Authorization before testing
Ownership, boundary, exclusions, and stop conditions come before any test run.
Methodology
Repeatability is valuable only when the product also exposes confidence, coverage, limitations, and the boundary of machine judgment.
Ownership, boundary, exclusions, and stop conditions come before any test run.
Test suite, configuration, grading method, scoring rule, and report version remain traceable.
Low-confidence machine judgments do not become confirmed vulnerabilities.
Use a test environment, synthetic data, and only the access needed for the approved assessment wherever possible.
Reports distinguish tested, passed, failed, inconclusive, not applicable, and not tested.
Manual judgment is optional, separately scoped, and attributed to the specialist providing it.
Layer 1
Each assessment selects the behaviors and customer-defined policies relevant to the approved scope. Evil AI runs non-destructive checks, captures the evidence needed to understand each result, and prepares supported findings for reporting.
Testing begins only after access and scope are confirmed. Submitting an inquiry does not authorize or initiate an assessment.
Layer 2
Findings retain their test-suite version, grading method, evidence, confidence, and limitations. Duplicate manifestations should not inflate risk, and historical reports do not silently change with new scoring rules.
Layer 3
Complex architectures, consequential findings, disputed results, and remediation decisions may warrant specialist review. That work is separately scoped and subject to finding a qualified independent specialist.
Alignment
Alignment does not imply endorsement, accreditation, compliance, or certification.
View dated framework versions, check-by-check coverage, and exclusions →
Private beta · authorized applications only