Automated evaluator · 21 behavioral checks

Automated AI security testing within your approved scope.

Evil AI applies 21 non-destructive behavioral checks to an explicitly authorized AI application. Sensitive access stays under the customer's control throughout the assessment.

Assessment safeguards

Useful evidence without unnecessary exposure.

Each assessment is restricted to the approved system and test boundary. Credentials remain in the customer's environment, and only the evidence needed to evaluate a result is retained.

Evidence is sanitized, limited, and evaluated consistently. Potential findings are reviewed before they appear in a customer report or affect the Evil Score.

01

Approved scope only

Testing is limited to the system and boundaries authorized for the engagement.

02

No external side effects

Every test explicitly prohibits browsing, messaging, purchases, permission changes, and other actions.

03

Consistent evaluation

Each potential failure is evaluated against consistent criteria, with evidence retained for review.

04

Reviewed before reporting

Automated findings are reviewed before they appear in your report or affect your Evil Score.

Current coverage · 21 checks

What Evil AI tests.

These checks examine observable behavior in controlled scenarios. They do not cover every aspect of infrastructure, access control, data isolation, or user-interface disclosure. Results without sufficient evidence remain inconclusive.

Review coverage and exclusions →

01Direct instruction hierarchy overrideInstruction hierarchy · high
02Fake developer instructionInstruction hierarchy · high
03System prompt disclosure requestPrompt confidentiality · moderate
04Internal policy extractionPrompt confidentiality · moderate
05Credential and secret requestSensitive data · critical
06Cross-user data isolationSensitive data · critical
07Unapproved messaging actionTool safety · high
08Unapproved purchase actionTool safety · critical
09Untrusted retrieved instructionRAG injection · high
10Fake admin note in retrieved contentRAG injection · high
11Unbounded autonomous taskAgency controls · moderate
12Self-escalation requestAgency controls · critical
13User privilege impersonationAuthorization · high
14Scope expansion by promptAuthorization · high
15False action-completion claimAction truthfulness · moderate
16Sensitive-value echo disciplineSensitive data · low
17Hidden context: memory, state and tool responsesHidden context exposure · high
18Token and cost amplification policyResource limits · moderate
19Recursive expansion policyResource limits · moderate
20AI identity disclosure responseAI transparency · moderate
21Sexual exploitation content refusal policyHarmful-output refusal · critical

What the evaluator delivers

From controlled test to actionable finding.

The evaluator applies a traceable set of checks to an authorized system, captures the relevant response evidence, evaluates failures consistently, and prepares supported findings for review.

Current limits: the evaluator does not perform browser-based testing, open-ended crawling, complex authenticated journeys, or destructive testing. These exclusions remain visible in every assessment.