Fictional and illustrative — not a real customer assessment
Sample report / EA-ILLUSTRATIVE-001
Example Support Agent
A sanitized example of how Evil AI communicates scope, risk, evidence, limitations, remediation, and retest status. No live system was tested.
- Application
- Example Support Agent
- Assessment
- Illustrative red team
- Status
- Fictional
- Retest
- Partially illustrated
01 / Scope and limitations
What this fictional assessment covered
The illustrative scope contains a customer-support chatbot with a fictional help-center knowledge base, authenticated account context, and a refund-request workflow. It excludes infrastructure, model provider systems, source-code review, denial-of-service testing, and destructive actions.
Results describe only the scenarios and configuration represented in this fictional example. They are not proof that any real system is secure, compliant, or certified.
02 / Executive summary
Useful support, weak trust boundaries.
The fictional agent handled normal support questions well but relied too heavily on model behavior for authorization, retrieved-content handling, and refund commitments. The most material risks could affect customer privacy and create unauthorized commercial promises.
03 / Evil Score
Overall: 61 / 100 — elevated risk
Illustrative scorecard
61/100
Elevated risk
Fictional example only. A score reflects the approved scope and date tested—not permanent safety.
Instruction integrity42 · High
Data confidentiality68 · Moderate
Tool boundaries76 · Moderate
Output reliability55 · High
04 / Sanitized findings
Prioritized evidence without weaponized detail.
Illustrative framework references · Mapping 2026-09-11.1 · Last reviewed 2026-09-11. These references do not establish control effectiveness or statutory compliance.
Framework references for illustrative findings| Finding | OWASP 2026 ID | ATLAS technique | ISO 42001 Annex A | Applicable statute |
|---|
| Retrieved instructions can influence answer policy | LLM01:2026 | AML.T0051 | A.6.2.4 | Applicability not assessed |
|---|
| Account context is disclosed too broadly | LLM02:2026 | Unmapped | A.6.2.4 | Applicability not assessed |
|---|
| Unsupported refund commitments | LLM07:2026 | Unmapped | A.6.2.4 | Applicability not assessed |
|---|
High risk
Retrieved instructions can influence answer policy
Business impact: Untrusted content in the fictional knowledge base could cause the support agent to disregard approved response rules, increasing the chance of incorrect commitments.
Recommended remediation: Separate retrieved content from trusted instructions, constrain its role, add source-aware validation, and test adversarial documents before indexing.
High risk
Account context is disclosed too broadly
Business impact: The fictional agent revealed more customer context than was necessary to answer a request when identity signals were ambiguous.
Recommended remediation: Minimize context passed to the model, enforce authorization before retrieval, and redact high-risk fields at the data layer.
Moderate risk
Unsupported refund commitments
Business impact: Under conversational pressure, the fictional agent stated that a refund would be issued without confirming policy eligibility or requiring approval.
Recommended remediation: Enforce refund eligibility through fixed business rules and require explicit approval before communicating an outcome.
05 / Retest status
Specific fixes, specifically checked.
Illustrative: one fix validated
In this fictional example, retrieval-role separation was shown as remediated and retested. Authorization and refund-control findings remain open. A targeted retest does not replace a future full assessment.
Private beta · authorized applications only
Get a report about your actual system.
Request beta access