Purpose
Prepare people, playbooks, and tooling to detect, contain, communicate, and learn from AI-specific incidents—including abuse, quality collapse, and unsafe tool actions.
Mandatory checks
Gate controls. Each check is pass/fail via artifact + pass condition. Expected from the annotated capability level when the system meets the minimum criticality tier.
Production AI systems must maintain incident playbooks for abuse, data leakage, bad agent/tool actions, and provider outage—each with a named owner and a review date within the last 12 months.
- Artifact
- Playbook set covering abuse, leakage, bad actions, and provider outage with owners + Review dates ≤12 months for each of the four playbooks
- Pass condition
- Four playbooks present (abuse, leakage, bad actions, provider outage), each with owner and review date ≤12 months (playbook evidence measuredAt ≤90 days). If no production AI system is in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapEnsure responders have prepared procedures for AI-specific incident classes.
Threats mitigated
Prompt InjectionData ExfiltrationHarmful Content GenerationAgent HijackingDenial of ServiceProtects
UsersDataAvailabilitySafetyMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
AI incidents such as leakage through model output or a misbehaving agent do not fit conventional infrastructure playbooks, so responders improvise exactly when speed matters. Prepared procedures reduce time to containment; readiness itself maps to no adversary technique.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
On-call must be able to pause agents, disable tools, and roll back prompts or models, proven by a drill in the last 90 days completed within documented time budgets.
- Artifact
- Containment runbook covering pause agents, disable tools, and prompt/model rollback + Drill record ≤90 days showing all three actions within documented time budgets
- Pass condition
- Drill in last 90 days successfully demonstrated pause agents, disable tools, and roll back prompt/model within documented time budgets (drill evidence measuredAt ≤90 days). If no production agents, tools, or prompt/model release units exist, score NOT_APPLICABLE.
Why this control exists
Threat mapEnsure on-call responders can actually stop AI activity during an incident.
Threats mitigated
Agent HijackingExcessive AgencyTool AbuseData ExfiltrationDenial of WalletProtects
RuntimeToolsDataCostAvailabilityMITRE: ATLAS AML.T0053 · ATLAS AML.T0086 · ATLAS AML.T0034.002
Containment capability turns an active compromise into a bounded one by cutting tool invocation and agent execution mid-incident. It limits impact and exfiltration duration rather than preventing initial access.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Evidence required
- Playbooks and ownership roster
- Alert configuration samples
- Post-incident review examples (redacted)
Recommended checks
Strengthen posture beyond the gate. Same measurable structure; non-blocking unless elevated by organizational policy.
On-call should receive page-worthy alerts for at least two non-infra AI safety/quality signals, each with a documented threshold and owner, and a policy reviewed within 90 days.
- Artifact
- On-call alert policy export listing safety/quality pages + Last 90 days of triggered incidents or drill tickets for those pages
- Pass condition
- At least two non-infra signals (e.g. refusal-rate spike, eval-score drop, toxicity/jailbreak hit rate) page an on-call; each has a documented threshold and owner; policy reviewed ≤90 days ago (alert evidence measuredAt ≤90 days). If no production AI system is in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapPage responders on safety and quality signals, not only on infrastructure health.
Threats mitigated
Harmful Content GenerationJailbreakMisinformationProtects
UsersSafetyMITRE: ATLAS AML.T0054
Infrastructure-only alerting leaves safety and quality failures to be discovered by customers. Paging on refusal-rate and safety-signal anomalies gives responders an early indication that jailbreak or abuse is succeeding at scale.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
SEV-eligible AI incidents should receive a post-incident review with at least one tracked action mapped to an APRF pillar—or an explicit no-action rationale.
- Artifact
- Post-incident review template requiring APRF pillar mapping + Last-90-day coverage: SEV-eligible AI reviews with tracked actions or no-action rationale
- Pass condition
- 100% of SEV-eligible AI incidents in last 90 days have a review with ≥1 tracked action mapped to an APRF pillar or explicit “no action” rationale (review evidence measuredAt ≤90 days). If no SEV-eligible AI incidents exist in the last 90 days, score NOT_APPLICABLE.
Why this control exists
Threat mapConvert incident learning into owned, tracked control improvements.
Threats mitigated
RepudiationProtects
Audit TrailSafetyMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
Reviews reduce recurrence only when they produce owned actions that are tracked to completion. This is a continual improvement control with no adversary technique mapping.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Teams should maintain criteria that map AI event types to notify / no-notify decisions, and show a drill or real notification ≤12 months that followed those criteria with timestamps.
- Artifact
- Customer notification criteria for AI-related events + Last drill or real notification sample ≤12 months with timestamps
- Pass condition
- Criteria map event types (safety incident, widespread quality fail, data exposure) to notify / no-notify; last drill or incident ≤12 months followed the criteria with timestamps (notification evidence measuredAt ≤90 days). If no customer-facing or externally disclosed AI system is in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapDefine in advance when an AI-related event requires customer notification.
Threats mitigated
RepudiationSensitive Information DisclosureProtects
UsersAudit TrailMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
Notification criteria decided during an incident are decided badly and late. Pre-agreed thresholds make disclosure consistent and defensible; no adversary technique maps.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Teams should complete an AI-focused incident tabletop at least every 180 days and retain an after-action report with actions and owners.
- Artifact
- Tabletop plan for an AI-specific incident scenario + Dated after-action report ≤180 days with retained actions and owners
- Pass condition
- An AI-focused tabletop completed ≤180 days with retained actions and owners (tabletop evidence measuredAt ≤90 days). If no production AI system is in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapRehearse AI-specific incident response before a real event tests it.
Threats mitigated
Prompt InjectionData ExfiltrationAgent HijackingProtects
AvailabilitySafetyMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
Tabletops validate that AI playbooks, containment authority, and escalation paths work under time pressure. This is readiness assurance with no adversary technique mapping.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
More detailPhilosophy, failures, practices, validations, examples, crosswalks, and evolution
Engineering philosophy
Incidents will happen. Readiness means AI-aware detection, clear ownership, containment that includes pausing agents, and blameless learning that improves pillars.
Why it matters
Traditional SEV playbooks miss prompt injection campaigns, model outages, and agentic damage. Without AI-specific readiness, mean time to contain balloons.
Common failures
- No playbook for model abuse or data leakage via chat
- On-call cannot pause agents or roll back prompts
- Customer communication templates ignore AI failure modes
- No post-incident action tracking into evals and controls
Severity & risk
- Severity
- high
- Impact if violated
- Risk level
- high
- Typical residual risk (impact × likelihood)
Engineering best practices
- Define SEV classifications for AI harm (financial, privacy, safety, reputation)
- Preserve traces for forensics under legal hold when needed
- Train support and on-call on AI failure literacy
- Link incidents to eval fixtures so regressions cannot silently return
Automatic validations
- Alert routing tests
- Automated creation of incident tickets from critical signals
- Verification that kill switches remain reachable
Manual validations
- Tabletop exercises
- After-action review quality checks
Examples
- A spike in tool denials pages on-call; the agent is paused while injection is investigated
- A postmortem adds a new adversarial eval case that fails CI until fixed
References
Crosswalks
MANAGE Manage
NIST AI Risk Management Framework · supports
§10 Improvement
ISO/IEC 42001 · aligns-with
C12.1 Request & Response Logging
OWASP AI Application Security Verification Standard (AISVS) · supports
C12.2 Detection and Alerting
OWASP AI Application Security Verification Standard (AISVS) · partial
C12.3 Model, Data, and Performance Drift Detection
OWASP AI Application Security Verification Standard (AISVS) · partial
C12.4 Proactive Security Behavior Monitoring
OWASP AI Application Security Verification Standard (AISVS) · partial
C12.5 Training Data & Model Lifecycle Audit
OWASP AI Application Security Verification Standard (AISVS) · partial
CWE-778 Insufficient Logging
OpenCRE (Open Common Requirements Enumeration) · aligns-with
CC7 System Operations
SOC 2 Trust Services Criteria · evidence-for
Operational Excellence Operational Excellence
AWS Well-Architected Framework · aligns-with
Reliability Reliability
AWS Well-Architected Framework · aligns-with
Future evolution
Shared AI incident taxonomies and anonymized industry learning feeds.