← Domains & pillars

Reliability & OperationsView domain

APRF-21

Incident Readiness

Detect, contain, and learn from AI-specific production incidents.

Purpose

Prepare people, playbooks, and tooling to detect, contain, communicate, and learn from AI-specific incidents—including abuse, quality collapse, and unsafe tool actions.

Mandatory checks

Gate controls. Each check is pass/fail via artifact + pass condition. Expected from the annotated capability level when the system meets the minimum criticality tier.

  • INC-M1L3 · DefinedTier 2 · ProductionhybridE3

    Production AI systems must maintain incident playbooks for abuse, data leakage, bad agent/tool actions, and provider outage—each with a named owner and a review date within the last 12 months.

    Artifact
    Playbook set covering abuse, leakage, bad actions, and provider outage with owners + Review dates ≤12 months for each of the four playbooks
    Pass condition
    Four playbooks present (abuse, leakage, bad actions, provider outage), each with owner and review date ≤12 months (playbook evidence measuredAt ≤90 days). If no production AI system is in scope, score NOT_APPLICABLE.

    Why this control exists

    Threat map

    Ensure responders have prepared procedures for AI-specific incident classes.

    Threats mitigated

    Prompt InjectionData ExfiltrationHarmful Content GenerationAgent HijackingDenial of Service

    Protects

    UsersDataAvailabilitySafety

    MITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.

    AI incidents such as leakage through model output or a misbehaving agent do not fit conventional infrastructure playbooks, so responders improvise exactly when speed matters. Prepared procedures reduce time to containment; readiness itself maps to no adversary technique.

    Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.

  • INC-M2L3 · DefinedTier 2 · ProductionhybridE3

    On-call must be able to pause agents, disable tools, and roll back prompts or models, proven by a drill in the last 90 days completed within documented time budgets.

    Artifact
    Containment runbook covering pause agents, disable tools, and prompt/model rollback + Drill record ≤90 days showing all three actions within documented time budgets
    Pass condition
    Drill in last 90 days successfully demonstrated pause agents, disable tools, and roll back prompt/model within documented time budgets (drill evidence measuredAt ≤90 days). If no production agents, tools, or prompt/model release units exist, score NOT_APPLICABLE.

    Why this control exists

    Threat map

    Ensure on-call responders can actually stop AI activity during an incident.

    Threats mitigated

    Agent HijackingExcessive AgencyTool AbuseData ExfiltrationDenial of Wallet

    Protects

    RuntimeToolsDataCostAvailability

    MITRE: ATLAS AML.T0053 · ATLAS AML.T0086 · ATLAS AML.T0034.002

    Containment capability turns an active compromise into a bounded one by cutting tool invocation and agent execution mid-incident. It limits impact and exfiltration duration rather than preventing initial access.

    Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.

Evidence required

  • Playbooks and ownership roster
  • Alert configuration samples
  • Post-incident review examples (redacted)
More detailPhilosophy, failures, practices, validations, examples, crosswalks, and evolution

Engineering philosophy

Incidents will happen. Readiness means AI-aware detection, clear ownership, containment that includes pausing agents, and blameless learning that improves pillars.

Why it matters

Traditional SEV playbooks miss prompt injection campaigns, model outages, and agentic damage. Without AI-specific readiness, mean time to contain balloons.

Common failures

  • No playbook for model abuse or data leakage via chat
  • On-call cannot pause agents or roll back prompts
  • Customer communication templates ignore AI failure modes
  • No post-incident action tracking into evals and controls

Severity & risk

Severity
high
Impact if violated
Risk level
high
Typical residual risk (impact × likelihood)

Engineering best practices

  • Define SEV classifications for AI harm (financial, privacy, safety, reputation)
  • Preserve traces for forensics under legal hold when needed
  • Train support and on-call on AI failure literacy
  • Link incidents to eval fixtures so regressions cannot silently return

Automatic validations

  • Alert routing tests
  • Automated creation of incident tickets from critical signals
  • Verification that kill switches remain reachable

Manual validations

  • Tabletop exercises
  • After-action review quality checks

Examples

  • A spike in tool denials pages on-call; the agent is paused while injection is investigated
  • A postmortem adds a new adversarial eval case that fails CI until fixed

References

Crosswalks

  • MANAGE Manage

    NIST AI Risk Management Framework · supports

  • §10 Improvement

    ISO/IEC 42001 · aligns-with

  • C12.1 Request & Response Logging

    OWASP AI Application Security Verification Standard (AISVS) · supports

  • C12.2 Detection and Alerting

    OWASP AI Application Security Verification Standard (AISVS) · partial

  • C12.3 Model, Data, and Performance Drift Detection

    OWASP AI Application Security Verification Standard (AISVS) · partial

  • C12.4 Proactive Security Behavior Monitoring

    OWASP AI Application Security Verification Standard (AISVS) · partial

  • C12.5 Training Data & Model Lifecycle Audit

    OWASP AI Application Security Verification Standard (AISVS) · partial

  • CWE-778 Insufficient Logging

    OpenCRE (Open Common Requirements Enumeration) · aligns-with

  • CC7 System Operations

    SOC 2 Trust Services Criteria · evidence-for

  • Operational Excellence Operational Excellence

    AWS Well-Architected Framework · aligns-with

  • Reliability Reliability

    AWS Well-Architected Framework · aligns-with

Future evolution

Shared AI incident taxonomies and anonymized industry learning feeds.