Purpose
Define SLIs/SLOs for AI latency, availability, and quality signals; manage error budgets so AI features meet user experience and capacity targets.
Mandatory checks
Gate controls. Each check is pass/fail via artifact + pass condition. Expected from the annotated capability level when the system meets the minimum criticality tier.
Every journey marked critical must have documented numeric availability and latency-percentile SLO targets in a maintained catalog.
- Artifact
- SLO catalog listing critical AI journeys with numeric availability and latency targets + Coverage showing 100% of marked-critical journeys have both target types
- Pass condition
- 100% of journeys marked critical have availability % and latency percentile targets recorded in an SLO catalog (catalog evidence measuredAt ≤90 days). If no AI journey is marked critical, score NOT_APPLICABLE.
Why this control exists
Threat mapDefine the availability and latency commitment for critical AI user journeys.
Threats mitigated
Denial of ServiceProtects
AvailabilityUsersMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
SLOs define the baseline against which degradation, abuse impact, and incident severity are measured. This is a reliability control with no adversary technique mapping.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Production AI services must collect and make available operational metrics for latency, error rate, and at least one AI-specific quality or task-success indicator.
- Artifact
- Metric definitions or exporters covering latency, error rate, and ≥1 AI quality/task-success signal + Proof metrics are queryable/available for operational monitoring
- Pass condition
- Metrics for latency, error rate, and at least one AI-specific quality or task-success indicator are collected and available for operational monitoring (metrics evidence measuredAt ≤90 days). If no production AI services are in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapExpose latency, error, and AI quality metrics so degradation is measurable.
Threats mitigated
Denial of ServiceMisinformationProtects
AvailabilityUsersLogsMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
Metrics that are not emitted cannot be alerted on or investigated later. Instrumentation is the prerequisite for PERF-M3; no adversary technique maps directly.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Every critical AI journey SLO must have burn-rate (or equivalent) alert policies, and a test or documented fire must prove the notification path works.
- Artifact
- Burn-rate or SLO alert policies covering each critical AI journey SLO + Alert test or documented fire proving notification path
- Pass condition
- Alert policies exist for each critical journey SLO; an alert test or documented fire demonstrates the notification path works (alert evidence measuredAt ≤90 days). If no critical AI journeys with SLOs are in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapAlert when SLO burn indicates the system is degrading beyond agreed limits.
Threats mitigated
Denial of ServiceDenial of WalletProtects
AvailabilityCostMITRE: ATLAS AML.T0029
Sustained SLO burn is often the first visible symptom of resource-exhaustion abuse against an AI endpoint. Alerting turns that into an actionable signal; it is detective and does not itself restore capacity.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Evidence required
- SLO definitions for AI journeys
- Dashboard and alert configuration
- Recent burn-rate or incident examples (redacted)
Recommended checks
Strengthen posture beyond the gate. Same measurable structure; non-blocking unless elevated by organizational policy.
When a critical AI journey's error budget is exhausted, release velocity should be blocked or require explicit risk acceptance, with ≥1 gated event or drill in the last 90 days.
- Artifact
- Error-budget policy linking AI SLOs to release freezes or risk acceptance + Last budget-burn gated release or drill ≤90 days
- Pass condition
- When error budget for a critical AI journey is exhausted, release velocity is blocked or requires explicit risk acceptance; ≥1 gated event or drill in 90 days (gate evidence measuredAt ≤90 days). If no critical AI journeys with error budgets are in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapUse error budgets to slow release velocity when AI reliability is degrading.
Threats mitigated
Denial of ServiceProtects
AvailabilityUsersMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
Error budgets make the trade-off between shipping speed and reliability explicit instead of implicit. This is a delivery governance control with no adversary technique mapping.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Capacity and load tests should include adversarial long-prompt and multi-step agent-loop scenarios within 90 days, with p95 latency and error rate staying within SLO under documented concurrency.
- Artifact
- Load-test plan including long-prompt and agent-loop scenarios + Latest capacity test report ≤90 days showing p95/error within SLO at documented concurrency
- Pass condition
- Last capacity test ≤90 days includes adversarial long prompts and multi-step agent loops; p95 latency and error rate stay within SLO under documented concurrency (capacity evidence measuredAt ≤90 days). If no production AI paths with capacity risk are in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapValidate that the system withstands deliberately expensive prompts and agent loops.
Threats mitigated
Denial of ServiceDenial of WalletExcessive AgencyProtects
AvailabilityCostRuntimeMITRE: ATLAS AML.T0029 · ATLAS AML.T0034.001 · ATLAS AML.T0034.002 · ATLAS AML.T0046 · ATT&CK T1499
Long-context and recursive agent workloads consume far more capacity per request than typical traffic, which makes them efficient denial-of-service and cost-harvesting vectors. Capacity testing against these shapes is what establishes the limits that COST-M1 and AGN-M2 then enforce.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Each streaming AI surface should expose TTFT and inter-token latency SLIs, with alerts on documented thresholds and metric series retained ≥30 days.
- Artifact
- Streaming SLI definitions/dashboard for TTFT and inter-token latency + Alert config on documented thresholds + retention ≥30 days
- Pass condition
- TTFT and inter-token latency SLIs exist for each streaming AI surface; alerts fire on documented thresholds; series retained ≥30 days (streaming evidence measuredAt ≤90 days). If no streaming AI surfaces (SSE/WebSocket/token streams) are in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapMeasure streaming-specific latency so perceived responsiveness is not invisible to SLOs.
Threats mitigated
Denial of ServiceProtects
AvailabilityUsersMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
Aggregate request latency hides time-to-first-token and inter-token stalls that dominate the streaming user experience. This is a reliability measurement control with no adversary technique mapping.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
Near-real-time operational dashboards should visualize latency, error rate, throughput, resource utilization, and AI quality metrics for production services.
- Artifact
- Dashboard URL/config covering latency, error rate, throughput, resource utilization, and AI quality + Near-real-time refresh evidence (e.g. panel freshness ≤15 minutes)
- Pass condition
- Near-real-time operational dashboards visualize latency, error rate, throughput, resource utilization, and AI quality metrics for production services (dashboard evidence measuredAt ≤90 days). If no production AI services with ops metrics are in scope, score NOT_APPLICABLE.
Why this control exists
Threat mapPresent AI operational metrics in near real time so degradation is seen as it happens.
Threats mitigated
Denial of ServiceProtects
AvailabilityMITRE: no technique mapped — this control addresses governance or assurance rather than a specific adversary technique.
Delayed dashboards extend time to detection during a live degradation. Near real-time presentation shortens that gap; no adversary technique maps.
Informative threat context — mappings reduce exposure and do not guarantee mitigation; not certification.
More detailPhilosophy, failures, practices, validations, examples, crosswalks, and evolution
Engineering philosophy
If it is not measured with an SLO, it is not operationally owned. Stochastic systems still need latency and quality budgets.
Why it matters
Unbounded TTFT, streaming stalls, and silent quality burn destroy UX and hide regressions that “availability” alone will not catch.
Common failures
- No p95/p99 latency targets for model calls
- Quality treated as a one-time eval, not an online SLO
- No error budget for AI feature releases
- Capacity planning ignores bursty agent tool loops
Severity & risk
- Severity
- high
- Impact if violated
- Risk level
- medium
- Typical residual risk (impact × likelihood)
Engineering best practices
- Separate SLOs for AI features vs core non-AI paths
- Budget tokens and wall-clock together
- Correlate performance regressions with model/prompt version changes
Automatic validations
- Synthetic latency probes
- Burn-rate alerts
- CI performance budgets for critical paths where feasible
Manual validations
- Quarterly SLO review with product owners
- Post-incident SLO recalibration
Examples
- Chat p95 TTFT < 1.5s with error-budget policy that freezes prompt experiments when burned
- Agent task-success rate tracked as a quality SLO alongside HTTP availability
References
Crosswalks
MEASURE Measure
NIST AI Risk Management Framework · supports
§9 Performance evaluation
ISO/IEC 42001 · supports
LLM10 Unbounded Consumption
OWASP Top 10 for Large Language Model Applications · supports
C9.1 Execution Budgets, Loop Control, and Circuit Breakers
OWASP AI Application Security Verification Standard (AISVS) · supports
L5 Evaluation & Observability
CSA MAESTRO (Multi-Agentic Threat Model) · supports
A1 Availability
SOC 2 Trust Services Criteria · evidence-for
Reliability Reliability
AWS Well-Architected Framework · aligns-with
Performance Efficiency Performance Efficiency
AWS Well-Architected Framework · partial
Future evolution
Standard GenAI SLI catalogs (TTFT, tool-loop duration, citation accuracy) across ecosystems.