{"$schema":"https://stackrail.io/aprf/spec-schema/0.7","id":"aprf","name":"AI Production Readiness Framework","question":"Can this AI application safely operate in production?","governance":{"version":"0.10.0","status":"working-draft","releaseName":"CHANGELOG + CI validators, Regulated Assess, WA/SLSA crosswalks, expanded N/A","publishedAt":"2026-07-25","publisher":{"name":"StackRail","role":"working-draft-publisher","url":"https://stackrail.io"},"intendedSteward":{"model":"neutral-working-group","description":"Multi-stakeholder technical oversight (vendors, practitioners, OWASP/CNCF-adjacent participants) with public RFCs — not a single commercial owner.","currentPhase":"working-draft","processPath":"/aprf/rfc/"},"versioning":{"scheme":"semver","rules":["MAJOR: breaking changes to domain IDs, check IDs, or gate semantics","MINOR: new domains, pillars, or checks; non-breaking field additions","PATCH: editorial clarifications to prose, artifacts, or pass conditions"],"stableIdPolicy":"Pillar IDs (APRF-NN) and check IDs (e.g. AUTHN-M1) are immutable once published in a MINOR+. Deprecate; do not reuse."},"schemaVersioning":{"frameworkSemVer":"APRF_GOVERNANCE.version — semantic version of the normative framework, check catalog, profiles, and scoring semantics.","specSchemaId":"https://stackrail.io/aprf/spec-schema/0.7 — JSON Schema for the machine-readable APRF document shape; may lag framework SemVer when only content changes.","attestationSchemaId":"https://stackrail.io/aprf/attestation-schema/0.6 — JSON Schema for aprf-self-attestation exports.","note":"Do not equate schema path versions (0.7 / 0.6) with APRF framework SemVer (e.g. 0.9.1). Validators and conformance claims should cite both."},"deprecation":{"policy":"Deprecated checks and pillars remain in the spec for at least one MINOR cycle (N−1), marked deprecated, with a replacement ID.","supportWindow":"N-1 minor versions"},"extension":{"lenses":"Formal lenses (RAG, Agents, Voice, Coding agents) add mandatory checks under existing pillars. Claim lenses in assessments and attestations when those system types apply.","customNamespace":"Organizations MAY add private checks with IDs prefixed x- (e.g. x-ACME-M1). Custom checks MUST NOT claim APRF conformance by themselves."},"certification":{"levels":[{"id":"self-attestation","name":"Self-attestation","description":"Organization publishes assessment results + evidence index against a pinned APRF version. No third-party validation."},{"id":"third-party","name":"Third-party assessment","description":"Independent assessor verifies gate blockers and samples evidence. Intended for Tier 3 and regulated use."}],"badges":"Conformance statements MUST cite APRF version, criticality tier, attained capability level, profile (if any), and list any open blockers. No single percentage badge.","referenceAssessment":{"path":"/aprf/assess/","exportFormat":"aprf-self-attestation","schema":"https://stackrail.io/aprf/attestation-schema/0.6","description":"Core or Regulated profile quiz with optional lenses: one question per gate, on-screen gate result, client download of full aprf-self-attestation JSON, optional save (POST /api/aprf/submit returns the full attestation) + email summary. AGN/HUM/TOL/MEM gates offer formal N/A. Unanswered checks count as failed."}},"compatibility":{"crosswalks":["NIST AI RMF","ISO/IEC 42001","SOC 2 Trust Services Criteria (evidence reuse)","OWASP LLM Top 10","AWS Well-Architected (conceptual)","SLSA / supply-chain"],"note":"Machine-readable maps for NIST AI RMF, ISO/IEC 42001, OWASP LLM Top 10, SOC 2, AWS Well-Architected, and SLSA ship in the APRF spec under crosswalks[]. Informative alignment only — not certification."},"stewardshipRef":"aprf-stewardship","draftCapabilityFloor":3},"stewardship":{"id":"aprf-stewardship","charter":{"purpose":"Evolve APRF into a ratifiable, vendor-neutral production-readiness standard for AI systems through open technical consensus—not through a single commercial owner.","principles":["Neutrality: no single vendor or publisher holds permanent veto over normative content.","Open process: substantive changes go through public RFCs with recorded decisions.","Evidence over marketing: checks stay measurable (artifact + pass condition); no vanity scores.","Stable identifiers: pillar and check IDs are immutable once published in a MINOR+; deprecate, do not reuse.","Separability: StackRail may publish implementations and services without owning the standard long-term.","Inclusive participation: vendors, practitioners, researchers, and OWASP/CNCF-adjacent contributors welcome."],"nonGoals":["Replacing NIST AI RMF, ISO/IEC 42001, SOC 2, or OWASP as peer frameworks.","Granting certification authority to the working-draft publisher by default.","Locking the taxonomy behind proprietary tooling or paywalled normative text."]},"phases":[{"id":"working-draft","name":"Working draft","summary":"StackRail publishes the draft, hosts the machine-readable spec, and accepts RFCs. Normative intent is public; stewardship is not yet transferred.","current":true},{"id":"interim-advisory","name":"Interim advisory board","summary":"Multi-org advisors review RFCs and release plans. Publisher still ships releases but commits to advisory consensus for MAJOR/MINOR.","current":false},{"id":"neutral-working-group","name":"Neutral working group","summary":"Standing technical WG owns the roadmap, RFC decisions, and release approvals. Publisher becomes one participant among peers.","current":false},{"id":"foundation-home","name":"Foundation / SDO home","summary":"Optional later home under a foundation or standards body with IP assignment, trademark, and formal membership rules.","current":false}],"interimAdvisory":{"status":"recruiting","targetSeats":5,"minimumIndependentOrgs":3,"seats":[{"category":"Independent security / AppSec practitioner"},{"category":"AI platform or model vendor (non-publisher)"},{"category":"Production AI operator (practitioner)"},{"category":"Standards / SDO-adjacent participant (e.g. OWASP/CNCF)"},{"category":"Researcher or academic"}],"openCall":"StackRail is recruiting an interim advisory board (≥3 independent organizations) before MAJOR/MINOR releases require advisory consensus. Volunteer via the interim contact; a seat does not grant commercial preference or veto over peer RFCs.","decisionLog":"RFC decisions and rationales are published on /aprf/rfc/ and in the machine-readable `rfcs` index in /aprf/spec/. Until seats are filled, working-draft quorum applies: publisher decides after the review window with public rationale.","credibilityNote":"Single-vendor publication is temporary and explicit. Normative text and the machine-readable catalog remain freely available; transfer triggers and requirements are published; certification programs (if any) stay separate from normative stewardship."},"transfer":{"triggers":["At least three independent organizations actively reviewing RFCs for two consecutive MINOR cycles.","Published interim charter signed by advisory participants.","Public RFC backlog with decisions recorded for ≥90 days.","Willingness of a neutral host (foundation, consortium, or multi-party WG) to accept stewardship."],"requirements":["Machine-readable APRF spec remains freely available under a permissive documentation license.","Check and pillar ID stability policy preserved.","StackRail (or any implementer) retains rights to build products against APRF without preferential gatekeeping.","Certification programs, if any, are separated from normative stewardship."],"ipIntent":"Normative APRF text and machine-readable catalogs are intended to transfer with stewardship. Product trademarks and commercial services remain with their owners."},"participation":{"who":["AI platform and model vendors","Practitioners running production AI systems","Security, SRE, and compliance specialists","Researchers and standards practitioners","Adjacent communities (e.g. OWASP GenAI, CNCF security TAG participants)"],"how":["File an RFC using the template and checklist.","Comment on open RFCs during the review window.","Propose lenses, crosswalks, or editorial fixes via PATCH/MINOR RFCs.","Volunteer for an interim advisory seat (open call while status is recruiting)."],"contact":{"emailHint":"prasoonanand@ymail.com (interim; transfers with steward)","processPath":"/aprf/rfc/","machineReadable":"stewardship in /aprf/spec/"}},"rfc":{"name":"APRF Request for Comments","numbering":"APRF-RFC-NNNN (zero-padded, monotonic; 0000 is the template)","minimumReviewDays":14,"decisionQuorum":"Working-draft phase: publisher decides after review window, with public rationale. Interim advisory+: majority of active advisors; ties escalate to published deadlock procedure.","stages":[{"id":"draft","name":"Draft","summary":"Author iterating privately or in a PR; not yet open for formal review."},{"id":"proposed","name":"Proposed","summary":"Submitted with checklist complete; awaiting scheduling into review."},{"id":"in-review","name":"In review","summary":"Public comment open for at least the minimum review window."},{"id":"accepted","name":"Accepted","summary":"Will ship in a named SemVer release; implementation PR linked."},{"id":"rejected","name":"Rejected","summary":"Will not ship as proposed; rationale recorded; may be revised and resubmitted."},{"id":"withdrawn","name":"Withdrawn","summary":"Author withdrew before decision."},{"id":"superseded","name":"Superseded","summary":"Replaced by a later RFC."}],"submissionChecklist":["Problem statement and who is affected","Proposed change (pillars, checks, profiles, lenses, crosswalks, scoring, or governance)","SemVer impact: MAJOR / MINOR / PATCH (with rationale)","Backward compatibility and deprecation plan if IDs change","Evidence that checks remain measurable (artifact + pass condition)","Crosswalk impact (if any)","Security / safety considerations","Open questions for reviewers"],"decisionCriteria":["Improves production readiness signal without diluting gate semantics","Does not introduce vanity scoring or unauditable prose-only controls","Fits taxonomy (prefer lenses/profiles over new top-level domains)","Preserves ID stability or provides a deprecation path","Feasible for adopters at the claimed criticality tier"],"template":{"id":"APRF-RFC-0000","fields":["Title","Status","Author(s)","Created","SemVer impact","Problem","Proposal","Alternatives considered","Compatibility","Security considerations","Open questions","Checklist"]}}},"rfcs":[{"id":"APRF-RFC-0001","number":1,"title":"Establish working-draft RFC process and public Open RFCs list","status":"in-review","created":"2026-07-24","semverImpact":"MINOR","summary":"Editorial RFC that ratifies the stewardship RFC stages on the site, publishes this index, and records the first open review window—no taxonomy changes.","slug":"0001-working-draft-rfc-process","href":"/aprf/rfc/0001-working-draft-rfc-process/","markdownPath":"/aprf/rfc/0001-working-draft-rfc-process.md"}],"scope":{"vendors":["OpenAI","Anthropic","Gemini","Llama","DeepSeek","Mistral","Self-hosted models"],"systemTypes":["Chatbots","AI Agents","MCP Servers","A2A Systems","RAG","Multi-agent systems","Coding Agents","Voice AI","Autonomous systems"],"peerStandards":["AWS Well-Architected Framework","OWASP Top 10","CIS Benchmarks","Google SRE Workbook","NIST AI RMF","ISO/IEC 42001"]},"taxonomy":{"domains":[{"id":"security","name":"Security","summary":"Adversarial resistance, identity, authorization, secrets, tool mediation, supply chain, and hardened runtime for AI systems.","pillarSlugs":["ai-security","authentication","authorization","secrets","tool-safety","supply-chain","infrastructure"]},{"id":"safety","name":"Safety & Responsible AI","summary":"Harm prevention, content safety, fairness, and transparency—NIST trustworthiness characteristics distinct from adversarial security.","pillarSlugs":["safety-responsible-ai","explainability"]},{"id":"data","name":"Data","summary":"Privacy, corpus and index governance, data quality, and memory integrity across AI pipelines.","pillarSlugs":["data-privacy","data-governance","memory-management"]},{"id":"model-lifecycle","name":"Model & Prompt Lifecycle","summary":"Model selection and versioning, prompt and context as production artifacts, and continuous evaluation gates.","pillarSlugs":["model-governance","prompt-engineering","context-engineering","evaluation"]},{"id":"agents","name":"Agents & Autonomy","summary":"Agent charters, autonomy limits, A2A trust, and human oversight for high-impact actions.","pillarSlugs":["agent-governance","human-approval"]},{"id":"reliability","name":"Reliability & Operations","summary":"Observability, performance SLOs, graceful degradation and continuity, change management with rollback, and incident readiness.","pillarSlugs":["observability","performance-slo","reliability-continuity","change-management","incident-readiness"]},{"id":"cost","name":"Cost","summary":"Spend bounds, attribution, caching, routing, and denial-of-wallet controls for AI workloads.","pillarSlugs":["cost-optimization"]},{"id":"governance","name":"Governance & Compliance","summary":"Organizational AI policy, ownership, risk acceptance, and auditable evidence of controls—without equating compliance with readiness.","pillarSlugs":["organizational-governance","compliance"]}],"crossCutting":{"id":"cross-cutting","name":"Cross-cutting concerns","summary":"Concerns that apply across every domain. They are not peer domains; they enable safe delivery of all other pillars.","pillarSlugs":["platform-engineering"]}},"maturity":{"capabilityLevels":[{"level":1,"name":"Initial","summary":"Ad-hoc practices. Works in demos; controls are informal or absent. Acceptable only for Tier 0 sandboxes.","entryCriteria":["No claim of production readiness","Changes may be undocumented console edits","Failures do not trigger organizational incident process","Spend and rate limits are owner-watched at best"]},{"level":2,"name":"Managed","summary":"Basic repeatable controls: auth, logging, spend caps, secret hygiene. Suitable baseline for Tier 1 (internal) systems.","entryCriteria":["Authenticated access for non-public surfaces","Request/cost logging exists","Hard spend or rate ceilings prevent runaway usage","Secrets are not hardcoded in production prompts or clients"]},{"level":3,"name":"Defined","summary":"Documented, enforced controls across critical domains. Mandatory checks for Tier 2 (customer production) are met.","entryCriteria":["Promotion path and rollback for prompts, models, and tools","Server-side authz, tool mediation, and eval gates on releases","Traces reconstruct model → tool → outcome paths","AI-specific incident playbooks and named owners exist"]},{"level":4,"name":"Quantitatively Managed","summary":"SLOs, audited evidence, formal governance, and measured quality/safety signals. Expected floor for many Tier 3 systems.","entryCriteria":["Published SLOs for latency, availability, and at least one quality/safety signal","Auditable evidence packs for critical controls","Change control and ownership across teams","Privacy/compliance obligations mapped to technical controls"]},{"level":5,"name":"Optimizing","summary":"Continuous improvement: automated gates, dual-control for irreversible actions, chaos/continuity drills, near-zero blast-radius ambiguity.","entryCriteria":["Dual-control (or equivalent) for irreversible high-impact actions","Continuous evaluation blocks unsafe releases automatically","Continuity and failure drills on a defined cadence with recorded results","Blast radius for every tool, agent, and data path is documented and enforced"]}],"criticalityTiers":[{"tier":0,"name":"Sandbox","summary":"Non-customer experiments. Failures are contained to the builder’s environment.","requiredCapability":1,"examples":["Hackathon prototype","Local notebook agent","Throwaway prompt playground"]},{"tier":1,"name":"Internal","summary":"Employee-only or tightly gated internal use. Organizational impact only.","requiredCapability":2,"examples":["Internal knowledge assistant","Employee coding agent on non-prod systems"]},{"tier":2,"name":"Production","summary":"Customer- or partner-facing. Material product, privacy, or financial impact if it fails.","requiredCapability":3,"examples":["Customer support chatbot","SaaS RAG feature","MCP tools touching customer data"]},{"tier":3,"name":"Mission Critical","summary":"High blast radius: safety, regulated decisions, large financial movement, or critical infrastructure dependency.","requiredCapability":4,"examples":["Medical triage assistant","Autonomous trading or payments agent","Life-safety or critical-infrastructure copilots"]}]},"profiles":[{"id":"aprf-profile-core","name":"Core (Tier 2 Production)","summary":"Minimum mandatory gates for customer- or partner-facing AI. Pass this gate before claiming production readiness. Tier 3 and regulated systems must still assess the Regulated profile or full catalog.","targetCriticality":2,"targetCapability":3,"mandatoryCheckIds":["AUTHN-M1","AUTHN-M2","AUTHZ-M1","AUTHZ-M2","SEC2-M1","SEC2-M2","SEC-M1","SEC-M3","TOL-M1","TOL-M2","TOL-M3","SCI-M2","INF-M1","SAF-M1","SAF-M2","SAF-M3","PRI-M1","PRI-M3","MEM-M1","PRM-M1","PRM-M2","MOD-M1","EVL-M1","EVL-M2","AGN-M2","HUM-M1","HUM-M3","OBS-M1","OBS-M2","PERF-M1","REL-M1","REL-M2","DEP-M1","CHG-M1","CHG-M3","INC-M1","INC-M2","COST-M1","COST-M3","ORG-M2"],"rationale":["Identity and authorization before any customer traffic (AUTHN/AUTHZ).","Secrets and injection/tool mediation to prevent common AI incidents (SEC/TOL).","Safety policy + eval gates so quality/harm regressions cannot silently ship (SAF/EVL).","Pinned prompts/models with promotion evidence (PRM/MOD).","Observability, timeouts, degraded mode, and tested rollback (OBS/REL/CHG).","Spend ceilings to prevent denial-of-wallet (COST).","Named owners so gates have stewards (ORG-M2)."]},{"id":"aprf-profile-regulated","name":"Regulated (Tier 3)","summary":"Core Profile plus Tier-3-only mandatories for mission-critical or regulated AI (residency/DPIA, fairness, dual control, signed supply chain, chaos/continuity, independent assessment). Target capability Level 5.","targetCriticality":3,"targetCapability":5,"mandatoryCheckIds":["AUTHN-M1","AUTHN-M2","AUTHZ-M1","AUTHZ-M2","SEC2-M1","SEC2-M2","SEC-M1","SEC-M3","TOL-M1","TOL-M2","TOL-M3","SCI-M2","INF-M1","SAF-M1","SAF-M2","SAF-M3","PRI-M1","PRI-M3","MEM-M1","PRM-M1","PRM-M2","MOD-M1","EVL-M1","EVL-M2","AGN-M2","HUM-M1","HUM-M3","OBS-M1","OBS-M2","PERF-M1","REL-M1","REL-M2","DEP-M1","CHG-M1","CHG-M3","INC-M1","INC-M2","COST-M1","COST-M3","ORG-M2","REL-M6","ORG-M3","CMP-M2","SEC-M5","AUTHN-M4","AUTHZ-M4","TOL-M5","SCI-M4","INF-M4","SAF-M4","EXP-M4","PRI-M4","PRI-M5","MEM-M4","EVL-M4","HUM-M4","REL-M7","REL-M8","CHG-M4","INC-M4","ORG-M4"],"rationale":["Includes every Core gate — regulated systems must still clear production minimums.","Adds residency/DPIA, fairness, dual control, and signed admission for regulated blast radius.","Requires chaos/continuity drills, independent assessment sampling, and automated quality rollback triggers.","Target capability Level 4 — Tier 3 is not Core with a different label."]}],"lenses":[{"id":"aprf-lens-rag","name":"RAG","summary":"Retrieval-augmented generation: corpus ownership, context labeling, memory isolation, and retrieval-quality gates.","appliesTo":["RAG","Chatbots","AI Agents"],"recommendedBaseProfileId":"aprf-profile-core","targetCapability":3,"additionalMandatoryCheckIds":["DG-M1","DG-M2","DG-M3","CTX-M1","CTX-M2","CTX-M3","MEM-M1","MEM-M2","MEM-M3","PRI-M1","EVL-M1","EVL-M2","SEC-M1","OBS-M1"],"rationale":["Retrieval corpora need owners, versioning, and promotion controls (DG).","Retrieved content must be sized, labeled, and access-controlled in context (CTX).","Vector/memory stores inherit tenant isolation and retention (MEM).","Eval gates must cover retrieval quality and grounding regressions (EVL)."]},{"id":"aprf-lens-agents","name":"Agents","summary":"Autonomous and tool-using agents: charters, step budgets, tool mediation, human gates, and kill switches.","appliesTo":["AI Agents","Multi-agent systems","MCP Servers","A2A Systems","Coding Agents","Autonomous systems"],"recommendedBaseProfileId":"aprf-profile-core","targetCapability":3,"additionalMandatoryCheckIds":["AGN-M1","AGN-M2","AGN-M3","AGN-M4","TOL-M1","TOL-M2","TOL-M3","TOL-M4","HUM-M1","HUM-M2","HUM-M3","AUTHZ-M1","AUTHZ-M2","AUTHN-M2","COST-M3","OBS-M1","REL-M1","REL-M3"],"rationale":["Every production agent needs a charter and hard step/time limits (AGN).","Tools fail closed with allowlists and schema validation (TOL).","High-impact actions require non-bypassable human approval (HUM).","Agent identities stay least-privilege; loops cannot burn unbounded spend (AUTHZ/COST)."]},{"id":"aprf-lens-voice","name":"Voice","summary":"Voice and telephony AI: session identity, privacy of recordings, latency SLOs, safety, and escalation paths.","appliesTo":["Voice AI","Chatbots"],"recommendedBaseProfileId":"aprf-profile-core","targetCapability":3,"additionalMandatoryCheckIds":["AUTHN-M1","AUTHN-M2","PRI-M1","PRI-M3","OBS-M1","OBS-M3","PERF-M1","PERF-M2","REL-M1","REL-M2","SAF-M1","SAF-M2","HUM-M1","INC-M1","COST-M1","TOL-M1"],"rationale":["Telephony sessions must authenticate before privileged tools (AUTHN).","Call audio and transcripts are sensitive personal data (PRI).","Latency and degraded mode matter more under real-time constraints (PERF/REL).","Safety refusals and human escalation must work on voice channels (SAF/HUM)."]},{"id":"aprf-lens-coding","name":"Coding agents","summary":"IDE and repo-connected coding agents: sandboxing, secret hygiene, tool allowlists, supply chain, and human gates for destructive changes.","appliesTo":["Coding Agents","AI Agents","MCP Servers"],"recommendedBaseProfileId":"aprf-profile-core","targetCapability":3,"additionalMandatoryCheckIds":["AGN-M1","AGN-M3","AGN-M4","TOL-M4","HUM-M2","SEC-M2","SEC-M4","SCI-M1","SCI-M3","PRM-M3","MOD-M2","AUTHZ-M3","DX-M1","DX-M2","INF-M2"],"rationale":["Coding agents need explicit charters, kill switches, and peer auth for multi-agent hops (AGN).","Shell/file tools require schema validation and fail-closed allowlists (TOL).","Repo and cloud credentials must stay out of prompts; model path stays bounded (SEC).","Dependency and model supply chain plus platform sandboxes reduce blast radius (SCI/DX/INF)."]}],"crosswalks":[{"id":"nist-ai-rmf","name":"NIST AI Risk Management Framework","peerVersion":"1.0 (2023)","url":"https://www.nist.gov/itl/ai-risk-management-framework","disclaimer":"Informative alignment only. Does not constitute certification, accreditation, or official endorsement.","controls":[{"id":"nist-ai-rmf:govern","ref":"GOVERN","title":"Govern","summary":"Culture, policies, accountability, and continuous improvement for AI risk."},{"id":"nist-ai-rmf:map","ref":"MAP","title":"Map","summary":"Context, intended use, impacts, and risk categorization."},{"id":"nist-ai-rmf:measure","ref":"MEASURE","title":"Measure","summary":"Quantitative and qualitative assessment of AI risks and trustworthiness."},{"id":"nist-ai-rmf:manage","ref":"MANAGE","title":"Manage","summary":"Prioritize, respond to, and recover from AI risks in operation."},{"id":"nist-ai-rmf:safe","ref":"Safe","title":"Safe","summary":"Trustworthiness characteristic — avoid harmful outcomes under intended use."},{"id":"nist-ai-rmf:secure-resilient","ref":"Secure & Resilient","title":"Secure and Resilient"},{"id":"nist-ai-rmf:explainable","ref":"Explainable","title":"Explainable and Interpretable"},{"id":"nist-ai-rmf:privacy","ref":"Privacy-Enhanced","title":"Privacy-Enhanced"},{"id":"nist-ai-rmf:fair","ref":"Fair","title":"Fair — Harmful Bias Managed"},{"id":"nist-ai-rmf:accountable","ref":"Accountable","title":"Accountable and Transparent"}],"mappings":[{"peerControlId":"nist-ai-rmf:govern","aprfPillarSlugs":["organizational-governance","compliance","model-governance","human-approval"],"aprfCheckIds":["ORG-M1","ORG-M2","ORG-M3","CMP-M1","HUM-M1"],"relation":"supports"},{"peerControlId":"nist-ai-rmf:map","aprfPillarSlugs":["model-governance","data-governance","context-engineering","explainability"],"aprfCheckIds":["MOD-M1","DG-M1","CTX-M1","EXP-M1"],"relation":"supports"},{"peerControlId":"nist-ai-rmf:measure","aprfPillarSlugs":["evaluation","observability","performance-slo","safety-responsible-ai"],"aprfCheckIds":["EVL-M1","EVL-M2","OBS-M1","PERF-M1","SAF-M2"],"relation":"supports"},{"peerControlId":"nist-ai-rmf:manage","aprfPillarSlugs":["incident-readiness","change-management","reliability-continuity","tool-safety","agent-governance"],"aprfCheckIds":["INC-M1","INC-M2","CHG-M1","REL-M1","TOL-M1","AGN-M2"],"relation":"supports"},{"peerControlId":"nist-ai-rmf:safe","aprfPillarSlugs":["safety-responsible-ai","evaluation","human-approval"],"aprfCheckIds":["SAF-M1","SAF-M2","SAF-M3","HUM-M3"],"relation":"aligns-with"},{"peerControlId":"nist-ai-rmf:secure-resilient","aprfPillarSlugs":["ai-security","authentication","authorization","secrets","tool-safety","supply-chain","infrastructure","reliability-continuity"],"relation":"aligns-with"},{"peerControlId":"nist-ai-rmf:explainable","aprfPillarSlugs":["explainability","observability"],"aprfCheckIds":["EXP-M1","EXP-M2","OBS-M2"],"relation":"aligns-with"},{"peerControlId":"nist-ai-rmf:privacy","aprfPillarSlugs":["data-privacy","memory-management","context-engineering"],"aprfCheckIds":["PRI-M1","PRI-M3","MEM-M1"],"relation":"aligns-with"},{"peerControlId":"nist-ai-rmf:fair","aprfPillarSlugs":["safety-responsible-ai","evaluation","data-governance"],"aprfCheckIds":["SAF-M3","EVL-M2","DG-M2"],"relation":"partial","note":"APRF covers fairness eval gates; organizational bias programs may need extra controls."},{"peerControlId":"nist-ai-rmf:accountable","aprfPillarSlugs":["organizational-governance","human-approval","explainability","change-management"],"aprfCheckIds":["ORG-M2","HUM-M1","EXP-M3","CHG-M3"],"relation":"supports"}]},{"id":"iso-42001","name":"ISO/IEC 42001","peerVersion":"2023 (clause-level conceptual)","url":"https://www.iso.org/standard/81230.html","disclaimer":"Informative alignment only. Does not constitute certification, accreditation, or official endorsement. Not a substitute for a certified AI management system.","controls":[{"id":"iso-42001:4","ref":"§4","title":"Context of the organization"},{"id":"iso-42001:5","ref":"§5","title":"Leadership"},{"id":"iso-42001:6","ref":"§6","title":"Planning"},{"id":"iso-42001:7","ref":"§7","title":"Support"},{"id":"iso-42001:8","ref":"§8","title":"Operation"},{"id":"iso-42001:9","ref":"§9","title":"Performance evaluation"},{"id":"iso-42001:10","ref":"§10","title":"Improvement"},{"id":"iso-42001:a","ref":"Annex A","title":"AI system controls (selected themes)","summary":"Mapped thematically to APRF domains — not clause-by-clause Annex A enumeration."}],"mappings":[{"peerControlId":"iso-42001:4","aprfPillarSlugs":["organizational-governance","compliance","data-governance"],"aprfCheckIds":["ORG-M1","CMP-M1","DG-M1"],"relation":"aligns-with"},{"peerControlId":"iso-42001:5","aprfPillarSlugs":["organizational-governance","human-approval"],"aprfCheckIds":["ORG-M2","HUM-M1"],"relation":"supports"},{"peerControlId":"iso-42001:6","aprfPillarSlugs":["organizational-governance","safety-responsible-ai","evaluation","cost-optimization"],"aprfCheckIds":["ORG-M3","SAF-M1","EVL-M1","COST-M1"],"relation":"partial","note":"Risk/objectives planning spans org process plus safety and eval gates."},{"peerControlId":"iso-42001:7","aprfPillarSlugs":["platform-engineering","secrets","infrastructure"],"aprfCheckIds":["DX-M1","SEC2-M1","INF-M1"],"relation":"partial","note":"Competence/awareness are org processes; APRF emphasizes platform & secrets support."},{"peerControlId":"iso-42001:8","aprfPillarSlugs":["prompt-engineering","model-governance","change-management","tool-safety","agent-governance","data-privacy"],"relation":"supports"},{"peerControlId":"iso-42001:9","aprfPillarSlugs":["evaluation","observability","performance-slo","compliance"],"aprfCheckIds":["EVL-M1","EVL-M2","OBS-M1","PERF-M1","CMP-M2"],"relation":"supports"},{"peerControlId":"iso-42001:10","aprfPillarSlugs":["incident-readiness","organizational-governance","change-management"],"aprfCheckIds":["INC-M2","ORG-M3","CHG-M1"],"relation":"aligns-with"},{"peerControlId":"iso-42001:a","aprfPillarSlugs":["ai-security","safety-responsible-ai","data-privacy","data-governance","model-governance","human-approval","supply-chain"],"relation":"partial","note":"Annex A themes → APRF security, safety, data, model, human oversight, supply chain."}]},{"id":"owasp-llm-top-10","name":"OWASP Top 10 for Large Language Model Applications","peerVersion":"2025","url":"https://genai.owasp.org/llm-top-10/","disclaimer":"Informative alignment only. Does not constitute certification, accreditation, or official endorsement.","controls":[{"id":"owasp-llm:01","ref":"LLM01","title":"Prompt Injection"},{"id":"owasp-llm:02","ref":"LLM02","title":"Sensitive Information Disclosure"},{"id":"owasp-llm:03","ref":"LLM03","title":"Supply Chain"},{"id":"owasp-llm:04","ref":"LLM04","title":"Data and Model Poisoning"},{"id":"owasp-llm:05","ref":"LLM05","title":"Improper Output Handling"},{"id":"owasp-llm:06","ref":"LLM06","title":"Excessive Agency"},{"id":"owasp-llm:07","ref":"LLM07","title":"System Prompt Leakage"},{"id":"owasp-llm:08","ref":"LLM08","title":"Vector and Embedding Weaknesses"},{"id":"owasp-llm:09","ref":"LLM09","title":"Misinformation"},{"id":"owasp-llm:10","ref":"LLM10","title":"Unbounded Consumption"}],"mappings":[{"peerControlId":"owasp-llm:01","aprfPillarSlugs":["ai-security","prompt-engineering","tool-safety"],"aprfCheckIds":["SEC-M1","SEC-M3","PRM-M1","TOL-M1"],"relation":"supports"},{"peerControlId":"owasp-llm:02","aprfPillarSlugs":["data-privacy","context-engineering","secrets","memory-management"],"aprfCheckIds":["PRI-M1","PRI-M3","SEC2-M1","MEM-M1"],"relation":"supports"},{"peerControlId":"owasp-llm:03","aprfPillarSlugs":["supply-chain","model-governance","infrastructure"],"aprfCheckIds":["SCI-M1","SCI-M2","MOD-M1"],"relation":"supports"},{"peerControlId":"owasp-llm:04","aprfPillarSlugs":["data-governance","memory-management","model-governance","evaluation"],"aprfCheckIds":["DG-M2","MEM-M2","MOD-M2","EVL-M1"],"relation":"supports"},{"peerControlId":"owasp-llm:05","aprfPillarSlugs":["ai-security","tool-safety","safety-responsible-ai"],"aprfCheckIds":["SEC-M3","TOL-M2","SAF-M1"],"relation":"supports"},{"peerControlId":"owasp-llm:06","aprfPillarSlugs":["tool-safety","agent-governance","human-approval","authorization"],"aprfCheckIds":["TOL-M1","TOL-M2","TOL-M3","AGN-M2","HUM-M1","AUTHZ-M1"],"relation":"supports"},{"peerControlId":"owasp-llm:07","aprfPillarSlugs":["prompt-engineering","ai-security","secrets"],"aprfCheckIds":["PRM-M2","SEC-M1","SEC2-M2"],"relation":"aligns-with"},{"peerControlId":"owasp-llm:08","aprfPillarSlugs":["memory-management","context-engineering","data-governance"],"aprfCheckIds":["MEM-M1","MEM-M3","CTX-M2","DG-M3"],"relation":"supports"},{"peerControlId":"owasp-llm:09","aprfPillarSlugs":["safety-responsible-ai","evaluation","explainability"],"aprfCheckIds":["SAF-M2","EVL-M2","EXP-M1"],"relation":"aligns-with"},{"peerControlId":"owasp-llm:10","aprfPillarSlugs":["cost-optimization","reliability-continuity","performance-slo"],"aprfCheckIds":["COST-M1","COST-M3","REL-M1","PERF-M1"],"relation":"supports","note":"Denial-of-wallet and resource exhaustion — spend ceilings + timeouts."}]},{"id":"soc2-tsc","name":"SOC 2 Trust Services Criteria","peerVersion":"2017 (with 2022 revisions) — evidence reuse","url":"https://www.aicpa-cima.com/topic/audit-assurance/audit-and-assurance-greater-than-soc-2","disclaimer":"Informative alignment only. Does not constitute certification, accreditation, or official endorsement. Passing APRF does not imply SOC 2 compliance.","controls":[{"id":"soc2:cc1","ref":"CC1","title":"Control Environment"},{"id":"soc2:cc3","ref":"CC3","title":"Risk Assessment"},{"id":"soc2:cc5","ref":"CC5","title":"Control Activities"},{"id":"soc2:cc6","ref":"CC6","title":"Logical and Physical Access"},{"id":"soc2:cc7","ref":"CC7","title":"System Operations"},{"id":"soc2:cc8","ref":"CC8","title":"Change Management"},{"id":"soc2:cc9","ref":"CC9","title":"Risk Mitigation"},{"id":"soc2:a1","ref":"A1","title":"Availability"},{"id":"soc2:c1","ref":"C1","title":"Confidentiality"},{"id":"soc2:pi1","ref":"PI1","title":"Processing Integrity"},{"id":"soc2:p","ref":"P-series","title":"Privacy (selected)"}],"mappings":[{"peerControlId":"soc2:cc1","aprfPillarSlugs":["organizational-governance","compliance"],"aprfCheckIds":["ORG-M1","ORG-M2","CMP-M1"],"relation":"evidence-for"},{"peerControlId":"soc2:cc3","aprfPillarSlugs":["organizational-governance","safety-responsible-ai","evaluation"],"aprfCheckIds":["ORG-M3","SAF-M1","EVL-M1"],"relation":"evidence-for"},{"peerControlId":"soc2:cc5","aprfPillarSlugs":["tool-safety","human-approval","platform-engineering"],"aprfCheckIds":["TOL-M1","HUM-M1","DX-M2"],"relation":"evidence-for"},{"peerControlId":"soc2:cc6","aprfPillarSlugs":["authentication","authorization","secrets","infrastructure"],"aprfCheckIds":["AUTHN-M1","AUTHN-M2","AUTHZ-M1","SEC2-M1","INF-M1"],"relation":"evidence-for"},{"peerControlId":"soc2:cc7","aprfPillarSlugs":["observability","incident-readiness","reliability-continuity"],"aprfCheckIds":["OBS-M1","OBS-M2","INC-M1","INC-M2","REL-M2"],"relation":"evidence-for"},{"peerControlId":"soc2:cc8","aprfPillarSlugs":["change-management","model-governance","prompt-engineering"],"aprfCheckIds":["DEP-M1","CHG-M1","CHG-M3","MOD-M1","PRM-M1"],"relation":"evidence-for"},{"peerControlId":"soc2:cc9","aprfPillarSlugs":["supply-chain","ai-security","reliability-continuity"],"aprfCheckIds":["SCI-M2","SEC-M1","REL-M1"],"relation":"evidence-for"},{"peerControlId":"soc2:a1","aprfPillarSlugs":["reliability-continuity","performance-slo","infrastructure"],"aprfCheckIds":["REL-M1","REL-M2","PERF-M1","INF-M2"],"relation":"evidence-for"},{"peerControlId":"soc2:c1","aprfPillarSlugs":["data-privacy","secrets","memory-management"],"aprfCheckIds":["PRI-M1","PRI-M3","SEC2-M1","MEM-M1"],"relation":"evidence-for"},{"peerControlId":"soc2:pi1","aprfPillarSlugs":["evaluation","observability","change-management"],"aprfCheckIds":["EVL-M1","EVL-M2","OBS-M1","CHG-M3"],"relation":"evidence-for"},{"peerControlId":"soc2:p","aprfPillarSlugs":["data-privacy","data-governance","compliance"],"aprfCheckIds":["PRI-M1","PRI-M2","DG-M1","CMP-M3"],"relation":"partial","note":"Privacy TSC is broader than APRF privacy checks; use as AI evidence pack only."}]},{"id":"aws-well-architected","name":"AWS Well-Architected Framework","peerVersion":"WA Framework pillars + Generative AI Lens (conceptual)","url":"https://aws.amazon.com/architecture/well-architected/","disclaimer":"Informative alignment only. Does not constitute certification, accreditation, or official endorsement. Not an AWS Well-Architected Lens review or AWS certification.","controls":[{"id":"aws-wa:ops","ref":"Operational Excellence","title":"Operational Excellence","summary":"Run and monitor systems; continually improve processes."},{"id":"aws-wa:sec","ref":"Security","title":"Security","summary":"Protect data, systems, and assets."},{"id":"aws-wa:rel","ref":"Reliability","title":"Reliability","summary":"Recover from failures and meet demand."},{"id":"aws-wa:perf","ref":"Performance Efficiency","title":"Performance Efficiency","summary":"Use resources efficiently and adapt to demand."},{"id":"aws-wa:cost","ref":"Cost Optimization","title":"Cost Optimization","summary":"Avoid unnecessary cost; manage spend."},{"id":"aws-wa:sus","ref":"Sustainability","title":"Sustainability","summary":"Minimize environmental impact of cloud workloads."},{"id":"aws-wa:genai","ref":"Generative AI Lens","title":"Generative AI Lens (themes)","summary":"Responsible AI, agent/tool safety, eval, and model lifecycle themes."}],"mappings":[{"peerControlId":"aws-wa:ops","aprfPillarSlugs":["observability","incident-readiness","change-management","platform-engineering","organizational-governance"],"aprfCheckIds":["OBS-M1","INC-M1","CHG-M1","ORG-M2"],"relation":"aligns-with"},{"peerControlId":"aws-wa:sec","aprfPillarSlugs":["authentication","authorization","secrets","ai-security","tool-safety","infrastructure","supply-chain"],"aprfCheckIds":["AUTHN-M1","AUTHZ-M1","SEC2-M1","SEC-M1","TOL-M1","INF-M1","SCI-M2"],"relation":"aligns-with"},{"peerControlId":"aws-wa:rel","aprfPillarSlugs":["reliability-continuity","performance-slo","incident-readiness"],"aprfCheckIds":["REL-M1","REL-M2","PERF-M1","INC-M2"],"relation":"aligns-with"},{"peerControlId":"aws-wa:perf","aprfPillarSlugs":["performance-slo","observability","context-engineering"],"aprfCheckIds":["PERF-M1","OBS-M2"],"relation":"partial"},{"peerControlId":"aws-wa:cost","aprfPillarSlugs":["cost-optimization"],"aprfCheckIds":["COST-M1","COST-M3"],"relation":"supports"},{"peerControlId":"aws-wa:sus","aprfPillarSlugs":["cost-optimization","model-governance"],"aprfCheckIds":["COST-M1","MOD-M1"],"relation":"partial","note":"APRF does not define sustainability metrics; cost/routing and model selection are nearest proxies."},{"peerControlId":"aws-wa:genai","aprfPillarSlugs":["safety-responsible-ai","evaluation","agent-governance","human-approval","prompt-engineering","model-governance","data-privacy"],"aprfCheckIds":["SAF-M1","EVL-M1","AGN-M2","HUM-M1","PRM-M1","MOD-M1","PRI-M1"],"relation":"aligns-with"}]},{"id":"slsa","name":"SLSA (Supply-chain Levels for Software Artifacts)","peerVersion":"v1.0 — conceptual levels / provenance","url":"https://slsa.dev/","disclaimer":"Informative alignment only. Does not constitute certification, accreditation, or official endorsement. APRF alignment does not confer a SLSA level attestation.","controls":[{"id":"slsa:l1","ref":"Level 1","title":"Build process documented","summary":"Scripted build with provenance documentation."},{"id":"slsa:l2","ref":"Level 2","title":"Hosted build + signed provenance","summary":"Version-controlled, hosted build service; authenticated provenance."},{"id":"slsa:l3","ref":"Level 3","title":"Hardened builds","summary":"Hardened build platform; non-falsifiable provenance."},{"id":"slsa:provenance","ref":"Provenance","title":"Artifact provenance","summary":"Who built what from which source, verifiable at deploy."},{"id":"slsa:verify","ref":"Verify-on-deploy","title":"Verification before use","summary":"Consumers verify signatures/provenance before production use."}],"mappings":[{"peerControlId":"slsa:l1","aprfPillarSlugs":["supply-chain","change-management","model-governance"],"aprfCheckIds":["SCI-M2","DEP-M1","MOD-M1"],"relation":"partial"},{"peerControlId":"slsa:l2","aprfPillarSlugs":["supply-chain","infrastructure","change-management"],"aprfCheckIds":["SCI-M2","SCI-R1","INF-M1","CHG-M1"],"relation":"aligns-with"},{"peerControlId":"slsa:l3","aprfPillarSlugs":["supply-chain","infrastructure","platform-engineering"],"aprfCheckIds":["SCI-M4","SCI-R1","INF-M1","DX-M2"],"relation":"partial","note":"APRF SCI-M4 / hardened admission approximate L3 themes; not a full SLSA L3 claim."},{"peerControlId":"slsa:provenance","aprfPillarSlugs":["supply-chain","model-governance"],"aprfCheckIds":["SCI-M2","SCI-R2","MOD-M1"],"relation":"supports"},{"peerControlId":"slsa:verify","aprfPillarSlugs":["supply-chain","infrastructure","change-management"],"aprfCheckIds":["SCI-R1","INF-M1","CHG-M3"],"relation":"supports"}]}],"scoring":{"methodology":{"name":"APRF gated evaluation","forbids":["A single averaged “readiness score” across all pillars","Trading a failed mandatory check against strong recommended scores","Conformance badges that omit open blockers or APRF version","Claiming Tier 3 / regulated readiness from Core Profile alone"],"steps":[{"id":"classify","title":"Classify criticality","detail":"Assign Tier 0–3. This sets the required capability floor."},{"id":"profile","title":"Choose catalog, profile, and lenses","detail":"Startups may assess Core Profile (Tier 2 minimum). Add formal lenses (RAG, Agents, Voice, Coding agents) when those system types apply. Full catalog for Tier 3 / enterprise / regulated."},{"id":"collect","title":"Collect evidence","detail":"For each applicable check, produce the named artifact and evaluate the pass condition."},{"id":"gate","title":"Evaluate the gate","detail":"All in-scope mandatory checks must pass or be formally marked notApplicable (N/A) with rationale. Failures and unanswered checks are blockers. Profile∪lens assessments only gate on the union of those check IDs. N/A is intended for agent, human-approval, tool-safety, and memory gates when those system types do not apply — not as a substitute for incomplete evidence."},{"id":"attainment","title":"Compute capability attainment","detail":"Per pillar: highest level L where all mandatory checks ≤ L pass. System attainment = minimum across pillars."},{"id":"recommended","title":"Score recommended controls per domain","detail":"Optional 0–100 domain scores from recommended checks, weighted by pillar severity. Never folded into the gate."},{"id":"report","title":"Report","detail":"Publish: APRF version, tier, profile (if any), lenses (if any), required capability, gate pass/fail, critical blockers, capability attained, per-domain recommended scores, evidence index."}]},"severityWeights":{"critical":4,"high":3,"medium":2,"low":1}},"pillars":[{"id":"APRF-01","slug":"ai-security","name":"Adversarial Security","summary":"Prevent prompt injection, jailbreaks, exfiltration, and model/tool abuse from becoming a production incident.","domain":"security","crossCutting":false,"severity":"critical","riskLevel":"critical","purpose":"Establish controls that prevent adversarial misuse of models, prompts, tools, and outputs from compromising confidentiality, integrity, or availability—distinct from content-safety and responsible-AI harms covered in the Safety domain.","engineeringPhilosophy":"Treat the model as an untrusted co-processor inside a zero-trust boundary. Security controls belong in the application and platform layers—never solely in prompt wording. Defense in depth spans input validation, output filtering, tool mediation, identity, and runtime isolation.","whyItMatters":"AI surfaces expand the attack surface: prompt injection, data exfiltration via tools, jailbreaks, model theft, and poisoned retrieval. A chatbot that 'mostly works' can still become an incident response event when an attacker turns natural language into a privileged API.","commonFailures":["Relying on system prompts alone to enforce security policy","Passing untrusted retrieval or user content into privileged tool contexts","Logging full prompts/completions that contain secrets or PII","No output schema enforcement—free-form text executes downstream logic","Missing abuse rate limits distinct from normal product quotas"],"mandatoryChecks":[{"id":"SEC-M1","requirement":"Untrusted input shall never authorize privileged actions without server-side policy","artifact":"Injection/privilege-escalation corpus (versioned) + CI gate report + policy-engine deny sample logs","passCondition":"≥95% of corpus cases that attempt privilege escalation via untrusted input are denied; 0 cases where model text alone granted a privileged tool call in the suite","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SEC-M2","requirement":"High-risk outputs shall be schema-validated or policy-filtered before side effects","artifact":"Schema/policy definitions for high-risk paths + contract test results","passCondition":"100% of high-risk side-effect paths (impact tier write/irreversible/financial) reject non-conforming model output in contract tests; coverage inventory lists every such path","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SEC-M3","requirement":"Abuse, jailbreak, and injection testing shall gate customer-facing releases","artifact":"CI gate configuration + last 30 days of release reports","passCondition":"100% of production releases in the last 30 days show security-suite gate = pass, or a time-boxed waiver with owner and expiry ≤ 30 days","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SEC-M4","requirement":"Network and identity boundaries shall prevent the model path from becoming a universal proxy to internal systems","artifact":"Architecture diagram of model/tool trust boundaries + network policy or egress allowlist export","passCondition":"Automated or reviewed probe shows the model/tool runtime can reach only allowlisted destinations; 0 unrestricted routes to internal admin APIs or data stores from the model identity","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"SEC-M5","requirement":"Canary tokens or tripwires shall detect exfiltration attempts in sensitive AI contexts","artifact":"Canary/tripwire design + detection test report with sample alerts for sensitive contexts","passCondition":"PASS if canaries/tripwires are deployed in production-sensitive paths and the latest detection test (≤90 days) shows expected alerts with 0 silent misses in the suite","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"SEC-R1","requirement":"Red-team exercises cover multi-turn and indirect injection via RAG/MCP","artifact":"Versioned red-team suite covering multi-turn and indirect injection via RAG/MCP + latest scored run report","passCondition":"Suite includes ≥10 multi-turn and ≥10 indirect RAG/MCP injection cases; latest run ≤90 days meets documented pass thresholds; report retained ≥90 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"SEC-R2","requirement":"Content safety and malware scanning on multimodal inputs where applicable","artifact":"Multimodal input safety/malware scanner config + latest CI or batch scan report for production-bound media","passCondition":"Where multimodal inputs are accepted, scanner runs before model ingest; last report ≤90 days shows coverage of image/file types in use and 0 unscanned production paths","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Threat model covering LLM-specific attack paths","Records of injection/jailbreak test suites and results","Architecture diagram showing trust boundaries around model and tools","Incident playbooks referencing AI-specific abuse scenarios"],"engineeringBestPractices":["Separate instruction channels from data channels; never concatenate untrusted data into system instructions without mediation","Apply least privilege to every tool and retrieval corpus the model can touch","Fail closed on policy violations; prefer refusal over silent degradation for security decisions","Version and review security policies the same way you review firewall rules"],"automaticValidations":["CI security tests for known injection payloads against critical flows","Runtime policy engines rejecting disallowed tool calls","Schema validation on structured model outputs","Anomaly detection on tool-call volume and destination"],"manualValidations":["Periodic adversarial review of new tools and prompts","Architecture review of trust boundaries before Level 3 launch","Tabletop exercises for data-exfiltration via the AI path"],"examples":["A support agent refuses to call refund APIs when the refund amount exceeds policy, regardless of user phrasing","RAG documents containing 'ignore previous instructions' cannot elevate privileges","MCP servers expose only scoped tools with server-side authorization independent of the model"],"references":[{"title":"OWASP Top 10 for Large Language Model Applications","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"},{"title":"NIST AI Risk Management Framework","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"CIS Controls (select mappings for AI workloads)","url":"https://www.cisecurity.org/controls"}],"futureEvolution":"Expect standardized machine-readable AI threat taxonomies, portable policy packs across providers, and attestation of model-path controls analogous to cloud security posture management."},{"id":"APRF-10","slug":"authentication","name":"Authentication","summary":"Strong identity for users, services, agents, and MCP callers.","domain":"security","crossCutting":false,"severity":"critical","riskLevel":"critical","purpose":"Ensure every actor that can invoke AI capabilities—humans, services, agents, MCP clients—is authenticated with appropriate assurance.","engineeringPhilosophy":"AI endpoints are privileged application surfaces. Anonymous or shared credentials for agents and MCP are incompatible with production readiness beyond Level 1.","whyItMatters":"Unauthenticated model and tool endpoints become public compute and data exfiltration oracles. Agent impersonation enables privilege abuse across systems.","commonFailures":["Public API keys embedded in clients calling powerful agents","MCP servers listening without auth","Service accounts shared across many agents","No MFA for admin consoles controlling prompts and tools"],"mandatoryChecks":[{"id":"AUTHN-M1","requirement":"Customer-facing AI APIs shall reject unauthenticated callers","artifact":"Automated auth probe report covering all production AI HTTP/RPC routes","passCondition":"100% of probed production AI endpoints return 401/403 without valid credentials; probe inventory matches production route catalog","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"AUTHN-M2","requirement":"Service-to-service and MCP connections shall use strong machine identity","artifact":"MCP/S2S connection inventory + auth config export","passCondition":"0 production MCP or AI S2S connections accept anonymous access or shared long-lived static keys; each connection has a named machine identity","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"AUTHN-M3","requirement":"Administrative access to AI control planes shall require strong authentication including MFA","artifact":"IdP policy export for AI admin roles + break-glass account inventory","passCondition":"100% of AI control-plane admin roles enforce MFA; break-glass accounts ≤ documented maximum and have monitoring enabled","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"AUTHN-M4","requirement":"End-user identity shall remain bound through agent and tool chains","artifact":"Identity-propagation design + sample traces showing end-user subject on tool calls","passCondition":"PASS if 100% of sampled privileged tool calls in the latest review carry an end-user (or documented service) subject; 0 anonymous privileged hops","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"AUTHN-R1","requirement":"Short-lived tokens for agents and tools; no long-lived static keys in prompts","artifact":"IdP/token TTL policy export + inventory of agent/tool credentials used in production prompts/config","passCondition":"100% of inventoried agent/tool credentials have TTL ≤1h (or a named exception ≤30 days); secret scan of prompts/config finds 0 long-lived static API keys","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"AUTHN-R2","requirement":"Workload identity for self-hosted model runtimes","artifact":"Workload-identity binding config for self-hosted model runtimes + sample authenticated call traces","passCondition":"Every self-hosted model runtime in production authenticates via workload identity (SPIFFE/IAM role/equivalent); 0 static shared keys in the runtime inventory","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Auth architecture for AI and MCP surfaces","Evidence that anonymous access is disabled in production","Admin access controls documentation"],"engineeringBestPractices":["Prefer OAuth2/OIDC or workload identity over static API keys","Issue distinct identities per agent class","Rotate credentials automatically; detect leaked keys in repos and prompts","Authenticate both directions for A2A where feasible"],"automaticValidations":["Integration tests rejecting unauthenticated calls","Secret scanning for exposed AI API keys","Config checks that MCP auth is enabled"],"manualValidations":["Identity design review for new agent entry points","Periodic access review of AI admin roles"],"examples":["An MCP server requires OAuth for every client; local-dev uses a distinct non-prod issuer","A voice AI channel authenticates the telephony session before invoking tools"],"references":[{"title":"OWASP Authentication Cheat Sheet","url":"https://cheatsheetseries.owasp.org/cheatsheets/Authentication_Cheat_Sheet.html"},{"title":"CIS Benchmarks — Identity controls","url":"https://www.cisecurity.org/cis-benchmarks"}],"futureEvolution":"Standard agent identity profiles and passport-like credentials for cross-org A2A."},{"id":"APRF-11","slug":"authorization","name":"Authorization","summary":"Enforce who/what may invoke which models, tools, data, and actions.","domain":"security","crossCutting":false,"severity":"critical","riskLevel":"critical","purpose":"Enforce fine-grained authorization over models, prompts, corpora, tools, and actions so authenticated actors only receive the capabilities they are entitled to.","engineeringPhilosophy":"Authentication without authorization is incomplete. Authorization decisions must be made outside the model and must apply to retrieved data and tool effects, not only to the HTTP route.","whyItMatters":"Confused-deputy and over-permissioned agents are primary causes of data breaches in AI systems. A model that can see everything can leak everything.","commonFailures":["Authorization checked at UI only; API and agent paths bypass it","Retrieval without document-level ACLs","Tools inheriting broad service roles","Tenant isolation bugs in multi-tenant AI products"],"mandatoryChecks":[{"id":"AUTHZ-M1","requirement":"Authorization shall be enforced server-side for AI features, tools, and retrieval","artifact":"Authz middleware/policy tests for AI feature, tool, and retrieval entry points","passCondition":"Automated tests cover 100% of production AI feature, tool, and retrieval entry points; unauthenticated or unauthorized callers are denied at 100% in the suite","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"AUTHZ-M2","requirement":"Multi-tenant isolation shall be tested for AI data and memory paths","artifact":"Cross-tenant attack test suite results for AI data stores and memory APIs","passCondition":"0 successful unauthorized cross-tenant reads/writes across ≥10 automated attack cases on AI data and memory paths","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"AUTHZ-M3","requirement":"Least-privilege roles shall be defined for agents and automation identities","artifact":"Role matrix for every production agent/automation identity + IAM/policy export","passCondition":"100% of production agent/automation identities appear in the role matrix with non-admin default roles; quarterly (or ≤90-day) access review recorded with 0 unexplained privilege escalations","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"AUTHZ-M4","requirement":"Attribute-based controls shall govern access to sensitive document classes","artifact":"ABAC/policy config + sample allow/deny decisions for sensitive classes","passCondition":"PASS if sensitive document classes are enumerated and policy denies unauthorized class access in tests; inventory matches production classes","method":"manual","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"AUTHZ-R1","requirement":"Policy-as-code for tool and model access","artifact":"Policy-as-code repo/config for tool and model access (OPA/Cedar/IAM-as-code or equivalent) + latest CI policy-check report","passCondition":"Tool and model access rules are expressed as code and enforced in CI or admission; last failing-to-passing policy change ≤90 days shows a deny for unauthorized tool/model","method":"manual","requiredFromLevel":4,"minCriticality":2},{"id":"AUTHZ-R2","requirement":"Continuous access reviews for high-privilege agents","artifact":"High-privilege agent inventory + last access-review spreadsheet/report with decisions (keep/revoke/modify)","passCondition":"Every high-privilege agent identity was reviewed ≤90 days ago; ≥1 revoke or scope-reduction appears in the last two cycles (or attestation that none were warranted with reviewer sign-off)","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Authorization model documentation","Tenant isolation test results","Role matrices for agents and tools"],"engineeringBestPractices":["Propagate user context into tool calls; avoid confused deputy","Deny by default for new tools and corpora","Re-check authorization at each hop in multi-agent workflows","Separate read and write capabilities aggressively"],"automaticValidations":["Automated cross-tenant access attempt tests","CI policy checks for IAM/tool scopes","Runtime denials logged and alertable"],"manualValidations":["Security review of new high-impact tools","Periodic privilege creep audits"],"examples":["A sales agent can query CRM records for the caller's accounts only","An internal coding agent cannot access production secrets vaults"],"references":[{"title":"OWASP Access Control Cheat Sheet","url":"https://cheatsheetseries.owasp.org/cheatsheets/Access_Control_Cheat_Sheet.html"},{"title":"NIST AI RMF — Manage","url":"https://www.nist.gov/itl/ai-risk-management-framework"}],"futureEvolution":"Portable authorization tokens that travel with agent tasks across organizational boundaries."},{"id":"APRF-12","slug":"secrets","name":"Secrets","summary":"Eliminate secret leakage into prompts, logs, tools, and client surfaces.","domain":"security","crossCutting":false,"severity":"critical","riskLevel":"critical","purpose":"Prevent API keys, credentials, tokens, and other secrets from entering prompts, completions, logs, training/fine-tuning corpora, or client-visible surfaces.","engineeringPhilosophy":"Secrets never belong in context windows as data. Treat any path that can echo or store model I/O as a potential secret exfiltration channel.","whyItMatters":"Models and logs are high-bandwidth leak channels. A single leaked provider key or cloud credential can escalate into full environment compromise and unbounded spend.","commonFailures":["API keys in frontend code or mobile apps","Secrets pasted into prompts for 'debugging'","Tool results returning credentials into context","Fine-tuning on datasets that include secrets"],"mandatoryChecks":[{"id":"SEC2-M1","requirement":"Production secrets shall live in a secrets manager and not in prompts, repos, or prod notebooks","artifact":"Secrets-manager config + CI/repo secret-scan report including prompt and fixture paths","passCondition":"0 privileged production secrets found in repos, prompt registries, or client bundles in the latest scan; 100% of production runtime secrets resolve from the secrets manager","method":"automated","requiredFromLevel":2,"minCriticality":1},{"id":"SEC2-M2","requirement":"Logging and tracing pipelines shall redact secret-like patterns","artifact":"Redaction config + sample of redacted traces + synthetic secret-injection test results","passCondition":"Synthetic secrets (API key/bearer/AWS-key patterns) injected into a canary request are redacted in persisted logs/traces at 100% detection in the test harness","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SEC2-M3","requirement":"Provider and cloud keys shall be rotatable and scoped; client apps shall not hold privileged keys","artifact":"Key inventory with rotation dates + client bundle/scan report","passCondition":"0 privileged provider/cloud keys embedded in client apps; every production key has a rotation date within policy (≤90 days or provider-managed short-lived credentials)","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"SEC2-R1","requirement":"Pre-commit and CI secret scanning including prompt/fixture files","artifact":"Pre-commit + CI secret-scan config including prompt/fixture globs + latest scan report with findings disposition","passCondition":"Secret scanning covers application code, prompts, and fixtures; blocking on high-confidence secrets; last green main-branch scan ≤7 days (or last PR merge evidence)","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SEC2-R2","requirement":"Egress controls limiting where runtime credentials can be used","artifact":"Egress allowlist/policy config for runtimes that hold credentials + sample deny/allow logs (24h)","passCondition":"Runtime credentials can only call documented destinations; ≥1 deny event is observed in test or production logs proving enforcement in the last 90 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"SEC2-R3","requirement":"Dataset scanning before fine-tuning or eval corpus publish","artifact":"Dataset secret/PII scan gate config before fine-tune or eval corpus publish + last scan report for a published corpus","passCondition":"100% of fine-tune/eval corpora published in the last 90 days have a linked scan report; publish is blocked when critical findings are open","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Secrets management design","Redaction configuration evidence","Key rotation records"],"engineeringBestPractices":["Use short-lived credentials for tools","Strip Authorization headers and env dumps from error payloads to models","Educate builders: never ask the model to 'remember' a key","Separate build-time and runtime secrets"],"automaticValidations":["Secret scanners in CI and image builds","Runtime detectors blocking known secret formats in prompts","Alerts on anomalous provider API key usage"],"manualValidations":["Review of new logging fields for secret risk","Incident drills for leaked model API keys"],"examples":["A developer tool refuses to send .env contents into an LLM chat","Tracing middleware redacts AWS keys and bearer tokens before persistence"],"references":[{"title":"OWASP Secrets Management Cheat Sheet","url":"https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html"},{"title":"CIS Benchmarks — credential hygiene","url":"https://www.cisecurity.org/cis-benchmarks"}],"futureEvolution":"Native provider APIs that never accept raw secrets in prompt fields, with automatic scrubbing attestations."},{"id":"APRF-05","slug":"tool-safety","name":"Tool Safety","summary":"Make tool/MCP invocation fail closed with least privilege and side-effect control.","domain":"security","crossCutting":false,"severity":"critical","riskLevel":"critical","purpose":"Ensure every tool, function call, MCP server, and external action invoked by a model is authorized, scoped, observable, and reversible or gated when impact is high.","engineeringPhilosophy":"The model proposes; the platform disposes. Tools are privileged APIs. Never let natural language become an unrestricted remote control over production systems.","whyItMatters":"Tool misuse converts language model errors into real-world damage—deleted data, fraudulent transactions, leaked secrets, or lateral movement through MCP and A2A bridges.","commonFailures":["Tools run with service-wide credentials instead of user-scoped tokens","No allowlist of tools per agent or environment","Destructive tools without confirmation or dry-run modes","MCP servers exposed without authentication"],"mandatoryChecks":[{"id":"TOL-M1","requirement":"Every tool invocation shall be authorized server-side independent of model output","artifact":"Tool gateway authz tests + deny logs","passCondition":"Automated tests show tool calls without a valid authz decision are denied at 100%; model-proposed tool name/args alone never bypass the gateway","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"TOL-M2","requirement":"Tools shall be allowlisted per agent/workload; unknown tools shall not be inventable at runtime","artifact":"Per-agent tool allowlist config + negative tests for unknown tool names","passCondition":"100% of production agents have an explicit tool allowlist; requests for tools outside the allowlist are denied at 100% in automated tests","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"TOL-M3","requirement":"High-impact tools shall require additional gates (approval, dual control, or policy engine)","artifact":"Impact-tiered tool inventory + gate configuration + bypass tests","passCondition":"100% of tools rated write/irreversible/financial have a configured gate; automated tests show ungated execution is impossible for those tools","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"TOL-M4","requirement":"Tool arguments shall be schema-validated and sanitized before execution","artifact":"JSON Schema (or equivalent) per tool + contract tests for invalid payloads","passCondition":"100% of production tools have a declared argument schema; invalid/malicious argument fixtures are rejected at 100% before side effects","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"TOL-M5","requirement":"MCP tool catalogs shall be signed and supply-chain reviewed before production use","artifact":"Signed tool catalog + supply-chain review record + verify-on-load config","passCondition":"PASS if production MCP runtimes reject unsigned/unapproved tool catalogs; last review ≤90 days or since last catalog change","method":"manual","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"TOL-R1","requirement":"Dry-run or simulation mode for destructive tools in lower environments","artifact":"Destructive-tool catalog marking dry-run/simulation support + sample lower-env dry-run logs","passCondition":"100% of tools classified destructive expose dry-run or simulation in non-prod; last promotion of a destructive tool includes a dry-run evidence link ≤90 days old","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"TOL-R2","requirement":"Per-tool rate limits and blast-radius budgets","artifact":"Per-tool rate-limit and blast-radius policy config + sample enforcement metrics (7 days)","passCondition":"Each high-impact tool has a documented QPS/daily cap and max-affected-entities budget; ≥1 limit hit or synthetic test proves enforcement in the last 30 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Tool inventory with owners, scopes, and impact ratings","Authorization policy examples","Logs of denied tool invocations"],"engineeringBestPractices":["Map each tool to an impact tier (read, write, irreversible, financial)","Pass user identity into tools; avoid god-mode service accounts where possible","Return structured errors; do not leak internal stack traces into the model context","Treat MCP and A2A endpoints as public-facing APIs with the same hardening"],"automaticValidations":["Contract tests for tool schemas","Policy-as-code denying unauthorized tool calls","CI scanning for overly broad tool permissions"],"manualValidations":["Impact rating review for new tools before production","Penetration testing focused on tool and MCP surfaces"],"examples":["A coding agent can run tests but cannot push to main without human approval","An MCP file server exposes only a workspace directory, not the host filesystem"],"references":[{"title":"Model Context Protocol security considerations","url":"https://modelcontextprotocol.io/"},{"title":"OWASP LLM — Excessive Agency","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"}],"futureEvolution":"Standardized tool impact taxonomies and portable capability tokens for agent tool use."},{"id":"APRF-24","slug":"supply-chain","name":"Supply Chain Integrity","summary":"Prove provenance and integrity of models, containers, MCP servers, and AI dependencies.","domain":"security","crossCutting":false,"severity":"critical","riskLevel":"high","purpose":"Ensure model artifacts, inference images, tools, MCP servers, and third-party AI dependencies are authenticated, reviewed, and verifiable—applying SLSA-style integrity to the AI stack.","engineeringPhilosophy":"An AI system is only as trustworthy as the artifacts it loads. Unsigned weights, unverified MCP hosts, and opaque tool catalogs are supply-chain defects, not product features.","whyItMatters":"Poisoned models, compromised MCP servers, and tampered tool plugins bypass application-layer controls and can exfiltrate data or execute attacker-chosen actions at scale.","commonFailures":["Downloading open-weight models from unverified mirrors without checksum or signature verification","MCP servers installed without source review or pinned versions","No SBOM/MBOM for inference images and agent runtimes","Fine-tuned adapters pulled from untrusted registries into production"],"mandatoryChecks":[{"id":"SCI-M1","requirement":"Production model and container artifacts shall have verified provenance/integrity","artifact":"Digest/signature verification logs for last production deploy","passCondition":"100% of production model/container pulls in the sample window verify against expected digest or signature; unverified pulls are blocked","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SCI-M2","requirement":"MCP servers and high-impact tools shall be inventoried, version-pinned, and reviewed before production use","artifact":"MCP/tool inventory with version pins, owners, and review dates","passCondition":"100% of production MCP servers and high-impact tools have pin + owner + review date ≤ 180 days; 0 unpinned “latest” entries in production","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"SCI-M3","requirement":"Dependency and image vulnerability scanning gates shall apply to AI runtimes","artifact":"Scanner config + CI gate reports for AI runtime images/deps","passCondition":"100% of AI runtime image builds in the last 30 days ran vulnerability scan; critical CVEs (CVSS ≥ 9.0 or org policy equivalent) block promote unless waived with expiry ≤ 14 days","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SCI-M4","requirement":"Admission controls shall block unsigned or unapproved AI artifacts","artifact":"Admission policy config + deny logs for unsigned/unapproved artifacts","passCondition":"PASS if admission is enforced in production deploy paths and the latest probe shows unsigned artifacts are blocked","method":"automated","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"SCI-R1","requirement":"Signed model/container artifacts with verify-on-deploy (SLSA-aligned)","artifact":"Admission/verify-on-deploy policy for signed model and container artifacts + last deploy verification log","passCondition":"Last production model/container deploy shows signature verification success; unsigned artifact is rejected in a recorded test or canary within 90 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"SCI-R2","requirement":"Machine-readable model bill of materials (MBOM) for production models","artifact":"Machine-readable MBOM (or SBOM+model metadata) for each production model pin + retention location","passCondition":"100% of production model pins have an MBOM/SBOM artifact retained ≥90 days and linked from the model registry entry","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Artifact signing/verification configuration","MCP and tool inventory with review records","Image/model scan reports for production releases"],"engineeringBestPractices":["Prefer private registries for production models and adapters","Treat MCP hosts as supply-chain surfaces equal to npm/PyPI packages","Separate build and serving identities; never use developer credentials in prod pulls","Record provenance in release evidence packs"],"automaticValidations":["CI verification of artifact digests against expected values","Admission controller or gateway denials for unsigned artifacts","SBOM generation and CVE gates on AI images"],"manualValidations":["Source review for new MCP servers and tool plugins","Periodic audit of model download sources"],"examples":["Inference pods only pull images signed by the org signing key","An MCP catalog requires peer review and a pinned commit before agents can call it"],"references":[{"title":"SLSA — Supply-chain Levels for Software Artifacts","url":"https://slsa.dev/"},{"title":"CNCF Software Supply Chain Best Practices","url":"https://www.cncf.io/"},{"title":"OWASP LLM — Supply Chain Vulnerabilities","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"}],"futureEvolution":"Portable MBOM formats and cross-vendor attestation for hosted and self-hosted models."},{"id":"APRF-20","slug":"infrastructure","name":"Infrastructure","summary":"Harden runtime, network, isolation, and supply chain for AI workloads.","domain":"security","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Provide secure, reliable runtime and network infrastructure for model serving, gateways, vector stores, agents, and MCP—including supply-chain integrity.","engineeringPhilosophy":"AI workloads inherit cloud security and SRE fundamentals, then add GPU/runtime and model-artifact supply-chain concerns. There is no AI exception to patching, isolation, or least privilege.","whyItMatters":"Compromised gateways, exposed vector DBs, or poisoned model artifacts undermine every application-layer control.","commonFailures":["Vector databases publicly reachable","Unpatched inference runtimes","Unsigned model artifacts from unverified sources","Overly broad network paths from agents to internal services"],"mandatoryChecks":[{"id":"INF-M1","requirement":"AI data stores and control planes shall not be publicly exposed without authenticated edge controls","artifact":"CSPM/network scan of AI data stores and control planes + edge auth config","passCondition":"0 AI data stores or control-plane endpoints publicly reachable without authentication in the latest scan; findings severity ≥ high closed or waived with expiry","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"INF-M2","requirement":"Base images and runtimes shall follow documented patching SLAs","artifact":"Patching SLA policy + image age/CVE backlog report for AI runtimes","passCondition":"100% of production AI runtime images are within patching SLA (e.g. critical fixes ≤ 14 days); report shows 0 SLA breaches or open waivers with expiry","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"INF-M3","requirement":"Network segmentation shall limit agent/tool reachability to required dependencies","artifact":"Network policy / security-group export for agent and tool runtimes + dependency allowlist","passCondition":"Egress/east-west allowlists for agent/tool identities match the documented dependency inventory; probe from agent identity to a non-allowlisted internal service fails at 100%","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"INF-M4","requirement":"GPU/accelerator isolation and noisy-neighbor controls shall protect multi-tenant inference","artifact":"Isolation design + capacity/isolation test report for shared accelerators","passCondition":"PASS if isolation controls are documented and the latest isolation/capacity test (≤90 days) meets stated limits","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"INF-R1","requirement":"Signed model and container artifacts with verify-on-deploy","artifact":"Cluster/admission controller config requiring signed images + last failed-unsigned admission event or drill","passCondition":"Unsigned model/container images cannot schedule in production namespaces; policy verified by a failed admission test within 90 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2,"deprecated":true,"replacedBy":"SCI-R1","deprecationNote":"Superseded by SCI-R1 (SLSA-aligned signed model/container verify-on-deploy). Retained for N−1 readers; new assessments should use SCI-R1."},{"id":"INF-R3","requirement":"Infrastructure-as-code with policy checks (CIS-aligned)","artifact":"IaC modules for AI infra + CIS-aligned policy scan config + latest scan report for production stacks","passCondition":"Production AI infra is declared in IaC; CIS (or equivalent) policy checks run on every apply/PR; critical findings for production stacks are 0 or have dated exceptions ≤90 days","method":"manual","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Network diagrams for AI components","Patch and image scan reports","Artifact signing/verification configuration"],"engineeringBestPractices":["Prefer private networking to providers where available","Harden MCP hosts as internet-facing services","Separate training/fine-tune clusters from serving","Apply CIS benchmarks to hosts and Kubernetes where used"],"automaticValidations":["CSPM/misconfiguration scanning","Image CVE gates in CI","Admission controllers verifying signatures"],"manualValidations":["Architecture review for new inference deployments","Supply-chain review for third-party model downloads"],"examples":["A self-hosted Llama deployment sits behind a private gateway with mTLS","Vector DB accepts connections only from the RAG service identity"],"references":[{"title":"CIS Benchmarks","url":"https://www.cisecurity.org/cis-benchmarks"},{"title":"CNCF security technical advisory guidance","url":"https://www.cncf.io/"},{"title":"AWS Well-Architected — Security","url":"https://aws.amazon.com/architecture/well-architected/"}],"futureEvolution":"Attested inference environments and standardized model artifact signing ecosystems."},{"id":"APRF-25","slug":"safety-responsible-ai","name":"Safety & Responsible AI","summary":"Prevent harmful content and unfair outcomes—NIST trustworthiness beyond adversarial security.","domain":"safety","crossCutting":false,"severity":"critical","riskLevel":"high","purpose":"Define and enforce content-safety, harm-prevention, fairness/bias, and user-facing AI disclosure requirements appropriate to the product domain and risk tier.","engineeringPhilosophy":"Safety is not a synonym for security. Adversarial controls stop attackers; safety controls stop harmful or unfair outcomes for users and society. Both are required for production readiness.","whyItMatters":"Systems can be “secure” yet produce toxic, discriminatory, or undisclosed AI-generated content that creates regulatory, brand, and user harm.","commonFailures":["No domain-specific harm taxonomy or refusal policy","Safety filters only on output, not on tool-mediated side effects","No fairness or disparity testing for high-stakes decisions","Users not informed when they interact with AI"],"mandatoryChecks":[{"id":"SAF-M1","requirement":"A documented harm taxonomy and refusal/escalation policy shall exist for the product domain","artifact":"Approved harm taxonomy + refusal/escalation policy (versioned, owned)","passCondition":"Policy document has version, owner, and review date ≤ 12 months; maps ≥ domain-minimum harm categories with explicit refuse vs escalate actions for each","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"SAF-M2","requirement":"Automated safety evaluation gates shall run on relevant releases","artifact":"Safety suite definition with numeric thresholds + CI gate reports","passCondition":"Safety gate executed on 100% of in-scope releases in last 30 days; fail blocks promote unless time-boxed waiver (expiry ≤ 14 days) with owner","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"SAF-M3","requirement":"Users shall be disclosed when interacting with AI where required by policy or law","artifact":"Disclosure UX inventory + screenshot/checklist audit of in-scope surfaces","passCondition":"AI-interaction disclosure present on 100% of in-scope user surfaces per policy checklist; 0 critical surfaces missing disclosure in the latest audit","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"SAF-M4","requirement":"Fairness/disparity evaluation shall cover high-stakes decision paths","artifact":"Fairness/disparity eval methodology + latest run report for in-scope high-stakes paths","passCondition":"PASS if in-scope high-stakes paths are inventoried and the latest eval (≤90 days) is retained with thresholds and owners","method":"hybrid","requiredFromLevel":4,"minCriticality":3}],"recommendedChecks":[{"id":"SAF-R2","requirement":"Red-team exercises covering jailbreak-to-harm scenarios (distinct from security red team)","artifact":"Jailbreak-to-harm red-team suite (distinct from security injection suite) + latest scored run report","passCondition":"Suite covers documented harm categories; latest run ≤90 days meets refusal/safety thresholds; findings feed the safety backlog with owners","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"SAF-R3","requirement":"Human review sampling of safety edge cases on a defined cadence","artifact":"Safety edge-case sampling plan + last review packet (sample size, labels, actions) with reviewer names","passCondition":"Defined sample size and cadence (≥ monthly or per release); last packet ≤90 days includes disposition for each fail/edge case and links to backlog items when needed","method":"manual","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Harm taxonomy and policy document","Safety eval suite definitions and recent gate results","Disclosure UX or policy evidence"],"engineeringBestPractices":["Separate safety eval datasets from adversarial security corpora","Fail closed on high-severity harm categories","Align safety thresholds to criticality tier, not headcount","Log safety refusals for measurement without storing unnecessary sensitive content"],"automaticValidations":["CI safety-gate failures block release","Online monitoring of refusal rates and safety classifier signals","Regression tests for known harmful prompt classes"],"manualValidations":["Domain expert review of harm taxonomy","Periodic calibration of automated safety graders vs human labels"],"examples":["A tutoring agent refuses self-harm content and escalates to a human protocol","A lending assistant reports group disparity metrics before each model promotion"],"references":[{"title":"NIST AI RMF — Trustworthiness characteristics","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"ISO/IEC 42001 — AI management systems","url":"https://www.iso.org/standard/81230.html"}],"futureEvolution":"Shared harm taxonomies and portable safety scorecards across model providers."},{"id":"APRF-26","slug":"explainability","name":"Explainability & Transparency","summary":"Make AI decisions reconstructable for operators, users, and auditors—not only for debugging.","domain":"safety","crossCutting":false,"severity":"high","riskLevel":"medium","purpose":"Provide appropriate explanations, citations, and decision traces so users, operators, and auditors can understand why the system produced an outcome.","engineeringPhilosophy":"Observability serves operators; explainability serves people affected by decisions and those who govern them. Production AI must support both audiences.","whyItMatters":"Opaque decisions block dispute resolution, regulatory review, and user trust—especially in high-stakes or regulated contexts.","commonFailures":["No citations for RAG answers that claim factual authority","Tool-driven actions with no human-readable rationale","Explanations that leak sensitive internal data","No audit-facing decision summary for contested outcomes"],"mandatoryChecks":[{"id":"EXP-M1","requirement":"High-stakes or factual RAG outputs shall include provenance (citations or source IDs)","artifact":"RAG eval suite measuring citation presence/accuracy + sample production traces","passCondition":"On the factual/high-stakes RAG eval set, ≥90% of answers include ≥1 valid source ID/citation that resolves to an authorized corpus document","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"EXP-M2","requirement":"Operators shall reconstruct the decision path for a sampled production outcome","artifact":"Operator reconstruction procedure + timed drill record on ≥3 sampled traces","passCondition":"On-call or operator successfully reconstructs model→retrieval/tools→outcome for 3/3 sampled production traces within documented time budget (e.g. ≤15 minutes each)","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"EXP-M3","requirement":"Explanation content shall not disclose secrets or unauthorized data","artifact":"Explanation redaction policy + automated secret/PII scan on explanation payloads","passCondition":"Synthetic secret/PII fixtures in explanation paths are redacted or blocked at 100% in tests; latest production explanation sample scan shows 0 privileged secret pattern hits","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"EXP-M4","requirement":"Counterfactual or change summaries shall be available for material model/prompt version diffs","artifact":"Change-summary tooling/config + sample summaries for recent promotions","passCondition":"PASS if the last material model/prompt promotion includes a retained change/counterfactual summary","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"EXP-R1","requirement":"User-facing rationale for material automated decisions","artifact":"Product UI/API screenshots or traces showing user-facing rationale for material automated decisions + decision catalog","passCondition":"100% of decision types marked material in the catalog return a rationale field or UI explanation in a 20-case sample; gaps tracked with owners","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"EXP-R3","requirement":"Formal explainability requirements mapped for regulated features","artifact":"Explainability requirements matrix (feature × regulation/obligation) + last compliance review record","passCondition":"Every regulated AI feature lists required explanation type and evidence; matrix reviewed ≤12 months ago with named owner","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Example cited answers (redacted)","Operator reconstruction procedure","Policy for explanation disclosure vs confidentiality"],"engineeringBestPractices":["Prefer source-grounded answers over free-form claims when stakes are high","Separate debug traces (privileged) from user explanations (minimized)","Version explanation formats with the product"],"automaticValidations":["Eval gates for citation presence/accuracy on RAG suites","Tests that explanation redaction strips secrets"],"manualValidations":["Legal/compliance review of explanation content for regulated flows","UX review of explanation clarity"],"examples":["A support bot cites knowledge-base article IDs with every policy answer","An underwriting assistant produces an auditor-facing feature contribution summary"],"references":[{"title":"NIST AI RMF — Explainability and Interpretability","url":"https://www.nist.gov/itl/ai-risk-management-framework"}],"futureEvolution":"Standard explanation schemas for agentic and tool-mediated decisions."},{"id":"APRF-18","slug":"data-privacy","name":"Data Privacy","summary":"Minimize, classify, and protect data flowing through AI pipelines.","domain":"data","crossCutting":false,"severity":"critical","riskLevel":"high","purpose":"Minimize collection and exposure of personal and sensitive data across prompts, retrieval, memory, logs, and third-party model providers.","engineeringPhilosophy":"Privacy by design for AI means data minimization, purpose limitation, residency awareness, and contractual/technical controls on processors—not only a privacy policy page.","whyItMatters":"AI pipelines concentrate sensitive data. Over-collection and uncontrolled provider sharing create regulatory, contractual, and trust failures even when 'security' looks fine.","commonFailures":["Sending full customer records to external models when aggregates would do","No data classification for prompt contents","Training or eval sets with uncleared personal data","Ignoring residency requirements when choosing providers"],"mandatoryChecks":[{"id":"PRI-M1","requirement":"Data flowing to models shall be classified; sensitive classes shall have documented handling rules","artifact":"Data classification scheme + handling rules for model-bound payloads","passCondition":"Classification scheme covers AI payload classes; ≥1 automated or sampled audit shows 100% of sampled production model requests tagged with a class, and sensitive-class handling matches policy","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"PRI-M2","requirement":"Third-party model processing terms shall be reviewed for training use and retention","artifact":"Vendor terms review record (DPA/training/retention) with date and owner","passCondition":"100% of production model providers have a review record ≤ 12 months covering training use and retention; 0 unreviewed providers in the inventory","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"PRI-M3","requirement":"Users or tenants shall have deletion/export paths covering AI memory and logs where required","artifact":"Deletion/export procedure + successful test execution record","passCondition":"Documented API/runbook covers AI memory and in-scope logs; test shows deletion/export completes within SLA for a sample tenant/user (measured duration recorded)","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"PRI-M4","requirement":"Residency-constrained routing shall be enforced for regulated workloads","artifact":"Routing policy config + sample routing decisions for regulated tenants/regions","passCondition":"PASS if regulated workloads are labeled and 100% of sampled regulated requests stay in approved regions","method":"hybrid","requiredFromLevel":4,"minCriticality":3},{"id":"PRI-M5","requirement":"DPIA/PIA equivalents shall exist for major AI features before production","artifact":"Completed DPIA/PIA (or equivalent) with sign-off for each major in-scope AI feature","passCondition":"PASS if 100% of major in-scope AI features have a completed assessment with owner sign-off before production traffic","method":"hybrid","requiredFromLevel":4,"minCriticality":3}],"recommendedChecks":[{"id":"PRI-R1","requirement":"Tokenization or redaction before model calls for high-sensitivity fields","artifact":"Pre-model tokenization/redaction pipeline config + sample of redacted payloads for high-sensitivity fields","passCondition":"Documented high-sensitivity fields never appear in cleartext in model request samples (≥50 sampled calls); pipeline fails closed when redaction errors","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Data flow diagrams for AI features","Provider DPA/terms review records","Deletion/export procedure evidence"],"engineeringBestPractices":["Default to least data in context","Separate privacy tiers of models (e.g., self-hosted for sensitive)","Ban pasting production data into consumer AI tools for debugging","Review eval datasets for personal data before publication"],"automaticValidations":["DLP scanning on outbound model payloads where feasible","Policy checks preventing disallowed data classes from leaving the VPC","Automated deletion job verification"],"manualValidations":["Privacy review for new AI features","Vendor assessments for new model providers"],"examples":["Support transcripts are redacted before being sent to an external LLM","A healthcare workload uses a residency-locked or self-hosted model path"],"references":[{"title":"NIST AI RMF — privacy considerations","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"OWASP LLM — Sensitive Information Disclosure","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"}],"futureEvolution":"Machine-readable privacy labels for model endpoints and agent capabilities."},{"id":"APRF-27","slug":"data-governance","name":"Data Governance & Quality","summary":"Govern corpora, indexes, labels, and feedback loops—distinct from privacy controls.","domain":"data","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Own lineage, quality, labeling, retrieval-index governance, and train/serve consistency for data that shapes AI behavior.","engineeringPhilosophy":"Privacy asks who may see data; governance asks whether the data is fit for purpose. Bad corpora and poisoned indexes create silent production failures.","whyItMatters":"Hallucinations, biased outcomes, and retrieval failures often originate in unmanaged data—not in the model alone.","commonFailures":["RAG indexes with no freshness or ownership","Eval and fine-tune sets without label quality review","Feedback loops that reinforce errors into memory or training","No lineage from production answer to source document version"],"mandatoryChecks":[{"id":"DG-M1","requirement":"Production retrieval corpora and indexes shall have owners, versioning, and refresh cadence","artifact":"Corpus/index inventory with owner, version ID, and refresh schedule","passCondition":"100% of production indexes have owner + version ID + refresh cadence; 0 indexes missing any field; stale beyond cadence flagged or rebuilt","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"DG-M2","requirement":"Data used for evals or fine-tuning shall have documented provenance and quality criteria","artifact":"Dataset cards or registry entries for eval/fine-tune sets","passCondition":"100% of datasets used in production gates or fine-tunes have provenance + quality criteria documented; promotion blocked if missing (CI or review checklist evidence)","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"DG-M3","requirement":"Feedback or memory promotion paths that affect future answers shall be governed","artifact":"Promotion policy + write-path controls for feedback→memory/training","passCondition":"100% of promotion paths require policy check or human approval; tests show ungated promotion to durable memory/training is denied","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"DG-R1","requirement":"Automated freshness and coverage metrics for critical corpora","artifact":"Corpus freshness dashboard (or job report) + alert config export covering each critical corpus ID","passCondition":"Each critical corpus has a documented freshness SLO (e.g. max age hours); ≥95% of sampled docs meet the SLO in the last 7 days; alert fires on freshness breach","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"DG-R2","requirement":"Train/serve skew monitoring for features or embeddings","artifact":"Skew monitor config + last weekly skew report comparing train vs serve feature/embedding distributions","passCondition":"Skew job ran within the last 7 days for every production embedding/feature pipeline; at least one documented threshold exists; breach creates a tracked ticket or page","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"DG-R3","requirement":"Dataset cards for major eval and fine-tune sets","artifact":"Dataset cards (or equivalent metadata) for each major eval and fine-tune corpus in the registry","passCondition":"100% of major eval/fine-tune sets used in production promotion have a dataset card with purpose, source, PII handling, and last-updated date ≤12 months","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Corpus/index inventory with owners","Data quality criteria and sample reviews","Lineage examples from answer to source version"],"engineeringBestPractices":["Treat indexes as release artifacts with version IDs","Separate experimental corpora from production indexes","Review user-generated feedback before it becomes training signal"],"automaticValidations":["Index build pipelines emit version and checksum","Alerts on stale critical corpora","CI checks that eval datasets meet schema/quality gates"],"manualValidations":["Periodic corpus content audits","Label quality sampling"],"examples":["A policy RAG index is rebuilt nightly with a new version ID and canary before cutover","Fine-tune data requires dual review before entering the training set"],"references":[{"title":"NIST AI RMF — Map and Measure (data)","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"ML systems engineering — data quality practices","url":"https://ml-ops.org/"}],"futureEvolution":"Portable dataset and index manifests with integrity proofs across vendors."},{"id":"APRF-04","slug":"memory-management","name":"Memory Management","summary":"Control retention, isolation, and poisoning of short- and long-term memory.","domain":"data","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Govern what the system remembers across turns and sessions—including user memory, agent state, and shared multi-agent memory—with retention, isolation, and integrity controls.","engineeringPhilosophy":"Memory is persistent attack surface and privacy liability. Default to minimal retention, strict tenancy isolation, and verifiable write paths. Memory that cannot be deleted or audited is not production-ready.","whyItMatters":"Poisoned or leaked memory causes lasting incorrect behavior, cross-tenant data exposure, and compliance failures. Autonomous agents amplify memory risk by writing their own state.","commonFailures":["Shared memory stores without tenant isolation","No user-facing or admin deletion path","Agents write unverified 'facts' into long-term memory","Memory reused across products without re-consent or reclassification"],"mandatoryChecks":[{"id":"MEM-M1","requirement":"Memory shall be isolated by tenant (and user where required) with tested boundaries","artifact":"Isolation tests for memory store APIs","passCondition":"0 successful cross-tenant (and cross-user where required) memory reads/writes across ≥10 automated attack cases","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"MEM-M2","requirement":"Retention and deletion policies shall exist and be executable","artifact":"Retention policy + TTL job config + deletion test record","passCondition":"TTL/deletion job succeeds in test; sample records older than retention are absent after job run; policy documents retention periods per memory class","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"MEM-M3","requirement":"Writes to long-term memory shall be validated against policy (who/what may write)","artifact":"Write-policy middleware config + deny tests","passCondition":"Unauthorized writers denied at 100% in tests; policy enumerates allowed writers and content classes for durable memory","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"MEM-M4","requirement":"Critical memory records shall have cryptographic or signed integrity protection","artifact":"Integrity/signing design + verification sample for critical memory stores","passCondition":"PASS if critical memory classes are inventoried and integrity verification succeeds in the latest check (≤90 days)","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"MEM-R1","requirement":"Memory poisoning tests in the eval suite","artifact":"Memory-poisoning cases in the adversarial eval suite + latest run report with pass/fail per case","passCondition":"≥5 poisoning scenarios (cross-tenant write, prompt-in-memory, stale trusted fact) are executed ≤90 days; critical fails block or require risk acceptance","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"MEM-R3","requirement":"Separate working memory from durable memory with promotion rules","artifact":"Memory architecture doc + config showing working vs durable stores and promotion rules + sample promotion audit","passCondition":"Working memory cannot silently become durable without a promotion rule; last 10 promotions show rule ID and actor; TTL differs by memory class","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"evidenceRequired":["Data retention schedule for AI memory stores","Isolation test results","Deletion and export procedure evidence"],"engineeringBestPractices":["Classify memory types: ephemeral session, user profile, agent scratchpad, org knowledge","Require confidence or human confirmation before promoting scratch facts to durable memory","Encrypt memory at rest; restrict operational access","Log memory writes for audit without logging raw sensitive values when avoidable"],"automaticValidations":["Automated tenant isolation tests","TTL enforcement jobs","Write-policy middleware on memory APIs"],"manualValidations":["Privacy review of memory schemas","Red-team attempts to poison shared agent memory"],"examples":["A personal assistant forgets user preferences on request within SLA","Multi-agent systems use namespaced memory; one agent cannot read another's private scratchpad"],"references":[{"title":"NIST AI RMF — Govern and Manage","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"OWASP LLM — Training Data Poisoning / Vector and Embedding Weaknesses","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"}],"futureEvolution":"Portable memory consent receipts and cross-vendor memory export formats with integrity proofs."},{"id":"APRF-07","slug":"model-governance","name":"Model Governance","summary":"Own model selection, versioning, deprecation, and capability boundaries.","domain":"model-lifecycle","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Control which models (hosted or self-hosted) may be used in which environments, how they are versioned, evaluated, deprecated, and constrained by capability.","engineeringPhilosophy":"Models are dependencies with behavioral contracts. Pin versions, document capabilities and known failure modes, and never silently change the production model under a customer-facing workload.","whyItMatters":"Provider default model changes, shadow upgrades, and unvetted open weights introduce quality, cost, security, and compliance regressions that teams discover only after incidents.","commonFailures":["Using 'latest' model aliases in production","No inventory of models and where they run","Self-hosted models without supply-chain or license review","Routing to cheaper models without eval coverage for those routes"],"mandatoryChecks":[{"id":"MOD-M1","requirement":"Production shall use pinned model identifiers; aliases like “latest” are forbidden on critical paths","artifact":"Production model pin config + lint/CI rule rejecting “latest” aliases","passCondition":"0 “latest”/floating aliases on critical production paths in config lint; 100% of critical paths reference immutable model IDs","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"MOD-M2","requirement":"Model inventory shall exist with owners, data residency, and intended use","artifact":"Model inventory registry export","passCondition":"100% of production models have owner, residency, and intended-use fields; query returns 0 incomplete rows","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"MOD-M3","requirement":"Model changes shall require evaluation evidence before promotion","artifact":"Promotion policy + CI/registry requiring eval artifact on model version bumps","passCondition":"100% of production model promotions in last 30 days have linked eval pass artifacts; promote without eval is blocked by gate","method":"automated","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"MOD-R1","requirement":"Deprecation and sunset policy for models and embeddings","artifact":"Model/embedding deprecation and sunset policy + registry entries showing sunset dates for superseded pins","passCondition":"Policy defines notice period and forced-sunset rules; ≥1 superseded production model/embedding has a sunset date in the registry; no undocumented pins past sunset without exception","method":"manual","requiredFromLevel":4,"minCriticality":2},{"id":"MOD-R2","requirement":"Capability allowlists (e.g., code execution, vision) per workload","artifact":"Per-workload capability allowlist config (code execution, vision, browsing, etc.) + deny logs","passCondition":"Each production workload has an explicit capability allowlist; a denied capability attempt is recorded in test or prod within 90 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"MOD-R3","requirement":"License and provenance review for open-weight and fine-tuned models","artifact":"License/provenance review checklist + completed reviews for each open-weight or fine-tuned production model","passCondition":"100% of open-weight/fine-tuned production models have a license+provenance review ≤12 months old; blocked licenses have documented exceptions with expiry","method":"manual","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Model inventory and version pins","Change records for model promotions","Eval comparisons for model swaps"],"engineeringBestPractices":["Abstract provider SDKs behind an internal model gateway with policy","Track embedding model versions separately—index rebuilds are part of change","Document fallback models and degraded modes","Separate experimental models into non-production projects"],"automaticValidations":["Config lint rejecting unpinned model IDs in production configs","Gateway deny rules for unapproved models","CI requiring eval artifact for model version bumps"],"manualValidations":["Architecture review when introducing a new model family","Legal/compliance review for data-processing terms of new providers"],"examples":["Production chat uses provider-model-2025-03-01; canaries test the next pin before cutover","A coding agent is restricted to models that meet internal code-exfiltration evals"],"references":[{"title":"NIST AI RMF — Map","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"AWS Well-Architected — Operational Excellence (change management parallels)","url":"https://aws.amazon.com/architecture/well-architected/"}],"futureEvolution":"Model bill of materials (MBOM) standards and portable evaluation attestations across vendors."},{"id":"APRF-02","slug":"prompt-engineering","name":"Prompt Engineering","summary":"Treat prompts as versioned production artifacts with regression control.","domain":"model-lifecycle","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Manage system, developer, and tool prompts as first-class production artifacts—versioned, reviewed, tested, and rolled back like application code.","engineeringPhilosophy":"Prompts are code. They change behavior, create regressions, and carry security and cost implications. Informal edits in dashboards without lineage are an anti-pattern for production AI.","whyItMatters":"Silent prompt drift is a leading cause of quality and safety regressions. Without versioning and eval gates, teams cannot answer what changed when production behavior breaks.","commonFailures":["Editing production prompts directly in a vendor console with no history","No golden-set regression tests when prompts change","Mixing business logic, safety policy, and tone in one unmaintainable blob","Different prompts per environment with no promotion path"],"mandatoryChecks":[{"id":"PRM-M1","requirement":"Every production prompt shall have an immutable version identifier and owner","artifact":"Prompt registry export","passCondition":"100% of production prompts have immutable version ID + owner; 0 unversioned production prompts","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"PRM-M2","requirement":"Prompt changes shall go through review and be tied to evaluation results before release","artifact":"Change records linking prompt versions to review + eval artifacts","passCondition":"100% of production prompt releases in last 30 days have review ID and eval pass artifact; promote without both is blocked","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"PRM-M3","requirement":"Rollback to a prior prompt version shall be possible without redeploying unrelated services","artifact":"Prompt rollback procedure + timed restore test","passCondition":"Prior prompt version restored in ≤ documented RTO without full app redeploy; demonstrated in last 90 days","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"PRM-R1","requirement":"Prompt templates are parameterized; secrets and PII never hardcoded","artifact":"Prompt template inventory showing parameters + static analysis or review finding zero hardcoded secrets/PII","passCondition":"100% of production prompt templates are parameterized for variable inputs; scan of templates finds 0 secrets and 0 hardcoded customer PII fields","method":"hybrid","requiredFromLevel":2,"minCriticality":1},{"id":"PRM-R2","requirement":"Prompt linting for length, forbidden patterns, and injection-prone constructs","artifact":"Prompt-lint CI config (length, forbidden patterns, injection-prone constructs) + latest lint report","passCondition":"Lint runs on every prompt change PR; blocking rules exist for secrets patterns and unbounded user concatenation; last failing lint example retained","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"PRM-R3","requirement":"A/B or shadow evaluation for high-traffic prompt changes","artifact":"A/B or shadow-eval config for high-traffic prompt changes + last experiment report with quality/safety deltas","passCondition":"Last high-traffic prompt change used A/B or shadow eval with pre-registered metrics; promotion required non-inferiority (or better) on safety and primary quality SLI","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Prompt registry with versions, owners, and change history","Eval reports linked to prompt releases","Rollback procedure documentation"],"engineeringBestPractices":["Store prompts in source control or a versioned registry with CI promotion","Separate safety/policy prompts from product/UX prompts where possible","Measure token cost impact of prompt changes alongside quality","Document intended failure modes and refusal behavior in the prompt contract"],"automaticValidations":["CI blocks deploy if prompt hash lacks associated eval pass","Diff reviews in pull requests for prompt files","Token budget checks against historical baselines"],"manualValidations":["Peer review of safety-critical prompt sections","Product review of tone and policy alignment for major releases"],"examples":["A coding agent system prompt is tagged v47; a quality drop triggers rollback to v46 in minutes","Marketing copy prompts and refund-policy prompts live in separate modules with different owners"],"references":[{"title":"Google SRE Workbook — Managing Risk (change management parallels)","url":"https://sre.google/workbook/"},{"title":"OWASP LLM Prompt Injection guidance","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"}],"futureEvolution":"Industry-standard prompt manifests (SBOM-like) and signed prompt releases for regulated environments."},{"id":"APRF-03","slug":"context-engineering","name":"Context Engineering","summary":"Bound and structure what the model is allowed to see and use.","domain":"model-lifecycle","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Design how context windows are filled—retrieval, history, tool results, and system state—so the model receives necessary, authorized, and bounded information.","engineeringPhilosophy":"Context is a scarce, trusted resource. Overstuffing, under-filtering, and leaking privileged context are engineering defects, not model quirks. Structure beats volume.","whyItMatters":"Poor context design causes hallucinations, privacy leaks, prompt injection via retrieved content, and unbounded cost. Agents fail when context is noisy or when critical constraints are buried.","commonFailures":["Dumping entire conversation history without summarization or TTL","Retrieving documents without ACL enforcement at query time","No provenance: model cannot distinguish user text from retrieved policy","Tool results concatenated without size or sensitivity limits"],"mandatoryChecks":[{"id":"CTX-M1","requirement":"Context assembly shall enforce maximum sizes and prioritization rules","artifact":"Context-budget config + unit/integration tests for overflow behavior","passCondition":"100% of production context builders enforce a max token/byte budget; tests show oversized inputs are truncated/rejected per priority rules with 0 silent overflows","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"CTX-M2","requirement":"Retrieved and tool-sourced content shall be labeled and access-checked before inclusion","artifact":"Context assembly code/tests showing source labels + ACL checks on retrieval/tool results","passCondition":"Automated tests: unauthorized retrieval chunks are excluded at 100%; included chunks carry a source label/type field in 100% of sampled assembled contexts","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"CTX-M3","requirement":"Sensitive context classes (secrets, regulated data) shall have explicit inclusion policies","artifact":"Data-class inclusion policy + enforcement tests or DLP hooks on context assembly","passCondition":"Policy enumerates sensitive classes and allow/deny rules; tests show disallowed classes are stripped or blocked at ≥95% on the sensitive-class fixture suite","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"CTX-R1","requirement":"Context budgets are monitored per request with alerts on saturation","artifact":"Per-request context token/budget metrics dashboard + alert rule export for saturation thresholds","passCondition":"≥99% of production requests in a 24h sample emit context-budget usage; alert fires when usage exceeds the documented % of max context (or hard truncate rate exceeds threshold)","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"CTX-R2","requirement":"Summarization or compaction strategies are tested for information loss on critical facts","artifact":"Compaction/summarization eval set of critical facts + latest information-loss report","passCondition":"Critical-fact retention ≥ documented threshold after compaction on the eval set; last run ≤90 days; regressions block context-pipeline releases","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"CTX-R3","requirement":"Structured context blocks (JSON/XML sections) separate instructions from data","artifact":"Context assembly schema/spec separating instruction vs data blocks + sample rendered request","passCondition":"Production assembler emits labeled structured sections (e.g. JSON/XML); untrusted data cannot overwrite the instruction section in a red-team or unit test","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"evidenceRequired":["Context assembly design document","ACL enforcement evidence on retrieval paths","Samples of labeled context structures (redacted)"],"engineeringBestPractices":["Design explicit context slots: instructions, user, memory, retrieval, tools","Prefer citations and IDs over pasting large documents when possible","Apply the same authorization checks to context as to API responses","Test adversarial documents in the corpus as first-class fixtures"],"automaticValidations":["Unit tests for context builders and ACL filters","Metrics on context token composition by source","Guards rejecting oversized tool payloads"],"manualValidations":["Review of retrieval corpora for injection and oversharing risk","Spot checks that production traces show expected context structure"],"examples":["A RAG system includes document IDs and ACL tags; the model only sees chunks the user may access","Agent tool output is truncated and classified before re-entering the context window"],"references":[{"title":"NIST AI RMF — Map and Measure functions","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"OWASP LLM — Sensitive Information Disclosure","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"}],"futureEvolution":"Standard context schemas across agent frameworks and portable ACL-aware retrieval protocols."},{"id":"APRF-08","slug":"evaluation","name":"Evaluation","summary":"Continuously prove quality, safety, and task success before and after release.","domain":"model-lifecycle","crossCutting":false,"severity":"critical","riskLevel":"high","purpose":"Institutionalize offline and online evaluation so releases are gated on measured quality, safety, and task success—not intuition or demo success.","engineeringPhilosophy":"If you cannot measure it, you cannot claim production readiness. Evaluation is a product dependency equal to tests in traditional software, extended for stochastic systems.","whyItMatters":"LLMs regress silently. Without evals, teams ship prompt or model changes that increase hallucination, toxicity, tool errors, or task failure rates under real traffic.","commonFailures":["Only manual spot checks before release","Eval sets that do not match production task distribution","No online monitoring of quality after deploy","Safety and quality evals owned by nobody"],"mandatoryChecks":[{"id":"EVL-M1","requirement":"Critical customer journeys shall have automated offline eval suites on every relevant change","artifact":"Eval suite registry mapped to journeys + CI runs on prompt/model/tool changes","passCondition":"100% of journeys marked critical have a versioned offline suite; 100% of relevant production changes in last 30 days triggered the suite (or documented waiver ≤ 14 days)","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"EVL-M2","requirement":"Release gates shall enforce numeric minimum thresholds for quality and safety metrics","artifact":"Gate config with metric names and thresholds + CI reports","passCondition":"Each critical journey has ≥1 quality and ≥1 safety metric with numeric threshold; failing gate blocks deploy in CI","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"EVL-M3","requirement":"Production shall have online signals for task success/failure and safety refusals","artifact":"Dashboard/metrics for task success and safety refusal rates","passCondition":"Online metrics exist and are updating for task success/failure and safety refusals; alert or review cadence defined; freshness ≤ 24 hours on the dashboard","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"EVL-M4","requirement":"Shadow deployment with eval comparison shall precede full cutover for high-risk AI changes","artifact":"Shadow/canary eval config + comparison report for the last high-risk cutover","passCondition":"PASS if the last high-risk AI cutover retained a shadow/canary comparison that met promotion criteria before 100% traffic","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"EVL-R1","requirement":"Separate eval tracks for regression, adversarial, and distribution-shift cases","artifact":"Eval catalog showing separate tracks for regression, adversarial, and distribution-shift + latest CI matrix","passCondition":"All three tracks exist with distinct corpora and owners; each track ran successfully on the last production model/prompt promotion (≤90 days)","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"EVL-R2","requirement":"Human preference or expert review sampling on a defined cadence","artifact":"Human preference / expert-review sampling protocol + last scored sample set with inter-rater notes","passCondition":"Cadence and sample size are defined; last sample ≤90 days covers production-like prompts; disagreements have adjudication recorded","method":"manual","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Eval suite definitions and ownership","Gate thresholds and recent pass/fail reports","Dashboards for online quality metrics"],"engineeringBestPractices":["Version eval datasets; treat label quality as seriously as code quality","Include tool-use and multi-turn scenarios, not only single-shot Q&A","Budget for eval compute as part of CI cost","Correlate eval failures to traces for fast debugging"],"automaticValidations":["CI pipelines blocking on eval gate failures","Scheduled regression runs against production prompts/models","Drift detection when online metrics diverge from offline baselines"],"manualValidations":["Periodic calibration of automated graders vs human labels","Product review of eval coverage gaps"],"examples":["A RAG support bot cannot ship a prompt change unless citation accuracy stays above threshold","An agent fails CI when tool-selection accuracy drops on the golden set"],"references":[{"title":"NIST AI RMF — Measure","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"Google SRE — Monitoring Distributed Systems (measurement culture)","url":"https://sre.google/sre-book/monitoring-distributed-systems/"}],"futureEvolution":"Shared public eval protocols for agentic and MCP workloads, with portable scorecards."},{"id":"APRF-06","slug":"agent-governance","name":"Agent Governance","summary":"Constrain goals, autonomy, loops, and escalation for agents and A2A.","domain":"agents","crossCutting":false,"severity":"critical","riskLevel":"critical","purpose":"Define and enforce how autonomous agents set goals, plan, loop, escalate, and collaborate with other agents so autonomy never exceeds authorized operational bounds.","engineeringPhilosophy":"Autonomy is a dial, not a virtue. Production agents have explicit charters: allowed goals, max steps, escalation paths, and kill switches. Multi-agent and A2A systems require governance of the mesh, not only of each node.","whyItMatters":"Unconstrained agents loop, spend unbounded budget, escalate privileges across A2A links, or pursue goals that conflict with business and safety policy. Governance failures look like 'the agent went rogue' but are usually missing product controls.","commonFailures":["No maximum step or time budget for agent loops","Agents that can redefine their own goals without oversight","A2A trust assumed by network presence alone","No kill switch or pause for runaway agents"],"mandatoryChecks":[{"id":"AGN-M1","requirement":"Each production agent shall have a documented charter: purpose, allowed tools, data scope, and autonomy limits","artifact":"Agent inventory with charter fields per production agent","passCondition":"100% of production agents have a charter with purpose, tool allowlist reference, data scope, and autonomy limits; inventory query returns 0 agents missing any required field","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"AGN-M2","requirement":"Hard limits shall exist on steps, wall-clock time, and recursive agent spawning","artifact":"Runtime config for max steps, wall-clock timeout, and spawn depth + enforcement tests","passCondition":"100% of production agents have finite max-steps, wall-clock timeout, and spawn-depth ≤ configured bound; tests show limits are enforced (run aborts) when exceeded","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"AGN-M3","requirement":"Operators shall be able to pause or terminate agent runs in production","artifact":"Kill-switch/pause control documentation + successful drill or production use record","passCondition":"Documented pause/terminate control exists for production agents; ≥1 successful pause or terminate action in drill or production in last 90 days with recorded time-to-effect ≤ documented SLO","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"AGN-M4","requirement":"A2A and multi-agent handoffs shall authenticate peers and carry scoped capabilities","artifact":"A2A auth config + capability-token/schema examples + negative auth tests","passCondition":"100% of production A2A/multi-agent handoff paths require authenticated peers; tests show unauthenticated or over-scoped handoffs denied at 100%","method":"automated","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"AGN-R1","requirement":"Goal conflict detection and policy checks before plan execution","artifact":"Pre-execution policy checks for agent plans (goal-conflict / disallowed-goal rules) + sample deny/allow traces","passCondition":"Agent planner runs policy checks before side-effecting tools; ≥1 synthetic conflict is denied in test or prod logs within 90 days; rules have a named owner","method":"manual","requiredFromLevel":4,"minCriticality":2},{"id":"AGN-R2","requirement":"Simulation/sandbox runs for new agent behaviors before production","artifact":"Sandbox/simulation environment config for agents + last pre-prod simulation report for a behavior change","passCondition":"Last new agent behavior promoted to production has a linked sandbox run ≤30 days before release with pass/fail criteria recorded","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"AGN-R3","requirement":"Formal RACI for agent ownership across teams","artifact":"RACI or ownership register entry for the system","passCondition":"Every production AI system ID has non-empty owner fields for required domains; inventory query returns 0 orphans","method":"manual","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Agent inventory with charters and owners","Loop/budget configuration evidence","Kill-switch runbook and test record"],"engineeringBestPractices":["Prefer narrow agents composed into workflows over one universal agent","Log plans and intermediate decisions for reconstructability","Escalate to humans on ambiguity or high-impact branches","Treat agent-to-agent messages as untrusted until verified"],"automaticValidations":["Runtime enforcement of step and cost budgets","Denial of unauthorized A2A peers","Alerts on loop detection and repeated identical tool calls"],"manualValidations":["Charter review for each new agent class","Tabletop of multi-agent failure cascades"],"examples":["A research agent may browse and summarize but cannot send email; a separate approved workflow handles outreach","A2A task delegation includes expiring capability tokens, not open network trust"],"references":[{"title":"OWASP LLM — Excessive Agency","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"},{"title":"NIST AI RMF — Govern","url":"https://www.nist.gov/itl/ai-risk-management-framework"}],"futureEvolution":"Interoperable agent identity and capability attestation standards for A2A ecosystems."},{"id":"APRF-17","slug":"human-approval","name":"Human Approval","summary":"Require human gates for high-impact or irreversible actions.","domain":"agents","crossCutting":false,"severity":"critical","riskLevel":"high","purpose":"Insert human approval and oversight where AI-proposed actions have high impact, irreversibility, or regulatory sensitivity.","engineeringPhilosophy":"Human-in-the-loop is not a product apology—it is a control. The goal is calibrated oversight: humans approve what machines should not decide alone.","whyItMatters":"Autonomous execution of irreversible actions (money movement, data deletion, external communications, code deploy) without approval creates unacceptable blast radius.","commonFailures":["Rubber-stamp UIs that hide risk details","Approval only in UI while API/agent path bypasses it","No dual control for mission-critical actions","Approvals without audit trail"],"mandatoryChecks":[{"id":"HUM-M1","requirement":"High-impact action classes shall be inventoried and gated by human approval in production","artifact":"High-impact action inventory + gate wiring evidence","passCondition":"100% of inventoried high-impact action classes have an approval gate in production; ungated execution tests fail at 100%","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"HUM-M2","requirement":"Approval decisions shall be logged with actor, context, and outcome","artifact":"Approval audit log schema + sample records","passCondition":"100% of sampled approvals in last 30 days include actor ID, action context, and approve/deny outcome; schema validation test passes","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"HUM-M3","requirement":"Approval shall not be bypassable via alternate agent or API paths","artifact":"Bypass-path threat tests across UI, API, and agent entry points","passCondition":"Automated or reviewed tests attempt bypass via alternate paths; 0 successful ungated high-impact executions","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"HUM-M4","requirement":"Dual control shall be required for Level 5 irreversible actions","artifact":"Dual-control workflow config + sample approval records for irreversible actions","passCondition":"PASS if Level 5 irreversible actions are inventoried and 100% of sampled executions show dual approval; 0 single-approver completes in the sample","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"HUM-R1","requirement":"Risk-based UI showing tool args, diffs, and confidence before approve","artifact":"Approval UI screenshots/spec showing tool args, diffs, and confidence + sample approval records","passCondition":"High-impact approvals display tool args, change diff (or equivalent), and confidence/risk; ≥10 sampled approvals in 90 days show those fields populated","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"HUM-R3","requirement":"SLA for approval queues to avoid unsafe workarounds","artifact":"Approval-queue SLA definition + queue age/metrics dashboard for the last 30 days","passCondition":"Documented SLA (e.g. p95 queue age) exists; measured p95 for the last 30 days is within SLA or open exceptions have owners and expiry","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["High-impact action inventory","Sample approval audit logs","Architecture proof that bypass paths are closed"],"engineeringBestPractices":["Classify actions by impact tier with default gates","Allow break-glass with heightened logging for emergencies","Never ask the model to 'pretend' approval was granted","Design approvals for voice and async channels, not only web UI"],"automaticValidations":["Tests that gated tools fail without approval tokens","Monitoring of approval bypass attempts","Alerts on abnormal approval velocity"],"manualValidations":["Process review of approval fatigue risk","Sample audits of approved vs executed actions"],"examples":["A coding agent opens a PR but cannot merge to main without human review","A finance agent queues wire transfers for dual approval above a threshold"],"references":[{"title":"NIST AI RMF — Govern and Manage (human oversight)","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"OWASP LLM — Excessive Agency","url":"https://owasp.org/www-project-top-10-for-large-language-model-applications/"}],"futureEvolution":"Standard approval protocols for cross-system agents with portable audit receipts."},{"id":"APRF-09","slug":"observability","name":"Observability","summary":"Make every decision path reconstructable: traces, prompts, tools, costs, outcomes.","domain":"reliability","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Provide end-to-end observability across prompts, model calls, retrieval, tools, agents, costs, and outcomes so incidents and regressions are diagnosable.","engineeringPhilosophy":"AI systems without reconstructable traces are not operable. Observability must cover the cognitive path, not only HTTP latency—while respecting privacy redaction.","whyItMatters":"When quality drops or a tool fires incorrectly, teams need to replay what the model saw and did. Blind production AI creates unresolvable incidents and slow learning loops.","commonFailures":["Logging only final user-visible answers","No correlation IDs across model, tool, and app spans","Storing raw sensitive prompts indefinitely","Cost metrics disconnected from product features"],"mandatoryChecks":[{"id":"OBS-M1","requirement":"Distributed traces shall link user request → model calls → tools → outcome","artifact":"Sample production traces (redacted) + tracing config","passCondition":"Canary or sampled traces show parent request ID linking model spans, tool spans, and outcome in ≥95% of canary requests over 24 hours","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"OBS-M2","requirement":"Token usage and cost shall be attributable per request/feature/tenant","artifact":"Cost/token metrics with request, feature, and tenant labels","passCondition":"≥95% of billed model calls in a 24h sample carry request ID + feature + tenant (or equivalent) attribution labels","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"OBS-M3","requirement":"Sensitive fields in traces shall be redacted or access-controlled","artifact":"Trace redaction/ACL config + synthetic sensitive-field test","passCondition":"Synthetic secrets/PII in traced fields are redacted or inaccessible to unauthorized roles at 100% in tests","method":"automated","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"OBS-R1","requirement":"Replay tooling for failed traces in a secure environment","artifact":"Secure replay tooling config + sample replay session for a failed production trace (redacted)","passCondition":"On-call can replay a failed AI trace in a restricted environment within the documented RTO; last drill or real replay ≤90 days old","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"OBS-R2","requirement":"Quality annotations attachable to traces for closed-loop eval","artifact":"Trace annotation schema + tooling config allowing quality labels on spans + sample annotated traces","passCondition":"Annotators can attach quality labels to production traces in a secure tool; ≥50 annotations in the last 90 days feed an eval or review loop","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"OBS-R3","requirement":"SLO dashboards for AI latency, error, and quality burn rates","artifact":"SLO/alert configuration export + dashboard screenshot or URL with retention note","passCondition":"Named SLOs with numeric targets exist for each critical AI journey; alert fires when burn exceeds documented threshold","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Example traces (redacted) showing full path","Cost attribution dashboards","Access control policy for prompt/completion logs"],"engineeringBestPractices":["Adopt OpenTelemetry-compatible tracing for LLM spans where possible","Capture tool names, latency, and status codes consistently","Separate debug retention from long-term analytics retention","Alert on silent quality degradation, not only 5xx rates"],"automaticValidations":["Synthetic checks that traces are emitted for canary requests","Alerts on missing cost telemetry","PII scanners on log pipelines"],"manualValidations":["Incident retrospectives verify traces were sufficient","Privacy review of observability data stores"],"examples":["An on-call engineer opens a trace and sees which retrieved chunk caused an incorrect refund recommendation","Finance reconciles model spend by product SKU using request-level cost tags"],"references":[{"title":"Google SRE Workbook — Observability","url":"https://sre.google/workbook/observability/"},{"title":"OpenTelemetry","url":"https://opentelemetry.io/"}],"futureEvolution":"Standard LLM span semantics across providers and agent frameworks."},{"id":"APRF-28","slug":"performance-slo","name":"Performance & SLO Engineering","summary":"Latency, throughput, and error budgets for AI features—SRE discipline applied to GenAI.","domain":"reliability","crossCutting":false,"severity":"high","riskLevel":"medium","purpose":"Define SLIs/SLOs for AI latency, availability, and quality signals; manage error budgets so AI features meet user experience and capacity targets.","engineeringPhilosophy":"If it is not measured with an SLO, it is not operationally owned. Stochastic systems still need latency and quality budgets.","whyItMatters":"Unbounded TTFT, streaming stalls, and silent quality burn destroy UX and hide regressions that “availability” alone will not catch.","commonFailures":["No p95/p99 latency targets for model calls","Quality treated as a one-time eval, not an online SLO","No error budget for AI feature releases","Capacity planning ignores bursty agent tool loops"],"mandatoryChecks":[{"id":"PERF-M1","requirement":"Critical AI user journeys shall have documented latency and availability SLOs","artifact":"SLO catalog with numeric targets per journey","passCondition":"100% of journeys marked critical have availability % and latency percentile targets recorded","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"PERF-M2","requirement":"Online dashboards shall track AI latency, error, and at least one quality or task-success signal","artifact":"Dashboard URLs/config showing latency, error, and quality/task-success panels","passCondition":"Dashboard panels for latency, error rate, and ≥1 quality/task-success signal are live with freshness ≤ 15 minutes","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"PERF-M3","requirement":"Alerting shall exist when SLOs burn beyond defined thresholds","artifact":"Burn-rate or SLO alert policies","passCondition":"Alert policies exist for each critical journey SLO; alert test or documented fire demonstrates notification path works","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"PERF-R1","requirement":"Error budgets formally gate release velocity for AI features","artifact":"Error-budget policy linking AI SLOs to release freezes + last budget burn report that gated a release","passCondition":"When error budget for a critical AI journey is exhausted, release velocity is blocked or requires explicit risk acceptance; ≥1 gated event or drill in 90 days","method":"automated","requiredFromLevel":4,"minCriticality":2},{"id":"PERF-R2","requirement":"Capacity and load tests include adversarial long-prompt and agent-loop scenarios","artifact":"Load-test plan including long-prompt and agent-loop scenarios + latest capacity test report","passCondition":"Last capacity test ≤90 days includes adversarial long prompts and multi-step agent loops; p95 latency and error rate stay within SLO under documented concurrency","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"PERF-R3","requirement":"Streaming-specific SLIs (TTFT, inter-token latency) for streaming UIs","artifact":"Streaming SLI dashboard for TTFT and inter-token latency + alert config for streaming UIs","passCondition":"TTFT and inter-token latency SLIs exist for each streaming AI surface; alerts fire on documented thresholds; series retained ≥30 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["SLO definitions for AI journeys","Dashboard and alert configuration","Recent burn-rate or incident examples (redacted)"],"engineeringBestPractices":["Separate SLOs for AI features vs core non-AI paths","Budget tokens and wall-clock together","Correlate performance regressions with model/prompt version changes"],"automaticValidations":["Synthetic latency probes","Burn-rate alerts","CI performance budgets for critical paths where feasible"],"manualValidations":["Quarterly SLO review with product owners","Post-incident SLO recalibration"],"examples":["Chat p95 TTFT < 1.5s with error-budget policy that freezes prompt experiments when burned","Agent task-success rate tracked as a quality SLO alongside HTTP availability"],"references":[{"title":"Google SRE — Service Level Objectives","url":"https://sre.google/sre-book/service-level-objectives/"},{"title":"AWS Well-Architected — Performance Efficiency","url":"https://aws.amazon.com/architecture/well-architected/"}],"futureEvolution":"Standard GenAI SLI catalogs (TTFT, tool-loop duration, citation accuracy) across ecosystems."},{"id":"APRF-14","slug":"reliability-continuity","name":"Reliability & Continuity","summary":"Survive provider outages and partial failures, and preserve critical AI capabilities under sustained disruption.","domain":"reliability","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Design graceful degradation for short failures and documented continuity options for sustained disruption—so AI-dependent journeys fail safely and recover within defined RTO/RPO.","engineeringPhilosophy":"Assume providers fail. Timeouts, circuit breakers, and degraded modes handle minutes; backups, failover, and human procedures handle hours and days. Safety properties must hold in every degraded mode.","whyItMatters":"Naive retries and single-provider dependency turn rate limits into outages; automation without continuity plans leaves the business stranded when AI is unavailable.","commonFailures":["No timeout on model calls","Infinite retries that worsen rate limits","No fallback when the primary model is unavailable","Agent continues after a failed critical tool as if it succeeded","No manual fallback for AI-automated workflows","Single provider with no contractual or technical alternative","Backups that exclude vector indexes and prompt registries","Untested continuity plans"],"mandatoryChecks":[{"id":"REL-M1","requirement":"All model and tool calls shall have timeouts and bounded retries","artifact":"Client config / static analysis report for timeouts and max retries","passCondition":"100% of model/tool client call sites have finite timeout and max-retry ≤ bound; verified by static check or integration test","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"REL-M2","requirement":"Critical user journeys shall define a degraded mode when AI is unavailable","artifact":"Degraded-mode spec per critical journey + feature-flag or fallback test","passCondition":"100% of critical journeys document degraded behavior; failover test shows non-AI or safe fallback activates when AI dependency fails","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"REL-M3","requirement":"Partial tool failures shall be detected and shall prevent false-success continuation","artifact":"Agent/tool error-handling tests for partial failure","passCondition":"Tests inject tool failure after partial success; agent does not report overall success in 100% of cases without explicit compensation/approval","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"REL-M4","requirement":"Critical AI-dependent processes shall have documented continuity options","artifact":"Continuity options doc per critical process (failover, manual, alternate provider)","passCondition":"100% of processes marked critical-AI-dependent have ≥1 documented continuity option with owner","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"REL-M5","requirement":"Backups shall include AI control-plane artifacts needed to restore service","artifact":"Backup job config covering prompt registry, policies, indexes as required + restore test","passCondition":"Backup inventory includes required AI control-plane artifacts; restore test in last 90 days succeeds for a sample artifact set","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"REL-M6","requirement":"RTO/RPO objectives shall exist for critical AI features","artifact":"RTO/RPO catalog for critical AI features","passCondition":"100% of critical AI features have numeric RTO and RPO documented and linked to a tested restore/failover procedure","method":"manual","requiredFromLevel":4,"minCriticality":3},{"id":"REL-M7","requirement":"Chaos experiments shall cover AI dependency failure modes","artifact":"Chaos experiment plan + dated after-action report including AI provider/tool failures","passCondition":"PASS if at least one AI-dependency chaos exercise completed in the last 180 days with retained actions","method":"hybrid","requiredFromLevel":5,"minCriticality":3},{"id":"REL-M8","requirement":"Contractual and technical multi-provider options shall exist for Level 5 continuity","artifact":"Provider contract summary + technical failover design and test evidence","passCondition":"PASS if Level 5 workloads have a documented alternate provider/path and a successful failover test ≤180 days","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"REL-R1","requirement":"Circuit breakers and bulkheads around provider clients","artifact":"Client library/config showing circuit breakers and bulkheads around model/provider calls + last trip log","passCondition":"Breaker opens under induced failure in test or observed prod trip within 90 days; bulkhead limits concurrent calls per provider dependency","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"REL-R2","requirement":"Multi-provider or multi-region fallback with eval coverage","artifact":"Multi-provider or multi-region fallback config + eval report comparing primary vs fallback quality/safety","passCondition":"Fallback path is configured and was exercised ≤90 days; fallback eval meets minimum quality/safety bars (not only connectivity)","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"REL-R4","requirement":"Periodic continuity drills including provider loss","artifact":"Continuity drill calendar + last provider-loss drill report with RTO/RPO results","passCondition":"Provider-loss (or equivalent) continuity drill completed ≤90 days; RTO/RPO met or exceptions have named owners and expiry","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"REL-R6","requirement":"Warm standby for self-hosted inference where required","artifact":"Warm-standby inference architecture + last failover test report with RTO result","passCondition":"Failover to warm standby completes within documented RTO in the last test (≤90 days); standby capacity covers declared peak for critical workloads","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"evidenceRequired":["Timeout/retry configuration","Degraded-mode UX and behavior documentation","Postmortems referencing resilience controls where applicable","Continuity plans for critical journeys","Backup/restore test records","RTO/RPO definitions"],"engineeringBestPractices":["Idempotent tools where retries are possible","Queue and backpressure for bursty AI workloads","Fail safe: prefer refusal over incorrect action under uncertainty","Separate availability SLOs for AI features vs core non-AI paths","Identify which AI features are critical vs optional","Design human-operated fallback procedures before full automation","Replicate prompt registries and policies, not only databases","Align continuity with resilience and rollback pillars"],"automaticValidations":["Synthetic canaries against providers","Alerts on elevated timeout and 429 rates","Integration tests for fallback paths","Scheduled backup success checks","Restore tests in isolated environments","Health checks on standby providers"],"manualValidations":["Game days simulating provider outage","Review of agent compensation logic after tool failure","Business impact analysis for AI features","Executive review of continuity acceptance risk"],"examples":["If the LLM provider is down, search falls back to keyword results with a clear banner","An agent aborts a payment flow when the ledger tool times out rather than guessing success","If the LLM provider fails, claims adjusters switch to a documented manual checklist","Prompt registry backups restore a prior known-good configuration in a DR region"],"references":[{"title":"Google SRE Book — Addressing Cascading Failures","url":"https://sre.google/sre-book/addressing-cascading-failures/"},{"title":"AWS Well-Architected — Reliability","url":"https://aws.amazon.com/architecture/well-architected/"},{"title":"Google SRE — Disaster Recovery","url":"https://sre.google/sre-book/disaster-recovery/"}],"futureEvolution":"Portable health/degradation protocols and multi-provider continuity playbooks for AI workloads."},{"id":"APRF-15","slug":"change-management","name":"Change Management & Release","summary":"Repeatable promotion of models, prompts, tools, and agents—with tested, fast rollback.","domain":"reliability","crossCutting":false,"severity":"critical","riskLevel":"high","purpose":"Promote AI artifacts (models, prompts, tools, agents, indexes) through controlled environments, and enable immediate rollback when production behavior regresses.","engineeringPhilosophy":"If you cannot roll back, you cannot safely roll forward. AI change management treats prompts and model pins as release units with the same discipline as application code.","whyItMatters":"Uncontrolled hot-edits and untested rollbacks create drift, unreproducible incidents, and long mean-time-to-recover when quality or safety collapses after a change.","commonFailures":["Hot-editing production prompts","Indexes rebuilt manually with no version tag","Tools registered in prod without staging soak","Environment config drift for model pins","No previous prompt version retained","Model pin changed with no path back","Schema migrations that break old tool clients irreversibly","Rollback untested until an incident"],"mandatoryChecks":[{"id":"DEP-M1","requirement":"A promotion path shall exist from non-prod to prod for prompts, models, and tools","artifact":"Pipeline/runbook documenting non-prod→prod promotion for prompts, models, and tools","passCondition":"100% of production prompt/model/tool releases in the last 30 days flowed through the documented promotion path; 0 production hot-edits without a linked change record","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"DEP-M2","requirement":"Production changes shall be recorded (who/what/when) and linked to reviews","artifact":"Change log or ticket export for AI artifact releases","passCondition":"100% of production AI artifact changes in the last 30 days have who/what/when fields and a review link (PR, ticket, or approval ID)","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"DEP-M3","requirement":"Infrastructure and AI config shall be defined as code or equivalent declarative config","artifact":"IaC/config repo paths for AI gateway, model pins, tool catalogs, and prompts","passCondition":"Drift check shows 0 unmanaged production AI config resources outside declarative sources; sample of live pins matches declared versions at 100%","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"CHG-M1","requirement":"Prior production versions of prompts and model pins shall be retained and restorable","artifact":"Version retention policy + registry listing of prior production versions","passCondition":"≥N prior production versions retained per policy (minimum N=2); restore dry-run successfully loads the immediate prior version in staging or prod-adjacent env","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"CHG-M2","requirement":"Rollback procedure shall be documented and operable by on-call","artifact":"Rollback runbook + on-call acknowledgment or drill checklist","passCondition":"Runbook lists exact commands/UI steps and owners; ≥1 on-call engineer completed a walkthrough or drill in the last 90 days with recorded time-to-execute","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"CHG-M3","requirement":"Rollback shall be tested in a drill or real event on a defined cadence","artifact":"Drill or incident record with timestamps and outcome","passCondition":"≥1 successful rollback (drill or real) in the last 90 days with measured time-to-restore ≤ documented RTO","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"CHG-M4","requirement":"Automated rollback shall trigger on quality SLO burn for AI releases","artifact":"Rollback automation config + sample trigger/test evidence tied to quality SLOs","passCondition":"PASS if quality SLO burn is wired to automated rollback (or page+runbook with measured MTTA) and a test/drill occurred ≤90 days","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"DEP-R1","requirement":"Canary or progressive delivery for high-traffic AI changes","artifact":"Progressive delivery/canary config for high-traffic AI changes + last canary promotion log","passCondition":"High-traffic AI changes use canary or progressive rollout with automated rollback criteria; last such release ≤90 days includes canary metrics link","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"DEP-R2","requirement":"Environment parity checks for model pins and tool catalogs","artifact":"Environment parity checklist for model pins and tool catalogs + last parity scan across envs","passCondition":"Prod vs staging model pins and tool catalogs match except documented deltas; last parity scan ≤30 days with 0 unexplained drifts","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"DEP-R3","requirement":"Automated migration for embedding/index version upgrades","artifact":"Embedding/index migration runbook + last automated migration job log and validation","passCondition":"Index/embedding version upgrades run via automated migration with validation gates; last upgrade ≤12 months (or N/A attestation if no upgrade) succeeded without dual-write gaps","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"CHG-R1","requirement":"One-click or single-command rollback for AI release units","artifact":"Rollback runbook or one-command script for AI release units + last rollback exercise log","passCondition":"AI release unit can be rolled back with a single documented command/action; exercise or real rollback ≤90 days completed within documented RTO","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"CHG-R2","requirement":"Feature flags for new agent behaviors","artifact":"Feature-flag config covering new agent behaviors + sample flag change audit trail","passCondition":"New agent behaviors ship behind flags; flag state changes are audited; kill/disable path tested ≤90 days ago","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"evidenceRequired":["Deployment pipeline documentation","Sample change records","Environment inventory","Rollback runbook","Evidence of tested rollback","Version retention policy"],"engineeringBestPractices":["Bundle related prompt/model/tool versions as a release unit when they interact","Keep staging data representative without copying unnecessary production secrets","Automate smoke tests post-deploy including one eval canary","Document freeze windows for peak business periods","Keep rollback independent of full application redeploy when possible","Version tool schemas with compatibility windows","Avoid one-way data migrations tied to prompt experiments","Measure time-to-rollback as an operational KPI"],"automaticValidations":["Pipeline gates requiring eval and security checks","Drift detection between declared and live config","Post-deploy smoke test automation","CI verifying previous artifacts remain fetchable","Synthetic rollback dry-runs in staging","Alerts recommending rollback on quality burn"],"manualValidations":["Release readiness review for major agent launches","Spot audits of console-only changes","On-call drills executing rollback","Post-incident verification that rollback path was used or improved"],"examples":["A release train ships prompt v12 + model pin + tool schema together after staging soak","Index rebuilds produce a new version ID switched via config flip","Quality drop after a prompt release is reversed in under five minutes via registry rollback","A new tool schema is dual-published so clients can fall back"],"references":[{"title":"Google SRE Workbook — Configuration Management","url":"https://sre.google/workbook/"},{"title":"AWS Well-Architected — Operational Excellence","url":"https://aws.amazon.com/architecture/well-architected/"},{"title":"Google SRE — Emergency Response","url":"https://sre.google/sre-book/emergency-response/"},{"title":"AWS Well-Architected — Reliability","url":"https://aws.amazon.com/architecture/well-architected/"}],"futureEvolution":"Unified release manifests covering prompts, models, tools, eval attestations, and automatic quality-triggered rollbacks."},{"id":"APRF-21","slug":"incident-readiness","name":"Incident Readiness","summary":"Detect, contain, and learn from AI-specific production incidents.","domain":"reliability","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Prepare people, playbooks, and tooling to detect, contain, communicate, and learn from AI-specific incidents—including abuse, quality collapse, and unsafe tool actions.","engineeringPhilosophy":"Incidents will happen. Readiness means AI-aware detection, clear ownership, containment that includes pausing agents, and blameless learning that improves pillars.","whyItMatters":"Traditional SEV playbooks miss prompt injection campaigns, model outages, and agentic damage. Without AI-specific readiness, mean time to contain balloons.","commonFailures":["No playbook for model abuse or data leakage via chat","On-call cannot pause agents or roll back prompts","Customer communication templates ignore AI failure modes","No post-incident action tracking into evals and controls"],"mandatoryChecks":[{"id":"INC-M1","requirement":"AI-specific incident playbooks shall exist for abuse, leakage, bad actions, and provider outage","artifact":"Playbook set covering the four scenarios with owners","passCondition":"Four playbooks present (abuse, leakage, bad actions, provider outage), each with owner and review date ≤ 12 months","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"INC-M2","requirement":"On-call shall be able to execute containment: pause agents, disable tools, roll back prompts/models","artifact":"Containment runbook + drill record exercising pause, disable, and rollback","passCondition":"Drill in last 90 days successfully demonstrated pause agents, disable tools, and roll back prompt/model within documented time budgets","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"INC-M3","requirement":"Post-incident reviews shall produce tracked actions against APRF pillars","artifact":"Post-incident review template + sample reviews with linked actions","passCondition":"100% of SEV-eligible AI incidents in last 90 days have a review with ≥1 tracked action mapped to an APRF pillar or explicit “no action” rationale","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"INC-M4","requirement":"Regular tabletop exercises shall cover AI-specific incidents","artifact":"Tabletop plan + dated after-action report for an AI incident scenario","passCondition":"PASS if an AI-focused tabletop completed ≤180 days with retained actions and owners","method":"hybrid","requiredFromLevel":4,"minCriticality":3}],"recommendedChecks":[{"id":"INC-R1","requirement":"Page-worthy alerts for safety and quality signals, not only infra","artifact":"On-call alert policy export listing safety/quality pages + last 90 days of triggered incidents (or drill tickets)","passCondition":"At least two non-infra signals (e.g. refusal-rate spike, eval-score drop, toxicity/jailbreak hit rate) page an on-call; each has a documented threshold and owner; policy reviewed ≤90 days ago","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"INC-R3","requirement":"Customer notification criteria for AI-related events","artifact":"Customer notification criteria for AI-related events + last drill or real notification sample","passCondition":"Criteria map event types (safety incident, widespread quality fail, data exposure) to notify / no-notify; last drill or incident ≤12 months followed the criteria with timestamps","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Playbooks and ownership roster","Alert configuration samples","Post-incident review examples (redacted)"],"engineeringBestPractices":["Define SEV classifications for AI harm (financial, privacy, safety, reputation)","Preserve traces for forensics under legal hold when needed","Train support and on-call on AI failure literacy","Link incidents to eval fixtures so regressions cannot silently return"],"automaticValidations":["Alert routing tests","Automated creation of incident tickets from critical signals","Verification that kill switches remain reachable"],"manualValidations":["Tabletop exercises","After-action review quality checks"],"examples":["A spike in tool denials pages on-call; the agent is paused while injection is investigated","A postmortem adds a new adversarial eval case that fails CI until fixed"],"references":[{"title":"Google SRE — Managing Incidents","url":"https://sre.google/sre-book/managing-incidents/"},{"title":"NIST AI RMF — Manage","url":"https://www.nist.gov/itl/ai-risk-management-framework"}],"futureEvolution":"Shared AI incident taxonomies and anonymized industry learning feeds."},{"id":"APRF-13","slug":"cost-optimization","name":"Cost Optimization","summary":"Bound spend with budgets, caching, routing, and abuse controls.","domain":"cost","crossCutting":false,"severity":"high","riskLevel":"high","purpose":"Keep AI spend predictable and attributable through budgets, caching, model routing, quota enforcement, and abuse prevention.","engineeringPhilosophy":"Cost is a reliability and security property. Unbounded token spend is a denial-of-wallet attack surface and a business continuity risk.","whyItMatters":"Agent loops, recursive tool use, and prompt bloat can create sudden five-figure bills. Without controls, AI features become financially unsafe to operate.","commonFailures":["No per-tenant or per-feature budgets","Retry storms that multiply completions","Sending huge contexts when retrieval would suffice","Using frontier models for trivial classification tasks"],"mandatoryChecks":[{"id":"COST-M1","requirement":"Production AI workloads shall have hard spend or rate limits","artifact":"Gateway/provider quota config","passCondition":"Enforced limit demonstrably denies or throttles when exceeded (automated test or production event log within last 90 days)","method":"automated","requiredFromLevel":2,"minCriticality":1},{"id":"COST-M2","requirement":"Cost shall be monitored with alerts on anomaly and budget burn","artifact":"Cost dashboard + alert policies for anomaly and budget burn","passCondition":"Alerts exist for budget burn and spend anomaly; synthetic or historical burn event would page/notify per runbook (alert test or documented fire in last 90 days)","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"COST-M3","requirement":"Retry and loop policies shall prevent unbounded cost amplification","artifact":"Retry/backoff and agent-loop budget config + amplification tests","passCondition":"Max retries and max agent loops are finite for 100% of production AI clients; tests show cost cannot grow without bound under forced failure/retry (bounded token or $ ceiling hit)","method":"automated","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"COST-R1","requirement":"Caching for idempotent or repeated prompts where safe","artifact":"Prompt/response cache config with safety exclusions + hit-rate report for 30 days","passCondition":"Cache enabled for documented idempotent paths; sensitive/personalized prompts are excluded; hit-rate and savings reported for ≥30 days","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"COST-R2","requirement":"Model routing: cheaper models for low-risk tasks with eval coverage","artifact":"Model-routing policy (cheap vs premium) + eval coverage report for low-risk routed tasks","passCondition":"Low-risk task classes route to cheaper models by default; eval shows quality within tolerance vs premium baseline; misroute rate monitored ≤30 days","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"COST-R3","requirement":"FinOps reviews of AI unit economics per product","artifact":"Per-product AI unit-economics report (cost per successful task / journey) + FinOps review minutes","passCondition":"Each customer-facing AI product has unit-cost metrics for the last quarter; review occurred ≤90 days with owners for outliers above documented thresholds","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["Budget and quota configuration","Cost dashboards and alert policies","Documented retry/loop limits"],"engineeringBestPractices":["Attribute cost to tenant and feature tags","Prefer embeddings + retrieval over stuffing large corpora","Set timeouts on model and tool calls","Load-test cost under adversarial long prompts"],"automaticValidations":["Gateway enforcement of quotas","Anomaly detection on token spend","CI checks estimating token cost of prompt changes"],"manualValidations":["Quarterly cost architecture review","Abuse scenario walkthroughs"],"examples":["A free-tier chatbot hard-stops after N tokens/day per user","An agent aborts when projected step cost exceeds remaining budget"],"references":[{"title":"AWS Well-Architected — Cost Optimization","url":"https://aws.amazon.com/architecture/well-architected/"},{"title":"FinOps Foundation practices","url":"https://www.finops.org/"}],"futureEvolution":"Standard cost telemetry schemas for multi-provider AI gateways."},{"id":"APRF-29","slug":"organizational-governance","name":"Organizational Governance","summary":"AI policy, ownership, risk acceptance, and continual improvement—ISO 42001-style management system.","domain":"governance","crossCutting":false,"severity":"high","riskLevel":"medium","purpose":"Establish organizational AI policy, clear ownership/RACI, risk-acceptance authority, and continual improvement so technical pillars have accountable stewards.","engineeringPhilosophy":"Technical controls without organizational governance drift. A production-ready AI program names owners, accepts residual risk explicitly, and improves from incidents and audits.","whyItMatters":"Enterprises fail AI readiness when no one owns a pillar, exceptions are informal, and leadership cannot see residual risk.","commonFailures":["AI features ship with no named owner for safety or eval gates","Exceptions granted in chat with no expiry","No AI policy covering acceptable use and prohibited applications","Compliance theater without leadership review of risk"],"mandatoryChecks":[{"id":"ORG-M1","requirement":"Documented AI policy covering acceptable use and prohibited applications shall exist","artifact":"Approved AI acceptable-use / prohibited-applications policy","passCondition":"Policy has version, owner, and review date ≤ 12 months; includes both acceptable-use and prohibited-application sections","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"ORG-M2","requirement":"Each production AI system shall have named owners for critical APRF domains","artifact":"System inventory with owner fields for critical domains","passCondition":"0 production systems missing required domain owners in inventory query","method":"hybrid","requiredFromLevel":3,"minCriticality":2},{"id":"ORG-M3","requirement":"Risk acceptance for known control gaps shall be recorded with owner and expiry","artifact":"Risk-acceptance register samples","passCondition":"100% of open control-gap waivers have owner + expiry date; 0 expired waivers without escalation record","method":"hybrid","requiredFromLevel":4,"minCriticality":3},{"id":"ORG-M4","requirement":"Internal audit or independent assessment shall sample APRF evidence on a defined cadence","artifact":"Independent assessment or internal-audit sampling report against APRF evidence","passCondition":"PASS if the last sampling report is ≤12 months old and lists sampled check IDs and findings","method":"manual","requiredFromLevel":5,"minCriticality":3}],"recommendedChecks":[{"id":"ORG-R1","requirement":"Periodic leadership review of AI risk posture and APRF maturity","artifact":"Leadership AI risk / APRF maturity review minutes (or board pack excerpt) + action log","passCondition":"Leadership reviewed AI risk posture and APRF capability attained ≤90 days ago; open actions have owners and due dates","method":"manual","requiredFromLevel":4,"minCriticality":2},{"id":"ORG-R3","requirement":"Continual improvement backlog fed by incidents and eval failures","artifact":"Improvement backlog (tickets) linked from incidents and eval failures + sample of closed items last quarter","passCondition":"≥80% of Sev-1/2 AI incidents and critical eval fails in the last quarter produced a backlog item; ≥50% of those items closed or have a dated plan","method":"hybrid","requiredFromLevel":4,"minCriticality":2}],"evidenceRequired":["AI policy document","Ownership/RACI for production AI systems","Risk-acceptance register samples"],"engineeringBestPractices":["Map APRF domains to teams; avoid orphan pillars","Time-box exceptions; auto-escalate expired waivers","Tie promotions and funding to measurable maturity for high-criticality systems"],"automaticValidations":["Inventory systems missing owners","Alerts on expired risk acceptances"],"manualValidations":["Annual policy review","Leadership tabletop on AI risk acceptance"],"examples":["A risk register entry accepts a missing multi-provider fallback until Q3 with CTO sign-off","Each agent product lists owners for Security, Safety, Evaluation, and Reliability"],"references":[{"title":"ISO/IEC 42001 — AI management systems","url":"https://www.iso.org/standard/81230.html"},{"title":"NIST AI RMF — Govern","url":"https://www.nist.gov/itl/ai-risk-management-framework"}],"futureEvolution":"Machine-readable organizational control catalogs mapped to APRF domains and ISO 42001 clauses."},{"id":"APRF-19","slug":"compliance","name":"Compliance","summary":"Produce auditable evidence of controls without equating compliance with readiness.","domain":"governance","crossCutting":false,"severity":"high","riskLevel":"medium","purpose":"Map regulatory and contractual obligations to concrete AI controls and produce auditable evidence—while recognizing compliance alone does not prove production readiness.","engineeringPhilosophy":"Compliance is necessary evidence, not sufficient engineering. APRF treats compliance as a pillar that packages proof of other pillars for auditors and customers.","whyItMatters":"Enterprises cannot buy or deploy AI systems that cannot demonstrate control evidence. Conversely, checkbox compliance without engineering depth still fails in production.","commonFailures":["Claiming 'we are compliant' without control-to-evidence mapping","Policies that do not match runtime behavior","No audit trail for model and prompt changes","Ignoring sector-specific AI obligations until late"],"mandatoryChecks":[{"id":"CMP-M1","requirement":"Applicable obligations for the AI system shall be identified and owned","artifact":"Obligations register with owner per obligation for each production AI system","passCondition":"Every production AI system ID has ≥1 mapped obligation entry or an explicit “none in scope” attestation with owner and review date ≤ 12 months","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"CMP-M2","requirement":"Control-to-evidence mapping shall exist for in-scope requirements","artifact":"Control→evidence matrix linking obligations to APRF checks or internal control IDs","passCondition":"100% of in-scope obligations map to ≥1 evidence artifact ID; matrix review date ≤ 12 months; 0 orphan obligations without evidence pointer","method":"manual","requiredFromLevel":4,"minCriticality":3},{"id":"CMP-M3","requirement":"Audit logs for critical AI control-plane changes shall be retained per policy","artifact":"Audit log retention config + sample of control-plane change events","passCondition":"Retention configured ≥ policy minimum (e.g. ≥ 365 days); synthetic control-plane change appears in audit log within ≤ 5 minutes and remains queryable after retention smoke check","method":"hybrid","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"CMP-R1","requirement":"Regular control testing with documented exceptions","artifact":"Control-testing schedule + latest test results pack with exceptions register","passCondition":"Documented controls were tested on schedule in the last cycle (≤90 days default); open exceptions have owner, expiry, and compensating control","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"CMP-R2","requirement":"Customer-facing trust documentation aligned to APRF pillars","artifact":"Customer-facing trust/security doc mapped to APRF pillars (or Core Profile) + published URL and last-updated date","passCondition":"Public trust doc covers identity, safety/eval, data handling, and incident contact at minimum; last-updated ≤12 months; pillar/Core mapping table is explicit","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"CMP-R3","requirement":"Independent assessment for Level 5 systems","artifact":"Independent assessment or internal-audit report sampling Level-5 AI systems against APRF gates","passCondition":"Last assessment ≤12 months old lists sampled check IDs, findings, and remediation owners; covers every system rated Level 5","method":"manual","requiredFromLevel":5,"minCriticality":3}],"evidenceRequired":["Obligations register for the AI product","Evidence packs per critical control","Audit log retention configuration"],"engineeringBestPractices":["Derive technical tickets from obligations; avoid orphan policies","Reuse APRF pillar evidence for multiple frameworks where mappings exist","Keep compliance language accurate: readiness ≠ certification","Version compliance artifacts with product releases"],"automaticValidations":["Continuous control monitoring where automatable","Alerts on audit log pipeline failure","CI attaching evidence artifacts to releases"],"manualValidations":["Internal audit sampling","Legal review of external AI claims"],"examples":["A SOC 2 narrative points to APRF eval gates and access reviews as evidence","Change tickets for prompt releases satisfy audit sampling requests"],"references":[{"title":"NIST AI RMF","url":"https://www.nist.gov/itl/ai-risk-management-framework"},{"title":"ISO/IEC 42001 (AI management systems) — conceptual alignment","url":"https://www.iso.org/standard/81230.html"}],"futureEvolution":"Common control catalogs mapping APRF pillars to regional AI regulations."},{"id":"APRF-23","slug":"platform-engineering","name":"Platform Engineering","summary":"Make the safe path the easy path—paved roads that apply across every APRF domain.","domain":null,"crossCutting":true,"severity":"medium","riskLevel":"medium","purpose":"Provide builders with paved roads—SDKs, templates, local evals, guardrails, and documentation—so secure and operable defaults are the path of least resistance across all domains.","engineeringPhilosophy":"Platform engineering is a cross-cutting control, not a peer readiness domain. Friction on unsafe paths and speed on safe ones determine whether other pillars are actually followed.","whyItMatters":"If the approved gateway is hard and the raw provider key is easy, builders will bypass controls. DX debt becomes security and reliability debt.","commonFailures":["No internal AI platform or golden paths","Docs that explain policy but not how to comply","Slow security review with no self-serve checklists","Local development that cannot run evals or policy checks"],"mandatoryChecks":[{"id":"DX-M1","requirement":"Builders shall have a documented golden path for deploying AI features to production","artifact":"Golden-path documentation with steps from scaffold to production","passCondition":"Doc exists with version/owner; covers auth, secrets, evals, and promote steps; reviewed ≤ 12 months","method":"manual","requiredFromLevel":3,"minCriticality":2},{"id":"DX-M2","requirement":"Local or CI self-serve checks shall cover critical mandatory controls (auth, secrets, basic evals)","artifact":"CI/local check config covering auth, secrets, and basic evals","passCondition":"Default AI pipeline runs auth, secret-scan, and basic eval checks; failing any blocks merge/promote in the golden-path template","method":"automated","requiredFromLevel":3,"minCriticality":2},{"id":"DX-M3","requirement":"Ownership and support channel shall exist for the AI platform/paved road","artifact":"Platform ownership record + support channel (on-call, Slack, ticket queue)","passCondition":"Named owner team and support channel documented; channel responds to a test ping within published SLA or has on-call rotation listed","method":"manual","requiredFromLevel":3,"minCriticality":2}],"recommendedChecks":[{"id":"DX-R1","requirement":"Scaffolding templates for agents, RAG, and MCP with safe defaults","artifact":"Scaffolding template repo/catalog for agents, RAG, and MCP with default security settings + adoption metrics","passCondition":"Templates exist for agents, RAG, and MCP with auth, secrets, and logging defaults on; ≥1 new service used a template in the last 90 days or adoption target is documented","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"DX-R2","requirement":"Inner-loop eval runners developers can run before PR","artifact":"Local/inner-loop eval runner docs + package + sample developer run log before a recent PR","passCondition":"Developers can run the core eval subset locally or in a one-command inner loop; last sampled AI PR (≤30 days) shows pre-PR eval evidence or documented waiver","method":"hybrid","requiredFromLevel":4,"minCriticality":2},{"id":"DX-R3","requirement":"DX metrics: time-to-safe-production and bypass rate","artifact":"Platform DX dashboard (or weekly report) with time-to-safe-production and policy-bypass rate definitions + last 30 days series","passCondition":"Both metrics are defined with formulas; published for ≥30 consecutive days; bypass rate has an alert or review threshold with a named owner","method":"hybrid","requiredFromLevel":5,"minCriticality":3}],"evidenceRequired":["Golden path documentation","Template/repository inventory","Support ownership for AI platform"],"engineeringBestPractices":["Default SDKs inject tracing, auth, and budget headers","Provide copy-paste threat model and APRF checklist stubs","Measure and reduce friction for approved tools","Celebrate teams that use paved roads in readiness reviews"],"automaticValidations":["Template CI that fails if safe defaults are removed","Telemetry on gateway vs direct-provider usage","Lint rules guiding builders to approved libraries"],"manualValidations":["Builder interviews on friction points","Periodic DX reviews with security and platform teams"],"examples":["create-agent CLI scaffolds an agent with budgets, authz hooks, and eval stubs","A PR template links to the APRF pillar checklist for the feature type"],"references":[{"title":"Google SRE Workbook — Soft skills / platform thinking parallels","url":"https://sre.google/workbook/"},{"title":"CNCF platforms white papers — paved roads concept","url":"https://www.cncf.io/"}],"futureEvolution":"Shared open-source APRF linters, project templates, and inner-loop conformance checks adopted across ecosystems."}],"stats":{"domainCount":8,"pillarCount":27,"mandatoryCheckCount":108,"recommendedCheckCount":69,"profileCount":2,"coreProfileCheckCount":40,"crosswalkCount":6,"crosswalkMappingCount":51,"lensCount":4,"lensCheckCount":63}}