Guides · APRF practice notes
Safety Guides
Refuse harmful content, disclose AI use, measure fairness, mediate tool side effects, and keep audit-ready traces. These guides support APRF Safety & Responsible AI and Explainability.
Harm prevention and transparency — APRF Safety & Responsible AI.
For injection and adversarial controls, see Security guides.
- AI Content Safety and Refusal Policies in Production
Alignment demos are not a safety program. Production needs written refusal policies, enforcement beyond the base model, and evals that prove them.
- AI Decision Traces and Citations for Audit
"The model said so" is not an audit trail. APRF Explainability expects reconstructable traces—citations, tool calls, and policy decisions—tied to versions you can roll back.
- AI Safety Eval Gates and Refusal Monitoring
A written refusal policy without release gates and live signals will rot. Gate on safety evals, separate them from security jailbreak corpora, and monitor refusals online.
- Disclose When Users Are Talking to AI
Hiding that a reply is AI-generated is a safety and trust failure—not a UX flourish. Production products need clear disclosure appropriate to channel and risk.
- Fairness and Disparity Testing for High-Stakes AI
A secure model can still be unfair. High-stakes AI needs defined protected attributes (where lawful), disparity metrics, and a gate that blocks promotion when thresholds break.
- Safety Filters Must Cover Tool Side Effects
Filtering the final chat reply leaves agents free to email, refund, post, or fetch harmful content via tools. Safety policy must mediate actions, not only text.
Next: Safety & Responsible AI
Open the related pillar specification for mandatory checks, artifacts, and pass conditions. Self-attest is optional.