Practical software assurance playbook
AI Assurance Playbook
A controlled path from context of use and human authority to evaluation evidence, release boundaries, monitoring, and stop conditions.
Professional interpretationEducational practitioner method · Quality review required for case-specific use
Decision and evidence workflow
Work through the assurance boundary
- 01
Intended Use
- Decision
- Specify the users, population, data, AI task, output, decision influenced, environment, exclusions, and prohibited actions.
- Evidence
- Feature-level intended use with permitted and prohibited reliance.
- 02
AI Role
- Decision
- Determine whether the model retrieves, extracts, predicts, drafts, recommends, decides, or executes.
- Evidence
- Authority map showing deterministic, AI, and human responsibilities.
- 03
Human Oversight
- Decision
- Define reviewer competence, time, source visibility, independence, escalation, and the consequences of automation bias.
- Evidence
- Oversight procedure and representative-user challenge results.
- 04
Failure Modes
- Decision
- Trace omission, unsupported content, wrong source/version, subgroup failure, instability, prompt injection, overreliance, and action errors.
- Evidence
- Failure chains linked to controls and critical acceptance criteria.
- 05
Evaluation Dataset
- Decision
- Represent the approved population, critical edge cases, realistic prevalence, source quality, subgroups, and prohibited contexts.
- Evidence
- Versioned dataset specification, provenance, exclusions, and leakage review.
- 06
Performance Metrics
- Decision
- Measure the error that matters: critical misses, unsupported claims, subgroup performance, calibration, consistency, and reviewer recovery.
- Evidence
- Predefined metrics with denominators, confidence, and acceptance rules.
- 07
Critical Misses
- Decision
- Examine every unacceptable event, not only the average score, and decide whether scope, control, or model changes are required.
- Evidence
- Critical-event review with disposition and retest traceability.
- 08
Controls
- Decision
- Constrain authority through source filters, citations, deterministic gates, permissions, confirmation, logging, and kill switches.
- Evidence
- AI Control Envelope with tested preventive and detective controls.
- 09
Release Decision
- Decision
- State the supported use, open uncertainty, conditions, fallback, authorized approver, and restart criteria.
- Evidence
- Signed release narrative and reconstructable evidence package.
- 10
Monitoring
- Decision
- Detect changes in data, use, supplier model, prompts, retrieval, performance, user response, and critical-event rate.
- Evidence
- Monitoring plan with thresholds, owners, stop conditions, and reassessment triggers.
Release gate
Before the decision is defended
- Critical errors are defined in operational terms and independently challenged.
- Human review is designed and tested as a control rather than assumed.
- The approved operating envelope and prohibited actions are explicit.
- Monitoring can locate affected records and trigger suspension or reassessment.
Scope and limitations
What this playbook does not establish
- A high aggregate accuracy score is not a release conclusion.
- This playbook does not replace clinical, safety, privacy, security, or legal evaluation.
- Model and supplier opacity may require narrower use, stronger controls, or a no-release decision.
Use applicable regulations, final guidance, organizational procedures, and subject-matter review for the actual system and jurisdiction.