PRACTITIONER METHOD

Practical software assurance playbook

AI Assurance Playbook

A controlled path from context of use and human authority to evaluation evidence, release boundaries, monitoring, and stop conditions.

Professional interpretationEducational practitioner method · Quality review required for case-specific use

Decision and evidence workflow

Work through the assurance boundary

  1. 01

    Intended Use

    Decision
    Specify the users, population, data, AI task, output, decision influenced, environment, exclusions, and prohibited actions.
    Evidence
    Feature-level intended use with permitted and prohibited reliance.
  2. 02

    AI Role

    Decision
    Determine whether the model retrieves, extracts, predicts, drafts, recommends, decides, or executes.
    Evidence
    Authority map showing deterministic, AI, and human responsibilities.
  3. 03

    Human Oversight

    Decision
    Define reviewer competence, time, source visibility, independence, escalation, and the consequences of automation bias.
    Evidence
    Oversight procedure and representative-user challenge results.
  4. 04

    Failure Modes

    Decision
    Trace omission, unsupported content, wrong source/version, subgroup failure, instability, prompt injection, overreliance, and action errors.
    Evidence
    Failure chains linked to controls and critical acceptance criteria.
  5. 05

    Evaluation Dataset

    Decision
    Represent the approved population, critical edge cases, realistic prevalence, source quality, subgroups, and prohibited contexts.
    Evidence
    Versioned dataset specification, provenance, exclusions, and leakage review.
  6. 06

    Performance Metrics

    Decision
    Measure the error that matters: critical misses, unsupported claims, subgroup performance, calibration, consistency, and reviewer recovery.
    Evidence
    Predefined metrics with denominators, confidence, and acceptance rules.
  7. 07

    Critical Misses

    Decision
    Examine every unacceptable event, not only the average score, and decide whether scope, control, or model changes are required.
    Evidence
    Critical-event review with disposition and retest traceability.
  8. 08

    Controls

    Decision
    Constrain authority through source filters, citations, deterministic gates, permissions, confirmation, logging, and kill switches.
    Evidence
    AI Control Envelope with tested preventive and detective controls.
  9. 09

    Release Decision

    Decision
    State the supported use, open uncertainty, conditions, fallback, authorized approver, and restart criteria.
    Evidence
    Signed release narrative and reconstructable evidence package.
  10. 10

    Monitoring

    Decision
    Detect changes in data, use, supplier model, prompts, retrieval, performance, user response, and critical-event rate.
    Evidence
    Monitoring plan with thresholds, owners, stop conditions, and reassessment triggers.

Release gate

Before the decision is defended

  • Critical errors are defined in operational terms and independently challenged.
  • Human review is designed and tested as a control rather than assumed.
  • The approved operating envelope and prohibited actions are explicit.
  • Monitoring can locate affected records and trigger suspension or reassessment.

Scope and limitations

What this playbook does not establish

  • A high aggregate accuracy score is not a release conclusion.
  • This playbook does not replace clinical, safety, privacy, security, or legal evaluation.
  • Model and supplier opacity may require narrower use, stronger controls, or a no-release decision.

Use applicable regulations, final guidance, organizational procedures, and subject-matter review for the actual system and jurisdiction.