AI Assurance Academy · Part 5

Monitoring, Drift, Incidents, and CAPA

Chapter 18 of 20 · Release testing establishes confidence at one point in time. Operational monitoring determines whether the AI-enabled system remains inside the approved envelope and whether users, data, suppliers, or models have changed the risk.

Author
Sandip Thorat
Published
September 4, 2026
Last reviewed
September 4, 2026
Category
AI Assurance
Reading time
6 min
Version
1.0
01Context02AI failure03Control envelope04Lifecycle evidence
A decision-focused assurance chain: every transition requires proportionate evidence.

Release testing establishes confidence at one point in time. Operational monitoring determines whether the AI-enabled system remains inside the approved envelope and whether users, data, suppliers, or models have changed the risk.

Published: September 4, 2026 | Version 1.0

Editorial owner: CSV to CSA Knowledge Hub | Review status: Open for practitioner peer review

MONITOR FOUR LEVELS

Service health

  • Availability, latency, errors, timeouts, queue depth, capacity
  • Provider incident and failover
  • Interface completeness and reconciliation

Input and data health

  • Missingness, ranges, categories, units, image quality, language
  • Population, product, site, instrument, and time distribution
  • Out-of-distribution rate
  • Retrieval corpus completeness and access state

Output and model health

  • Confidence distribution
  • Abstention and refusal
  • Class distribution
  • Performance on delayed ground truth
  • Subgroup performance
  • Hallucination, citation, and critical-error samples
  • Agent tool failures or unauthorized attempts

Process and human health

  • Overrides, disagreement, acceptance, escalation
  • Time-to-review and workload
  • Downstream deviations, complaints, rejects, or misses
  • Use outside approved scope
  • Benefit and unintended workflow adaptation

DRIFT TYPES

Data drift

Input distribution changes.

Concept drift

The relationship between input and correct output changes.

Label drift

Outcome definition or reviewer practice changes.

Population drift

Sites, products, users, languages, or instruments change.

Operational drift

Workflow, review behavior, alert volume, or fallback changes.

Supplier drift

Model routing, safety, training, or infrastructure changes without a clear version.

Prompt or knowledge drift

System prompts, templates, approved sources, or index contents change.

Drift is a signal for investigation, not automatic proof of unacceptable performance. Performance can degrade without an obvious distribution metric, so outcome monitoring remains important.

MONITORING SPECIFICATION

For every metric, define:

  • Purpose and linked failure mode
  • Data source and calculation
  • Population and stratification
  • Baseline and expected variation
  • Warning and action limits
  • Review frequency
  • Owner and backup
  • Investigation workflow
  • Required record
  • Escalation, restriction, rollback, or retraining action

Avoid dashboards without decision rules.

GROUND TRUTH DELAY

Some outcomes become known days or months later. Use layered monitoring:

Immediate proxies

Schema failure, confidence, missingness, source retrieval, deterministic rule conflict.

Short-term review

Human overrides, adjudicated samples, rejected outputs, operational errors.

Long-term outcomes

Confirmed defects, complaints, investigation conclusions, process performance, or clinical outcomes where applicable.

Link later outcomes back to the exact model and system version.

INCIDENT CLASSIFICATION

Examples:

  • Critical wrong output affects or could affect a regulated decision
  • Repeated subgroup underperformance
  • Unauthorized data exposure
  • Prompt injection or tool misuse
  • Silent model or configuration change
  • Inability to reconstruct a decision
  • Monitoring failure
  • Widespread service outage or corrupted queue

Containment may include suspending AI, reverting to manual process, blocking a model, disabling tools, preserving evidence, notifying users, and reviewing affected records.

AI INCIDENT INVESTIGATION

Preserve:

  • Inputs and authoritative sources
  • Model, prompt, index, application, and configuration version
  • Tool calls and permissions
  • Output and confidence
  • User actions and workflow state
  • Monitoring signals
  • Supplier status and recent changes
  • Similar prior and subsequent events

Investigate systemically. “The model hallucinated” is a symptom, not a root cause. Causes may include unsupported use, weak retrieval, misleading interface, insufficient training, silent change, missing guardrail, poor data, or inadequate review.

WORKED EXAMPLE: COMPLAINT TRIAGE DRIFT

Urgent-class recall was acceptable at launch. Six months later, override rate rises at one site.

Investigation:

  • Site introduced a new abbreviation set
  • Recent serious cases use free text absent from development data
  • The model confidence remains high, so confidence monitoring did not alert
  • Reviewers learned to recognize the pattern and override

Actions:

  • Add deterministic rules for critical new abbreviations
  • Reassess affected historical complaints
  • Create an adjudicated challenge set
  • Retrain or update only through controlled change
  • Add site-language distribution and override reason monitoring
  • Evaluate whether users at other sites detect the same failure

CAPA FOR AI

Corrective actions can address model, data, prompt, interface, supplier, process, training, monitoring, or governance. Preventive action may include dependency mapping, automated version checks, stronger challenge sets, access restrictions, or revised intended use.

Effectiveness checks should verify the original failure chain. A generic “no recurrence” check may be weak if the event is rare.

PERIODIC REVIEW

Assess:

  • Intended use and actual use
  • Inventory and versions
  • Performance and subgroup evidence
  • Drift and monitoring effectiveness
  • Overrides and human review
  • Incidents, deviations, complaints, and CAPA
  • Supplier and service changes
  • Security, privacy, access, and data retention
  • Open limitations and residual risk
  • Business benefit and non-AI alternative
  • Retirement readiness

PROFESSIONAL INTERPRETATION

Monitoring is part of validation for a changing, data-dependent system. The strongest program defines what signal leads to what decision and can rapidly reconstruct which records and uses were exposed.

PRIMARY SOURCES

FDA and EMA, Guiding Principles of Good AI Practice in Drug Development:

www.fda.gov/about-fda/artificial-intelligence-drug-development/guiding-principles-good-ai-practice-drug-development

NIST AI RMF:

www.nist.gov/itl/ai-risk-management-framework

EMA AI reflection paper:

www.ema.europa.eu/en/use-artificial-intelligence-ai-medicinal-product-lifecycle-scientific-guideline

ISPE, GAMP Guide: Artificial Intelligence. Industry guidance; not a regulation:

ispe.org/publications/guidance-documents/gamp-guide-artificial-intelligence