AI Assurance Academy · Part 5

Change Control and Revalidation Triggers

Chapter 17 of 20 · AI behavior can change when the model, data, prompt, retrieval corpus, software, supplier routing, or operating population changes. A mature change-control system identifies these dependencies and scales revalidation to the affected risk.

Author
Sandip Thorat
Published
September 4, 2026
Last reviewed
September 4, 2026
Category
AI Assurance
Reading time
6 min
Version
1.0
01Context02AI failure03Control envelope04Lifecycle evidence
A decision-focused assurance chain: every transition requires proportionate evidence.

AI behavior can change when the model, data, prompt, retrieval corpus, software, supplier routing, or operating population changes. A mature change-control system identifies these dependencies and scales revalidation to the affected risk.

Published: September 4, 2026 | Version 1.0

Editorial owner: CSV to CSA Knowledge Hub | Review status: Open for practitioner peer review

THE AI CHANGE INVENTORY

Model changes

  • New weights or model family
  • Retraining or fine-tuning
  • Hyperparameter or threshold change
  • Quantization, distillation, or compression
  • Safety layer or moderation change
  • Provider routing or fallback model

Data changes

  • New source, site, product, instrument, language, or population
  • Label definition or adjudication change
  • Extraction, transformation, imputation, or feature change
  • Training-window or prevalence shift

RAG changes

  • Source repository or approval metadata
  • Parser, OCR, chunking, embedding, index, ranking, or top-k
  • Access filter or re-index schedule

Prompt and application changes

  • System prompt, template, examples, tool definition, or output schema
  • Business rule, interface, role, integration, or user experience
  • Confidence display, warning, fallback, or override

Infrastructure and supplier changes

  • Cloud region, hardware, library, API, latency, availability, or subprocessor
  • Retention, logging, privacy, or security control

Use changes

  • New user, decision, authority, workflow, product, language, or record
  • Advisory output becomes an automated action

CHANGE CLASSIFICATION

Class 0 — No relevant effect

No component inside the approved boundary changes. Record the screening conclusion.

Class 1 — Low-impact controlled change

Behavior is not expected to affect critical performance; evidence may include supplier review, configuration comparison, and targeted smoke or regression testing.

Class 2 — Material change within approved use

Behavior, data, prompt, retrieval, integration, or human workflow changes. Perform targeted risk reassessment and revalidation of affected requirements and failure modes.

Class 3 — New or expanded intended use

The operating envelope, decision authority, user, population, or regulated purpose changes. Revisit scope, risk, data, evaluation, oversight, records, monitoring, and approval as a new use case.

Class 4 — Emergency or uncontrolled change

Silent supplier change, critical defect, security event, unexpected drift, or forced update. Contain, identify affected use, apply fallback, assess impact, and perform retrospective evidence under emergency procedures.

THE IMPACT ASSESSMENT

Ask:

  • What component changed and can it be uniquely identified?
  • Which intended uses depend on it?
  • Which failure modes or controls can change?
  • Does the evaluation data remain representative?
  • Are acceptance thresholds still appropriate?
  • Does human review remain effective?
  • Are records, signatures, or audit trails affected?
  • Can change be detected in production?
  • Is rollback possible and tested?

REVALIDATION SELECTION

Supplier-evidence review

Appropriate when credible vendor evidence covers a standard-product change and customer-specific behavior is not affected.

Configuration confirmation

Compare controlled configuration and enabled features.

Targeted regression

Exercise affected requirements and linked critical paths.

Model reevaluation

Run representative and challenge sets when behavior may change.

Human-factors reevaluation

Required when confidence display, workflow, alert volume, explanation, or authority changes.

Prospective silent mode

Useful for material changes when outputs can be compared before influencing decisions.

Full reassessment

Required for new intended use, major architecture change, or invalidated prior evidence.

WORKED EXAMPLE: HOSTED LLM UPDATE

The vendor announces a new default model with improved reasoning and a larger context window. The application uses it to draft laboratory-investigation chronologies.

Potential impact:

  • Different summarization and refusal behavior
  • New tokenization affects units and symbols
  • Longer context includes attachments previously excluded
  • Different safety layer changes allowed questions
  • Latency and cost change user behavior
  • Prior prompt tuning may no longer perform as expected

Assurance response:

  • Pin old version if contract allows while assessing
  • Compare outputs on locked representative and critical challenge cases
  • Reevaluate harmful omission, fabricated events, numeric accuracy, citation, and repeated variability
  • Test long-context and attachment handling
  • Confirm prompt and output schema
  • Run a user review study on intentionally wrong outputs
  • Approve, restrict, or reject with rollback plan

RETRAINING CONTROL

Predefine:

  • Trigger and authorized initiator
  • Eligible data and cutoff
  • Label and quality review
  • Training code and environment
  • Independent test set protection
  • Acceptance criteria and statistical comparison
  • Subgroup and regression limits
  • Model registry, approval, deployment, and rollback
  • Monitoring baseline reset

Do not allow automatic production retraining to bypass release evidence for critical use.

CONFIGURATION DRIFT

Automate comparison of model endpoint, prompt hash, enabled features, retrieval sources, permissions, thresholds, and tool definitions. Alert on unapproved difference. Periodically test that the deployed artifact matches the approved registry.

CHANGE RECORD

Change ID and trigger:

Component/version before and after:

Affected uses and records:

Risk impact:

Supplier evidence:

Assurance selected:

Results and issues:

Monitoring change:

Rollback or fallback:

Approval and deployment:

Post-implementation review:

PROFESSIONAL INTERPRETATION

AI change control should be dependency-aware, not document-heavy. The organization must know which uses depend on a changed model, prompt, source, or tool and generate evidence proportional to the altered failure risk.

PRIMARY SOURCES

FDA CSA final guidance, lifecycle and SaaS automatic-update example:

www.fda.gov/media/188844/download

FDA and EMA, Guiding Principles of Good AI Practice in Drug Development:

www.fda.gov/about-fda/artificial-intelligence-drug-development/guiding-principles-good-ai-practice-drug-development

NIST AI RMF:

www.nist.gov/itl/ai-risk-management-framework