AI Assurance Academy · Part 4

Generative AI, Hallucination, and Content Risk

Chapter 14 of 20 · Generative AI can draft, summarize, translate, compare, and search at scale, but fluent language can hide unsupported claims, harmful omission, and invented evidence. Assurance must evaluate claims, sources, workflow reliance, and user behavior—not simply whether an answer sounds good.

Author
Sandip Thorat
Published
September 4, 2026
Last reviewed
September 4, 2026
Category
AI Assurance
Reading time
6 min
Version
1.0
01Context02AI failure03Control envelope04Lifecycle evidence
A decision-focused assurance chain: every transition requires proportionate evidence.

Generative AI can draft, summarize, translate, compare, and search at scale, but fluent language can hide unsupported claims, harmful omission, and invented evidence. Assurance must evaluate claims, sources, workflow reliance, and user behavior—not simply whether an answer sounds good.

Published: September 4, 2026 | Version 1.0

Editorial owner: CSV to CSA Knowledge Hub | Review status: Open for practitioner peer review

GENAI FAILURE TAXONOMY

Fabrication

Creates a fact, reference, event, or source that does not exist.

Unsupported inference

Moves beyond available evidence without making the uncertainty clear.

Harmful omission

Leaves out a condition, exception, unit, negation, or required step.

Source conflict

Combines incompatible instructions from different products, sites, versions, or jurisdictions.

Instruction failure

Does not follow required format, scope, or prohibition.

Over-refusal

Declines a safe and supported task, reducing process utility.

Under-refusal

Answers when evidence is missing or the request is outside approved use.

Privacy or confidentiality leakage

Exposes sensitive data through output, logs, training, retrieval, or memorization.

Prompt injection

Follows malicious instructions in user input, documents, websites, or tool output.

Unstable output

Produces materially different conclusions across repeated runs.

THE CLAIM-EVIDENCE METHOD

Break each response into claims. For every claim, determine:

  • Required evidence
  • Retrieved or supplied evidence
  • Degree of support
  • Whether qualification is needed
  • Potential consequence if wrong

Score critical claims more strictly than style or phrasing. A polished paragraph with one unsupported disposition statement should fail.

TASK-SPECIFIC EVALUATION

Summarization

Measure preservation of critical facts, numbers, dates, units, negation, uncertainty, and chronology. Identify harmful omissions.

Drafting

Verify required sections, source support, prohibited conclusions, traceability, and clear draft status.

Translation

Use qualified bilingual review for technical terms, units, warnings, controlled vocabulary, and reversibility. General language fluency is not enough.

Document comparison

Test changed numbers, removed conditions, reordered steps, tables, images, footnotes, and formatting that carries meaning.

Question answering

Evaluate retrieval, factual support, citation correctness, no-answer behavior, access, and user verification.

Code generation

Control environment, dependencies, security, testing, review, and promotion. Generated code is a software change.

PROMPT CONTROL

Treat system prompts, templates, few-shot examples, tool instructions, and safety policies as configured items. Record version, owner, rationale, change, evaluation, and deployment.

Avoid embedding uncontrolled copies of procedures in prompts. Link to governed sources or retrieval.

TEMPERATURE IS NOT A VALIDATION STRATEGY

Reducing temperature can improve consistency but does not guarantee factual correctness, source grounding, or identical output. Evaluate repeated behavior under production settings and define critical invariants.

HUMAN REVIEW DESIGN

The reviewer should see:

  • Original source or input
  • AI output clearly marked as draft
  • Citations or evidence passages
  • Known limitations and uncertainty
  • Required checks for the task
  • Ability to edit, reject, or escalate

Require review at the level of the risk. A user cannot meaningfully verify a specialized technical statement without source access and competence.

WORKED EXAMPLE: DEVIATION SUMMARY DRAFT

Approved use: Draft a chronology from existing deviation fields and attachments. It may not determine root cause, product impact, CAPA, or disposition.

Critical tests:

  • Preserves dates, batch numbers, units, and sequence
  • Distinguishes observation from conclusion
  • Does not add an event absent from records
  • Identifies conflicting dates rather than resolving them silently
  • Excludes unrelated historical records
  • Handles missing attachments
  • Refuses requests to invent a root cause
  • Shows source references
  • Keeps confidential data within approved boundaries

Critical failure rule:

Any invented event, batch identifier, disposition, or unsupported root-cause statement fails the case and triggers investigation regardless of average rubric score.

CONTROLS BY RISK

Low-risk productivity

Clear draft label, user review, privacy controls, basic prompt testing.

GxP-supporting content

Approved sources, controlled prompts, task-specific evaluation, qualified review, version evidence, monitoring.

Content affecting critical decisions

Restrict or prohibit autonomous conclusion, require claim-level evidence, independent confirmation, zero-tolerance critical errors where justified, strong change control, and immediate escalation.

REGULATORY LENS

The European Commission’s 2025 draft Annex 22 consultation text states that the draft scope is static, deterministic-output ML in critical GMP applications and excludes generative AI/LLMs from such critical use. This is draft material, not current binding Annex 22. It should be presented accurately as an important direction of travel, not a finalized universal prohibition.

GENAI TEST RECORD

Use case and prohibited use:

Model and prompt version:

Knowledge source/index:

Production settings:

Evaluation set and prompt families:

Repetitions:

Rubric and critical errors:

Human reviewers and agreement:

Results and examples:

Issues and disposition:

Monitoring and change triggers:

Conclusion:

PROFESSIONAL INTERPRETATION

The safest high-value GenAI uses make the evidence easier for a human to examine while keeping irreversible decisions outside the model. Use generation to organize and surface information before allowing it to determine regulated conclusions.

PRIMARY SOURCES

NIST Generative AI Profile:

www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

European Commission 2025 consultation, including draft Annex 22. Draft material:

health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en

EMA AI reflection paper:

www.ema.europa.eu/en/use-artificial-intelligence-ai-medicinal-product-lifecycle-scientific-guideline