AI audits for GxP

Audit what the AI actually does.

Nobody disputes that AI in a regulated process should be audited. The difficulty is that a conventional audit checks whether a system was tested, and a model can pass every test it was given while being wrong about the population you deployed it on.

An audit lens trained on a model core, tracing back through evaluation data, oversight, and monitoring, with one deviating signal marked.

Perspective

The risks AI introduces into a GxP environment are distributed across governance, data, development, system controls, human behaviour, and the passage of time — and the failures that matter are rarely confined to one of them. An audit narrow enough to examine a single document category certifies each area in isolation and misses the failure in the gap. One broad enough to touch everything, without structure, produces motion without direction. The seven-domain model exists to resolve both.

Audit principle

The failure lives in the seams.

A system can look controlled at every individual point and still fail: unrepresentative training data, a reviewer whose scrutiny has decayed into confirmation, and a vendor model refreshed without notice will each pass their own domain. The value of a structured model is not seven things to check — it is the ability to trace a weakness from where it surfaces to the domain that produced it.

Capabilities

Specialist support, connected to the whole system.

Scope is tailored to the engagement; these are the core areas in which QA4Tech can contribute.

01

AI supplier audits

Audit a vendor’s model lifecycle, data governance, evaluation practice, release control, and their obligation to notify you of model change.

02

Internal AI use audits

Examine your own AI register, risk classification, approvals, and whether the oversight your policy requires is genuinely being performed.

03

AI inside validated systems

Audit AI features added to an existing GxP platform, where the original validation never contemplated statistical behaviour.

04

The human–AI team

Audit the reviewer and the system as one unit, because that pairing — not the model alone — is what produces the regulated decision.

05

Data provenance review

Trace training, tuning, and evaluation data for origin, permitted use, representativeness, leakage, and fitness for the deployed population.

06

For-cause AI audits

Establish scope, cause, and exposure after a model produced a wrong or unexplained output that oversight failed to catch.

The seven-domain model

Seven domains, read across rather than worked through.

A reusable, proportionate structure for organizing an AI audit around the areas where risk arises, controls should exist, and evidence can be gathered. The domains are coupled, not independent: a deficiency in one is evidence about the others. Depth within each tracks the risk the system actually carries — this is a structure, not a checklist.

01

Governance & accountability

Confirm that accountable ownership, policy, and oversight exist and bite. Ask who answers if the system proves unfit — a list of teams is not an answer.

02

Intended use & context of use

Confirm the purpose, users, population, setting, inputs, outputs, and exclusions are documented and honest. This is the measuring stick for everything downstream.

03

Data & data integrity

Confirm training, validation, and operational data are fit for purpose, controlled, and traceable — including how set independence was verified rather than assumed.

04

Model development & performance

Confirm the model was developed rigorously, evaluated against criteria fixed in advance, and its limitations understood by the people relying on it.

05

System validation & technical controls

Confirm the system as deployed is validated and controlled. A validated model inside an uncontrolled system is not assured; assurance is a property of the whole assembly.

06

Human oversight & operational use

Confirm review is substantive, competence is demonstrated, and accountability is assigned. The most over-claimed control in AI assurance, and the easiest to test by interview.

07

Lifecycle, change & monitoring

Confirm the system is monitored, change is controlled, and fitness is maintained rather than assumed. Correct at deployment is not correct forever.

Evidence requested

What the organization has to be able to produce.

The evidence request goes out before the audit, specific enough that a curated substitute is obvious. The instruments below are the ones organizations typically hold — a domain is deficient when its question has no adequate answer, not when the answer arrives under an unfamiliar title.

  • An AI inventory with GxP classification, named ownership, and review dates
  • A context-of-use statement: users, population, setting, inputs, outputs, exclusions
  • Data lineage and datasheets, with set independence demonstrated rather than asserted
  • A development plan pre-dating the result, plus model cards and subgroup performance
  • Validation evidence, a model registry, and audit trails that tie an output to a version
  • Review records, competence assessments, and a route by which a reviewer can disagree
  • Monitoring thresholds, drift analyses, and records of what happened when one was crossed
  • Vendor notification records, and the organization’s assessment of supplier model updates

How it is commissioned

The same subject, three kinds of engagement.

An AI audit is not a separate category from the audit types — it is a subject that can be examined under any of them. Which one applies changes the access you have, the tone in the room, and what the report has to support.

01 · As an internal audit

Auditing your own AI use against your own policy, as part of a self-inspection programme or ahead of an inspection.

02 · As a supplier audit

Auditing an AI or AI-enabled provider before qualification, on a risk-based cycle, or after they add AI to a service you already use.

03 · As a for-cause audit

Triggered by a wrong output, a missed review, a regulatory question, or a pattern that only became visible in aggregate.

Approach

Context first. Evidence throughout.

A clear sequence keeps the work rigorous while avoiding unnecessary process.

  1. 01

    Fix the measuring stick

    Establish intended use and context of use precisely enough to be falsifiable, because every other domain is judged against it.

  2. 02

    Weight the domains

    Use a pre-audit risk assessment to decide where to press. Even a low-priority domain gets enough examination to confirm the risk was recognized.

  3. 03

    Follow the evidence

    Test each assertion against lineage, evaluation records, live demonstration, reviewer interviews, and samples chosen by the auditor.

  4. 04

    Trace across the seams

    Ask what a weakness in one domain implies about the others, and report the structural cause rather than a scattered list of findings.

Deliverables

What you receive.

An audit record complete enough to support a qualification decision, an inspection response, or a board-level assurance statement about AI.

  • An audit plan with scope, risk rationale, and a structured evidence request
  • A written report with graded findings and the basis for each conclusion
  • An explicit conclusion on fitness for the stated intended use, with conditions
  • An assessment of residual risk where evidence was unavailable or inconclusive
  • Review and challenge of the auditee’s response and corrective action plan
  • Formal closure documentation once actions are verified as effective

Reference frameworks

The criteria an AI audit is set against.

AI does not have its own GxP regulation. Criteria are assembled from computerized system expectations, quality risk management, and the AI-specific instruments that now apply, then stated in the audit plan before the audit begins.

FDA & EMA, Guiding Principles of Good AI Practice in Drug Development (January 2026)
Ten joint guiding principles from the two agencies — the closest thing to a shared transatlantic statement of what good AI practice in drug development looks like.
FDA draft guidance on AI to support regulatory decision-making (January 2025)
Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products. Draft guidance; its credibility framing around context of use is durable, its status is not final.
EMA reflection paper on AI in the medicinal product lifecycle (September 2024)
EMA/CHMP/CVMP/83833/2023. European thinking on AI across the product lifecycle, including expectations around change and continued fitness.
GAMP 5 (2nd edition)
Risk-based lifecycle and critical thinking, including its treatment of AI and machine learning within computerized system validation.
EU AI Act & ISO/IEC 42001
Obligations by role and risk class, and the AI management system requirements used where a formal AIMS exists to audit against.
ICH E6(R3), 21 CFR Part 11 & EU Annex 11
Sponsor oversight, data governance, electronic records, and the audit trail obligations that AI-generated and AI-influenced content does not escape.

These are examples, not a complete list. The frameworks and criteria that apply to a particular engagement are identified and agreed as part of defining its scope.

Typical applications

Where this work can apply.

  • Auditing an AI or AI-enabled supplier before qualification
  • AI features added to a platform you have already validated
  • Independent audit of your own AI use against your own policy
  • Testing whether human oversight is a real control or a documented one
  • For-cause audit after a wrong output or a missed review
  • EU AI Act or ISO/IEC 42001 readiness examined against evidence
  • AI providers preparing to be audited by regulated customers

Start a conversation

Bring the right level of assurance to the next decision.

Begin with a focused discussion about context, risk, evidence, and the outcome you need.

Request an AI Audit