AI supplier audits
Audit a vendor’s model lifecycle, data governance, evaluation practice, release control, and their obligation to notify you of model change.
AI audits for GxP
Nobody disputes that AI in a regulated process should be audited. The difficulty is that a conventional audit checks whether a system was tested, and a model can pass every test it was given while being wrong about the population you deployed it on.
Perspective
The risks AI introduces into a GxP environment are distributed across governance, data, development, system controls, human behaviour, and the passage of time — and the failures that matter are rarely confined to one of them. An audit narrow enough to examine a single document category certifies each area in isolation and misses the failure in the gap. One broad enough to touch everything, without structure, produces motion without direction. The seven-domain model exists to resolve both.
Audit principle
A system can look controlled at every individual point and still fail: unrepresentative training data, a reviewer whose scrutiny has decayed into confirmation, and a vendor model refreshed without notice will each pass their own domain. The value of a structured model is not seven things to check — it is the ability to trace a weakness from where it surfaces to the domain that produced it.
Capabilities
Scope is tailored to the engagement; these are the core areas in which QA4Tech can contribute.
Audit a vendor’s model lifecycle, data governance, evaluation practice, release control, and their obligation to notify you of model change.
Examine your own AI register, risk classification, approvals, and whether the oversight your policy requires is genuinely being performed.
Audit AI features added to an existing GxP platform, where the original validation never contemplated statistical behaviour.
Audit the reviewer and the system as one unit, because that pairing — not the model alone — is what produces the regulated decision.
Trace training, tuning, and evaluation data for origin, permitted use, representativeness, leakage, and fitness for the deployed population.
Establish scope, cause, and exposure after a model produced a wrong or unexplained output that oversight failed to catch.
The seven-domain model
A reusable, proportionate structure for organizing an AI audit around the areas where risk arises, controls should exist, and evidence can be gathered. The domains are coupled, not independent: a deficiency in one is evidence about the others. Depth within each tracks the risk the system actually carries — this is a structure, not a checklist.
Confirm that accountable ownership, policy, and oversight exist and bite. Ask who answers if the system proves unfit — a list of teams is not an answer.
Confirm the purpose, users, population, setting, inputs, outputs, and exclusions are documented and honest. This is the measuring stick for everything downstream.
Confirm training, validation, and operational data are fit for purpose, controlled, and traceable — including how set independence was verified rather than assumed.
Confirm the model was developed rigorously, evaluated against criteria fixed in advance, and its limitations understood by the people relying on it.
Confirm the system as deployed is validated and controlled. A validated model inside an uncontrolled system is not assured; assurance is a property of the whole assembly.
Confirm review is substantive, competence is demonstrated, and accountability is assigned. The most over-claimed control in AI assurance, and the easiest to test by interview.
Confirm the system is monitored, change is controlled, and fitness is maintained rather than assumed. Correct at deployment is not correct forever.
Evidence requested
The evidence request goes out before the audit, specific enough that a curated substitute is obvious. The instruments below are the ones organizations typically hold — a domain is deficient when its question has no adequate answer, not when the answer arrives under an unfamiliar title.
How it is commissioned
An AI audit is not a separate category from the audit types — it is a subject that can be examined under any of them. Which one applies changes the access you have, the tone in the room, and what the report has to support.
Auditing your own AI use against your own policy, as part of a self-inspection programme or ahead of an inspection.
Auditing an AI or AI-enabled provider before qualification, on a risk-based cycle, or after they add AI to a service you already use.
Triggered by a wrong output, a missed review, a regulatory question, or a pattern that only became visible in aggregate.
Approach
A clear sequence keeps the work rigorous while avoiding unnecessary process.
Establish intended use and context of use precisely enough to be falsifiable, because every other domain is judged against it.
Use a pre-audit risk assessment to decide where to press. Even a low-priority domain gets enough examination to confirm the risk was recognized.
Test each assertion against lineage, evaluation records, live demonstration, reviewer interviews, and samples chosen by the auditor.
Ask what a weakness in one domain implies about the others, and report the structural cause rather than a scattered list of findings.
Deliverables
An audit record complete enough to support a qualification decision, an inspection response, or a board-level assurance statement about AI.
Reference frameworks
AI does not have its own GxP regulation. Criteria are assembled from computerized system expectations, quality risk management, and the AI-specific instruments that now apply, then stated in the audit plan before the audit begins.
These are examples, not a complete list. The frameworks and criteria that apply to a particular engagement are identified and agreed as part of defining its scope.
Typical applications
Start a conversation
Begin with a focused discussion about context, risk, evidence, and the outcome you need.