Insights
Audit trail review: from sampling to full coverage without breaking Part 11
What 21 CFR Part 11, EU Annex 11 and the MHRA data integrity guidance actually require for audit trail review, why review by exception is the accepted route to full coverage, and where a human must still decide.
The problem
Audit trails are switched on in most GxP systems and reviewed in few. Where a review exists it is often a sample: ten records a month, read by the system owner, signed off with "no issues." The volume makes anything else feel impossible.
The risk
21 CFR 11.10(e) requires secure, computer-generated, time-stamped audit trails that record who created, modified or deleted a record and when, retained as long as the record and available to FDA. EU Annex 11 (2011) §9 requires that audit trails be available, convertible to an intelligible form, and regularly reviewed. A sample with no documented logic is not a risk-based review, it is a hope. And the draft Annex 11 revision (not in force) points where expectations are heading: review by someone not involved in the activity, and review before batch release unless a later detection risk is justified (draft §12.6, §12.8).
What the regulators actually allow
The MHRA GxP Data Integrity Guidance (2018) is the most practical text on this. §6.13 says it is not necessary for audit trail review to include every system activity such as log-ons or keystrokes, and that audit trails may be reviewed as a list of relevant data or by an exception reporting process, where an exception report is a validated search tool that identifies predetermined abnormal data. §3.6 sets the principle: effort should be commensurate with the risk and impact of a data integrity failure. §3.5 adds that a routine forensic approach is not expected. §6.15 requires a documented review procedure with a positive statement of whether issues were found, signed and dated.
The GAMP Good Practice Guide on Data Integrity by Design (§6.3.2) distinguishes three reviews: audit trail review as part of normal data review before data release, review during an investigation, and verification that the audit trail function works, as part of periodic review. Its §6.3.2.1 states the principle plainly: use software to screen all results and flag suspect data, then use people to investigate what is flagged. That is review by exception, and it requires a validated search tool, configuration management of its limits, and an audit trail on the tool itself.
Sampling to full coverage
So the route is not "read everything" and not "sample and hope." It is: define the abnormal patterns by risk (changes after approval, deletions, edits outside working hours, repeated re-analysis, privilege changes), build or configure a validated exception report that scans every record, and have a qualified reviewer decide on every flag. Coverage becomes total, human effort becomes proportional to what the screen finds.
This is also where automation and AI agents enter a quality routine legitimately. The scanning step is a computerized function and is validated like one (Annex 11 §4). The decision stays with the reviewer. The ISPE GAMP AI Guide (2025, Appendix M9) frames the balance: AI should complement rather than replace cognitive decision-making, with human control points kept throughout, and with attention to the tendency for human vigilance to drop once a tool is trusted.
What a defensible program looks like
An audit trail review SOP that names the systems in scope, the risk basis for what is screened, the exception logic, the reviewer's independence, the frequency (tied to batch release for batch-critical systems), and the record of each review with a positive statement. A validated exception tool with its own change control. Periodic verification that the audit trail function still works. And training for reviewers on what an anomaly looks like, because a review is only as good as the person who reads the flag.
Want this applied to your systems?
Book a discovery call. We will map where manual review is costing you the most, and whether CSV/CSA, AI governance, or an AI tool assessment is the right place to start.