Insights

Choosing an AI tool for a GxP process: four questions before the demo

Four questions a QA or IT manager should answer before an AI vendor demo: supplier maturity, model category, intended use, and the human control point. Anchored to the ISPE GAMP AI Guide (2025), the draft Annex 22 and Annex 11 §3.

The problem

AI tools arrive through a sales deck. The demo is impressive, the price is attractive, and QA is asked to sign off with whatever the vendor chose to disclose. By the time a validation question is raised, the tool is already in someone's daily routine.

The risk

Annex 11 (2011) §3 makes the regulated user responsible for its suppliers: formal agreements, a risk-based decision on whether to audit, and review of supplier documentation against user requirements. The draft Annex 22 (not in force) adds that model documentation must be reviewed by the regulated user whether the model was built in-house or bought (§2.2), and that acceptance criteria must be at least as high as the process the model replaces (§4.3). A tool chosen on a demo alone leaves all of that undone.

Four questions before the demo

1. How mature is the supplier, and what exactly are they delivering?

The ISPE GAMP AI Guide (2025, Appendix M2 §13.2.1) ties supplier risk to the supplier's impact on the intended use and to the scope of what is delivered: a data set, a pre-trained model, a product with AI sub-systems. Suppliers are categorized by their maturity and their product's maturity. Ask for the software bill of materials and for evidence of the supplier's own testing. Where a foundation model sits inside the product, the guide (Appendix S6 §28.4.4) says supplier controls matter most, including evaluation of software of unknown provenance.

2. Which model category is it?

The guide's Appendix M11 (§22.4) distinguishes rule-based systems, models trained by the regulated company, pre-trained models with fine-tuning, and pre-trained models used without fine-tuning. Each carries different change-control ownership. A foundation model is a starting point, not a finished component; the guide names integration complexity, comprehensibility, reliability, bias and fine-tuning maintenance as its risks (§2.4.2), and states that rigorous testing remains necessary with or without fine-tuning (§9.7.3.3). If the vendor cannot say which category applies, the demo is premature.

3. What is the intended use, in your process, with your data?

Quality risk management for AI starts from a shared understanding of impact on patient safety, product quality and data integrity, the user requirements, and the system functions (GAMP AI Guide Chapter 5). The draft Annex 22 makes intended use concrete: a documented input space covering common and rare variations, approved by subject-matter experts before testing (§3), and per-subgroup acceptance criteria (§4). Write the intended use before the demo and ask the vendor to show the tool against it, not against their sample data.

4. Where is the human control point, and what may the tool never decide?

The guide's trustworthy-AI appendix (M9 §20.2) expects an adequate level of human control and the ability to regain control, with AI complementing rather than replacing cognitive decision-making. The draft Annex 22 goes further: where a human decides on the model's output and testing was reduced on that basis, the human's performance is monitored like any other manual process (§3.3, §10.5). Decide before the demo which decisions stay human (batch disposition, deviation and CAPA closure, complaint classification, any electronic signature) and how the record will show both the tool's proposal and the person's decision. Privacy belongs here too: the guide (M9 §20.5) expects local data protection law, such as GDPR, to be considered, and consent to be adequate where personal data is involved.

After the demo

The outcome should be a documented, auditor-defensible decision: adopt, adopt with conditions, or walk away, with the four answers on file. That document is the first artifact of the tool's validation record and the one an inspector will ask to see.

Want this applied to your systems?

Book a discovery call. We will map where manual review is costing you the most, and whether CSV/CSA, AI governance, or an AI tool assessment is the right place to start.