Back to case studiesAI QAHealthcareIndia / Middle East
Red-teaming a clinical documentation assistant before launch
Before a clinical summarisation assistant went live, we ran independent adversarial testing across safety, privacy and robustness.
- Client
- A hospital network piloting a clinical LLM assistant
- Service
- AI QA
- Frameworks
- NIST AI RMF · ISO/IEC 42001
The challenge
- A vendor-supplied LLM assistant drafting clinical notes, tested only for accuracy
- Clinical leadership concerned about hallucinated findings and patient-data leakage
- No independent evidence to support a go/no-go decision
What we did
- 01
Threat modelling
Mapped realistic misuse and failure paths specific to clinical documentation.
- 02
Adversarial testing
Prompt injection, jailbreaks, data-exfiltration attempts and unsafe-output probing.
- 03
Robustness checks
Behaviour under ambiguous, incomplete and adversarially phrased clinical inputs.
- 04
Independent report
Findings rated by severity with concrete mitigations for the vendor and the hospital.
Outcomes
- Safety-relevant failure modes identified and mitigated before clinical rollout
- Guardrails, prompt controls and human-review checkpoints agreed with clinicians
- An independent test report for the clinical governance committee
- A repeatable test suite to re-run at each model or vendor update
Facing something similar?
A short conversation is usually enough to scope the right assurance work.
Book a free consulting