AI model monitoring: drift, performance and bias in production
An AI model that passed every test at launch can still fail six months later, quietly, as the world changes around it. AI model monitoring is how you catch that. This guide covers what to monitor and why testing is a lifecycle activity. Part of our guide to AI testing.
AramGRC Team·AI Testing & Assurance·September 11, 2026·8 min read
Why AI testing doesn't stop at launch
AI models are trained on a snapshot of the world, but the world keeps moving. Customer behaviour shifts, fraud patterns evolve, language changes — and a model that was accurate at launch slowly stops being accurate. Monitoring is how you detect that decay before it becomes a customer-facing failure or a compliance breach.
What is model drift?
Drift is the gap that opens up between the world a model was trained on and the world it now operates in. Two kinds matter: data drift (the input data changes distribution) and concept drift (the relationship the model learned changes). Both degrade performance without any change to the model itself.
What to monitor
Performance metrics — accuracy, error rates and the task metrics you evaluated on.
Data drift — changes in the distribution of input features versus training data.
Prediction drift — changes in the model's output distribution over time.
Bias — fairness metrics tracked continuously, not just at launch.
Operational signals — latency, error rates and usage anomalies.
Incidents — harmful outputs, complaints and near-misses.
How to monitor
Set baselines from your evaluation, define thresholds that trigger alerts, dashboard the key signals, and schedule regular reviews. When a threshold is breached, it should trigger investigation and, if needed, retraining or rollback.
Monitoring and re-assessment
Monitoring feeds risk management: a drift alert or a new incident should trigger a fresh AI risk assessment and re-testing of the affected system. Monitoring is the sensor; assessment and testing are the response.
Monitoring and regulation
Continuous monitoring is increasingly required: the EU AI Act mandates post-market monitoring for high-risk systems, and ISO/IEC 42001 requires performance evaluation. See the EU AI Act and ISO 42001 certification.
How AramGRC helps
AramGRC's continuous assurance monitors deployed AI for drift, performance decay and emerging bias, and re-tests when the signals move — because models drift.
Frequently asked questions
What is AI model monitoring?+
Continuously tracking a deployed AI model's performance, drift, bias and operational signals to detect degradation or harm before it becomes a failure.
What is model drift?+
The gap between the world a model was trained on and the world it now operates in — data drift (inputs change) and concept drift (the learned relationship changes) — which degrades performance over time.
What should you monitor in an AI model?+
Performance metrics, data and prediction drift, fairness/bias metrics, operational signals like latency and error rates, and incidents.
How often should you monitor AI?+
Continuously for operational signals, with scheduled reviews and automatic alerts on threshold breaches; high-risk systems need the closest watch.