There are now more than 1,100 FDA-cleared AI applications in radiology. That number represents years of clinical validation, regulatory review, and significant investment by application developers and health systems alike and it reflects genuine progress in what medical imaging AI can do.
Regulatory clearance is a snapshot of AI performance. It tells you how a model performs under controlled conditions, on a carefully curated dataset, at a specific point in time. As AI becomes part of routine clinical care, a new question is emerging: how does any particular model perform over time, in your hospital, on your patients, and with your diagnostic equipment? Increasingly the evidence suggests the answer deserves closer attention.
In a recent study published by Emory Healthcare1, researchers assessed the performance of a commercial AI triage application used for intracranial haemorrhage (ICH) detection across a range of criteria. The team measured algorithm performance across more than 100,000 CT scans from over 74,000 patients across two years, setting a performance target of 96.15% sensitivity.2 based on the application’s FDA clearance figure.
Overall, real-world aggregate sensitivity came in at 82.2%, lower than the performance target, but not alarmingly so. The authors of the study considered this performance clinically acceptable for an assistive tool where the final call always rests with a physician.
The researchers went on to evaluate the performance of the algorithm by various demographic and clinical subgroups, and found outputs varied across different clinical use cases. For example, algorithm performance was very good for large acute ICH (95%) but not as good for subacute ICH (45%). Algorithm performance for outpatients (72.2%) was still good but less than aggregate performance levels.
The algorithm was performing in ways that aggregate figures would never reveal, and that without systematic monitoring, nobody would know.
This isn’t an isolated finding. AI models can perform differently across various patient populations and different clinical settings. The training datasets used to build AI models don’t always reflect the full diversity of real-world hospital environments and global patient populations. When a hospital purchases a new CT scanner with a better detector there could be changes to what an AI algorithm can detect and therefore how it behaves. Workflows shift. Protocols get updated. A new technologist joins and uses the system differently. Each of those changes can move model performance in ways that are completely invisible without systematic measurement.
AI models are being created all over the world and are being trained on different populations. A model trained predominantly on one demographic group will not necessarily perform with the same accuracy on others. For health systems serving diverse communities, that matters. Hospitals are increasingly demanding AI governance tools that give them the ability to detect and analyse model performance across a range of criteria.
What was until recently only a clinical concern is rapidly becoming a compliance one.
The EU AI Act3 is bringing specific requirements to medical imaging AI this year, including bias testing, transparency obligations, and audit logging. For health systems and model developers operating across European markets, post-market monitoring is no longer optional infrastructure; it is a regulatory baseline.
The evidence suggests that baseline is not yet being met. At the European Congress of Radiology earlier this year, researcher Kicky Van Leeuwen4 highlighted a striking finding: none of the 13 manufacturers visited by the Dutch Health and Youth Care Inspectorate in 2023/2024 met post-market surveillance requirements. As she put it: "If we want to ensure long-term safety of AI, in a world where the only constant is change, we need post-deployment monitoring."
In the US, the picture is moving in the same direction. The FDA's introduction of Predetermined Change Control Plans5 means that developers who want to update their models over time (without seeking a new clearance for every change) must include performance monitoring as part of that process.
And as foundation models -- a newer, more powerful and potentially more drift-prone generation of AI -- begin entering the medical imaging market, monitoring requirements are likely to intensify further. The direction of travel is clear. What has been a gap in good practice is becoming a gap in compliance.
We know that AI will play a significant role in delivering healthcare now and in the future. AI is already having an extremely positive impact and helping clinicians to save time, improve accuracy and take better care of their patients.
The question isn’t “Does AI work”? The question is, “Where does AI work best”? Like any tool in a clinician’s hands, understanding the clinical conditions where a tool works best is the ultimate goal.
If a tool cleared at 96.15% sensitivity is performing at 45.5% sensitivity for a specific patient group in your department, would you know?
For most health systems right now, the honest answer is no because data to answer that question isn't being captured. Measuring AI performance at the level of granularity that matters clinically is genuinely difficult, and most departments are working without the tools to do it.
Algorithm developers are in the same position. Understanding how any particular AI model works in the real world provides insights and opportunities for improvements and, right now, that feedback loop is complex. Post-market monitoring isn't just a compliance obligation; it’s understanding how to use and improve each of the parts and how the whole ecosystem gets better.
Clearance marks the beginning of an application’s life in clinical practice, not the end of the evaluation process.
The Emory study is a reminder of how important it is to understand how any AI model is working post deployment, and to know which AI can most benefit your own patient population.
That's the conversation worth having.
1. Trivedi H, et al. Real-world performance evaluation of a commercial deep learning model for intracranial haemorrage detection. npj Digital Medicine. 2025;9(1):66. https://doi.org/10.1038/s41746-025-02244-3
2. US Food and Drug Administration. 510(k) Premarket Notification K221240: https://www.accessdata.fda.gov/cdrh_docs/pdf22/K221240.pdf
3. Regulation (EU) 2024/1689 of the European Parliament and of the Council (EU AI Act), Article 26. In force August 2024, applied from August 2026. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689
4. Van Leeuwen K. Ethical AI in Radiology - Performance, People, and Post-Market Responsibility. European Medical Journal. April 2026. https://www.emjreviews.com/radiology/congress-review/congress-feature-ethical-ai-in-radiology-performance-people-and-post-market-responsibility-j14126/
5. US Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions. Final Guidance, December 2024. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence