As medical intelligence/" class="gmj-dict-autolink" title="Dictionary: Artificial Intelligence">artificial intelligence evolves from single-task diagnostic tools toward complex multimodal agents that navigate entire clinical workflows, the field faces a critical measurement problem: current benchmarks assess only whether AI arrives at the right answer, not whether it reasons soundly, acts safely, or uses resources wisely. According to a perspective published in PLOS Medicine, this gap threatens to deploy AI systems into clinical practice that perform well on narrow tasks but fail in real-world complexity.
Key takeaways
- Current AI benchmarks measure final diagnostic or treatment outputs, ignoring clinical reasoning quality and safety decisions
- Modern medical AI is shifting from single-task models to multimodal large language model–based agents capable of handling multi-step clinical workflows
- New benchmarking frameworks must evaluate process safety, resource stewardship, and clinical reasoning—not just accuracy
From single tasks to clinical agents
For over a decade, medical AI researchers have focused on narrow problems: identifying pneumonia on a chest X-ray, predicting sepsis from vital signs, or recommending drug interactions. These single-task models are easy to benchmark—feed in data, measure if the answer is correct. But real clinical work is not compartmentalized. A patient arrives with chest pain; the clinician must take a history, order tests, interpret results, revise hypotheses, coordinate with specialists, and make moment-by-moment resource decisions. According to Silas Ruhrberg Estévez and colleagues at the University of Cambridge and collaborators, large language model–based agents can now operate across these multi-step workflows, but benchmarking tools have not evolved to match.
This mismatch creates risk. An AI system that correctly diagnoses pneumonia 95% of the time might still order redundant imaging, miss contraindications, or fail to escalate appropriately—and standard benchmarks would never capture these failures. The authors argue that clinical practice requires evaluation frameworks that measure the entire decision pathway, not just endpoints.
Why process safety cannot be an afterthought
Clinical reasoning is not a black box that can be ignored once outputs are validated. When a physician treats a patient, every step carries safety implications: which tests to order (avoiding harm from unnecessary radiation or invasive procedures), in what sequence (recognizing that early decisions shape later options), with what communication to other providers, and under what resource constraints. According to the PLOS Medicine perspective, benchmarks that ignore process essentially grade AI on luck—systems that arrive at correct diagnoses through flawed reasoning may pass current tests and then fail unpredictably on novel cases in practice.
This is especially critical for multimodal agents handling complex workflows. These systems must integrate text (clinical notes, literature), structured data (lab values, imaging measurements), and unstructured information (radiology reports, audio from patient conversations). If benchmarks do not explicitly measure how an agent weighs conflicting signals, justifies exclusions, or escalates uncertainty, clinicians deploying these tools have no assurance they are safe beyond the narrow training distribution.
Dimensions of Clinical AI Benchmarking
Current benchmarks assess only final outputs; comprehensive evaluation requires measuring process quality across three dimensions
Illustrative framework based on Estévez et al. perspective, PLOS Medicine 2025
Building frameworks for multimodal agent evaluation
The authors propose that benchmarks for next-generation medical AI must explicitly measure three elements currently neglected. First, clinical reasoning quality: Does the agent explain its logic transparently? Does it recognize uncertainty and avoid overconfidence? Does it update hypotheses when new data arrives? Second, process safety: Does the agent avoid redundant or harmful interventions? Does it flag contraindications and drug interactions? Does it escalate appropriately when facing ambiguity? Third, resource stewardship: Does the system minimize unnecessary tests, imaging, or procedures? Does it allocate resources fairly and efficiently? These are not peripheral concerns—they are central to whether an AI tool improves or harms clinical care.
Implementing such benchmarks requires different data structures and evaluation metrics than current systems. Rather than simple accuracy scores, researchers will need to build synthetic clinical scenarios with multiple valid pathways, explicit safety hazards, and resource trade-offs. Evaluators would score not just whether the AI reaches the right diagnosis but how it reasons along the way, where it missteps, and how it responds to new information. This is more complex and time-intensive than current benchmarking—but, as the authors argue, it is the only approach that actually predicts real-world clinical performance.
Medical AI benchmarks must assess clinical reasoning quality, process safety, and resource stewardship—not just diagnostic accuracy—to ensure systems are safe and effective in real clinical workflows.
— Silas Ruhrberg Estévez and colleagues, University of Cambridge (PLOS Medicine, 2025)
Implications for AI deployment in healthcare
The shift from output-only benchmarking to process-focused evaluation has immediate consequences for how hospitals and health systems adopt AI tools. Institutions currently piloting large language model–based clinical agents should demand comprehensive safety evaluations before integration into patient care pathways. Regulatory bodies like the U.S. Food and Drug Administration and the European Medicines Agency will need updated guidance on what constitutes adequate evidence of AI safety and effectiveness for complex workflows—a gap the authors highlight as urgent.
Additionally, this framework raises questions about transparency and accountability. If an AI agent makes a clinical decision, who is responsible when that decision is later questioned—the system’s developers, the clinician who acted on it, or the institution that deployed it? Process-based benchmarking, by making AI reasoning explicit and auditable, supports clearer accountability structures and enables learning from errors. Health policy frameworks will need to evolve in parallel.
What this means
Frequently asked questions
Why does clinical reasoning matter if an AI system gets the diagnosis right?
Because a correct answer can arrive through flawed logic. An AI system might diagnose pneumonia correctly but recommend redundant imaging, miss a critical contraindication, or fail to recognize when a patient needs urgent escalation. These process failures only appear in real clinical use—not in benchmarks that measure final outputs alone. Moreover, systems that reason soundly are more likely to generalize safely to new situations and populations beyond their training data.
What would a process-focused benchmark for medical AI actually look like?
Rather than presenting isolated cases and grading accuracy, researchers would evaluate AI agents on multi-step clinical scenarios with embedded safety hazards, ambiguous data, and resource constraints. Evaluators would score how transparently the system reasons, whether it flags risks appropriately, how it updates thinking when new information arrives, and whether its resource use is justifiable. This requires richer evaluation datasets and longer assessment times than current benchmarks—but it predicts real-world performance.
Should hospitals stop using AI tools until these new benchmarks exist?
Not necessarily, but they should demand comprehensive safety evaluations beyond accuracy claims. Current deployments should focus on lower-stakes applications (decision support, documentation assistance) where errors are caught by human oversight, and require ongoing monitoring for unintended harms. As benchmarking frameworks mature, hospitals can confidently expand AI use to higher-stakes clinical decisions with greater assurance of safety and effectiveness.
The evolution of medical AI benchmarking reflects a maturation of the field itself. As these systems move from research tools to clinical agents integrated into patient care, the criteria for success must broaden from narrow accuracy metrics to comprehensive assessment of safety, reasoning quality, and resource stewardship. This shift requires investment in new evaluation frameworks, but it is the foundation upon which trustworthy, deployable clinical AI will be built.
Source: Medical AI agents require benchmarks measuring clinical reasoning, safety, and resource use—not outputs alone, PLOS Medicine 2025
Was this article helpful?
Disclaimer. This article is health journalism intended for general information and education. It is not medical advice and is not a substitute for professional diagnosis or treatment. Always consult a qualified healthcare provider about your individual circumstances. Full disclaimer →
Related Coverage




Editorial standards. This article was produced under the GMJ News editorial process, with oversight by the GMJ Editorial Board. Our editorial process. Spotted an error? Contact the editorial team.






