ARCHIVES
Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases
Published Online: May-August 2026
Pages: 997-1004
Cite this article
↗ https://www.doi.org/10.59256/indjcst.20260502110Abstract
Large language models show potential for clinical diagnostic support, but their diagnostic accuracy across diverse real-world patient presentations remains uncertain. We evaluated diagnostic retrieval and ranking using multi-system emergency-department narratives from MIMIC-IV-Ext version 1.0.2, a deidentified research dataset derived from MIMIC-IV and curated for research involving referral, triage and diagnostic prediction. The dataset was selected because it links early clinical information, including presenting complaints, history and initial vital signs, with documented diagnoses derived from routine care. A locked cohort of 995 diagnosis-free vignettes was used, with one protected index primary diagnosis per case. GPT-5.6 Thinking, Claude Sonnet 5, Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3 independently generated exactly three ranked differential diagnoses for every vignette. The principal outcome was concept-equivalent Top-3 accuracy; Top-1 accuracy, mean reciprocal rank, omission rate, strict text matching, system-wise performance and paired statistical comparisons were secondary outcomes. GPT-5.6 achieved the highest Top-1 accuracy (41.7%), Top-3 accuracy (61.1%) and mean reciprocal rank (0.504). Claude ranked second at 39.8%, 55.3% and 0.467, respectively. Llama reached 25.5% Top-1 and 41.3% Top-3 accuracy, while Mistral reached 22.8% and 38.7%. Overall Top-3 outcomes differed significantly across models (Cochran Q=271.73, df=3, p=1.31×10⁻⁵⁸). Under identical clinical inputs and scoring rules, the proprietary models retrieved the documented index diagnosis more often and ranked it higher than the two open-weight models. These findings provide a reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations and establish a baseline for further clinical validation.
Related Articles
2026
Artificial Intelligence in Learning and Teaching
2026
Admin Assist: An AI – Driven Configuration and Orchestration for Enterprise Application
2026
Enhancing Blood Group Identification using pigeon inspired optimization: An Innovative Approach
2026
Eco-Genius: Power Up Smart, Power Down Waste
2026
Crowd-Sourced Disaster Response and Rescue Assistant
2026
Unveiling Deepfake Detection Using Vision Transformers: A Survey and Experimental Study
Share Article
Or copy link
*Instagram doesn't support direct link sharing from web. Copy the link and share it in your Instagram story or post.