ARCHIVES

Original Article

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

Saurabh Lalwani1 Dr. Vishala Bodetti2 Kishan Gor3 Nrip Nihalani4 Aditya Patkar5
1 2 3 4 5 Plus91 Technologies Pvt. Ltd., Pune, Maharashtra, India.

Published Online: May-August 2026

Pages: 997-1004

Abstract

Large language models show potential for clinical diagnostic support, but their diagnostic accuracy across diverse real-world patient presentations remains uncertain. We evaluated diagnostic retrieval and ranking using multi-system emergency-department narratives from MIMIC-IV-Ext version 1.0.2, a deidentified research dataset derived from MIMIC-IV and curated for research involving referral, triage and diagnostic prediction. The dataset was selected because it links early clinical information, including presenting complaints, history and initial vital signs, with documented diagnoses derived from routine care. A locked cohort of 995 diagnosis-free vignettes was used, with one protected index primary diagnosis per case. GPT-5.6 Thinking, Claude Sonnet 5, Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3 independently generated exactly three ranked differential diagnoses for every vignette. The principal outcome was concept-equivalent Top-3 accuracy; Top-1 accuracy, mean reciprocal rank, omission rate, strict text matching, system-wise performance and paired statistical comparisons were secondary outcomes. GPT-5.6 achieved the highest Top-1 accuracy (41.7%), Top-3 accuracy (61.1%) and mean reciprocal rank (0.504). Claude ranked second at 39.8%, 55.3% and 0.467, respectively. Llama reached 25.5% Top-1 and 41.3% Top-3 accuracy, while Mistral reached 22.8% and 38.7%. Overall Top-3 outcomes differed significantly across models (Cochran Q=271.73, df=3, p=1.31×10⁻⁵⁸). Under identical clinical inputs and scoring rules, the proprietary models retrieved the documented index diagnosis more often and ranked it higher than the two open-weight models. These findings provide a reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations and establish a baseline for further clinical validation.

Related Articles

2026

Artificial Intelligence in Learning and Teaching

2026

Admin Assist: An AI – Driven Configuration and Orchestration for Enterprise Application

2026

Enhancing Blood Group Identification using pigeon inspired optimization: An Innovative Approach

2026

Eco-Genius: Power Up Smart, Power Down Waste

2026

Crowd-Sourced Disaster Response and Rescue Assistant

2026

Unveiling Deepfake Detection Using Vision Transformers: A Survey and Experimental Study

Share Article

X
LinkedIn
Facebook
WhatsApp

Or copy link

https://www.indjcst.com/archives/evaluation-of-diagnostic-accuracy-of-open-source-and-proprietary-large-language-models-across-multi-system-clinical-cases

*Instagram doesn't support direct link sharing from web. Copy the link and share it in your Instagram story or post.