ARCHIVES
Year 2026 · Volume 5 · Issue 2
Original Article
Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases
Saurabh Lalwani1
Dr. Vishala Bodetti2
Kishan Gor3
Nrip Nihalani4
Aditya Patkar5
1 2 3 4 5 Plus91 Technologies Pvt. Ltd., Pune, Maharashtra, India.
Published Online: May-August 2026
Pages: 997-1004
Cite this article
↗ https://www.doi.org/10.59256/indjcst.20260502110References
1. A. E. W. Johnson, L. Bulgarelli, L. Shen, et al., “MIMIC-IV, a freely accessible electronic health record dataset,” Scientific Data, vol. 10, art.
1, 2023. doi: 10.1038/s41597-022-01899-x.
2. F. Gaber and A. Akalin, “MIMIC-IV-Ext clinical decision support for referral, triage and diagnosis,” PhysioNet, version 1.0.2, 2025, doi:
10.13026/stnm-qx35. Available: https://physionet.org/content/mimic-iv-ext-cds/1.0.2/ (accessed Jul. 22, 2026).3. F. Gaber, M. Shaik, F. Allega, et al., “Evaluating large language model workflows in clinical decision support for triage and referral and
diagnosis,” npj Digital Medicine, vol. 8, art. 263, 2025. doi: 10.1038/s41746-025-01684-1.
4. T. A. Buckley, B. Crowe, R.-E. E. Abdulnour, A. Rodman, and A. K. Manrai, “Comparison of frontier open-source and proprietary large
language models for complex diagnoses,” JAMA Health Forum, vol. 6, no. 3, e250040, 2025. doi: 10.1001/jamahealthforum.2025.0040.
5. D. McDuff, M. Schaekermann, T. Tu, et al., “Towards accurate differential diagnosis with large language models,” Nature, vol. 642, pp. 451-
457, 2025. doi: 10.1038/s41586-025-08869-4.
6. K. Singhal, S. Azizi, T. Tu, et al., “Large language models encode clinical knowledge,” Nature, vol. 620, pp. 172-180, 2023. doi:
10.1038/s41586-023-06291-2.
7. W. G. Cochran, “The comparison of percentages in matched samples,” Biometrika, vol. 37, no. 3/4, pp. 256-266, 1950.
8. Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, pp. 153-
157, 1947.
9. S. Holm, “A simple sequentially rejective multiple test procedure,” Scandinavian Journal of Statistics, vol. 6, no. 2, pp. 65-70, 1979.
10. W. Kwon, Z. Li, S. Zhuang, et al., “Efficient memory management for large language model serving with PagedAttention,” in Proceedings of
the 29th ACM Symposium on Operating Systems Principles, pp. 611-626, 2023.
11. Anthropic, “Introducing Claude Sonnet 5,” Jun. 30, 2026. Available: https://www.anthropic.com/news/claude-sonnet-5 (accessed Jul. 22,
2026).
12. OpenAI, “GPT-5.6: Frontier intelligence that scales with your ambition,” Jul. 9, 2026. Available: https://openai.com/index/gpt-5-6/ (accessed
Jul. 22, 2026).
1, 2023. doi: 10.1038/s41597-022-01899-x.
2. F. Gaber and A. Akalin, “MIMIC-IV-Ext clinical decision support for referral, triage and diagnosis,” PhysioNet, version 1.0.2, 2025, doi:
10.13026/stnm-qx35. Available: https://physionet.org/content/mimic-iv-ext-cds/1.0.2/ (accessed Jul. 22, 2026).3. F. Gaber, M. Shaik, F. Allega, et al., “Evaluating large language model workflows in clinical decision support for triage and referral and
diagnosis,” npj Digital Medicine, vol. 8, art. 263, 2025. doi: 10.1038/s41746-025-01684-1.
4. T. A. Buckley, B. Crowe, R.-E. E. Abdulnour, A. Rodman, and A. K. Manrai, “Comparison of frontier open-source and proprietary large
language models for complex diagnoses,” JAMA Health Forum, vol. 6, no. 3, e250040, 2025. doi: 10.1001/jamahealthforum.2025.0040.
5. D. McDuff, M. Schaekermann, T. Tu, et al., “Towards accurate differential diagnosis with large language models,” Nature, vol. 642, pp. 451-
457, 2025. doi: 10.1038/s41586-025-08869-4.
6. K. Singhal, S. Azizi, T. Tu, et al., “Large language models encode clinical knowledge,” Nature, vol. 620, pp. 172-180, 2023. doi:
10.1038/s41586-023-06291-2.
7. W. G. Cochran, “The comparison of percentages in matched samples,” Biometrika, vol. 37, no. 3/4, pp. 256-266, 1950.
8. Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, pp. 153-
157, 1947.
9. S. Holm, “A simple sequentially rejective multiple test procedure,” Scandinavian Journal of Statistics, vol. 6, no. 2, pp. 65-70, 1979.
10. W. Kwon, Z. Li, S. Zhuang, et al., “Efficient memory management for large language model serving with PagedAttention,” in Proceedings of
the 29th ACM Symposium on Operating Systems Principles, pp. 611-626, 2023.
11. Anthropic, “Introducing Claude Sonnet 5,” Jun. 30, 2026. Available: https://www.anthropic.com/news/claude-sonnet-5 (accessed Jul. 22,
2026).
12. OpenAI, “GPT-5.6: Frontier intelligence that scales with your ambition,” Jul. 9, 2026. Available: https://openai.com/index/gpt-5-6/ (accessed
Jul. 22, 2026).
Related Articles
2026
Artificial Intelligence in Learning and Teaching
2026
Admin Assist: An AI – Driven Configuration and Orchestration for Enterprise Application
2026
Enhancing Blood Group Identification using pigeon inspired optimization: An Innovative Approach
2026
Eco-Genius: Power Up Smart, Power Down Waste
2026
Crowd-Sourced Disaster Response and Rescue Assistant
2026
Unveiling Deepfake Detection Using Vision Transformers: A Survey and Experimental Study
Share Article
Or copy link
https://www.indjcst.com/archives/evaluation-of-diagnostic-accuracy-of-open-source-and-proprietary-large-language-models-across-multi-system-clinical-cases
*Instagram doesn't support direct link sharing from web. Copy the link and share it in your Instagram story or post.