CALM-RAD: Calibrated GPT-4o Confidence-Based Triage Enables 96–100% Accuracy in Automated TNM Staging of Head and Neck Cancer Reports

Bhayana R. Chatbots and Large Language Models in Radiology: A Practical Primer for Clinical and Research Applications. Radiology. 2024 Jan;310(1):e232756.

Bhayana R, Biswas S, Cook TS, Kim W, Kitamura FC, Gichoya J, et al. From Bench to Bedside With Large Language Models: AJR Expert Panel Narrative Review. AJR Am J Roentgenol. 2024 Sep;223(3):e2430928.

Meng X, Yan X, Zhang K, Liu D, Cui X, Yang Y, et al. The application of large language models in medicine: A scoping review. iScience. 2024 Apr 23;27(5):109713.

Zhao H, Chen H, Yang F, Liu N, Deng H, Cai H, et al. Explainability for Large Language Models: A Survey [Internet]. arXiv; 2023 [cited 2025 May 15]. Available from: http://arxiv.org/abs/2309.01029

Byrd C, Ajawara U, Laundry R, Radin J, Bhandari P, Leung A, et al. Performance of a rule-based semi-automated method to optimize chart abstraction for surveillance imaging among patients treated for non-small cell lung cancer. BMC Med Inform Decis Mak. 2022 Jun 3;22:148.

Hanna GJ, Patel N, Tedla SG, Baugnon KL, Aiken A, Agrawal N. Personalizing Surveillance in Head and Neck Cancer. Am Soc Clin Oncol Educ Book. 2023 Jan;43:e389718.

Zhou L, Blackley SV, Kowalski L, Doan R, Acker WW, Landman AB, et al. Analysis of Errors in Dictated Clinical Documents Assisted by Speech Recognition Software and Professional Transcriptionists. JAMA Netw Open. 2018 Jul 6;1(3):e180530.

Using logprobs | OpenAI Cookbook [Internet]. [cited 2025 May 15]. Available from: https://cookbook.openai.com/examples/using_logprobs

Jiang X, Osl M, Kim J, Ohno-Machado L. Smooth Isotonic Regression: A New Method to Calibrate Predictive Models. AMIA Jt Summits Transl Sci Proc. 2011 Mar 7;2011:16–20.

Lydiatt WM, Patel SG, O’Sullivan B, Brandwein MS, Ridge JA, Migliacci JC, et al. Head and Neck cancers-major changes in the American Joint Committee on cancer eighth edition cancer staging manual. CA Cancer J Clin. 2017 Mar;67(2):122–37.

API Reference - OpenAI API [Internet]. [cited 2025 Jul 3]. Available from: https://platform.openai.com

Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010 Jan;21(1):128–38.

Baran E, Lee M, Aviv S, Weiss J, Pettengell C, Karam I, et al. Oropharyngeal Cancer Staging Health Record Extraction Using Artificial Intelligence. JAMA Otolaryngol Head Neck Surg. 2024 Dec 1;150(12):1051–7.

Kayaalp M, Bölek H, Yaşar HA. Confirmation of Large Language Models in Head and Neck Cancer Staging. Diagnostics (Basel). 2025 Sep 18;15(18):2375.

Hamroun A, Amouyel P, Bentegeac R, Kuchcinski G, Le Guellec B. Unveiling Inner Doubts beyond ChatGPT’s Apparent Overconfidence. Radiology. 2025 Jan;314(1):e241766.

Krishna S, Bhambra N, Bleakney R, Bhayana R. Evaluation of Reliability, Repeatability, Robustness, and Confidence of GPT-3.5 and GPT-4 on a Radiology Board-style Examination. Radiology. 2024 May;311(2):e232715.

Omar M, Agbareia R, Glicksberg BS, Nadkarni GN, Klang E. Benchmarking the Confidence of Large Language Models in Answering Clinical Questions: Cross-Sectional Evaluation Study. JMIR Med Inform. 2025 May 16;13:e66917.

Geng J, Cai F, Wang Y, Koeppl H, Nakov P, Gurevych I. A Survey of Confidence Estimation and Calibration in Large Language Models [Internet]. arXiv; 2024 [cited 2025 May 15]. Available from: http://arxiv.org/abs/2311.08298

Pawitan Y, Holmes C. Confidence in the Reasoning of Large Language Models [Internet]. arXiv; 2024 [cited 2025 Jun 27]. Available from: http://arxiv.org/abs/2412.15296

Wang C, Zhou W, Ghosh S, Batmanghelich K, Li W. Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation [Internet]. arXiv; 2025 [cited 2025 Jun 27]. Available from: http://arxiv.org/abs/2412.04606

Ma H, Chen J, Zhou JT, Wang G, Zhang C. Estimating LLM Uncertainty with Evidence [Internet]. arXiv; 2025 [cited 2025 Jun 27]. Available from: http://arxiv.org/abs/2502.00290

Stangel P, Bani-Harouni D, Pellegrini C, Özsoy E, Zaripova K, Keicher M, et al. Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models [Internet]. arXiv; 2025 [cited 2025 Jun 27]. Available from: http://arxiv.org/abs/2503.02623

Gu B, Desai RJ, Lin KJ, Yang J. Probabilistic medical predictions of large language models. NPJ Digit Med. 2024 Dec 19;7(1):367.

Savage T, Wang J, Gallo R, Boukil A, Patel V, Safavi-Naini SAA, et al. Large language model uncertainty proxies: discrimination and calibration for medical diagnosis and treatment. J Am Med Inform Assoc. 2025 Jan 1;32(1):139–49.

Gao Y, Myers S, Chen S, Dligach D, Miller T, Bitterman DS, et al. Uncertainty estimation in diagnosis generation from large language models: next-word probability is not pre-test probability. JAMIA Open. 2025 Feb;8(1):ooae154.

Comments (0)

No login
gif