{"id":150521,"date":"2025-05-27T00:00:00","date_gmt":"2025-05-27T00:00:00","guid":{"rendered":"https:\/\/medgoo.com\/index.php\/2025\/05\/27\/fine-tuned-large-language-models-enhance-error-id-in-radiology-reports\/"},"modified":"2025-05-29T15:10:24","modified_gmt":"2025-05-29T15:10:24","slug":"fine-tuned-large-language-models-enhance-error-id-in-radiology-reports","status":"publish","type":"post","link":"https:\/\/medgoo.com\/index.php\/2025\/05\/27\/fine-tuned-large-language-models-enhance-error-id-in-radiology-reports\/","title":{"rendered":"Fine-Tuned Large Language Models Enhance Error ID in Radiology Reports"},"content":{"rendered":"<h3>\n<p>Fine-tuned Llama-3-70B-Instruct model achieved the best performance using zero-shot prompting<\/p>\n<\/h3>\n<p><b>By Elana Gotkine HealthDay Reporter<\/b><br \/>\n<b><\/b><\/p>\n<p>TUESDAY, May 27, 2025 (HealthDay News) &#8212; Large language models (LLMs), fine-tuned on radiology reports, enhance error detection in radiology reports, according to a study published online May 20 in <em>Radiology<\/em>.<\/p>\n<p>Cong Sun, Ph.D., from Weill Cornell Medicine in New York City, and colleagues developed and evaluated generative LLMs for detecting errors in radiology reports pertaining to health care proofreading in a retrospective study. A dataset was constructed with two parts: The first included 1,656 synthetic chest radiology reports generated by GPT-4 (OpenAI) with 828 error-free synthetic reports and 828 containing errors. A total of 614 reports were included in the second part: 307 error-free from the MIMIC chest radiograph (MIMIC-CXR) database and 307 synthetic reports with errors generated by GPT-4. Using zero-shot prompting, few-shot prompting, or fine-tuning strategies, several models were refined, and the performance of these models was assessed.<\/p>\n<p>The researchers found that the fine-tuned Llama-3-70B-Instruct model achieved the best performance using zero-shot prompting, with F1 scores of 0.769, 0.772, 0.750, 0.828, and 0.780 for negation errors, left\/right errors, interval change errors, transcription errors, and overall, respectively. Two radiologists reviewed 200 randomly selected reports output by the model in a real-world evaluation phase; 99 were confirmed by both radiologists to contain errors detected by the models and 163 were confirmed by at least one radiologist.<\/p>\n<p>&#8220;The findings show that fine-tuning is crucial for enabling local deployment of LLMs while also demonstrating the importance of prompt design in optimizing performance for specific medical tasks,&#8221; the authors write.<\/p>\n<p>One author has patents planned, issued, or pending with Weill Cornell Hospital.<\/p>\n<p><a href=\"https:\/\/pubs.rsna.org\/doi\/10.1148\/radiol.242575\">Abstract\/Full Text<\/a><\/p>\n<p><a href=\"https:\/\/pubs.rsna.org\/doi\/10.1148\/radiol.251259\">Editorial (subscription or payment may be required)<\/a><\/p>\n<p><i><\/i><br \/>\n<i>Copyright &#169; 2025 <a href=\"https:\/\/consumer.healthday.com\/\" target=\"_new\" rel=\"noopener\">HealthDay<\/a>. All rights reserved.<\/i><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Fine-tuned Llama-3-70B-Instruct model achieved the best performance using zero-shot prompting<\/p>\n","protected":false},"author":6,"featured_media":151473,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[11],"class_list":["post-150521","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-news","tag-news"],"_links":{"self":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/posts\/150521","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/comments?post=150521"}],"version-history":[{"count":0,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/posts\/150521\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/media\/151473"}],"wp:attachment":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/media?parent=150521"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/categories?post=150521"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/tags?post=150521"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}