{"id":145657,"date":"2025-05-13T00:00:00","date_gmt":"2025-05-13T00:00:00","guid":{"rendered":"https:\/\/medgoo.com\/index.php\/2025\/05\/13\/openai-releases-healthbench-dataset-to-test-ai-in-health-care\/"},"modified":"2025-05-14T16:10:17","modified_gmt":"2025-05-14T16:10:17","slug":"openai-releases-healthbench-dataset-to-test-ai-in-health-care","status":"publish","type":"post","link":"https:\/\/medgoo.com\/index.php\/2025\/05\/13\/openai-releases-healthbench-dataset-to-test-ai-in-health-care\/","title":{"rendered":"OpenAI Releases HealthBench Dataset to Test AI in Health Care"},"content":{"rendered":"<h3><\/h3>\n<p><b>By I. Edwards HealthDay Reporter<\/b><br \/>\n<b><\/b><\/p>\n<p>TUESDAY, May 13, 2025 (HealthDay News) \u2014 OpenAI has unveiled a large dataset to help test how well artificial intelligence (AI) models answer health care questions.<\/p>\n<p>Experts call it a major step forward, but they also say more work is needed to ensure safety.<\/p>\n<p>The dataset \u2014 called HealthBench \u2014 is OpenAI&#8217;s first major independent health care project. It includes 5,000 \u201crealistic health conversations,\u201d each with detailed grading tools to evaluate AI responses, <em>STAT News <\/em>reported.<\/p>\n<p>\u201cOur mission as OpenAI is to ensure AGI is beneficial to humanity,\u201d <a href=\"https:\/\/www.linkedin.com\/in\/karan1149\">Karan Singhal<\/a>, head of the San Francisco-based company&#8217;s health AI team, said. AGI is shorthand for artificial general intelligence.<\/p>\n<p>\u201cOne part of that is building and deploying technology,&#8221; Singhal said. &#8220;Another part of it is ensuring that positive applications like health care have a place to flourish and that we do the right work to ensure that the models are safe and reliable in these settings.\u201d<\/p>\n<p>The dataset was created with help from 262 doctors who have worked in 60 countries. They provided more than 57,000 unique criteria to judge how well AI models answer health questions.<\/p>\n<p>HealthBench aims to fix a common problem: Comparing different AI models fairly.<\/p>\n<p>\u201cWhat OpenAI has done is they have provided this in a scalable way from a really big, reputable brand that\u2019s going to enable people to use this very easily,\u201d <a href=\"https:\/\/www.medstarhealth.org\/innovation-and-research\/medstar-health-research-institute\/principal-investigators\/raj-ratwani\">Raj Ratwani<\/a>, a health AI researcher at MedStar Health, said.<\/p>\n<p>The 5,000 examples in HealthBench were made using synthesized conversations designed by physicians.<\/p>\n<p>\u201cWe wanted to balance the benefits of being able to release the data with, of course, the privacy constraints of using realistic data,&#8221; Singhal told <em>STAT News.<\/em><\/p>\n<p>The dataset also includes a special group of 1,000 hard examples where AI models struggled. OpenAI hopes this group \u201cprovides a worthy target for model improvements for months to come,&#8221; <em>STAT News<\/em> reported.<\/p>\n<p>OpenAI also tested its own models as well as models from Google, Meta, Anthropic and xAI. OpenAI\u2019s o3 model scored the best, especially in communication quality, <em>STAT News<\/em> reported.<\/p>\n<p>But models performed poorly in areas like context awareness and completeness, experts said.<\/p>\n<p>Some warned about OpenAI grading its own models.<\/p>\n<p>&#8220;In sensitive contexts like health care, where we are discussing life and death, that level of opacity is unacceptable,&#8221; Hao explained.<\/p>\n<p>Others noted that AI itself was used to grade some of the responses, which could result in errors being overlooked.<\/p>\n<p>It \u201cmay hide errors shared by both model and grader,\u201d <a href=\"https:\/\/profiles.mountsinai.org\/girish-n-nadkarni\">Girish Nadkarni<\/a>, head of artificial intelligence and human health at the Icahn School of Medicine at Mount Sinai in New York City, told <em>STAT News.<\/em><\/p>\n<p>He and others called for more reviews to ensure models work well in different countries and among different demographics.<\/p>\n<p>\u201cHealthBench improves large language model&nbsp;health care evaluation but still needs subgroup analysis and wider human review before it can support safety claims,\u201d Nadkarni said.<\/p>\n<p><strong>More information<\/strong><\/p>\n<p>The National Institutes of Health has more on <a href=\"https:\/\/pmc.ncbi.nlm.nih.gov\/articles\/PMC8285156\/\">artificial intelligence in health care<\/a>.<\/p>\n<p>SOURCE: <em>STAT News<\/em>, May 12, 2025<\/p>\n<p><i><\/i><br \/>\n<i>Copyright &#169; 2025 <a href=\"https:\/\/consumer.healthday.com\/\" target=\"_new\" rel=\"noopener\">HealthDay<\/a>. All rights reserved.<\/i><\/p>\n","protected":false},"excerpt":{"rendered":"<p>By I. Edwards HealthDay Reporter TUESDAY, May 13, 2025 (HealthDay News) \u2014 OpenAI has unveiled a large dataset to help test how well artificial intelligence (AI) models answer health care questions. Experts call it a major step forward, but they also say more work is needed to ensure safety. The dataset \u2014 called HealthBench \u2014 [&hellip;]<\/p>\n","protected":false},"author":6,"featured_media":146090,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[11],"class_list":["post-145657","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-news","tag-news"],"_links":{"self":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/posts\/145657","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/comments?post=145657"}],"version-history":[{"count":0,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/posts\/145657\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/media\/146090"}],"wp:attachment":[{"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/media?parent=145657"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/categories?post=145657"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/medgoo.com\/index.php\/wp-json\/wp\/v2\/tags?post=145657"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}