| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:A brief review on benchmarking for large language models evaluation in healthcare
Authors:ID Cilar Budler, Leona (Author)
ID Chen, Hongyu (Author)
ID Chen, Aokun (Author)
ID Topaz, Maxim (Author)
ID Tam, Wilson (Author)
ID Bian, Jiang (Author)
ID Štiglic, Gregor (Author)
Files:URL https://wires.onlinelibrary.wiley.com/doi/10.1002/widm.70010
 
.pdf WIREs-Data-Min-Knowl---2025---Budler---A-Brief-Review-on-Ben.pdf (943,83 KB)
MD5: 132DFD2ED2C2E99251CEF9B5CEE42FF8
 
Language:English
Work type:Scientific work
Typology:1.02 - Review Article
Organization:FZV - Faculty of Health Sciences
FERI - Faculty of Electrical Engineering and Computer Science
Abstract:This paper reviews benchmarking methods for evaluating large language models (LLMs) in healthcare settings. It highlights the importance of rigorous benchmarking to ensure LLMs' safety, accuracy, and effectiveness in clinical applications. The review also discusses the challenges of developing standardized benchmarks and metrics tailored to healthcare-specific tasks such as medical text generation, disease diagnosis, and patient management. Ethical considerations, including privacy, data security, and bias, are also addressed, underscoring the need for multidisciplinary collaboration to establish robust benchmarking frameworks that facilitate LLMs' reliable and ethical use in healthcare. Evaluation of LLMs remains challenging due to the lack of standardized healthcare-specific benchmarks and comprehensive datasets. Key concerns include patient safety, data privacy, model bias, and better explainability, all of which impact the overall trustworthiness of LLMs in clinical settings.
Keywords:artificial intelligence, benchmarking, chatbots, healthcare, large language models, natural language processing
Publication status:Published
Publication version:Version of Record
Submitted for review:17.02.2025
Article acceptance date:11.03.2025
Publication date:01.01.2025
Publisher:Wiley
Year of publishing:2025
Number of pages:str. 1-16
Numbering:Letn. 15, išt. 2, št. članka e70010
PID:20.500.12556/DKUM-92741 New window
UDC:004.89:614
ISSN on article:1942-4795
COBISS.SI-ID:234990851 New window
DOI:10.1002/widm.70010 New window
Publication date in DKUM:12.05.2025
Views:184
Downloads:21
Metadata:XML DC-XML DC-RDF
Categories:Misc.
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Wiley interdisciplinary reviews : Data mining and knowledge discovery
Publisher:John Wiley & Sons
ISSN:1942-4795
COBISS.SI-ID:15994646 New window

Document is financed by a project

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:GC-0001
Name:Artificial Intelligence for Science

Licences

License:CC BY-NC-ND 4.0, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
Link:http://creativecommons.org/licenses/by-nc-nd/4.0/
Description:The most restrictive Creative Commons license. This only allows people to download and share the work for no commercial gain and for no other purposes.
Licensing start date:01.01.2025

Secondary language

Language:Slovenian
Keywords:umetna inteligenca, primerjalna analiza, zdravstvena oskrba, veliki jezikovni modeli, obdelava naravnega jezika


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica