<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://dk.um.si/IzpisGradiva.php?id=95932"><dc:title>Human-led and artificial intelligence-automated critical appraisal of systematic reviews</dc:title><dc:creator>Gosak,	Lucija	(Avtor)
	</dc:creator><dc:creator>Štiglic,	Gregor	(Avtor)
	</dc:creator><dc:creator>Tam,	Wilson	(Avtor)
	</dc:creator><dc:creator>Vrbnjak,	Dominika	(Avtor)
	</dc:creator><dc:subject>artificial intelligence in healthcare</dc:subject><dc:subject>multimodal large language models</dc:subject><dc:subject>nursing</dc:subject><dc:subject>evidence-based practice</dc:subject><dc:description>Aim To evaluate and compare human-led and artificial intelligence-automated critical appraisal of evidence. Background Critical appraisal is essential in evidence-based practice, yet many nurses lack the skills to perform it. Large language models offer potential support, but their role in critical appraisal remains underexplored. Design We conducted a comparative study to evaluate the performance of five commonly used large language models versus two human reviewers in appraising four systematic reviews on interventions to reduce medication administration errors. Methods We compared large language models and two human reviewers in independently appraising four systematic reviews using the JBI Critical Appraisal Checklist. These models were Perplexity Sonar (Pro), Claude 3.7 Sonnet, Gemini 2.0 Flash, GPT-4.5 and Grok-2. All models received identical full texts and standardized prompts. Responses were analyzed descriptively and agreement was assessed using Cohen’s Kappa. Results Large language models showed full agreement with human reviewers on five of 11 JBI items. Most disagreements occurred in appraising search strategy, inclusion criteria and publication bias. The agreement between human reviewers and large language models ranged from slight to moderate. The highest level of agreement was observed with Claude (κ = 0.732), while the lowest level was observed with Gemini (κ = 0.394). Conclusion Large language models can support aspects of critical appraisal evidence but lack contextual reasoning and methodological insight required for complex judgments. While Claude 3.7 Sonnet aligned most closely with human reviewers, human oversight remains essential. Large language models should serve as adjuncts and not substitutes for evidence-based practice.</dc:description><dc:publisher>Elsevier</dc:publisher><dc:date>2025</dc:date><dc:date>2025-11-12 12:23:49</dc:date><dc:type>Znanstveno delo</dc:type><dc:identifier>95932</dc:identifier><dc:language>sl</dc:language><dc:rights>© 2025 The Authors. Published by Elsevier Ltd.</dc:rights></rdf:Description></rdf:RDF>
