| Title: | Development and evaluation of a generative AI chatbot for database searching in systematic review |
|---|
| Authors: | ID Tam, Wilson (Author) ID Tung, Neo (Author) ID Lee, Shi Xuan (Author) ID Štiglic, Gregor (Author) ID Huynh, Tom (Author) ID Tang, Arthur (Author) |
| Files: | J_of_Nursing_Scholarship_-_2026_-_Tam_-_Development_and_Evaluation_of_a_Generative_AI_Chatbot_for_Database_Searching_in.pdf (1,09 MB) MD5: 8498E1E2F58433264ACFA44F9AE8B076
https://sigmapubs.onlinelibrary.wiley.com/doi/10.1111/jnu.70076
|
|---|
| Language: | English |
|---|
| Work type: | Scientific work |
|---|
| Typology: | 1.01 - Original Scientific Article |
|---|
| Organization: | FZV - Faculty of Health Sciences
|
|---|
| Abstract: | Introduction Systematic reviews (SRs) require comprehensive, reproducible searches, yet developing search strategies is resource-intensive and demands specialized expertise. Generative AI offers potential to streamline this process, but empirical evaluations for GAI-assisted SR searching remain scarce. The objectives of this study are to: demonstrate a step-by-step process for developing a custom ChatGPT-based chatbot to support SR search strategy development, and evaluate its performance. Design A cross-sectional evaluation study. Methods We used ChatGPT-4.0 to create a chatbot designed to mimic a medical librarian, generating PICO-informed searches. Its knowledge base was augmented with two methodological references. After piloting testing, we refined its instructions. For evaluation, we randomly sampled 50 Cochrane SRs published in 2024. Standardized P–I–O prompts produced database-ready queries for PUBMED and EMBASE. The primary outcome was per-review success rate, summarized by median and inter-quartile range. A sensitivity analysis was conducted. Results Pilot testing achieved a retrieval rate of 41/49 (83.7%). In the main sample (1169 studies; median 13.5 studies per SR), the chatbot identified a median of 67.4% of included studies (IQR: 43.1%–88.4%). When limited to indexed studies (n = 1114), retrieval rose to 72.0% (IQR: 46.0%–92.5%). Lower performance was observed when outcomes were absent from the abstracts or interventions had many lexical variants. Conclusions A GAI-based chatbot can rapidly generate SR searches (~67%–72% identification), serving as a useful starting point but not a replacement for expert-led approaches. Integration of librarian expertise, structured prompts, and controlled vocabularies may improve performance. Further benchmarking and transparent reporting are needed to guide adoption. |
|---|
| Keywords: | database searching, generative artificial intelligence, large language model, systematic review |
|---|
| Publication status: | Published |
|---|
| Publication version: | Version of Record |
|---|
| Submitted for review: | 13.10.2025 |
|---|
| Article acceptance date: | 17.02.2026 |
|---|
| Publication date: | 10.03.2026 |
|---|
| Publisher: | Wiley |
|---|
| Year of publishing: | 2026 |
|---|
| Number of pages: | str. 1-7 |
|---|
| Numbering: | Letn. 58, št. 2, št. članka e70076 |
|---|
| PID: | 20.500.12556/DKUM-97514  |
|---|
| UDC: | 007.52+004.8:025.4.036:004.65 |
|---|
| ISSN on article: | 1547-5069 |
|---|
| COBISS.SI-ID: | 271552515  |
|---|
| DOI: | 10.1111/jnu.70076  |
|---|
| Publication date in DKUM: | 19.03.2026 |
|---|
| Views: | 225 |
|---|
| Downloads: | 2 |
|---|
| Metadata: |  |
|---|
| Categories: | Misc.
|
|---|
|
:
|
Copy citation |
|---|
| | | | Average score: | (0 votes) |
|---|
| Your score: | Voting is allowed only for logged in users. |
|---|
| Share: |  |
|---|
Hover the mouse pointer over a document title to show the abstract or click
on the title to get all document metadata. |