<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://dk.um.si/IzpisGradiva.php?id=97620"><dc:title>Exploration of augmentation strategies in multi-modal retrieval-augmented generation for the biomedical domain</dc:title><dc:creator>Kocbek,	Primož	(Avtor)
	</dc:creator><dc:creator>Frkatović-Hodžić,	Azra	(Avtor)
	</dc:creator><dc:creator>Lalić,	Dora	(Avtor)
	</dc:creator><dc:creator>Hui,	Vivian	(Avtor)
	</dc:creator><dc:creator>Lauc,	Gordan	(Avtor)
	</dc:creator><dc:creator>Štiglic,	Gregor	(Avtor)
	</dc:creator><dc:creator>,	IEEE	(Lastnik avtorskih pravic)
	</dc:creator><dc:subject>multimodal retrieval-augmentation generation</dc:subject><dc:subject>evaluation</dc:subject><dc:subject>large language models</dc:subject><dc:description>Multi-modal retrieval-augmented generation (MMRAG) promises grounded biomedical QA, but it is unclear when to (i) convert figures/tables into text versus (ii) use optical character recognition (OCR)-free visual retrieval that returns page images and leaves interpretation to the generator. We study this trade-off in glycobiology, a visually dense domain. We built a benchmark of 120 multiple-choice questions (MCQs) from 25 papers, stratified by retrieval difficulty (easy text, medium figures/tables, hard cross-evidence). We implemented four augmentations—None, Text RAG, Multi-modal conversion, and late-interaction visual retrieval (ColPali)—using Docling parsing and Qdrant indexing. We evaluated mid-size opensource and frontier proprietary models (e.g., Gemma-3-27BIT, GPT-4o family). Additional testing used the GPT-5 family and multiple visual retrievers (ColPali/ColQwen/ColFlor). Accuracy with Agresti–Coull 95% confidence intervals (CIs) was computed over 5 runs per configuration. With Gemma-3-27BIT, Text and Multi-modal augmentation outperformed OCR-free retrieval (0.722-0.740 vs. 0.510 average accuracy). With GPT-4o, Multi-modal achieved 0.808, with Text 0.782 and ColPali 0.745 close behind; within-model differences were small. In follow-on experiments with the GPT-5 family, the best results with ColPali and ColFlor improved by 2% to 0.828 in both cases. In general across the GPT-5 family, ColPali, ColQwen, and ColFlor were statistically indistinguishable; ColFlor matched ColPali while being far smaller. GPT-5-nano trailed larger GPT-5 variants by roughly 8-10%. Pipeline choice is capacity-dependent: converting visuals to text lowers the reader burden and is more reliable for mid-size models, whereas OCR-free visual retrieval becomes competitive under frontier models. Among retrievers, ColFlor offers parity with heavier options at a smaller footprint, making it an efficient default when strong generators are available.</dc:description><dc:date>2026</dc:date><dc:date>2026-03-25 14:21:27</dc:date><dc:type>Znanstveno delo</dc:type><dc:identifier>97620</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
