<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Evaluating Proprietary and Open-Weight Large Language Models as Universal Decimal Classification Recommender Systems</dc:title><dc:creator>Borovič,	Mladen	(Avtor)
	</dc:creator><dc:creator>Tomovski,	Eftimije	(Avtor)
	</dc:creator><dc:creator>Li Dobnik,	Tom	(Avtor)
	</dc:creator><dc:creator>Majninger,	Sandi	(Avtor)
	</dc:creator><dc:subject>universal decimal classification</dc:subject><dc:subject>large language models</dc:subject><dc:subject>conversational systems</dc:subject><dc:subject>recommender systems</dc:subject><dc:subject>prompt engineering</dc:subject><dc:subject>zero-shot classification</dc:subject><dc:subject>hierarchical similarity</dc:subject><dc:description>Manual assignment of Universal Decimal Classification (UDC) codes is time-consuming and inconsistent as digital library collections expand. This study evaluates 17 large language models (LLMs) as UDC classification recommender systems, including ChatGPT variants (GPT-3.5, GPT-4o, and o1-mini), Claude models (3-Haiku and 3.5-Haiku), Gemini series (1.0-Pro, 1.5-Flash, and 2.0-Flash), and Llama, Gemma, Mixtral, and DeepSeek architectures. Models were evaluated zero-shot on 900 English and Slovenian academic theses manually classified by professional librarians. Classification prompts utilized the RISEN framework, with evaluation using Levenshtein and Jaro–Winkler similarity, and a novel adjusted hierarchical similarity metric capturing UDC’s faceted structure. Proprietary systems consistently outperformed open-weight alternatives by 5–10% across metrics. GPT-4o achieved the highest hierarchical alignment, while open-weight models showed progressive improvements but remained behind commercial systems. Performance was comparable between languages, demonstrating robust multilingual capabilities. The results indicate that LLM-powered recommender systems can enhance library classification workflows. Future research incorporating fine-tuning and retrieval-augmented approaches may enable fully automated, high-precision UDC assignment systems.</dc:description><dc:publisher>MDPI AG</dc:publisher><dc:date>2025</dc:date><dc:date>2025-07-09 12:11:32</dc:date><dc:type>Članek v reviji</dc:type><dc:identifier>93644</dc:identifier><dc:identifier>UDK: 004.8</dc:identifier><dc:identifier>eISSN: 2076-3417</dc:identifier><dc:identifier>COBISS_ID: 243245571</dc:identifier><dc:identifier>DOI: 10.3390/app15147666</dc:identifier><dc:identifier>ISSN pri članku: 2076-3417</dc:identifier><dc:language>sl</dc:language><dc:rights>© 2025 by the authors</dc:rights></metadata>
