| Title: | Evaluating Proprietary and Open-Weight Large Language Models as Universal Decimal Classification Recommender Systems |
|---|
| Authors: | ID Borovič, Mladen, Faculty of Electrical Engineering and Computer Science, University of Maribor, 2000 Maribor, Slovenia (Author) ID Tomovski, Eftimije, Faculty of Electrical Engineering and Computer Science, University of Maribor, 2000 Maribor, Slovenia (Author) ID Li Dobnik, Tom, Faculty of Electrical Engineering and Computer Science, University of Maribor, 2000 Maribor, Slovenia (Author) ID Majninger, Sandi, Faculty of Electrical Engineering and Computer Science, University of Maribor, 2000 Maribor, Slovenia (Author) |
| Files: | applsci-15-07666-v2.pdf (447,50 KB) MD5: 72ACD70840709FDBA57806D0FE92EBAF
https://www.mdpi.com/2076-3417/15/14/7666/pdf
|
|---|
| Language: | English |
|---|
| Work type: | Article |
|---|
| Typology: | 1.01 - Original Scientific Article |
|---|
| Organization: | FERI - Faculty of Electrical Engineering and Computer Science
|
|---|
| Abstract: | Manual assignment of Universal Decimal Classification (UDC) codes is time-consuming and inconsistent as digital library collections expand. This study evaluates 17 large language models (LLMs) as UDC classification recommender systems, including ChatGPT variants (GPT-3.5, GPT-4o, and o1-mini), Claude models (3-Haiku and 3.5-Haiku), Gemini series (1.0-Pro, 1.5-Flash, and 2.0-Flash), and Llama, Gemma, Mixtral, and DeepSeek architectures. Models were evaluated zero-shot on 900 English and Slovenian academic theses manually classified by professional librarians. Classification prompts utilized the RISEN framework, with evaluation using Levenshtein and Jaro–Winkler similarity, and a novel adjusted hierarchical similarity metric capturing UDC’s faceted structure. Proprietary systems consistently outperformed open-weight alternatives by 5–10% across metrics. GPT-4o achieved the highest hierarchical alignment, while open-weight models showed progressive improvements but remained behind commercial systems. Performance was comparable between languages, demonstrating robust multilingual capabilities. The results indicate that LLM-powered recommender systems can enhance library classification workflows. Future research incorporating fine-tuning and retrieval-augmented approaches may enable fully automated, high-precision UDC assignment systems. |
|---|
| Keywords: | universal decimal classification, large language models, conversational systems, recommender systems, prompt engineering, zero-shot classification, hierarchical similarity |
|---|
| Publication status: | Published |
|---|
| Publication version: | Version of Record |
|---|
| Submitted for review: | 20.06.2025 |
|---|
| Article acceptance date: | 07.07.2025 |
|---|
| Publication date: | 08.07.2025 |
|---|
| Publisher: | MDPI AG |
|---|
| Year of publishing: | 2025 |
|---|
| Number of pages: | 7666-7689 |
|---|
| Numbering: | let. 15, št. 14 |
|---|
| PID: | 20.500.12556/DKUM-93644  |
|---|
| UDC: | 004.8 |
|---|
| ISSN on article: | 2076-3417 |
|---|
| eISSN: | 2076-3417 |
|---|
| COBISS.SI-ID: | 243245571  |
|---|
| DOI: | 10.3390/app15147666  |
|---|
| Copyright: | © 2025 by the authors |
|---|
| Publication date in DKUM: | 21.07.2025 |
|---|
| Views: | 219 |
|---|
| Downloads: | 18 |
|---|
| Metadata: |  |
|---|
| Categories: | Misc.
|
|---|
|
:
|
Copy citation |
|---|
| | | | Average score: | (0 votes) |
|---|
| Your score: | Voting is allowed only for logged in users. |
|---|
| Share: |  |
|---|
Hover the mouse pointer over a document title to show the abstract or click
on the title to get all document metadata. |