| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:A hybrid NLLB and large language model pipeline for diachronic intralingual translation of 16th-century slovene literary heritage
Authors:ID Šprajc, Žan Tomaž (Author)
ID Sekirnik, Rok (Author)
ID Kučiš, Vlasta (Author)
ID Jesenko, David (Author)
Files:URL https://www.mdpi.com/2076-3417/16/14/7317
 
.pdf applsci-16-07317-v2_(2).pdf (1,60 MB)
MD5: 50FD6676AB65DF2F49D397BB567C349D
 
Language:English
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Modernising historical literature into contemporary language is a form of diachronic intralingual translation that supports access to written cultural heritage. For low-resource languages such as Slovene, this task is hindered by orthographic, lexical and syntactic shifts, as well as the scarcity of parallel data. We present a two-stage pipeline that combines a fine-tuned No Language Left Behind (NLLB) model with Claude Opus 4.8 post-editing for the modernisation of 16th-century Slovene literature, retaining archaic words and phrases while normalising the alphabet, orthography and grammar. Using Jurij Dalmatin’s 1584 Bible and its 2017 modernised edition, we constructed an aligned parallel corpus of 14,876 sentence and clause-level pairs through Bohorič-to-Gaj normalisation and LaBSE-based embedding alignment. The hybrid pipeline achieved the best overall scores, reaching BLEU 45.78, CHRF 67.68 and METEOR 71.72, with TER 43.25 and CER 34.00. Its gains over the standalone LLM baseline were large and statistically significant across all the metrics, while its improvement over the fine-tuned NLLB model was smaller and significant mainly for overlap-based measures. We applied the pipeline further to Tulščak’s Kerszhanske leipe molitve (1579), producing the first preliminary modernisation of the earliest known Slovene prayer book assessed qualitatively and tested the generalisation on additional 16th-century texts, including an out-of-domain legal text. The results demonstrate that combining task-specific neural machine translation with controlled LLM post-editing offers a practical strategy for modernising low-resource historical texts and contributes a reusable methodology for digital cultural heritage preservation.
Keywords:cultural heritage, digital humanities, diachronic intralingual translation, historical text modernisation, low-resource language, neural machine translation, large language model, NLLB
Publication status:Published
Publication version:Version of Record
Submitted for review:17.06.2026
Article acceptance date:16.07.2026
Publication date:01.01.2026
Publisher:MDPI
Year of publishing:2026
Number of pages:25 str.
Numbering:Vol. 16, no. 14, [article no.] 7317
PID:20.500.12556/DKUM-98985 New window
UDC:004.45
ISSN on article:2076-3417
COBISS.SI-ID:285726211 New window
DOI:10.3390/app16147317 New window
Copyright:© 2026 by the authors
Publication date in DKUM:23.07.2026
Views:350
Downloads:8
Metadata:XML DC-XML DC-RDF
Categories:Misc.
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Applied sciences
Shortened title:Appl. sci.
Publisher:MDPI
ISSN:2076-3417
COBISS.SI-ID:522979353 New window

Document is financed by a project

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:J7-70247-2026
Name:SPHERE - Vrednotenje okoljskih pojavov z informiranim globokim učenjem na podlagi podatkov opazovanja Zemlje

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P2-0041-2020
Name:Računalniški sistemi, metodologije in inteligentne storitve

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:kulturna dediščina, digitalna humanistika, diahroni intralingualni prevod, modernizacija zgodovinskih besedil, jezik z omejenimi viri, nevronsko strojno prevajanje, veliki jezikovni modeli


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica