| Title: | A hybrid NLLB and large language model pipeline for diachronic intralingual translation of 16th-century slovene literary heritage |
|---|
| Authors: | ID Šprajc, Žan Tomaž (Author) ID Sekirnik, Rok (Author) ID Kučiš, Vlasta (Author) ID Jesenko, David (Author) |
| Files: | https://www.mdpi.com/2076-3417/16/14/7317
applsci-16-07317-v2_(2).pdf (1,60 MB) MD5: 50FD6676AB65DF2F49D397BB567C349D
|
|---|
| Language: | English |
|---|
| Work type: | Article |
|---|
| Typology: | 1.01 - Original Scientific Article |
|---|
| Organization: | FERI - Faculty of Electrical Engineering and Computer Science
|
|---|
| Abstract: | Modernising historical literature into contemporary language is a form of diachronic intralingual translation that supports access to written cultural heritage. For low-resource languages such as Slovene, this task is hindered by orthographic, lexical and syntactic shifts, as well as the scarcity of parallel data. We present a two-stage pipeline that combines a fine-tuned No Language Left Behind (NLLB) model with Claude Opus 4.8 post-editing for the modernisation of 16th-century Slovene literature, retaining archaic words and phrases while normalising the alphabet, orthography and grammar. Using Jurij Dalmatin’s 1584 Bible and its 2017 modernised edition, we constructed an aligned parallel corpus of 14,876 sentence and clause-level pairs through Bohorič-to-Gaj normalisation and LaBSE-based embedding alignment. The hybrid pipeline achieved the best overall scores, reaching BLEU 45.78, CHRF 67.68 and METEOR 71.72, with TER 43.25 and CER 34.00. Its gains over the standalone LLM baseline were large and statistically significant across all the metrics, while its improvement over the fine-tuned NLLB model was smaller and significant mainly for overlap-based measures. We applied the pipeline further to Tulščak’s Kerszhanske leipe molitve (1579), producing the first preliminary modernisation of the earliest known Slovene prayer book assessed qualitatively and tested the generalisation on additional 16th-century texts, including an out-of-domain legal text. The results demonstrate that combining task-specific neural machine translation with controlled LLM post-editing offers a practical strategy for modernising low-resource historical texts and contributes a reusable methodology for digital cultural heritage preservation. |
|---|
| Keywords: | cultural heritage, digital humanities, diachronic intralingual translation, historical text modernisation, low-resource language, neural machine translation, large language model, NLLB |
|---|
| Publication status: | Published |
|---|
| Publication version: | Version of Record |
|---|
| Submitted for review: | 17.06.2026 |
|---|
| Article acceptance date: | 16.07.2026 |
|---|
| Publication date: | 01.01.2026 |
|---|
| Publisher: | MDPI |
|---|
| Year of publishing: | 2026 |
|---|
| Number of pages: | 25 str. |
|---|
| Numbering: | Vol. 16, no. 14, [article no.] 7317 |
|---|
| PID: | 20.500.12556/DKUM-98985  |
|---|
| UDC: | 004.45 |
|---|
| ISSN on article: | 2076-3417 |
|---|
| COBISS.SI-ID: | 285726211  |
|---|
| DOI: | 10.3390/app16147317  |
|---|
| Copyright: | © 2026 by the authors
|
|---|
| Publication date in DKUM: | 23.07.2026 |
|---|
| Views: | 350 |
|---|
| Downloads: | 8 |
|---|
| Metadata: |  |
|---|
| Categories: | Misc.
|
|---|
|
:
|
Copy citation |
|---|
| | | | Average score: | (0 votes) |
|---|
| Your score: | Voting is allowed only for logged in users. |
|---|
| Share: |  |
|---|
Hover the mouse pointer over a document title to show the abstract or click
on the title to get all document metadata. |