| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Sequence-to-Sequence models and their evaluation for spoken language normalization of Slovenian
Authors:ID Sepesy Maučec, Mirjam (Author)
ID Verdonik, Darinka (Author)
ID Donaj, Gregor (Author)
Files:.pdf applsci-14-09515.pdf (437,99 KB)
MD5: 19D67562D60D137292969407CD398506
 
URL https://www.mdpi.com/2076-3417/14/20/9515
 
Language:English
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Sequence-to-sequence models have been applied to many challenging problems, including those in text and speech technologies. Normalization is one of them. It refers to transforming non-standard language forms into their standard counterparts. Non-standard language forms come from different written and spoken sources. This paper deals with one such source, namely speech from the less-resourced highly inflected Slovenian language. The paper explores speech corpora recently collected in public and private environments. We analyze the efficiencies of three sequence-to-sequence models for automatic normalization from literal transcriptions to standard forms. Experiments were performed using words, subwords, and characters as basic units for normalization. In the article, we demonstrate that the superiority of the approach is linked to the choice of the basic modeling unit. Statistical models prefer words, while neural network-based models prefer characters. The experimental results show that the best results are obtained with neural architectures based on characters. Long short-term memory and transformer architectures gave comparable results. We also present a novel analysis tool, which we use for in-depth error analysis of results obtained by character-based models. This analysis showed that systems with similar overall results can differ in the performance for different types of errors. Errors obtained with the transformer architecture are easier to correct in the post-editing process. This is an important insight, as creating speech corpora is a time-consuming and costly process. The analysis tool also incorporates two statistical significance tests: approximate randomization and bootstrap resampling. Both statistical tests confirm the improved results of neural network-based models compared to statistical ones.
Keywords:low-resource language, applications, spoken language, normalization, character unit, subword unit, statistical model, long short-term memory, transformer, error analysis
Publication status:Published
Publication version:Version of Record
Submitted for review:06.09.2024
Article acceptance date:16.10.2024
Publication date:18.10.2024
Publisher:MDPI
Year of publishing:2024
Number of pages:24 str.
Numbering:let. 14, št. 20, št. članka 9515
PID:20.500.12556/DKUM-91741 New window
UDC:004.8
ISSN on article:2076-3417
COBISS.SI-ID:213048067 New window
DOI:10.3390/app14209515 New window
Copyright:© 2024 by the authors
Publication date in DKUM:31.01.2025
Views:125
Downloads:17
Metadata:XML DC-XML DC-RDF
Categories:Misc.
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Applied sciences
Shortened title:Appl. sci.
Publisher:MDPI
ISSN:2076-3417
COBISS.SI-ID:522979353 New window

Document is financed by a project

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:J7-4642-2022
Name:Temeljne raziskave za razvoj govornih virov in tehnologij za slovenščino

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:aplikacije, govorjeni jeziki, statistični modeli


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica