| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Influence of highly inflected word forms and acoustic background on the robustness of automatic speech recognition for human–computer interaction
Authors:ID Žgank, Andrej (Author)
Files:.pdf mathematics-10-00711.pdf (1,12 MB)
MD5: 822271573C9F339CAF19ED6CD073B307
 
URL https://www.mdpi.com/2227-7390/10/5/711
 
Language:English
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Automatic speech recognition is essential for establishing natural communication with a human–computer interface. Speech recognition accuracy strongly depends on the complexity of language. Highly inflected word forms are a type of unit present in some languages. The acoustic background presents an additional important degradation factor influencing speech recognition accuracy. While the acoustic background has been studied extensively, the highly inflected word forms and their combined influence still present a major research challenge. Thus, a novel type of analysis is proposed, where a dedicated speech database comprised solely of highly inflected word forms is constructed and used for tests. Dedicated test sets with various acoustic backgrounds were generated and evaluated with the Slovenian UMB BN speech recognition system. The baseline word accuracy of 93.88% and 98.53% was reduced to as low as 23.58% and 15.14% for the various acoustic backgrounds. The analysis shows that the word accuracy degradation depends on and changes with the acoustic background type and level. The highly inflected word forms’ test sets without background decreased word accuracy from 93.3% to only 63.3% in the worst case. The impact of highly inflected word forms on speech recognition accuracy was reduced with the increased levels of acoustic background and was, in these cases, similar to the non-highly inflected test sets. The results indicate that alternative methods in constructing speech databases, particularly for low-resourced Slovenian language, could be beneficial.
Keywords:human–computer interaction, automatic speech recognition, acoustic modeling, highly inflected word forms, acoustic background
Publication status:Published
Publication version:Version of Record
Submitted for review:30.12.2021
Article acceptance date:22.02.2022
Publication date:24.02.2022
Publisher:MDPI AG
Year of publishing:2022
Number of pages:16 str.
Numbering:Vol. 10, no. 5
PID:20.500.12556/DKUM-92311 New window
UDC:621.39
ISSN on article:2227-7390
COBISS.SI-ID:98760451 New window
DOI:10.3390/math10050711 New window
Copyright:© 2022 by the author
Publication date in DKUM:28.03.2025
Views:220
Downloads:12
Metadata:XML DC-XML DC-RDF
Categories:Misc.
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Mathematics
Shortened title:Mathematics
Publisher:MDPI AG
ISSN:2227-7390
COBISS.SI-ID:523267865 New window

Document is financed by a project

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P2-0069-2018
Name:Napredne metode interakcij v telekomunikacijah

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:interakcija človek - stroj, avtomatsko razpoznavanje govora, akustično ozadje


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica