| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Razvoj sistema za pretvorbo besedil v govor z globokimi nevronskimi mrežami : magistrsko delo
Authors:ID Bratina, Matevž (Author)
ID Rojc, Matej (Mentor) More about this mentor... New window
Files:.pdf MAG_Bratina_Matevz_2021.pdf (3,01 MB)
MD5: F8735A0E177C83E13AF7AB07D532B402
PID: 20.500.12556/dkum/02827375-9289-4938-88cc-61a50c5b4a89
 
Language:Slovenian
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:V magistrski nalogi smo razvili sistem pretvorbe besedila v govor PLATTOS za več jezikov. Sistem bazira na osnovi globokih nevronskih mrež. Osnovni cilj naloge je bil razviti in testirati sistem sinteze govora na osnovi globokega učenja, ki bo čim bolje generiral govor v več jezikih, pri čemer je tudi pomemben čas generiranja. Prvi del naloge tako predstavlja pregled tehnologij sistemov sinteze govora in njihova podrobnejša analiza. Zanimala nas je namreč arhitektura sistema sinteze govora, medsebojna primerjava zmogljivosti sistemov, njihov razvoj in kvaliteta sintetiziranega signala, ki ga določen TTS lahko generira. Sledila je izbira tehnologije globokega učenja, in razvoj novega TTS sistema. Izbrali smo tisto, ki je izkazovala največji potencial, da izpolni vse zastavljene cilje. Sledil je razvoj TTS sistema. Za prvo stopnjo (pretvorba vhodnega besedila v spektrogram) smo izbrali Tacotron globoki model. Ta je namenjen pretvorbi spektrogramov v pripadajoči govorni signal. V drugi stopnji, smo izbrali vokoder Waveglow. Pred izbiro komponent sistema, smo različne tipe vokoderjev in rekonstrukcijskih algoritmov tudi testirali. Sistem TTS na osnovi globokih nevronskih mrež PLATTOS smo testirali na različnih prosto dostopnih bazah govornih podatkov večih jezikov. Ocenjevali in primerjali smo tudi kvaliteto sinteze govora različnih arhitektur z globokimi nevronskimi mrežami. Kot kriterij kvalitete sinteze govora, smo bili predvsem pozorni na naravnost in razumljivost sintetiziranega govora. Pri ocenjevanju kvalitete smo tako uporabili subjektivne MUSHRA teste. Pokazalo se je, da kombinacija globokih nevronskih modelov Tacotron in Waveglow zagotovi najboljše rezultate v večih jezikih, kar se tiče kvalitete sintetiziranega govora in hitrosti generiranja odziva.
Keywords:globoko učenje, nevronska mreža, sinteza govora, umetna inteligenca, Pytorch, Tensorflow, Tacotron, Waveglow, Wavenet, WaveRNN
Place of publishing:Maribor
Place of performance:Maribor
Publisher:[M. Bratina]
Year of publishing:2021
Number of pages:XII, 99 str.
PID:20.500.12556/DKUM-79655 New window
UDC:004.7(043.2)
COBISS.SI-ID:86391299 New window
Publication date in DKUM:18.10.2021
Views:1050
Downloads:128
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-NC-ND 4.0, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
Link:http://creativecommons.org/licenses/by-nc-nd/4.0/
Description:The most restrictive Creative Commons license. This only allows people to download and share the work for no commercial gain and for no other purposes.
Licensing start date:09.08.2021

Secondary language

Language:English
Title:Development of a text-to-speech system with deep neural networks
Abstract:In this master 's thesis, we developed a text-to-speech synthesis system PLATTOS for several languages based on deep neural networks. The main goal of the task was to develop and test a text-to-speech synthesis system based on deep learning techniques, which will best synthesise speech in several languages, where also speech generation time is important. The first part of the paper presents an overview of technologies regarding speech synthesis systems, and their detailed analysis. Namely, we were interested in the architecture of the text-to-speech synthesis system, the mutual comparison of several systems’ capabilities, their development and the speech quality that a certain TTS system is capable to generate. This was followed by the selection of deep learning technology, and the development of a novel TTS system PLATTOS, in which we recognized the potential to meet all our goals. This was followed by the development of the TTS system, where the Tacotron deep model for the first stage (conversion of input text into a spectrogram), and the Waveglow vocoder for the conversion of spectrograms into the corresponding speech signal in the second stage were finally selected, after testing several vocoders and reconstruction algorithms. The TTS system PLATTOS based on deep neural networks was tested on several freely accessible speech databases in several languages.We also evaluated and compared the quality of synthesized speech of several deep neural architectures. As a criterion for the quality of synthesized speech, we paid particular attention to the naturalness and intelligibility of the synthesized speech. Therefore, subjective MUSHRA tests were used to assess this quality. The combination of the Tacotron and Waveglow neural models has been shown to provide the best results in several languages in terms of the quality of synthesized speech and the speed of response generation.
Keywords:deep learning, neural network, speech synthesis, artificial intelligence, Pytorch, Tensorflow, Tacotron, Waveglow, Wavenet, WaveRNN


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica