| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Uporaba nevronskih jezikovnih modelov za prepoznavanje imenskih entitet iz nestrukturiranih dokumentov : diplomsko delo
Authors:ID Knupleš, Urban (Author)
ID Holobar, Aleš (Mentor) More about this mentor... New window
ID Ferme, Marko (Comentor)
Files:.pdf UN_Knuples_Urban_2021.pdf (1,56 MB)
MD5: 9593ED9635CEAAFC9FED4A1724E47A91
PID: 20.500.12556/dkum/285cb4cc-b93c-40f3-a17c-bbbfb9e1f2b5
 
Language:Slovenian
Work type:Bachelor thesis/paper
Typology:2.11 - Undergraduate Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Nestrukturirani dokumenti zajemajo informacije v oblikah in postavitvah, ki se lahko od enega primerka do drugega razlikujejo, kar lahko oteži in podraži nalogo pridobivanja informacij. Kot rešitev se je v zadnjih letih za razumevanje dokumentov na področju dokumentne inteligence pričela uporaba nevronskih jezikovnih modelov, usposobljenih na učnih množicah dokumentov. V diplomskem delu za pridobivanje informacij iz skeniranih trgovinskih računov uporabljamo prehodno učeni nevronski jezikovni model, zgrajen iz transformatorjev. Model je natančno učen z uporabo učne množice SROIE za izluščitev štirih kategorij, tj. imen in naslovov trgovin, datumov in skupnih cen. Za pridobivanje informacij smo uporabili prepoznavo imenskih entitet. Za primerjavo izvajamo poskuse s spreminjanem hiperparametrov modela. S spremembo nevronskega jezikovnega modela smo pri poskusih dosegli največjo natančnost klasifikacije: 96,7 %.
Keywords:Dokumentna inteligenca, obdelava naravnih jezikov, prepoznava imenskih entitet, jezikovni modeli, transformatorji
Place of publishing:Maribor
Place of performance:Maribor
Publisher:[U. Knupleš]
Year of publishing:2021
Number of pages:IX, 35 str.
PID:20.500.12556/DKUM-80444 New window
UDC:004.652.8(043.2)
COBISS.SI-ID:95975171 New window
Publication date in DKUM:18.10.2021
Views:1155
Downloads:52
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-NC 4.0, Creative Commons Attribution-NonCommercial 4.0 International
Link:http://creativecommons.org/licenses/by-nc/4.0/
Description:A creative commons license that bans commercial use, but the users don’t have to license their derivative works on the same terms.
Licensing start date:13.09.2021

Secondary language

Language:English
Title:Named entity recognition on unstructured documents using neural language models
Abstract:Layouts and formats of information, in unstructured documents, can differ from one another and can make the extraction of information difficult and costly. Therefore, in recent years, the field of document intelligence began with the usage of neural language models trained on datasets of documents for document understanding. In the thesis, we adopt a pre-trained neural language model based on transformers, for information extraction out of scanned store invoices. The model is fine-tuned, using the SROIE dataset, based on four categories to extract store names and addresses, dates and total prices. For information extraction we used named entity recognition to classify tokens into the four prementioned categories. We conducted experiments using altered hyperparameters of the model for comparison. With the usage of the fine-tuned, altered neural language model, we achieved a maximum classification accuracy score of 96.7 %.
Keywords:Document intelligence, natural language processing, named entity recognition, langauge models, transformers


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica