| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Zajemanje in obdelava besedila iz slik z uporabo tehnologij optičnega prepoznavanja znakov in transformerskih modelov za tolmačenje dokumentov : diplomsko delo
Authors:ID Kajba, Tjan (Author)
ID Novak, Damijan (Mentor) More about this mentor... New window
Files:.pdf VS_Kajba_Tjan_2025.pdf (3,39 MB)
MD5: A4FDF8693C7C6DD599A6647E2F5AEE82
 
Language:Slovenian
Work type:Bachelor thesis/paper
Typology:2.11 - Undergraduate Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Informacije imajo dandanes ključno vlogo v digitalnem svetu. Tehnologije zajemanja besedila s slik so ključna rešitev za pretvorbo fizičnih medijev v digitalno obliko. Teoretični del diplomskega dela predstavlja, kaj so omenjene tehnologije in kako jih razvrščamo. V nadaljevanju smo prikazali delovanje posameznih podvrst orodij, tako tistih, ki temeljijo na velikih jezikovnih modelih, kot tudi tistih brez njih ter njihove priporočene uporabe. Proti koncu teoretičnega dela smo izvedli primerjalno analizo različnih tehnologij. Na podlagi rezultata analize je bil izbran transformerski model za razumevanje dokumentov kot osnova aplikacije. Praktični del prikazuje ustvarjeno aplikacijo, namenjeno branju računov.
Keywords:multimodalni veliki jezikovni modeli, optično prepoznavanje znakov, transformerski modeli za razumevanje dokumentov, veliki jezikovni modeli, zajemanje besedila s slik
Place of publishing:Maribor
Place of performance:Maribor
Publisher:[T. Kajba]
Year of publishing:2025
Number of pages:1 spletni vir (1 datoteka PDF (VI, 66 str.))
PID:20.500.12556/DKUM-94618 New window
UDC:004.932(043.2)
COBISS.SI-ID:271053315 New window
Publication date in DKUM:23.09.2025
Views:215
Downloads:46
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-NC-SA 4.0, Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International
Link:http://creativecommons.org/licenses/by-nc-sa/4.0/
Description:A Creative Commons license that bans commercial use and requires the user to release any modified works under this license.
Licensing start date:22.08.2025

Secondary language

Language:English
Title:Text extraction and processing from images using optical character recognition technologies and document interpretation transformer models
Abstract:Information plays a crucial role in today's digital world. Technologies for text extraction from images are a crucial solution for the digitalization of data. The theoretical segment of my thesis introduces the mentioned technologies and their categorizations. The following chapters describe the inner workings of each subcategory whether it uses large language models or not and their recommended use cases. The theoretical segment is concluded with a comparative analysis of the different technologies. Based on the analysis, document understanding transformer models were selected as a foundation for our app. The practical segment displays the app for reading receipts.
Keywords:multimodal large language models, optical character recognition, document understanding transformer models, large language models, text extraction from images


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica