| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:NADZOROVANO ODKRIVANJE PREDMETA TEKSTOVNIH VSEBIN Z UPORABO SELEKCIJSKIH IN STATISTIČNIH METOD
Authors:ID Hrnčić, Sašo (Author)
ID Kosar, Tomaž (Mentor) More about this mentor... New window
ID Podgorelec, Vili (Comentor)
Files:.pdf VS_Hrncic_Saso_2016.pdf (2,31 MB)
MD5: 087986506A000E5CC992A3E222C1C702
 
Language:Slovenian
Work type:Undergraduate thesis
Typology:2.11 - Undergraduate Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Cilj diplomske naloge je izdelati preprost kategorizacijski sistem, ki zna nov tekstovni dokument čim natančneje uvrstiti v naprej definirane kategorije. Ena izmed funkcionalnosti sistema je prepoznavanje jezika, ki je bilo testirano na podatkovnih korpusih dokumentov Wikipedije, Europarla in jezikovnih modelov projekta LibTextCat. Kategorizacijski sistem je bil razširjen še na prepoznavanje v naprej definiranih tematikah korpusa 20 Newsgroups in Reuters-21578. Za predstavitev dokumentov smo uporabili n-gramsko tehniko, ki smo jo kombinirali s selekcijskimi in statističnimi metodami. Dosežene rezultate smo analizirali ter dokumentirali. Podrobneje smo predstavili problematiko, lastne izkušnje, lastnosti uporabljenih metod ter obstoječe raziskave.
Keywords:tekstovno kategoriziranje, n-grami, strojno učenje, teorija informacij, odmik od najpomembnejšega elementa
Place of publishing:[Maribor
Publisher:S. Hrnčić
Year of publishing:2016
PID:20.500.12556/DKUM-61990 New window
UDC:004.05:004.5(043.2)
COBISS.SI-ID:19991318 New window
NUK URN:URN:SI:UM:DK:GT5AAOMU
Publication date in DKUM:16.09.2016
Views:1219
Downloads:111
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Secondary language

Language:English
Title:SUPERVISED TOPICS' DETECTION BASED ON FEATURE SELECTION AND STATISTICAL METHODS
Abstract:The main goal of diploma work is to develop simple text classification system that is able to automatically classify a document into predefined categories as accurately as possible. One of the functionalities of the system is language detection that has been tested on documents of Wikipedia, Europarl and language models of project LibTextCat. Classification system has been expanded to identify predefine topics of the corpus 20 Newsgroups and Reuters-21578. For document presentation we used n-grams technique, which was combined with feature selection methods and statistical methods. The obtained results were analyzed and documented. We also present text classification problem, our experiences, features of used methods and some existing research.
Keywords:text classification, n-grams, machine learning, information theory, out of place


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica