| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:PORAZDELJENA POMENSKA ANALIZA DOKUMENTOV V PROGRAMSKEM OGRODJU APACHE HADOOP
Authors:ID Starina, David (Author)
ID Ojsteršek, Milan (Mentor) More about this mentor... New window
Files:.pdf UN_Starina_David_2016.pdf (1,33 MB)
MD5: 741BBE3BB24F58B70B46E807F6D0C71F
 
Language:Slovenian
Work type:Undergraduate thesis
Typology:2.11 - Undergraduate Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:V diplomskem delu obravnavamo porazdeljeno pomensko analizo dokumentov v programskem ogrodju Apache Hadoop. Opišemo sestavo in delovanje Hadoopa, predvsem porazdeljenega datotečnega sistema HDFS in pogajalca za vire YARN. Predstavimo različne metode za pomensko analizo besedil, osredotočimo se na linearno Dirichletovo razporeditev (LDA) in podamo različne metrike za ugotavljanje podobnosti med vektorji. Predstavimo implementacijo rešitve za iskanje podobnih dokumentov s pomočjo programske knjižnice Apache Mahout in razpravljamo o primerih z LDA-jem generiranih tem. Predstavimo rezultate meritev na porazdeljeni in ne-porazdeljeni različici in predstavimo nekaj predlogov za hitrejšo analizo.
Keywords:pomenska analiza, porazdeljena obdelava, Hadoop, linearna Dirichletova razporeditev, procesiranje naravnega jezika
Place of publishing:[Maribor
Publisher:D. Starina
Year of publishing:2016
PID:20.500.12556/DKUM-62419 New window
UDC:004.6:004.728.8(043.2)
COBISS.SI-ID:20190742 New window
NUK URN:URN:SI:UM:DK:Y1GNB3MD
Publication date in DKUM:08.09.2016
Views:1668
Downloads:182
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Secondary language

Language:English
Title:DISTRIBUTED SEMANTIC ANALYSIS OF DOCUMENTS USING APACHE HADOOP
Abstract:In this thesis we deal with distributed semantic analysis of documents in the Apache Hadoop programming framework. We describe the composition and operation of Hadoop, mainly of the distributed file system, HDFS and resource negotiator, YARN. We present different methods of semantic text analysis, focusing on the linear Dirichlet allocation (LDA), and describe different metrics to determine vector similarity. We present implementation of the solution for searching similar documents using Apache Mahout software library and discuss examples of LDA-generated topics. We presented the measurement results on distributed and non-distributed version and present some suggestions for faster analysis.
Keywords:semantic analysis, distributed processing, Hadoop, linear Dirichlet allocation, natural language processing


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica