| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:STROJNO UČENJE NA VELIKIH PODATKIH Z UPORABO MONGODB, R IN HADOOP
Authors:ID Adanza Dopazo, Daniel (Author)
ID Podgorelec, Vili (Mentor) More about this mentor... New window
Files:.pdf MAG_Adanza_Dopazo_Daniel_2016.pdf (5,92 MB)
MD5: 33A2622B2BDC944541B2DF60AB466FAA
 
Language:Slovenian
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Osrednji namen tega magistrskega dela je testiranje različnih pristopov in izvedba več eksperimentov na različnih podatkovnih zbirkah, razporejenih na infrastrukturi za obdelavo velikih podatkov. Da bi dosegli ta cilj, smo magistrsko nalogo strukturirali v tri glavne dele. Najprej smo pridobili deset javno dostopnih podatkovnih zbirk z različnih področij, ki so dovolj kompleksne (glede na obseg podatkov in število atributov) za namen izvajanja analize velikih podatkov na ustrezen način. Zbrane podatke smo najprej predhodno obdelali, da bi bili združljivi s podatkovno bazo MongoDB. V drugem delu smo analizirali zbrane podatke in izvedli različne poskuse s pomočjo orodja R, ki omogoča izvedbo statistične obdelave podatkov. Orodje R smo pri tem povezali s podatkovno bazo MongoDB. V zadnjem delu smo uporabili še ogrodje Hadoop, s pomočjo katerega smo dokončali načrtovano infrastrukturo za obdelavo in analizo velikih podatkov. Za namen tega magistrskega dela smo vzpostavili sistem v načinu enega vozlišča v gruči. Analizirali smo razlike z vidika učinkovitosti vzpostavljene infrastrukture in delo zaključili z razpravo o prednostih in slabostih uporabe predstavljenih tehnologij za obdelavo velikih podatkov.
Keywords:veliki podatki, strojno učenje, analiza podatkov
Place of publishing:[Maribor
Publisher:D. Adanza Dopazo
Year of publishing:2016
PID:20.500.12556/DKUM-64710 New window
UDC:004.8:004.65(043.2)
COBISS.SI-ID:20345622 New window
NUK URN:URN:SI:UM:DK:1BP9VDDA
Publication date in DKUM:08.12.2016
Views:1625
Downloads:204
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Secondary language

Language:English
Title:MACHINE LEARNING ON BIG DATA USING MOGODB, R AND HADOOP
Abstract:The main purpose of this master thesis is to test different approaches and perform several experiments on different datasets, deployed on a big data infrastructure. In order to achieve that goal we will structure the thesis in three different parts. First of all, we will obtain ten publicly available datasets from different domains, which are complex enough (in terms of size and number of attributes) in order to perform the big data analysis in the proper way. Once they are gathered, we will pre-process them in order to be compatible with the MongoDB database. Second of all, we will analyse the data and perform various experiments using the R statistical and data analysis tool, which at the same time will be linked to the MongoDB database. Finally, we will use Hadoop for deploying this structure on big data. For the purpose of this master thesis, we will use it in a single node cluster mode. We will analyse the differences from the performance point of view and discuss the advantages and disadvantages of using the presented big data technologies.
Keywords:big data, machine learning, data analysis


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica