| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Analiza uporabe in postavitve podatkovnega jezera : magistrsko delo
Authors:ID Koren, Marcel (Author)
ID Kamišalić Latifić, Aida (Mentor) More about this mentor... New window
ID Šestak, Martina (Comentor)
Files:.pdf MAG_Koren_Marcel_2021.pdf (2,31 MB)
MD5: 32BDF1E1D8F573F36E54D8BEDC28D86E
PID: 20.500.12556/dkum/f0afb304-7d28-4491-82b1-173fed7f0e33
 
Language:Slovenian
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Velepodatki in podatkovna jezera sta pojma, ki jih v zadnjih letih vedno pogosteje uporabljamo v povezavi s porastom količine ustvarjenih podatkov. V magistrskem delu predstavljamo lastnosti podatkovnih jezer, čemu so namenjena, kako jih lahko vzpostavimo ter kako so povezana z velepodatki. Podrobno opišemo odprtokodno rešitev Apache Hadoop in oblačno rešitev Microsoft Azure Data Lake. Pri tem smo spoznali tudi orodja, ki jih rešitvi ponujata, med katerimi sta pomembnejši Apache Spark in Azure Databricks. V nadaljevanju predstavljamo, kako ju vzpostavimo ter izvedemo eksperiment, kjer na podlagi hitrosti izvajanja in stroškov spoznamo njune prednosti in slabosti.
Keywords:velepodatki, podatkovna jezera, Hadoop, Spark, Azure Data Lake
Place of publishing:Maribor
Place of performance:Maribor
Publisher:[M. Koren]
Year of publishing:2021
Number of pages:1 spletni vir (1 datoteka PDF (VIII, 61 f.))
PID:20.500.12556/DKUM-80869 New window
UDC:004.65+004.76(043.2)
COBISS.SI-ID:97445635 New window
Publication date in DKUM:16.12.2021
Views:1387
Downloads:145
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-SA 4.0, Creative Commons Attribution-ShareAlike 4.0 International
Link:http://creativecommons.org/licenses/by-sa/4.0/
Description:This Creative Commons license is very similar to the regular Attribution license, but requires the release of all derivative works under this same license.
Licensing start date:03.11.2021

Secondary language

Language:English
Title:An analysis of the usage and deployment approaches for data lakes
Abstract:Big data and data lakes are terms that have grown in popularity in the last couple of years, because of the increase in the amount of data produced. In this master thesis, we present the properties of data lakes, what they are used for, how they can be set up and how they are linked to big data. We describe the open-source solution Apache Hadoop and the cloud solution Microsoft Azure Data Lake. We learn about the tools they offer, the most important of which are Apache Spark and Azure Databricks. Later we present how to set them up and run an experiment to see their advantages and disadvantages based on execution speed and cost.
Keywords:big data, data lakes, Hadoop, Spark, Azure Data Lake


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica