| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Porazdeljeno generiranje poročil detektorja plagiatov
Authors:ID Fartek, Jože (Author)
ID Ojsteršek, Milan (Mentor) More about this mentor... New window
Files:.pdf VS_Fartek_Joze_2018.pdf (1,23 MB)
MD5: 359D64350CD4BDA7544D7062073AA4F4
PID: 20.500.12556/dkum/c0f47491-51d9-45c3-973f-68bf958e0bd8
 
Language:Slovenian
Work type:Bachelor thesis/paper
Typology:2.11 - Undergraduate Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Ker je relacijske podatkovne baze za hranjenje velike količine izvlečkov iz besedil in generiranje poročil detektorja podobnih vsebin težko horizontalno razširiti, smo za ta namen raziskali možnost uporabe podatkovnih baz NoSQL. Preizkusili smo več podatkovnih baz in izbrali najprimernejšo. Implementirali smo tudi nekaj algoritmov, ki so primerni za ugotavljanje podobnosti v parafraziranih besedilih in temeljijo na tvorjenju izvlečkov iz besedil s pomočjo normaliziranih n-gramov. Te algoritme smo primerjali z algoritmom za tvorjenje izvlečkov, ki se na Univerzi v Mariboru uporablja za detekcijo podobnih dokumentov. Po izbiri najustreznejše podatkovne baze NoSQL in algoritma za tvorjenje izvlečkov, smo implementirali prototip porazdeljenega sistema za ugotavljanje podobnih dokumentov in generiranje poročil detektorja podobnih vsebin.
Keywords:porazdeljeno procesiranje, koncept »MapReduce«, NoSQL, detekcija podobnih vsebin
Place of publishing:[Maribor
Publisher:J. Fartek
Year of publishing:2018
PID:20.500.12556/DKUM-71768 New window
UDC:004.4'415:7.061(043.2)
COBISS.SI-ID:21849878 New window
NUK URN:URN:SI:UM:DK:NZGMAOES
Publication date in DKUM:19.10.2018
Views:1247
Downloads:130
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-NC-ND 4.0, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
Link:http://creativecommons.org/licenses/by-nc-nd/4.0/
Description:The most restrictive Creative Commons license. This only allows people to download and share the work for no commercial gain and for no other purposes.
Licensing start date:29.08.2018

Secondary language

Language:English
Title:Distributed generation of plagiarism detection reports
Abstract:Due to the fact that relational databases for storing large quantities of calculated hashes from documents and generations of plagiarism detection reports of similar content have difficulties extending horizontally, we have explored the possibility of using NoSQL databases for this purpose. We have tested several NoSQL databases and selected the most appropriate one. Furthermore, we have implemented several algorithms that are suitable for searching similarities in paraphrased documents, that are based on generating hashes from documents using normalized n-grams. These algorithms were compared with a hash generation algorithm, used at the University of Maribor to detect similar documents. After selecting the most suitable NoSQL database and hash generation algorithm, we implemented a prototype of distributed computer system for identifying similar documents and generating the detector reports of similar content.
Keywords:distributed processing, MapReduce, NoSQL, text matching


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica