| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Relevantnost informacijskega priklica pri strojnem učenju za binarno besedilno klasifikacijo
Authors:ID Marijan, Robert (Author)
ID Leskovar, Robert (Mentor) More about this mentor... New window
Files:.pdf DOK_Marijan_Robert_2019.pdf (6,50 MB)
MD5: 7C3EDD46F0FDE6B4874A6557AB3BF5C7
PID: 20.500.12556/dkum/d9f431bf-3135-4e12-9735-77f418119f83
 
Language:Slovenian
Work type:Doctoral dissertation
Organization:FOV - Faculty of Organizational Sciences in Kranj
Abstract:Raziskava obravnava relevantnost v procesih informacijskega priklica oziroma iskanja in prenosa informacij pri algoritmih strojnega učenja za binarno klasifikacijo besedilnih primerkov. Predstavljena je metodologija dela z opisom problema in okolja, metod, orodij in podatkov raziskovanja, ter hipotezi. Izvedli smo pojmovno razgradnjo gradnikov disertacije: agent, informacijska znanost, iskanje in prenos informacij, klasifikacija, komunikacija, podatkovno rudarjenje in prepoznavanje vzorcev, povezave, reševanje problemov, hevristika in intuicija, strojno učenje, svet, umetna inteligenca, umetne nevronske mreže, znanje. Pri ključnem pojmu raziskave, relevantnosti, smo opredelili njegov izvor, izhodiščne definicije, opredelitve pojma kot povezave, ter predstavili tri modele: Mizzarov model štirih dimenzij relevantnosti, Dervinino sense-making teorijo, ter Greisdorfov konjunktivno-disjunktivni model. Proučili smo stopnje relevantnosti in zadostnost dokazov, tipe relevantnosti, pogoje za relevantnost, predpostavke in uporabnikove izbore konteksta ter sodbe relevantnosti. Izvedli smo dve skupini eksperimentov, v katerih smo (1) merili učinkovitost iskanja in prenosa informacij, in (2) analizirali dejavnike, ki vplivajo na učinkovitost iskanja in prenosa informacij: vpliv naključja, vpliv števila atributov, vpliv števila primerkov in tipa vektorizacije, vpliv izbora algoritmov strojnega učenja, vpliv izbora atributov z razredno napovednimi lastnostmi, vpliv lematizacije, vpliv razredne neenakosti, vpliv učinka pretiranega prilagajanja modela podatkom. V disertaciji smo testirali hipotezi zamenjave (kognitivnih) sodb relevantnosti človeških ekspertov z računalniško ustvarjenimi. Z uporabo odprtokodnih in brezplačno dostopnih programskih orodij smo s postopki podatkovnega rudarjenja in algoritmi strojnega učenja merili učinkovitost iskanja in prenosa informacij agenta z zaznano informacijsko potrebo. Korenski pojem relevantnosti smo analizirali s sistemsko in uporabniško usmerjenim pristopom. Izvedli smo dve skupini eksperimentov, ter dokazali, da lahko ob danih predpostavkah odprtokodna programska orodja (tako delno kot v celoti) nadomestijo človeškega eksperta kot ocenjevalca v postopkih binarne klasifikacije relevantnosti. Založniki lahko, v primeru uporabe dokumentacijskih oziroma knjižničnih informacijskih sistemov, ki tehnično in pravno omogočajo modularno nadgrajevanje, zamenjajo programske module iskanja in prenosa informacij z odprtokodnimi orodji strojnega učenja, ki so bili predmet disertacije, in s tem začnejo postopek uporabe le-teh kot nadomestilo ali dopolnilo človeškim sodbam v postopkih binarne klasifikacije besedilnih primerkov.
Keywords:Relevantnost, iskanje in prenos informacij, strojno učenje
Place of publishing:Maribor
Year of publishing:2019
PID:20.500.12556/DKUM-72984 New window
COBISS.SI-ID:302092032 New window
NUK URN:URN:SI:UM:DK:0N1SK0LB
Publication date in DKUM:20.01.2020
Views:1737
Downloads:133
Metadata:XML DC-XML DC-RDF
Categories:FOV
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-NC-ND 4.0, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
Link:http://creativecommons.org/licenses/by-nc-nd/4.0/
Description:The most restrictive Creative Commons license. This only allows people to download and share the work for no commercial gain and for no other purposes.
Licensing start date:13.01.2019

Secondary language

Language:English
Title:Relevance of information retrieval in machine learning binary text classification
Abstract:In research we studied the notion of relevance in the field of information retrieval using the machine learning algorithms for binary text classification. We described the methodology, used in research, which included definition of problem and environment, methods, tools and research data, and hypothesis. We analyzed building blocks of dissertation, which included the notions of agent, information science, information retrieval, classification, communication, data mining and pattern recognition, relations, problem solving, heuristics and intuition, machine learning, world, artificial intelligence, artificial neural networks and knowledge. For the key research notion, relevance, we defined concept’s origin and basic definitions. We analyzed different relevance models, including Mizzaro’s four dimensions of relevance, Dervin’s sense-making theory and Greisdorf’s conjunctive/disjunctive model. We studied the levels, types and conditions of relevance, premises and user’s context selection and relevance judgments. We conducted two sets of experiments, where we (1) measured the performance of information retrieval and (2) analyzed the factors that influence the performance of information retrieval: the number of attributes, text records, types of vectorization, selection of machine learning algorithms, selection of attributes with predictive properties, the impact of selected lemmatization and others. In dissertation we tested the hypothesis of replacing cognitive relevance judgments, created by human experts, with relevance judgments, created solely by computer algorithms. We were measuring the performance of information retrieval in order to satisfy agent’s perceived information need using open source and freely available data mining procedures and machine learning algorithms. We analyzed the root notion of relevance using systems-centered and user-centered approach. We conducted the number of experiments and showed that in comply with certain premises open source software can (fully and in part) substitute the human expert as a gold standard creator. Publishers can, if their integrated library systems technically and legally allow, replace parts of ILS with open source software modules for information retrieval, discussed in this dissertation, as s step towards replacing human experts with machine learning algorithms performing tasks of binary text classification.
Keywords:Relevance, information retrieval, machine learning


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica