| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Metoda hierarhične večznačne klasifikacije na osnovi ekstrakcije značilnic s tekstovno analizo mikrobiotskih podatkov : doktorska disertacija
Authors:ID Brezočnik, Lucija (Author)
ID Podgorelec, Vili (Mentor) More about this mentor... New window
Files:.pdf DOK_Brezocnik_Lucija_2025.pdf (22,60 MB)
MD5: E053B64C51AD54343E42B6AAB8BE3645
 
Language:Slovenian
Work type:Doctoral dissertation
Typology:2.08 - Doctoral Dissertation
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Zanesljiva identifikacija kompleksnih vsebinskih struktur v primerih, kjer posamezni primerki podatkovnega nabora niso homogeni, pač pa združujejo informacije več virov, predstavlja enega izmed ključnih metodoloških izzivov sodobne podatkovne analitike. Relativno enostavna je namreč naloga, kjer je določen primerek homogen in ga z uporabo večrazredne klasifikacije znamo relativno preprosto razvrstiti v enega izmed ponujenih razredov. Kompleksnost pa se drastično poveča, ko se v istem primerku skriva več virov. V tem primeru osnovne metode analize ne zadostujejo več in potrebujemo naprednejše pristope, ki so sposobni razbrati soobstoj več razredov oziroma oznak, kar je tudi domena večznačne klasifikacije. V predloženi doktorski disertaciji obravnavamo omenjeni problem na področju metagenomike, ki med drugim omogoča raziskovanje mikrobiote, raznolike skupnosti bakterij in drugih mikroorganizmov v določenem okolju. Z naprednimi tehnikami sekvenciranja iz njih pridobimo zaporedja DNK celotne mikrobne združbe, ki jih lahko opišemo kot izjemno dolga besedila, zapisana z abecedo štirih nukleotidov: A, T, G in C. Naš cilj je v teh besedilih poiskati t. i. označevalne gene, ki so izključno ali močno povezani z gostiteljem. V ta namen smo na podlagi optimizacijskih pristopov in domenskih pravil predlagali metodo ekstrakcije značilnic, temelječo na osnovi k-merov, tj. krajših delov DNK. Pristop na osnovi k-merov se je izkazal za zelo učinkovitega, zato smo ga uporabili tudi pri sintetičnem generiranju vzorcev mikrobnih oziroma mikrobiotskih podatkov. Metoda temelji na pripravi profilov k-merov in na nanje osnovanih grafih prehodov. Ker smo v doktorski disertaciji analizirali lokacijsko specifične vzorce, smo morali njihov manjši nabor čistih vzorcev ustrezno razširiti. Še več, sintetično smo razširili tudi nabor mešanih vzorcev, kar predstavlja še večji izziv v realnih okoljih. Obe predlagani metodi sta se združili v konceptualno najzahtevnejšem delu doktorske naloge, predlagani metodi hierarhične večznačne klasifikacije na osnovi ekstrakcije značilnic, imenovani MLB. Z njo smo na osnovi vhodnih podatkov, tj. čistih ali sintetično ustvarjenih vzorcev, napovedovali deleže gostiteljev v mešanih mikrobnih vzorcih. Rezultate metode MLB smo primerjali s tistimi, pridobljenimi z orodjem SourceTracker, vodilnim orodjem za natančno identifikacijo in kvantifikacijo gostiteljev mikrobov v mešanih vzorcih. Metodi smo ovrednotili z uveljavljenimi metrikami na področju večznačne klasifikacije, ki razkrivajo, da metoda MLB učinkovito rešuje problem določitve gostiteljev in njihovih deležev ter poda primerljive, večinoma pa boljše rezultate kot orodje SourceTracker.
Keywords:strojno učenje, večznačna klasifikacija, ekstrakcija značilnic, obdelava naravnega jezika, mikrobiotski podatki
Place of publishing:Maribor
Place of performance:Maribor
Publisher:[L. Brezočnik]
Year of publishing:2025
Number of pages:XV, 157 str.
PID:20.500.12556/DKUM-94160 New window
UDC:004.5:004.8(043.3)
COBISS.SI-ID:253974787 New window
Publication date in DKUM:20.10.2025
Views:163
Downloads:111
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-NC-ND 4.0, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
Link:http://creativecommons.org/licenses/by-nc-nd/4.0/
Description:The most restrictive Creative Commons license. This only allows people to download and share the work for no commercial gain and for no other purposes.
Licensing start date:06.08.2025

Secondary language

Language:English
Title:Hierarchical multilabel classification method with feature extraction based on textual analysis of microbial data
Abstract:Reliable identification of complex content structures in cases where individual dataset instances are not homogeneous, but consist of information from multiple sources, represents one of the key methodological challenges in modern data analytics. The task is relatively straightforward when an instance is homogeneous, and can be easily classified into one of the available classes using multiclass classification. However, the complexity increases drastically when multiple sources are present within the same instance. In such cases, basic analytical methods are no longer sufficient, and advanced approaches capable of determining the coexistence of multiple classes or labels are required, which is the domain of multilabel classification. In the presented doctoral dissertation, we address this challenge within the field of metagenomics, which enables the exploration of the microbiome, a diverse community of bacteria and other microorganisms in a given environment. Using advanced sequencing techniques, we obtain DNA sequences of the entire microbial community, which can be described as extremely long texts written with the alphabet of four nucleotides: A, T, G, and C. Our goal is to identify so-called marker genes in these texts that are exclusively or strongly associated with a specific host. Hence, we proposed a feature extraction method based on optimization approaches and domain rules, relying on k-mers, i.e., shorter segments of DNA. The k-mer-based approach proved to be very effective, and we therefore applied it to the synthetic generation of microbial data samples. The method relies on creating k-mer profiles and corresponding transition graphs. Since we analyzed location-specific samples, we needed to expand their smaller set of pure samples accordingly. Furthermore, we synthetically expanded the set of mixed samples, which poses an even greater challenge in real-world settings. Both proposed methods were combined in the conceptually most demanding part of the dissertation: a proposed hierarchical multilabel classification method with feature extraction, named MLB. This method predicts the proportions of hosts in mixed microbial samples based on input data, i.e., pure or synthetically generated samples. The results of the MLB method were compared with those obtained by the SourceTracker, a leading tool for predicting and quantifying microbial hosts in mixed samples. Both methods were evaluated using established metrics, revealing that the MLB method effectively addresses the problem of identifying hosts and their proportions, delivering comparable, and in most cases, superior results to SourceTracker.
Keywords:Machine Learning, Multilabel Classification, Feature Extraction, Natural Language Processing, Microbial data


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica