| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Razvoj spletne aplikacije in napovednega modela za napoved nediagnosticirane sladkorne bolezni tipa 2
Authors:ID Fajfar, Andrej (Author)
ID Štiglic, Gregor (Mentor) More about this mentor... New window
Files:.pdf MAG_Fajfar_Andrej_2017.pdf (1,38 MB)
MD5: 75C92B03E527E3A89BC74D64FDE2C94E
PID: 20.500.12556/dkum/75f01c6d-58ca-41f0-a4e8-3922b54d7485
 
Language:Slovenian
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FZV - Faculty of Health Sciences
Abstract:V magistrski nalogi smo s pomočjo metode strojnega učenja »Random Forest« skušali napovedati stopnjo tveganja za nastanek sladkorne bolezni oz. verjetnost prisotnosti nediagnosticirane sladkorne bolezni na podlagi podatkov iz Slovenije. Za izbrano metodo smo določili optimalno število in vrsto spremenljivk za posamezni model. Za evalvacijo modela smo uporabili povprečno območje pod krivuljo (AUC), točnost in F-mero. Za model populacije s povečanim tveganjem smo dosegli povprečno AUC 0,823, točnost 0,824 in F-mero 0,804. V modelu za napoved nediagnostirane sladkorne bolezni smo dosegli povprečno AUC in točnost 0,749 in F-mero 0,654. Na podlagi podatkov smo pokazali, da je možno z veliko uspešnostjo določiti osebe z visokim tveganjem, ki predstavljajo preddiabetike in nediagnostirane diabetike oz. skupino s tveganjem za nastanek sladkorne bolezni tipa 2. Pokazali smo uporabo tehnik uravnoteženja odločitvenega razreda in rezultate primerjali z neuravnoteženim razredom. Uravnoteženje razreda zviša klasifikacijsko uspešnost modela. Rezultate smo primerjali z rezultati drugih znanstvenih objav in zasledili podobnost med rezultati. Tuje raziskave navajajo, da je klasifikator Random Forest najpogosteje izbran model, v primerjavi z drugimi modeli za napovedovanje kroničnih bolezni. S korelacijskim testom smo pokazali, da napovedna uspešnost modela ne korelira s številom dreves v ansamblu (p = 0,00015).
Keywords:strojno učenje, Random Forest, neuravnoteženi podatki, diabetes mellitus
Place of publishing:Maribor
Publisher:[A. Fajfar]
Year of publishing:2017
PID:20.500.12556/DKUM-68620 New window
UDC:616.379-008.64:004
COBISS.SI-ID:2361764 New window
NUK URN:URN:SI:UM:DK:4EO1TWZI
Publication date in DKUM:19.10.2017
Views:1486
Downloads:177
Metadata:XML DC-XML DC-RDF
Categories:FZV
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY-NC-ND 4.0, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
Link:http://creativecommons.org/licenses/by-nc-nd/4.0/
Description:The most restrictive Creative Commons license. This only allows people to download and share the work for no commercial gain and for no other purposes.
Licensing start date:30.09.2017

Secondary language

Language:English
Title:Development of web aplication and predictive model for predicting undiagnosed type 2 diabetes
Abstract:The master’s thesis develops predictive models based on machine learning method Random Forest for prediction of diabetes type 2 risk population group and undiagnosed diabetes group using data collected in Slovenia. The optimal set of features for optimal results was determined. We used mean area under the curve (AUC), accuracy and F-measure for evalvation of the model. We achieved mean AUC of 0.823, accuracy 0.824 and F- measure 0.804 for risk in the preddiabetes group and mean AUC and accuracy of 0.749 and F-measure 0.654 for undiagnosed diabetes model. We achieved the best predictive performance for high-risk population representing preddiabetes and undiagnosed diabetes patients. We show techniques for balancing class variable and compare results with unbalance data. Balancing improved overall accuracy of classification model. Comparison to related scientific papers shows similarities between results and techniques and that Random Forest is a preferred model of choice in the field of prediction of chronic diseases with high accuracy rate. Correlation test shows no correlation between mean AUC and number of trees in ensemble (p = 0,00015).
Keywords:machine learning, Random Forest, unbalance data, diabetes mellitus


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica