| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Utilizing large language models for cyber attack data generation : magistrsko delo
Authors:ID Zgaga, Bine (Author)
ID Beranič, Tina (Mentor) More about this mentor... New window
ID Kaymaz, Sevtap Duman (Comentor)
Files:.pdf MAG_Zgaga_Bine_2026.pdf (3,40 MB)
MD5: 22AB2BECFA652F6BBDBC0C6F41A89C7D
 
Language:English
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:This master’s thesis addresses the challenge of limited, imbalanced, and restricted-access datasets in the field of cybersecurity, which significantly hinder the development and evaluation of intrusion detection systems. To overcome these limitations, the thesis explores the use of synthetic data generation as a viable and effective alternative for producing realistic and diverse cyber-attack data while preserving privacy and improving dataset usability. The work presents a systematic review of existing approaches to synthetic data generation based on Generative Adversarial Networks and Large Language Models, analysing their fundamental characteristics, advantages, and limitations in the context of cyber-attack generation. Particular emphasis is placed on their suitability for creating realistic and diverse attack patterns relevant to intrusion detection research. The experimental part of the thesis investigates and compares multiple approaches to synthetic data generation, including standalone generative models and a combined methodology. The generated data are evaluated with respect to syntactic correctness, diversity, and practical applicability, supported by the introduction of a novel metric for assessing attack diversity. Additionally, the generated attacks are validated in a controlled test environment to assess their effectiveness in realistic scenarios. The results demonstrate that while individual generative approaches exhibit distinct strengths, their integration enables the most balanced and effective generation of synthetic cyber-attack data. The thesis concludes that combining Generative Adversarial Networks and Large Language Models provides a robust framework for producing realistic, diverse, and practically useful datasets, thereby contributing to the development of more reliable cybersecurity systems.
Keywords:large language models, attack data generation, artificial intelligence, cybersecurity
Place of publishing:Maribor
Place of performance:Maribor
Publisher:[B. Zgaga]
Year of publishing:2026
Number of pages:1 spletni vir (1 datoteka PDF ([XXI], 67 str.))
PID:20.500.12556/DKUM-97077 New window
UDC:004.6:004.8(043.2)
COBISS.SI-ID:281650179 New window
Publication date in DKUM:03.06.2026
Views:315
Downloads:13
Metadata:XML DC-XML DC-RDF
Categories:KTFMB - FERI
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.
Licensing start date:16.02.2026

Secondary language

Language:Slovenian
Title:Uporaba velikih jezikovnih modelov za generiranje testnih podatkov
Abstract:Magistrsko delo obravnava problem omejene dostopnosti kakovostnih, uravnoteženih in reprezentativnih podatkovnih zbirk na področju kibernetske varnosti, ki pomembno vpliva na razvoj in vrednotenje sistemov za zaznavanje vdorov. Zaradi zaupnosti, redkosti in neuravnoteženosti realnih podatkov se kot učinkovita alternativa uveljavlja uporaba sintetičnih podatkov, ki omogočajo varnejše raziskave, večjo raznolikost primerov ter boljše pokrivanje redkih, a za varnost kritičnih vrst napadov. V magistrskem delu je predstavljen pregled obstoječih metod generiranja sintetičnih podatkov z uporabo generativnih adversarialnih mrež in velikih jezikovnih modelov. Analizirane so njihove temeljne značilnosti, prednosti in omejitve pri ustvarjanju realističnih vzorcev kibernetskih napadov, s poudarkom na njihovi uporabnosti v sistemih za zaznavanje vdorov. Eksperimentalni del vključuje primerjavo različnih pristopov k generiranju sintetičnih podatkov, pri čemer so bili ti ovrednoteni z vidika sintaktične pravilnosti, raznolikosti in praktične uporabnosti. Razvita je bila tudi nova metrika za ocenjevanje raznolikosti napadov, ki omogoča objektivnejšo primerjavo med posameznimi metodami. Dodatno je bila izvedena praktična validacija v testnem okolju, ki je omogočila preverjanje učinkovitosti generiranih napadov v realističnih pogojih. Rezultati raziskave kažejo, da posamezni pristopi izkazujejo različne prednosti, pri čemer kombinacija generativnih adversarialnih mrež in velikih jezikovnih modelov omogoča najbolj uravnoteženo generiranje sintetičnih podatkov. Tak pristop omogoča ustvarjanje realističnih, raznolikih in uporabnih podatkovnih zbirk ter predstavlja pomemben prispevek k razvoju zanesljivejših sistemov za kibernetsko varnost.
Keywords:veliki jezikovni modeli, generiranje podatkov za napade, umetna inteligenca, kibernetska varnost


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica