<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://dk.um.si/IzpisGradiva.php?id=97077"><dc:title>Utilizing large language models for cyber attack data generation</dc:title><dc:creator>Zgaga,	Bine	(Avtor)
	</dc:creator><dc:creator>Beranič,	Tina	(Mentor)
	</dc:creator><dc:creator>Kaymaz,	Sevtap Duman	(Komentor)
	</dc:creator><dc:subject>large language models</dc:subject><dc:subject>attack data generation</dc:subject><dc:subject>artificial intelligence</dc:subject><dc:subject>cybersecurity</dc:subject><dc:description>This master’s thesis addresses the challenge of limited, imbalanced, and restricted-access datasets in the field of cybersecurity, which significantly hinder the development and evaluation of intrusion detection systems. To overcome these limitations, the thesis explores the use of synthetic data generation as a viable and effective alternative for producing realistic and diverse cyber-attack data while preserving privacy and improving dataset usability.
The work presents a systematic review of existing approaches to synthetic data generation based on Generative Adversarial Networks and Large Language Models, analysing their fundamental characteristics, advantages, and limitations in the context of cyber-attack generation. Particular emphasis is placed on their suitability for creating realistic and diverse attack patterns relevant to intrusion detection research.
The experimental part of the thesis investigates and compares multiple approaches to synthetic data generation, including standalone generative models and a combined methodology. The generated data are evaluated with respect to syntactic correctness, diversity, and practical applicability, supported by the introduction of a novel metric for assessing attack diversity. Additionally, the generated attacks are validated in a controlled test environment to assess their effectiveness in realistic scenarios.
The results demonstrate that while individual generative approaches exhibit distinct strengths, their integration enables the most balanced and effective generation of synthetic cyber-attack data. The thesis concludes that combining Generative Adversarial Networks and Large Language Models provides a robust framework for producing realistic, diverse, and practically useful datasets, thereby contributing to the development of more reliable cybersecurity systems.</dc:description><dc:publisher>[B. Zgaga]</dc:publisher><dc:date>2026</dc:date><dc:date>2026-02-16 20:24:24</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>97077</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
