| Title: | The Cert dataset decade : a systematic review of methodological evolution and performance bias |
|---|
| Authors: | ID Jurišić, Marko (Author) ID Tomičić, Igor (Author) |
| Files: | RAZ_Jurisic_Marko_2026.pdf (928,41 KB) MD5: 9B30B98F26AE5E189F52541180F55E79
https://www.fvv.um.si/rv/arhiv/2026/2026-02-Jurisic-Tomicic-E.html
|
|---|
| Language: | English |
|---|
| Work type: | Scientific work |
|---|
| Typology: | 1.02 - Review Article |
|---|
| Organization: | FVV - Faculty of Criminal Justice and Security
|
|---|
| Abstract: | Purpose: The purpose of this paper is to identify methodological biases and limitations in machine learning–based insider threat detection using the Computer Emergency Response Team [CERT] dataset, in order to guide the development of more realistic, robust, and operationally relevant detection approaches. Design/Methods/Approach: The objectives are achieved through a systematic literature analysis of 131 peer-reviewed studies published between 2013 and 2025 that apply machine learning to insider threat detection using the CERT dataset, employing a Preferred Reporting Items for Systematic Reviews and Meta-Analyses [PRISMA]-guided selection process and a structured comparative framework to examine dataset versions, feature engineering strategies, model architectures, and evaluation metrics from a methodological and empirical perspective. Findings: The analysis shows that most studies rely on the less realistic CERT v4.2 dataset, resulting in inflated performance that does not generalize to operational settings. It also finds that feature engineering is a stronger determinant of detection performance than model complexity, while inconsistent evaluation practices hinder meaningful comparison across studies. Research Limitations / Implications: The study is limited by its reliance on published research using a single synthetic dataset, which constrains generalization to real-world environments. Practical Implications: The findings indicate that practitioners should be cautious when adopting models validated on simplified benchmark settings, and instead prioritize solutions tested under extreme class imbalance. Emphasis should be placed on robust feature engineering, unsupervised or hybrid detection approaches, and evaluation metrics. Originality/Value: This paper provides the first large-scale, methodologically focused analysis of insider threat detection research that explicitly exposes performance inflation caused by dataset version bias and evaluation inconsistency, offering concrete, evidence-based guidance for improving the realism, comparability, and operational value of future studies in the field. |
|---|
| Keywords: | insider threat detection, CERT dataset, machine learning, anomaly detection, dataset bias, evaluation metrics |
|---|
| Publication status: | Published |
|---|
| Publication version: | Version of Record |
|---|
| Publication date: | 19.03.2026 |
|---|
| Year of publishing: | 2026 |
|---|
| Number of pages: | str. 1-24 |
|---|
| Numbering: | Vol. 28 |
|---|
| PID: | 20.500.12556/DKUM-97658  |
|---|
| UDC: | 004.056 |
|---|
| ISSN on article: | 2232-2981 |
|---|
| COBISS.SI-ID: | 273493251  |
|---|
| Publication date in DKUM: | 30.03.2026 |
|---|
| Views: | 310 |
|---|
| Downloads: | 4 |
|---|
| Metadata: |  |
|---|
| Categories: | Misc.
|
|---|
|
:
|
Copy citation |
|---|
| | | | Average score: | (0 votes) |
|---|
| Your score: | Voting is allowed only for logged in users. |
|---|
| Share: |  |
|---|
Hover the mouse pointer over a document title to show the abstract or click
on the title to get all document metadata. |