| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Fairness evaluation paradox : how biased test data masks true group fairness assessment
Authors:ID Karakatič, Sašo (Author)
ID Colakovic, Ivona (Author)
ID Heričko, Tjaša (Author)
Files:URL https://www.mdpi.com/2227-7390/14/16/2894
 
.pdf mathematics-14-02894.pdf (2,37 MB)
MD5: 2F4F92A129B0D74F8FDF3F31D8636251
 
Language:English
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:The EU AI Act makes fairness metrics for high-risk AI systems’ compliance evidence, turning their trustworthiness into a safety and accountability concern. Fairness audits assume that test data reflects the properties of real-world conditions, whereas standard evaluation protocols use test data drawn from the same biased records as the training data. Studies measuring how strongly this bias in test data distorts fairness metrics between the validation phase and real-world deployment are still very rare. We conduct an experiment on synthetic and real data, measuring this discrepancy across five fairness interventions on four unfairness types (1000 repetitions per combination, 20,000 total runs). Synthetic data lets us encode human bias in labels and compare fairness metrics on biased test labels (data available during development) against clean labels (conditions models face in deployment). We find that a systematic evaluation bias is present across all metrics, so the same models on the same test data can support opposite fairness conclusions and mask the mistreatment of the most disadvantaged groups. This pattern of fairness misevaluation is confirmed by a real-world validation on the Adult Census Income dataset. We conclude that trustworthy fairness auditing and regulatory standards should require bias-aware evaluation protocols, in which observed labels are not treated as ground truth.
Keywords:algorithmic fairness, evaluation bias, fairness metrics, fairness interventions, machine learning, EU AI Act, trustworthy AI, model auditing
Publication status:Published
Publication version:Version of Record
Submitted for review:06.07.2026
Article acceptance date:27.07.2026
Publication date:11.08.2026
Publisher:MDPI
Year of publishing:2026
Number of pages:28 str.
Numbering:Vol. 14, iss. 16, [article no.] 2894
PID:20.500.12556/DKUM-99735 New window
UDC:004.8
ISSN on article:2227-7390
COBISS.SI-ID:288823043 New window
DOI:10.3390/math14162894 New window
Copyright:© 2026 by the authors
Publication date in DKUM:25.08.2026
Views:161
Downloads:6
Metadata:XML DC-XML DC-RDF
Categories:Misc.
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Mathematics
Shortened title:Mathematics
Publisher:MDPI AG
ISSN:2227-7390
COBISS.SI-ID:523267865 New window

Document is financed by a project

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P2-0057-2018
Name:Informacijski sistemi

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:J5-50176-2023
Name:Razumevanje hrepenenja po hrani in vnosa hrane: Od skupinskih povprečij k personaliziranemu pristopu

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:algoritemska pravičnost, metrike pravičnosti, pristranskost pri vrednotenju, strojno učenje


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica