| Title: | Fairness evaluation paradox : how biased test data masks true group fairness assessment |
|---|
| Authors: | ID Karakatič, Sašo (Author) ID Colakovic, Ivona (Author) ID Heričko, Tjaša (Author) |
| Files: | https://www.mdpi.com/2227-7390/14/16/2894
mathematics-14-02894.pdf (2,37 MB) MD5: 2F4F92A129B0D74F8FDF3F31D8636251
|
|---|
| Language: | English |
|---|
| Work type: | Article |
|---|
| Typology: | 1.01 - Original Scientific Article |
|---|
| Organization: | FERI - Faculty of Electrical Engineering and Computer Science
|
|---|
| Abstract: | The EU AI Act makes fairness metrics for high-risk AI systems’ compliance evidence, turning their trustworthiness into a safety and accountability concern. Fairness audits assume that test data reflects the properties of real-world conditions, whereas standard evaluation protocols use test data drawn from the same biased records as the training data. Studies measuring how strongly this bias in test data distorts fairness metrics between the validation phase and real-world deployment are still very rare. We conduct an experiment on synthetic and real data, measuring this discrepancy across five fairness interventions on four unfairness types (1000 repetitions per combination, 20,000 total runs). Synthetic data lets us encode human bias in labels and compare fairness metrics on biased test labels (data available during development) against clean labels (conditions models face in deployment). We find that a systematic evaluation bias is present across all metrics, so the same models on the same test data can support opposite fairness conclusions and mask the mistreatment of the most disadvantaged groups. This pattern of fairness misevaluation is confirmed by a real-world validation on the Adult Census Income dataset. We conclude that trustworthy fairness auditing and regulatory standards should require bias-aware evaluation protocols, in which observed labels are not treated as ground truth. |
|---|
| Keywords: | algorithmic fairness, evaluation bias, fairness metrics, fairness interventions, machine learning, EU AI Act, trustworthy AI, model auditing |
|---|
| Publication status: | Published |
|---|
| Publication version: | Version of Record |
|---|
| Submitted for review: | 06.07.2026 |
|---|
| Article acceptance date: | 27.07.2026 |
|---|
| Publication date: | 11.08.2026 |
|---|
| Publisher: | MDPI |
|---|
| Year of publishing: | 2026 |
|---|
| Number of pages: | 28 str. |
|---|
| Numbering: | Vol. 14, iss. 16, [article no.] 2894 |
|---|
| PID: | 20.500.12556/DKUM-99735  |
|---|
| UDC: | 004.8 |
|---|
| ISSN on article: | 2227-7390 |
|---|
| COBISS.SI-ID: | 288823043  |
|---|
| DOI: | 10.3390/math14162894  |
|---|
| Copyright: | © 2026 by the authors
|
|---|
| Publication date in DKUM: | 25.08.2026 |
|---|
| Views: | 161 |
|---|
| Downloads: | 6 |
|---|
| Metadata: |  |
|---|
| Categories: | Misc.
|
|---|
|
:
|
Copy citation |
|---|
| | | | Average score: | (0 votes) |
|---|
| Your score: | Voting is allowed only for logged in users. |
|---|
| Share: |  |
|---|
Hover the mouse pointer over a document title to show the abstract or click
on the title to get all document metadata. |