Federated Learning (FL) enables collaborative training while preserving privacy, yet it introduces a critical challenge: the “illusion of fairness”. A global model, usually evaluated on the server, appears fair on average while keeping persistent unfairness at the client level. Current fairness-enhancing FL solutions often fall short, as they typically mitigate biases for a single, usually binary, sensitive attribute, while ignoring two realistic and conflicting scenarios: Attribute-Bias (where clients are unfair toward different sensitive attributes) and Value-Bias (where clients exhibit conflicting biases toward different values of the same attribute). To support more robust and reproducible fairness research in FL, we introduce FeDa4Fair, the first benchmarking framework designed to stress-test FL fairness methods under these heterogeneous conditions. Our contributions are three-fold: (1) We introduce FeDa4Fair, a library designed to create datasets tailored to evaluating fair FL methods under heterogeneous client bias; (2) we release a benchmark suite generated by the FeDa4Fair library to standardize the evaluation of fair FL methods; (3) we provide ready-to-use functions for evaluating fairness outcomes for these datasets.

FeDa4Fair: Client-Level Federated Datasets for Fairness Evaluation

Luca Corbucci;
2026-01-01

Abstract

Federated Learning (FL) enables collaborative training while preserving privacy, yet it introduces a critical challenge: the “illusion of fairness”. A global model, usually evaluated on the server, appears fair on average while keeping persistent unfairness at the client level. Current fairness-enhancing FL solutions often fall short, as they typically mitigate biases for a single, usually binary, sensitive attribute, while ignoring two realistic and conflicting scenarios: Attribute-Bias (where clients are unfair toward different sensitive attributes) and Value-Bias (where clients exhibit conflicting biases toward different values of the same attribute). To support more robust and reproducible fairness research in FL, we introduce FeDa4Fair, the first benchmarking framework designed to stress-test FL fairness methods under these heterogeneous conditions. Our contributions are three-fold: (1) We introduce FeDa4Fair, a library designed to create datasets tailored to evaluating fair FL methods under heterogeneous client bias; (2) we release a benchmark suite generated by the FeDa4Fair library to standardize the evaluation of fair FL methods; (3) we provide ready-to-use functions for evaluating fairness outcomes for these datasets.
File in questo prodotto:
File Dimensione Formato  
3805689.3812291.pdf

accesso aperto

Licenza: Creative commons
Dimensione 4.01 MB
Formato Adobe PDF
4.01 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11582/373548
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
social impact