Fed-Focal Loss: Federated Learning with Imbalanced Data

Dipankar Sarkar
Dipankar Sarkar · · 2 min read

Federated learning keeps training data with the organisations or devices that produced it. That helps with privacy and data ownership, but it creates a difficult statistical problem: each client may have a different and highly imbalanced local dataset.

Fed-Focal Loss changes the local training objective so that examples the model already classifies confidently contribute less to the loss, while difficult examples contribute more. In an imbalanced classification problem, those difficult examples often include the minority class that ordinary averaging learns poorly.

Why ordinary averaging struggles

Federated Averaging combines model updates from multiple clients. When common examples dominate local training, the combined model can achieve apparently strong accuracy while performing poorly on the rare class.

Centralised rebalancing is often unavailable because the server cannot inspect or redistribute the clients’ raw data. The adjustment therefore has to work during local training.

What Fed-Focal Loss changes

Focal loss scales the contribution of each example according to how confidently it is classified. Easy examples receive less weight; hard examples receive more. Fed-Focal Loss applies that mechanism inside each federated client’s local optimisation step before the server aggregates the updates.

The method does not require the server to reconstruct the global class distribution or collect raw examples from participating clients.

Experimental result

The paper evaluated Fed-Focal Loss across MNIST, FEMNIST, Vehicle Sensor Network, and Human Activity Recognition datasets. On the unbalanced MNIST benchmark, it improved performance by more than nine absolute percentage points in the reported comparison.

That is a result from the paper’s experimental setup. A production deployment would still need its own client sampling, drift analysis, privacy controls, minority-class metrics, and evaluation against the cost of false positives and false negatives.

Publication

  • Title: Fed-Focal Loss for Imbalanced Data Classification in Federated Learning
  • Authors: Dipankar Sarkar, Ankur Narang, and Sumit Rai
  • Status: Accepted at the FL-IJCAI 2020 workshop
  • Primary record: arXiv:2011.06283

The companion preprint CatFedAvg studies communication efficiency and classification accuracy in federated aggregation.