Skip to yearly menu bar Skip to main content


Poster

Improving Subgroup Robustness via Data Selection

Saachi Jain · Kimia Hamidieh · Kristian Georgiev · Andrew Ilyas · Marzyeh Ghassemi · Aleksander Madry

[ ]
Fri 13 Dec 11 a.m. PST — 2 p.m. PST

Abstract:

Machine learning models can often fail on subgroups that are underrepresentedduring training. While dataset balancing can improve performance onunderperforming groups, it requires access to training group annotations and canend up removing large portions of the dataset. In this paper, we introduceData Debiasing with Datamodels (D3M), a debiasing approachwhich isolates and removes specific training examples that drive the model'sfailures on minority groups. Our approach enables us to efficiently traindebiased classifiers while removing only a small number of examples, and doesnot require training group annotations or additional hyperparameter tuning.

Live content is unavailable. Log in and register to view live content