Teaching a GAN About Fairness

A Project on Mitigating Racial Bias in Face Aging Models

Authors: Cédric Caruzzo, Juchul Shin, Minshik Choi [GitHub Repository]
← Back to Projects & Writing Hub

The Problem

Facial age progression models predict how someone may look as they age. Many are trained on datasets that do not represent the population well. FFHQ, for example, is heavily skewed toward white faces.

Chart showing racial bias in the FFHQ dataset
The racial distribution in the popular FFHQ dataset. Truncation, a common technique in GANs, further amplifies this bias, increasing the ratio of generated white faces.

Models trained on this data may change the perceived race of non-white subjects as they age. We wanted to reduce that failure.

Baseline: Style-Based Age Manipulation

We started with Style-based Age Manipulation (SAM), an image-to-image model built on StyleGAN. SAM encodes a face into a latent representation, changes its age, and tries to preserve the person's identity.

SAM model architecture diagram
The baseline SAM architecture, which maps an input image and a target age to a set of style vectors to generate the aged face.

Adding a Race-Preservation Loss

We added a penalty when the generated face's predicted race differed from the input. A pretrained DeepFace classifier supplied that signal during training.

DeepFace compared the original and generated faces after each SAM pass. A mismatch contributed a race-preservation loss alongside the existing identity and age objectives.

Our modified SAM architecture with DeepFace for race loss
Our proposed architecture. We add a DeepFace classifier that compares the input and output images, calculating a "race loss" to guide the model towards fairer results.

Results

We did not yet have complete quantitative metrics, so these results are qualitative. The video compares the baseline and modified models directly.

First half: baseline. Second half: our model. Our model better preserves the subject's perceived race, while the baseline often shifts toward white features.

Side-by-Side Image Comparisons

We also compared still images from our model, trained from scratch with a race-loss weight of 15, against the original SAM.

Young Age

Comparison of our model vs baseline for a young face
Comparison at a younger age. Our model (left) vs. the baseline model (right).

Mature Age

Comparison of our model vs baseline for a mature face
Comparison at a mature age. Our model retains racial characteristics more faithfully.

Elderly Age

Comparison of our model vs baseline for an elderly face
Comparison at an elderly age. The baseline model shows a significant shift in features.

What We Learned

Adding a classifier signal directly to the loss reduced this bias without requiring a perfectly balanced training set. Compute limits kept the evaluation preliminary, but two points stood out:

Try It Yourself: Notebooks

These Colab notebooks run the model without local setup.

← Back to Projects & Writing Hub