Semester

Summer

Date of Graduation

2026

Document Type

Problem/Project Report

Degree Type

MS

College

Statler College of Engineering and Mineral Resources

Department

Not Listed

Committee Chair

Prashnna K. Gyawali

Committee Co-Chair

Anthony Sicilia

Committee Member

Jeremy Dawson

Abstract

Vision–language models (VLMs) and the embedding models that underlie them have become the default representation layer for multimodal artificial intelligence. Trained by contrastive alignment over enormous, loosely curated image–text corpora, they inherit the demographic skew of that data, and they encode it in the geometry of the shared embedding space itself. The consequence is a model that performs unevenly across demographic groups even when no protected attribute is ever supplied as an input. In consumer applications this is an equity problem; in medicine it is a safety problem, because an embedding that is systematically less discriminative for one subpopulation translates directly into missed or delayed diagnoses for that subpopulation.

This report studies demographic bias in medical VLMs in the setting of automated glaucoma screening from retinal fundus images. Glaucoma is well suited to this purpose: it is a leading cause of irreversible blindness, it is asymptomatic until late, it disproportionately burdens underserved populations, and it is one of the few medical imaging domains with a public dataset (Harvard–FairVLMed) that pairs images with clinical notes and with a rich set of demographic annotations.

Existing debiasing methods for VLMs assume that protected attributes are available at training time. In real clinical deployments they typically are not: demographic labels may be legally restricted, unrecorded, or simply unreliable. This report develops an attribute-agnostic debiasing framework, Debiased CLIP, that requires no demographic supervision. The method relaxes the Rawlsian worst-group objective into a differentiable top-k loss over the hardest image–text pairs, computes the gradient alignment between that top-k CLIP loss and a per-sample image–image contrastive loss, and uses the resulting alignment scores as surrogate weights that up-weight exactly those visual samples whose corrective updates reduce the worst-case multimodal loss. Because the hardest pairs are disproportionately drawn from under-represented subpopulations, the weighting acts as an implicit reweighting of latent demographic groups.

Evaluated zero-shot on the Harvard–FairVLMed glaucoma subset across six protected attributes, Debiased CLIP reduces Equalized Odds Difference on four of six attributes (ethnicity 17.10 to 3.50; language 22.28 to 16.00; age 27.72 to 16.24; marital status 46.88 to 28.57) and improves Equalized Subgroup AUC on five of six, with the largest group-wise gains concentrated in the smallest subgroups: Spanish-speaking patients (1.7% of the cohort) improve from 58.24 to 70.17 AUC, and Hispanic patients (4.0%) from 57.92 to 63.09. These gains are obtained without any protected-attribute labels and with a single model rather than one model per attribute, which makes the approach compatible with the privacy constraints under which clinical systems actually operate.

Share

COinS