Semester
Summer
Date of Graduation
2026
Document Type
Problem/Project Report
Degree Type
MS
College
Statler College of Engineering and Mineral Resources
Department
Not Listed
Committee Chair
Prashnna K. Gyawali
Committee Co-Chair
Anthony Sicilia
Committee Member
Jeremy Dawson
Abstract
Vision–language models (VLMs) and the embedding models that underlie them have become the default representation layer for multimodal artificial intelligence. Trained by contrastive alignment over enormous, loosely curated image–text corpora, they inherit the demographic skew of that data, and they encode it in the geometry of the shared embedding space itself. The consequence is a model that performs unevenly across demographic groups even when no protected attribute is ever supplied as an input. In consumer applications this is an equity problem; in medicine it is a safety problem, because an embedding that is systematically less discriminative for one subpopulation translates directly into missed or delayed diagnoses for that subpopulation.
This report studies demographic bias in medical VLMs in the setting of automated glaucoma screening from retinal fundus images. Glaucoma is well suited to this purpose: it is a leading cause of irreversible blindness, it is asymptomatic until late, it disproportionately burdens underserved populations, and it is one of the few medical imaging domains with a public dataset (Harvard–FairVLMed) that pairs images with clinical notes and with a rich set of demographic annotations.
Existing debiasing methods for VLMs assume that protected attributes are available at training time. In real clinical deployments they typically are not: demographic labels may be legally restricted, unrecorded, or simply unreliable. This report develops an attribute-agnostic debiasing framework, Debiased CLIP, that requires no demographic supervision. The method relaxes the Rawlsian worst-group objective into a differentiable top-k loss over the hardest image–text pairs, computes the gradient alignment between that top-k CLIP loss and a per-sample image–image contrastive loss, and uses the resulting alignment scores as surrogate weights that up-weight exactly those visual samples whose corrective updates reduce the worst-case multimodal loss. Because the hardest pairs are disproportionately drawn from under-represented subpopulations, the weighting acts as an implicit reweighting of latent demographic groups.
Evaluated zero-shot on the Harvard–FairVLMed glaucoma subset across six protected attributes, Debiased CLIP reduces Equalized Odds Difference on four of six attributes (ethnicity 17.10 to 3.50; language 22.28 to 16.00; age 27.72 to 16.24; marital status 46.88 to 28.57) and improves Equalized Subgroup AUC on five of six, with the largest group-wise gains concentrated in the smallest subgroups: Spanish-speaking patients (1.7% of the cohort) improve from 58.24 to 70.17 AUC, and Hispanic patients (4.0%) from 57.92 to 63.09. These gains are obtained without any protected-attribute labels and with a single model rather than one model per attribute, which makes the approach compatible with the privacy constraints under which clinical systems actually operate.
Recommended Citation
Akash, Ahsan Habib, "Fairness Without Demographic Attributes in Medical Vision–Language Models" (2026). Graduate Theses, Dissertations, and Problem Reports (ETD). 13470.
https://researchrepository.wvu.edu/etd/13470