Author ORCID Identifier
Semester
Summer
Date of Graduation
2026
Document Type
Dissertation
Degree Type
PhD
College
Statler College of Engineering and Mineral Resources
Department
Lane Department of Computer Science and Electrical Engineering
Committee Chair
Jeremy Dawson
Committee Member
Matthew Valenti
Committee Member
Gianfranco Doretto
Committee Member
Prashnna Gyawali
Committee Member
Jacqueline Speir
Abstract
The rapid advancement of intelligent surveillance systems and the increasing demand for reliable biometric identification in border security, public safety, and digital forensics require robust face understanding under unconstrained conditions, including low resolution, pose variation, and occlusion. While Vision Transformer (ViT)-based foundation models have greatly improved visual representation learning, their patch-based tokenization and lack of spatial inductive bias limit their ability to capture fine-grained details in low-resolution inputs. This dissertation investigates multimodal representation learning for face understanding through natural language supervision, large-scale face-caption pre-training, and parameter-efficient foundation model adaptation. It hypothesizes that textual supervision provides complementary semantic cues that improve the robustness, interpretability, and generalizability of facial representations across diverse downstream tasks. To achieve this goal, the dissertation introduces a unified set of multimodal learning frameworks that address several fundamental challenges, including modality heterogeneity and cross-modal alignment, the scarcity and noise of face-caption data, generic and non-discriminative caption generation, and reward hacking in reinforcement learning-based optimization. First, a caption-guided face recognition framework is proposed that utilizes natural language supervision to improve recognition performance in low-resolution surveillance scenarios. Second, a face-caption pre-training framework is developed to learn universal facial representations with strong transferability across both cross-modal understanding and face analysis tasks. Third, an attribute-driven face captioning framework is introduced with a novel reward optimization strategy that generates identity-discriminative descriptions, substantially improving text-based face retrieval. Collectively, these contributions establish multimodal representation learning through caption supervision as a scalable and generalizable framework for face understanding. Extensive experiments on multiple face verification, retrieval, and captioning benchmarks, substantiate these contributions: the caption-guided recognition framework improves 1:1 verification rate by up to 1.8 percentage points (TAR@FAR=10−6) over state-of-the-art text-guided face recognition methods; the face-caption pre-training framework improves zero-shot text-to-face retrieval (R@5) by up to 2.1 percentage points over strong facial representation learning baselines while reaching 92.12% facial attribute recognition accuracy on CelebA; and the attribute-driven captioning framework improves zero-shot text-to-face self-retrieval (R@10) by 5.0–15.8 percentage points over the strongest reward-based baseline across three out-of-domain benchmarks.
Recommended Citation
Hasan, Md Mahedi, "Multimodal Representation Learning for Face Understanding: From Caption Supervision to Foundation Model Adaptation" (2026). Graduate Theses, Dissertations, and Problem Reports (ETD). 13478.
https://researchrepository.wvu.edu/etd/13478