Author ORCID Identifier

https://orcid.org/0000-0002-6337-3571

Semester

Summer

Date of Graduation

2026

Document Type

Dissertation

Degree Type

PhD

College

Statler College of Engineering and Mineral Resources

Department

Lane Department of Computer Science and Electrical Engineering

Committee Chair

Jeremy Dawson

Committee Member

Matthew Valenti

Committee Member

Gianfranco Doretto

Committee Member

Prashnna Gyawali

Committee Member

Jacqueline Speir

Abstract

The rapid advancement of intelligent surveillance systems and the increasing demand for reliable biometric identification in border security, public safety, and digital forensics require robust face understanding under unconstrained conditions, including low resolution, pose variation, and occlusion. While Vision Transformer (ViT)-based foundation models have greatly improved visual representation learning, their patch-based tokenization and lack of spatial inductive bias limit their ability to capture fine-grained details in low-resolution inputs. This dissertation investigates multimodal representation learning for face understanding through natural language supervision, large-scale face-caption pre-training, and parameter-efficient foundation model adaptation. It hypothesizes that textual supervision provides complementary semantic cues that improve the robustness, interpretability, and generalizability of facial representations across diverse downstream tasks. To achieve this goal, the dissertation introduces a unified set of multimodal learning frameworks that address several fundamental challenges, including modality heterogeneity and cross-modal alignment, the scarcity and noise of face-caption data, generic and non-discriminative caption generation, and reward hacking in reinforcement learning-based optimization. First, a caption-guided face recognition framework is proposed that utilizes natural language supervision to improve recognition performance in low-resolution surveillance scenarios. Second, a face-caption pre-training framework is developed to learn universal facial representations with strong transferability across both cross-modal understanding and face analysis tasks. Third, an attribute-driven face captioning framework is introduced with a novel reward optimization strategy that generates identity-discriminative descriptions, substantially improving text-based face retrieval. Collectively, these contributions establish multimodal representation learning through caption supervision as a scalable and generalizable framework for face understanding. Extensive experiments on multiple face verification, retrieval, and captioning benchmarks, substantiate these contributions: the caption-guided recognition framework improves 1:1 verification rate by up to 1.8 percentage points (TAR@FAR=10−6) over state-of-the-art text-guided face recognition methods; the face-caption pre-training framework improves zero-shot text-to-face retrieval (R@5) by up to 2.1 percentage points over strong facial representation learning baselines while reaching 92.12% facial attribute recognition accuracy on CelebA; and the attribute-driven captioning framework improves zero-shot text-to-face self-retrieval (R@10) by 5.0–15.8 percentage points over the strongest reward-based baseline across three out-of-domain benchmarks.

Share

COinS