Author ORCID Identifier

https://orcid.org/0009-0007-3948-2608

Semester

Summer

Date of Graduation

2026

Document Type

Problem/Project Report

Degree Type

MS

College

Statler College of Engineering and Mineral Resources

Department

Lane Department of Computer Science and Electrical Engineering

Committee Chair

Prashnna Gyawali

Committee Co-Chair

Jeremy Dawson

Committee Member

Anthony Sicilia

Abstract

The safe clinical deployment of deep learning models for high-stakes medical imaging tasks requires more than high average accuracy; it requires demonstrable, per-case reliability. Uncertainty quantification (UQ) provides the missing signal that tells a clinician when a model prediction can be trusted and when a case should be escalated for expert review. Among UQ approaches, conformal prediction (CP) is especially attractive because it produces prediction sets that are guaranteed, under the assumption of exchangeability, to contain the true label with a user chosen probability, and it does so without assumptions about the model or the data distribution. This report first develops why UQ is a central problem for trustworthy medical imaging and why conformal prediction is a principled way to address it. It then examines a practical tension that has received little attention: data augmentation is a near universal ingredient of modern training pipelines, yet by design it alters the training distribution and may therefore interact with the exchangeability assumption that underpins conformal guarantees. To study this interaction in a concrete, clinically meaningful setting, we use diabetic retinopathy (DR) grading as our experimental case study. Using the publicly available DDR dataset, we evaluate two backbone architectures, ResNet-50 and a Co-Scale Conv-Attentional Transformer (CoaT), each trained under five augmentation regimes: no augmentation, standard geometric transforms, Contrast Limited Adaptive Histogram Equalization (CLAHE), Mixup, and CutMix. We analyze the downstream effects on conformal metrics, including empirical coverage, average prediction set size, and correct efficiency. Our results show that sample mixing strategies such as Mixup and CutMix not only improve predictive accuracy but also tend to yield more reliable and efficient uncertainty estimates, whereas contrast enhancement through CLAHE can degrade model certainty. These findings support a central recommendation: augmentation strategies should be co-designed with downstream uncertainty quantification in mind, rather than optimized for accuracy alone, if we are to build genuinely trustworthy artificial intelligence systems for medical imaging.

Share

COinS