Uncertainty Quantification in Structured High Dimensional Regression Tasks with Missing Data

Nathan Hwangbo
Ph.D., 2026
MCKINNON, KAREN A
Modern statistical analyses often involve inference in the setting where the number of parameters exceeds the number of datapoints. Valid inference in this high dimensional setting relies on modeling assumptions that constrain the dimensionality of the parameter space. The set of assumptions appropriate for a given task should ideally make use of domain knowledge about the underlying structure of the data. This thesis develops statistical models for high dimensional regression tasks in two scientific domains, with particular attention paid to quantifying uncertainty due to these structural assumptions. Chapters 1 and 2 discuss the field of paleoclimate reconstruction, whereby climate of the past is inferred through natural archives like tree-ring widths (TRW). The paleoclimate reconstruction task is best formulated as an inverse problem, where the unobserved climate of the past is viewed as missing data that must be estimated by inverting a forward model which maps climate to tree-ring widths. Without constraints, this problem is ill-posed, in the sense that it is impossible to disentangle the contribution of climatic and non-climatic influences on variation in the paleoclimate proxy. We consider the Bayesian approach to the inverse problem, which estimates the probability distribution of the inverse solution and constrains the solution space by imposing prior distributions on both past climate and the climate-proxy relationship. Chapter 2 proposes a Bayesian hierarchical model for TRW-based paleoclimate reconstructions at a single spatial location, with emphasis placed on propagating uncertainties in the forward model. One source of uncertainty in existing paleoclimate reconstruction workflows is the temporal change of support problem, whereby TRWs produce a single measurement per year, but are typically not equally sensitive to climate in all seasons. The proposed model resolves the change-of-support problem by modeling climate at monthly temporal resolution, simultaneously estimating the season which the proxies are sensitive to alongside the reconstruction. Chapter 3 considers uncertainty in paleoclimate field reconstructions, where proxies at multiple spatial locations are used to generate a gridded spatial reconstruction of the past. The high dimensional spatiotemporal nature of this task make computational feasability a limiting factor, so that popular existing methods for climate field reconstruction make use of simplifying assumptions for practical use. In Chapter 3, Bayesian hierarchical models for climate field reconstruction are developed which include these existing methods as special cases. This framework is used alongside pseudoproxy experiments to highlight the simplifying assumptions implicit in existing climate field reconstruction methods, and to assess the influence of these assumptions on reconstruction skill. Insights gained from these experiments are used to develop a simplified model for the common paleoclimate task of estimating a large-scale spatial average. Chapter 4 discusses the application of high dimensional regression to a different scientific domain, the identification of biomarkers in metabolomic profiles of cerebrospinal fluid (CSF). Recent advances in high-throughput biology have allowed for the measurement of concentrations of all small molecules in a biological sample (known as the metabolome). These metabolomic profiles are influenced by both the genome and the environment, and can be used to identify biomarkers of different phenotypes. We focus on the identification of biomarkers for Alzhiemer's and Parkinson's disease using metabolomic profiles of CSF. Due to the invasive procedure required to extract CSF, very few samples are available for analysis, leading to a “small $n$, large $p$” setting where the number of metabolites greatly exceeds the sample size. In addition, multiple technologies and processing techniques are available for annotating and identifying metabolites, so that the results of any metabolomic analysis are sensitive to the particular data collection methods used. These metabolomic profiles are subject to measurement error and frequently feature missing data, so that the imputation of missing regressors (here, metabolite concentrations) is a common theme shared across the thesis. We compare biomarker identification across metabolomic profiles and multiple statistical analyses to assess the robustness of our findings to these data and methodological choices. Chapter 5 contains a discussion of possible extensions to the methods developed here to obtain more reliable inference in each of these scientific domains.
2026