Active Subsampling for Estimation and Inference of Individualized Thresholds
Access to this document is restricted. Some items have been embargoed at the request of the author, but will be made publicly available after the "No Access Until" date.
During the embargo period, you may request access to the item by clicking the link to the restricted file(s) and completing the request form. If we have contact information for a Cornell author, we will contact the author and request permission to provide access. If we do not have contact information for a Cornell author, or the author denies or does not respond to our inquiry, we will not be able to provide access. For more information, review our policies for restricted content.
In this thesis, we study individualized thresholds, where our goal is to learn a high-dimensional parameter θ in a linear threshold θᵀZ for a continuous variable X, such that the discrepancy between whether X exceeds the threshold θᵀZ and a binary outcome Y is minimized. This framework arises naturally in a variety of applications, including the individualized minimal clinically important difference, which aims to quantify clinically meaningful changes at the individual level in biomedical studies. On the estimation side, in many modern applications, labeled data may be expensive or difficult to obtain, leading to measurement-constrained settings in which large amounts of covariate data are available but obtaining labeled outcomes is costly and time-consuming. As a result, only a limited number of observations can be labeled within a given budget. This raises a fundamental question: how can we efficiently utilize a limited labeling budget to estimate individualized thresholds? To address this question, we propose a novel active subsampling framework for M-estimation. The key idea is to adaptively select the most informative observations for labeling, rather than sampling uniformly at random. We introduce a K-step procedure that iteratively refines the sampling strategy and solves a regularized M-estimation problem. Our theoretical analysis reveals a sharp phase transition phenomenon governed by the smoothness of the conditional density, and shows that the proposed method can achieve the parametric convergence rate under suitable conditions, strictly improving over the minimax rate attainable in the i.i.d. setting. Furthermore, we formulate an N-budget minimax framework for the measurement-constrained M-estimation problem and prove that our estimator is minimax rate optimal up to a logarithmic factor. We also develop Lepski's methods to achieve adaptation to the unknown smoothness and sparsity respectively, and provide practical guidelines for the implementation of our algorithm. Finally, we demonstrate the superior performance of our method in simulation studies and apply the method to analyze a large diabetes dataset from 130 US hospitals. On the inference side, we study the problem of testing the significance of components of the individualized threshold parameter in a high-dimensional setting without measurement constraints, where labeled data are fully observed. The difficulty dues to the high-dimensional nuisance in developing such a testing procedure, and also stems from the fact that this high-dimensional threshold model is nonregular and the limiting distribution of the corresponding estimator is nonstandard. To deal with these challenges, we construct a test statistic via a new bias-corrected smoothed decorrelated score approach, and establish its asymptotic distributions under both null and local alternative hypotheses. We propose a double-smoothing approach to select the optimal bandwidth in our test statistic and provide theoretical guarantees for the selected bandwidth. We conduct simulation studies to demonstrate how our proposed procedure can be applied in empirical studies. We apply the proposed method to a clinical trial where the scientific goal is to assess the clinical importance of a surgery procedure. Together, these results provide a unified framework for individualized thresholds, combining data-efficient estimation under measurement constraints with valid high-dimensional inference in fully observed settings.