Essays on Latent Variable Models
This thesis consists of three essays on latent variable models, with a central focus on how unobserved heterogeneity affects identification, estimation, and inference in econometric settings. The essays develop methodological advances for handling latent structures in the presence of missing data, text-based latent representations, and discrete choice heterogeneity, with applications in causal inference, text analysis, and industrial organization. The first chapter studies causal inference when key confounders are missing not at random. In such settings, standard conditional independence assumptions fail, rendering treatment effects generally unidentified. We consider a class of missingness mechanisms in which confounder missingness depends on the confounders themselves but is independent of the outcome, conditional on treatment and confounders. Under this structure, we develop a doubly robust and efficient estimator for the average treatment effect that corrects for missing confounders through novel propensity scores and conditional outcome models. Identification is established via a system of integral equations, including one previously studied in the literature and another that appears novel. We further introduce a low-rank structure for the missingness mechanism to accommodate settings with limited outcome variation, such as binary outcomes. The resulting estimator is both doubly robust and rate-doubly robust, achieving square-root-n consistency under weak convergence rates and attaining the semiparametric efficiency bound under outcome-independent missingness. We validate the method through simulations and empirical applications. The second chapter studies the statistical foundations of topic models in text analysis. It is a joint work with Simon Freyaldenhoven, Barry Ke, and José Luis Montiel Olea. Topic models are widely used for uncovering latent structure in documents, typically relying on the assumption of anchor words—words that uniquely identify topics—to ensure identification. We show that this assumption is not merely a normalization, but is statistically testable. Specifically, we construct a valid hypothesis test for the existence of anchor words with nontrivial power. This result implies that separability conditions commonly imposed in topic modeling are empirically restrictive and can be assessed using data. We apply the test to datasets of Federal Reserve monetary policy discussions and find evidence rejecting the existence of anchor words in one of the corpora, highlighting potential limitations of standard topic model assumptions in applied settings. The third chapter analyzes finite mixtures of multinomial logit models used in discrete choice and industrial organization applications. We study a setting in which each mixture component—representing latent consumer heterogeneity—satisfies a “pure alternatives” condition, requiring that each component has at least two alternatives uniquely associated with it. This assumption generalizes separability conditions commonly used in nonnegative matrix factorization and enables identification of the mixture structure. Building on this result, we propose a three-step estimation procedure and establish consistency when either the number of markets or the number of consumers per market grows large. The framework provides a flexible approach to modeling latent heterogeneity in multinomial choice environments while maintaining tractable identification and estimation.