<?xml version='1.0' encoding='UTF-8'?><?xml-stylesheet href='static/style.xsl' type='text/xsl'?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-19T18:14:29Z</responseDate><request verb="GetRecord" identifier="oai:ecommons.cornell.edu:1813/59414" metadataPrefix="dim">https://ecommons.cornell.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:ecommons.cornell.edu:1813/59414</identifier><datestamp>2026-05-15T19:50:42Z</datestamp><setSpec>com_1813_35</setSpec><setSpec>col_1813_47</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="author">Lee, Moontae</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="chair">Mimno, David</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="committeeMember">Frazier, Peter</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="committeeMember">Bindel, David S.</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2018-10-23T13:22:40Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2020-06-04T06:00:32Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2018-05-30</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="other">ProQuest Submission ID: 10890</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="other">ProQuest Publication ID: 10823238</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">https://hdl.handle.net/1813/59414</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="doi">https://doi.org/10.7298/X4FB514V</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="bibid">10489499</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract">Co-occurrence information is powerful statistics that can model various discrete objects by their joint instances with other objects. Transforming unsupervised problems of learning low-dimensional geometry into provable decompositions of co-occurrence information, spectral inference provides fast algorithms and optimality guarantees for non-linear dimensionality reduction or latent topic analysis. Spectral approaches reduce the dependence on the original training examples and produce substantial gain in efficiency, but at costs: a) The algorithms perform poorly on real data that does not necessarily follow underlying models; b) Users can no longer infer information about individual examples, which is often important for real-world applications; c) Model complexity rapidly grows as the number of objects increases, requiring a careful curation of the vocabulary. The first issue is called model-data mismatch, which is a fundamental problem common in every spectral inference method for latent variable models. As real data never follows any particular computational model, this issue must be ad- dressed for practicality of the spectral inference beyond synthetic settings. For the second issue, users could revisit probabilistic inference to infer information about individual examples, but this brings back all the drawbacks of traditional approaches. One method is recently developed for spectral inference, but it works only on tiny models, quickly losing its performance for the datasets whose underlying structures exhibit realistic correlations. While probabilistic inference also suffers from the third issue, the problem is more serious for spectral inferences because co-occurrence information easily exceeds storable capacity as the size of vocabulary becomes larger. We cast the learning problem in the framework of Joint Stochastic Matrix Factorization (JSMF), showing that existing methods violate the theoretical conditions necessary for a good solution to exist. Proposing novel rectification paradigms for handling the model-data mismatch, the Rectified Anchor Word Algorithm (RAWA) is able to learn quality latent structures and their interactions even on small noisy data. We also propose the Prior Aware Dual Decomposition (PADD) that is capable of considering the learned interactions as well as the learned latent structures to robustly infer example- specific information. Beyond the theoretical guarantees, our experimental results show that RAWA recovers quality low-dimensional geometry on various textual/non-textual datasets comparable to probabilistic Gibbs sampling, and PADD substantially outperforms the recently developed method for learning low-dimensional representations of individual examples. Although this thesis does not address the complexity issue for large vocabulary, we have developed new methods that can drastically compress co-occurrence information and learn only with the compressed statistics without losing much precision. Providing rich capability to operate on millions of objects and billions of examples, we complete all the necessary tools to make spectral inference robust and scalable competitor to probabilistic inference for unsupervised latent structure learning. We hope our research serves an initial basis for a new perspective that combines the benefits of both spectral and probabilistic worlds.</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso">en_US</dim:field>
   <dim:field mdschema="dc" element="subject">Statistics</dim:field>
   <dim:field mdschema="dc" element="subject">Applied mathematics</dim:field>
   <dim:field mdschema="dc" element="subject">Anchor Word Algorithm</dim:field>
   <dim:field mdschema="dc" element="subject">Co-occurrence Modeling</dim:field>
   <dim:field mdschema="dc" element="subject">Joint-Stochastic Matrix Factorization</dim:field>
   <dim:field mdschema="dc" element="subject">Rectification</dim:field>
   <dim:field mdschema="dc" element="subject">Spectral Inference</dim:field>
   <dim:field mdschema="dc" element="subject">Topic Modeling</dim:field>
   <dim:field mdschema="dc" element="subject">Artificial intelligence</dim:field>
   <dim:field mdschema="dc" element="title">Joint-stochastic Spectral Inference for Robust Co-occurrence Modeling and Latent Topic Analysis</dim:field>
   <dim:field mdschema="dc" element="type">dissertation or thesis</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="discipline">Computer Science</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="grantor">Cornell University</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="level">Doctor of Philosophy</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="name">Ph. D., Computer Science</dim:field>
   <dim:field mdschema="dcterms" element="license">https://hdl.handle.net/1813/59810</dim:field>
   <dim:field mdschema="dspace" element="entity" qualifier="type">Publication</dim:field>
   <dim:field mdschema="cris" element="virtual" qualifier="collection" authority="https://cornell-ecommons.eks.prod.4science.cloud/handle/1813/47" confidence="600">Cornell Theses and Dissertations</dim:field>
   <dim:field mdschema="cris" element="virtual" qualifier="author">Lee, Moontae</dim:field>
   <dim:field mdschema="cris" element="virtualsource" qualifier="collection">5893a6ea-7af3-41d7-abc6-04bcd26ab5df</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
   <dim:field mdschema="cerif" element="openaire" authority="" confidence="-1">&lt;Publication xmlns="https://www.openaire.eu/cerif-profile/1.1/" id="0e43dfb5-d421-4fd4-9163-978c125adf9d">
	&lt;Type xmlns="https://www.openaire.eu/cerif-profile/vocab/COAR_Publication_Types">http://purl.org/coar/resource_type/c_1843&lt;/Type>
	&lt;Language>en_US&lt;/Language>
   	&lt;Title>Joint-stochastic Spectral Inference for Robust Co-occurrence Modeling and Latent Topic Analysis&lt;/Title>
   	&lt;PublishedIn>
    	&lt;Publication>
      	&lt;/Publication>
   	&lt;/PublishedIn>
   	&lt;PublicationDate>2018-05-30&lt;/PublicationDate>
   	&lt;DOI>https://doi.org/10.7298/X4FB514V&lt;/DOI>
   	&lt;Authors>
      	&lt;Author>
        	&lt;DisplayName>Lee, Moontae&lt;/DisplayName>
         	&lt;Affiliation>
         		&lt;OrgUnit>
         		&lt;/OrgUnit>
         	&lt;/Affiliation>
      	&lt;/Author>
	&lt;/Authors>
   	&lt;Editors>
	&lt;/Editors>
    &lt;Publishers>
        &lt;Publisher>
            &lt;OrgUnit />
        &lt;/Publisher>
    &lt;/Publishers>
    &lt;Keyword>Statistics&lt;/Keyword>
    &lt;Keyword>Applied mathematics&lt;/Keyword>
    &lt;Keyword>Anchor Word Algorithm&lt;/Keyword>
    &lt;Keyword>Co-occurrence Modeling&lt;/Keyword>
    &lt;Keyword>Joint-Stochastic Matrix Factorization&lt;/Keyword>
    &lt;Keyword>Rectification&lt;/Keyword>
    &lt;Keyword>Spectral Inference&lt;/Keyword>
    &lt;Keyword>Topic Modeling&lt;/Keyword>
    &lt;Keyword>Artificial intelligence&lt;/Keyword>
   	&lt;Abstract>Co-occurrence information is powerful statistics that can model various discrete objects by their joint instances with other objects. Transforming unsupervised problems of learning low-dimensional geometry into provable decompositions of co-occurrence information, spectral inference provides fast algorithms and optimality guarantees for non-linear dimensionality reduction or latent topic analysis. Spectral approaches reduce the dependence on the original training examples and produce substantial gain in efficiency, but at costs: a) The algorithms perform poorly on real data that does not necessarily follow underlying models; b) Users can no longer infer information about individual examples, which is often important for real-world applications; c) Model complexity rapidly grows as the number of objects increases, requiring a careful curation of the vocabulary. The first issue is called model-data mismatch, which is a fundamental problem common in every spectral inference method for latent variable models. As real data never follows any particular computational model, this issue must be ad- dressed for practicality of the spectral inference beyond synthetic settings. For the second issue, users could revisit probabilistic inference to infer information about individual examples, but this brings back all the drawbacks of traditional approaches. One method is recently developed for spectral inference, but it works only on tiny models, quickly losing its performance for the datasets whose underlying structures exhibit realistic correlations. While probabilistic inference also suffers from the third issue, the problem is more serious for spectral inferences because co-occurrence information easily exceeds storable capacity as the size of vocabulary becomes larger. We cast the learning problem in the framework of Joint Stochastic Matrix Factorization (JSMF), showing that existing methods violate the theoretical conditions necessary for a good solution to exist. Proposing novel rectification paradigms for handling the model-data mismatch, the Rectified Anchor Word Algorithm (RAWA) is able to learn quality latent structures and their interactions even on small noisy data. We also propose the Prior Aware Dual Decomposition (PADD) that is capable of considering the learned interactions as well as the learned latent structures to robustly infer example- specific information. Beyond the theoretical guarantees, our experimental results show that RAWA recovers quality low-dimensional geometry on various textual/non-textual datasets comparable to probabilistic Gibbs sampling, and PADD substantially outperforms the recently developed method for learning low-dimensional representations of individual examples. Although this thesis does not address the complexity issue for large vocabulary, we have developed new methods that can drastically compress co-occurrence information and learn only with the compressed statistics without losing much precision. Providing rich capability to operate on millions of objects and billions of examples, we complete all the necessary tools to make spectral inference robust and scalable competitor to probabilistic inference for unsupervised latent structure learning. We hope our research serves an initial basis for a new perspective that combines the benefits of both spectral and probabilistic worlds.&lt;/Abstract>
	&lt;Access xmlns="http://purl.org/coar/access_right" 
    >
    &lt;/Access>
&lt;/Publication>
</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>