Time Series, Clustering, Network Models and Tensor Analysis
Access to this document is restricted. Some items have been embargoed at the request of the author, but will be made publicly available after the "No Access Until" date.
During the embargo period, you may request access to the item by clicking the link to the restricted file(s) and completing the request form. If we have contact information for a Cornell author, we will contact the author and request permission to provide access. If we do not have contact information for a Cornell author, or the author denies or does not respond to our inquiry, we will not be able to provide access. For more information, review our policies for restricted content.
High dimensional data comes in many forms and with modern technologies we are able to capture and store ever larger amounts of data. These increasing amounts of high dimensional data lead to a need for statistical methods which are able analyze and digest such data in a computationally feasible manner. Toward this end, this thesis develops methods for analyzing some of the types of high dimensional data that are observed in practice: time series, networks, and higher-order tensors with an emphasis on clustering methods applied to these data types.First we look at analyzing corpora of time series data. We develop a model-based approach to clustering time series data using Autoregressive Moving Average (ARMA) models. Using these models we are able to not only identify groups of related time series, but also develop goodness of fit tests for models applied to the resulting clusters. Next, we look at analyzing sets of networks. To analyze these, we apply Hutchinson Trace Estimation to estimate quantiles of the spectrum of a graph to use as an embedding. With these embeddings we show how sets of graphs can be clustered using both the K-means algorithm and dynamic time warping. Finally, we look at analyzing higher-order tensor data via a decomposition which interpolates between the CP-decomposition and the Tucker decomposition. This decomposition, which we call the grouped-tucker decomposition, groups together the different dimensions of the tensor and allows for different behavior within the groups and between the groups. We also develop a method for identifying the groups via a hierarchical clustering inspired algorithm.