Distributed Estimation for Principal Component Analysis: An Enlarged Eigenspace Analysis

Chen, Xi; Lee, Jason D; Li, He; Yang, Yun

Distributed Estimation for Principal Component Analysis: An Enlarged Eigenspace Analysis

Author(s): Chen, Xi; Lee, Jason D; Li, He; Yang, Yun

Download

To refer to this page use: http://arks.princeton.edu/ark:/88435/pr16h4cq69

Full metadata record

DC Field	Value	Language
dc.contributor.author	Chen, Xi	-
dc.contributor.author	Lee, Jason D	-
dc.contributor.author	Li, He	-
dc.contributor.author	Yang, Yun	-
dc.date.accessioned	2024-01-20T17:49:03Z	-
dc.date.available	2024-01-20T17:49:03Z	-
dc.date.issued	2021-04-06	en_US
dc.identifier.citation	Chen, Xi, Lee, Jason D, Li, He, Yang, Yun. (2022). Distributed Estimation for Principal Component Analysis: An Enlarged Eigenspace Analysis. Journal of the American Statistical Association, 117 (540), 1775 - 1786. doi:10.1080/01621459.2021.1886937	en_US
dc.identifier.issn	0162-1459	-
dc.identifier.uri	http://arks.princeton.edu/ark:/88435/pr16h4cq69	-
dc.description.abstract	The growing size of modern datasets brings many challenges to the existing statistical estimation approaches, which calls for new distributed methodologies. This article studies distributed estimation for a fundamental statistical machine learning problem, principal component analysis (PCA). Despite the massive literature on top eigenvector estimation, much less is presented for the top-L-dim (L > 1) eigenspace estimation, especially in a distributed manner. We propose a novel multi-round algorithm for constructing top-L-dim eigenspace for distributed data. Our algorithm takes advantage of shift-and-invert preconditioning and convex optimization. Our estimator is communication-efficient and achieves a fast convergence rate. In contrast to the existing divide-and-conquer algorithm, our approach has no restriction on the number of machines. Theoretically, the traditional Davis–Kahan theorem requires the explicit eigengap assumption to estimate the top-L-dim eigenspace. To abandon this eigengap assumption, we consider a new route in our analysis: instead of exactly identifying the top-L-dim eigenspace, we show that our estimator is able to cover the targeted top-L-dim population eigenspace. Our distributed algorithm can be applied to a wide range of statistical problems based on PCA, such as principal component regression and single index model. Finally, we provide simulation studies to demonstrate the performance of the proposed distributed estimator.	en_US
dc.format.extent	1775 - 1786	en_US
dc.language	en	en_US
dc.language.iso	en_US	en_US
dc.relation.ispartof	Journal of the American Statistical Association	en_US
dc.rights	Author's manuscript	en_US
dc.title	Distributed Estimation for Principal Component Analysis: An Enlarged Eigenspace Analysis	en_US
dc.type	Journal Article	en_US
dc.identifier.doi	doi:10.1080/01621459.2021.1886937	-
dc.identifier.eissn	1537-274X	-
pu.type.symplectic	http://www.symplectic.co.uk/publications/atom-terms/1.0/journal-article	en_US

Files in This Item:

File	Description	Size	Format
2004.02336.pdf		5.01 MB	Adobe PDF	View/Download

Show Simple Item Record