IP Library › Granted Patent US 12,608,444
Granted Patent B2
US 12,608,444 · App. 17/816,288 · Granted Apr 21, 2026

Automated selection of principal component analysis variants for large-scale datasets

Inventors: Xi Cheng (Kirkland, WA); Mingge Deng (Kirkland, WA); Amir Hossein Hormati (Seattle, WA)
Assignee: Google LLC
G06F18/2135G06F16/2433G06N20/00G05B23/024G06F16/7343G06N7/00G06V10/77H04L41/024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,444
App. No.
17/816,288
Granted
Apr 21, 2026
Kind
B2
Abstract

A method for principal component analysis includes receiving a principal component analysis (PCA) request from a user requesting data processing hardware to perform PCA on a dataset, the dataset including a plurality of input features. The method further includes training a PCA model on the plurality of input features of the dataset. The method includes determining, using the trained PCA model, one or more principal components of the dataset. The method also includes generating, based on the plurality of input features and the one or more principal components, one or more embedded features of the dataset. The method includes returning the one or more embedded features to the user.

Claims (32)

1 . A computer-implemented method comprising:

receiving, by data processing hardware, a principal component analysis request from a user requesting the data processing hardware to perform principal component analysis on a dataset, the dataset comprising a plurality of input features;

determining, by the data processing hardware, whether a number of the plurality of input features satisfies a threshold number of input features, wherein the threshold number of input features corresponds to a number of the input features that fit into a memory of a single server that implements the data processing hardware;

responsive to determining that the number of the plurality of input features satisfies the threshold number of input features, selecting, by the data processing hardware, a basic principal component analysis variant as a selected principal component analysis variant;

responsive to determining that the number of the plurality of input features fails to satisfy the threshold number of input features, selecting, by the data processing hardware, a randomized principal component analysis variant as the selected principal component analysis variant;

training, by the data processing hardware and using the selected principal component analysis variant, a principal component analysis model on the plurality of input features of the dataset;

determining, by the data processing hardware and using the trained principal component analysis model, one or more principal components of the dataset;

generating, by the data processing hardware and based on the plurality of input features and the one or more principal components, one or more embedded features of the dataset; and

returning, by the data processing hardware, the one or more embedded features to the user.

2 . The method of claim 1 , wherein a number of the one or more embedded features of the dataset is less than a number of the plurality of input features of the dataset.

3 . The method of claim 1 , wherein the principal component analysis request comprises a single Structured Query Language query.

4 . The method of claim 1 , wherein determining, using the trained principal component analysis model, the one or more principal components of the dataset comprises using an economy sized quadrature remainder decomposition algorithm.

5 . The method of claim 1 , wherein determining, using the trained principal component analysis model, the one or more principal components of the dataset comprises determining, by the data processing hardware and using a quadratic programming with non-decreasing constraints algorithm, the one or more principal components of the dataset.

6 . The method of claim 1 , wherein the randomized principal component analysis variant is configured to process a matrix having a number of rows in a defined range of 100,000 to 1,000,000.

7 . The method of claim 1 , wherein the randomized principal component analysis variant is a transposed randomized principal component analysis algorithm.

8 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to:

receive a principal component analysis request from a user requesting the data processing hardware to perform principal component analysis on a dataset, the dataset comprising a plurality of input features;

determine whether a number of the plurality of input features satisfies a threshold number of input features, wherein the threshold number of input features corresponds to a number of the input features that fit into a memory of a single server that implements the data processing hardware;

responsive to determining that the number of the plurality of input features satisfies the threshold number of input features, select a basic principal component analysis variant as a selected principal component analysis variant;

responsive to determining that the number of the plurality of input features fails to satisfy the threshold number of input features, select a randomized principal component analysis variant as the selected principal component analysis variant;

train, using the selected principal component analysis variant, a principal component analysis model on the plurality of input features of the dataset;

determine, using the trained principal component analysis model, one or more principal components of the dataset;

generate, based on the plurality of input features and the one or more principal components, one or more embedded features of the dataset; and

return the one or more embedded features to the user.

9 . The system of claim 8 , wherein a number of the one or more embedded features of the dataset is less than a number of the plurality of input features of the dataset.

10 . The system of claim 8 , wherein the principal component analysis request comprises a single Structured Query Language query.

11 . The system of claim 8 , wherein the instructions cause the data processing hardware to determine the one or more principal components of the dataset using an economy sized quadrature remainder decomposition algorithm.

12 . The system of claim 8 , wherein the instructions cause the data processing hardware to determine the one or more principal components of the dataset using a quadratic programming with non-decreasing constraints algorithm.

13 . The system of claim 8 , wherein the randomized principal component analysis variant is configured to process a matrix having a number of rows in a defined range of 100,000 to 1,000,000.

14 . The system of claim 8 , wherein the randomized principal component analysis variant is a transposed randomized principal component analysis algorithm.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2022
From: CHENG, XI; DENG, MINGGE; HORMATI, AMIR HOSSEIN
To: GOOGLE LLC
Reel/Frame 061057/0372 →
Continuity (2)
Provisional Application 63203934 · Aug 4, 2021
Related Publication 20230045139A1 · Feb 9, 2023
References Cited (40)
US 20040260682A1 · Herley · 2004 [cited by examiner]
US 20150186775A1 · Cruz Mota et al. · 2015 [cited by applicant]
US 20190311552A1 · Zhang · 2019 [cited by examiner]
US 20200034740A1 · Yang et al. · 2020 [cited by applicant]
US 20200381084A1 · Kawas · 2020 [cited by examiner]
US 20200402660A1 · Chakravarthy · 2020 [cited by examiner]
US 20210174207A1 · Parker · 2021 [cited by examiner]
US 20210264201A1 · Pandey · 2021 [cited by examiner]
US 20210374553A1 · Li · 2021 [cited by examiner]
US 20220374765A1 · Wu · 2022 [cited by examiner]
US 20220394629A1 · Lau · 2022 [cited by examiner]
US 20240086745A1 · Gemp · 2024 [cited by examiner]
US 20250086652A1 · Wang · 2025 [cited by examiner]
WO WO2022198680A1 · 2022 [cited by examiner]
Inelus, Gabriel-Robert “Quadratically Regularised Principal Component Analysis over Multi-Relational Databases” Jul. 5, 2019, pp. i-82. (Year: 2019). [cited by examiner]
Erichson et al., “Randomized Matrix Decomposition Using R” Nov. 26, 2019, arXiv: 1608.02148v5, pp. i-47. (Year: 2019). [cited by examiner]
Sharma et al., “Principal component analysis using QR decomposition” Sep. 25, 2012, pp. 679-683. (Year: 2012). [cited by examiner]
Rolinek et al., “Variational Autoencoders Pursue PCA Directions (by Accident)” Apr. 16, 2019, arXiv: 1812.06775v2, pp. 1-19. (Year: 2019). [cited by examiner]
Dou et al., “PCA-SRGAN: Incremental Orthogonal Projection Discrimination for Face Super-resolution” Aug. 28, 2020, arXiv: 2005.00306v2, pp. 1-8. (Year: 2020). [cited by examiner]
He et al., “EigenGAN: Layer-wise Eigen-Learning for GANs” Apr. 26, 2021, arXiv: 2104.12476v1, pp. 1-35. (Year: 2021). [cited by examiner]
Balcan et al., “Communication Efficient Distributed Kernel Principal Component Analysis” Aug. 2016, pp. 725-734. (Year: 2016). [cited by examiner]
Ballabio et al., “A MATLAB toolbox for Principal Component Analysis and unsupervised exploration of data structure” Oct. 19, 2015, pp. 1-9. (Year: 2015). [cited by examiner]
Fan et al., “Distributed estimation of principal eigenspaces” Jan. 10, 2018, arXiv: 1702.06488v4, pp. 1-47. (Year: 2018). [cited by examiner]
Ghojogh et Crowley, “Unsupervised and Supervised Principal Component Analysis: Tutorial” Jun. 1, 2019, arXiv: 1906.03148v1, pp. 1-25. (Year: 2019). [cited by examiner]
Golinko et Zhu, “Generalized Feature Embedding for Supervised, Unsupervised, and Online Learning Tasks” Apr. 16, 2018, pp. 125-142. (Year: 2018). [cited by examiner]
Hertrich et al., “PCA Reduced Gaussian Mixture Models with Applications in Superresolution” May 6, 2021, arXiv: 2009.07520v3, pp. 1-31. (Year: 2021). [cited by examiner]
Wu et al., “A Review of Distributed Algorithms for Principal Component Analysis” 2018, pp. 1-19. (Year: 2018). [cited by examiner]
International Search Report and Written Opinion for the related Application No. PCT/US2022/074341 dated Nov. 24, 2022, 231 pages. [cited by applicant]
Benjamin Erichson N et al: “Randomized Matrix Decompositions using R”, arxj:v. org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 6, 2016 (Aug. 6, 2016), XP081560626 , DOI: 10.18… [cited by applicant]
Zhang Yiqun et al: “Big Data Analytics, Integrating a Parallel Columnar DBMS and the R. Language”, 2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing CCGRID), IEEE, May 16, 2016 (May 16, 201… [cited by applicant]
Nathan Halko et al: “An algorithm for the principal component analysis of large data sets,” arxiv.org, Cornell University Library, 201, Olin Library Cornell University Ithaca, NY 14853, Jul. 30, 2010 (Jul. 30, 2010), XP… [cited by applicant]
Elgamal Tarek Tgamal@QF ORG QA et al: “sPCA Scalable Principal Component Analysis for Big Data on Distributed Platforms”, Proceedings of the 27th Annual International Conference on Mobile Computing and Networking, ACM, … [cited by applicant]
Andreas Grammenos et al: “Federated Principal Component Analysis”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 22, 2020 (Oct. 22, 2020), XP081795533, 36 pages. [cited by applicant]
Boutsidis@gmail com et al: “Optimal principal component analysis in distributed and streaming models”, Theory of Computing, ACM, 2 Penn Plaza, Sute 701 New York NY 10121-0701 USA, Jun. 19, 2016 (Jun. 19, 2016), pp. 236-… [cited by applicant]
“PCA”, scikit-learn 1.7.1 documentation, 11 pp., Retrieved from the Internet on Jul. 18, 2025 from URL: https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.PCA.html. [cited by applicant]
“PCA: Principal Component Analysis (PCA)”, RDocumentation, 4 pp., Retrieved from the Internet on Jul. 18, 2025 from URL: https://www.rdocumentation.org/packages/FactoMineR/versions/2.4/topics/PCA. [cited by applicant]
Response to Communication Pursuant to Rules 161(1) and 162 EPC dated Mar. 13, 2024, from counterpart European Application No. 22761867.5, filed Sep. 12, 2024, 18 pp. [cited by applicant]
First Examination Report from counterpart Indian Application No. 202447013919 dated Oct. 14, 2025, 11 pp. [cited by applicant]
Communication pursuant to Article 94(3) EPC from counterpart European Application No. 22761867.5 dated Jan. 22, 2026, 9 pp. [cited by applicant]
Response to First Examination Report dated Oct. 14, 2025 from IN Application No. 202447013919, filed Jan. 30, 2026, 16 pp. [cited by applicant]