IP Library › Granted Patent US 12,639,573
Granted Patent B2
US 12,639,573 · App. 18/656,024 · Granted May 26, 2026

Method, system, and computer program product for embedding compression and regularization

Inventors: Haoyu Li (Columbus, OH); Junpeng Wang (Santa Clara, CA); Liang Wang (San Jose, CA); Yan Zheng (Los Gatos, CA); Wei Zhang (Fremont, CA)
Assignee: Visa International Service Association
G06N3/08G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,573
App. No.
18/656,024
Granted
May 26, 2026
Kind
B2
Abstract

A method, system, and computer program product is provided for embedding compression and reconstruction. The method includes receiving embedding vector data comprising a plurality of embedding vectors. A beta-variational autoencoder is trained based on the embedding vector data and a loss equation. The method includes determining a respective entropy of a respective mean and a respective variance of each respective dimension of a plurality of dimensions. A first subset of the plurality of dimensions is determined based on the respective entropy of the respective mean and the respective variance for each respective dimension of the plurality of dimensions. A second subset of the plurality of dimensions is discarded based on the respective entropy of the respective mean and the respective variance for each respective dimension of the plurality of dimensions. The method includes generating a compressed representation of the embedding vector data based on the first subset of dimensions.

Claims (357)

1 . A computer-implemented method for generating a compressed representation of embedding vector data, comprising:

receiving, with at least one processor, the embedding vector data comprising a plurality of embedding vectors;

training, with at least one processor, a beta-variational autoencoder based on the embedding vector data and a loss equation, the beta-variational autoencoder comprising an encoder network, a latent layer, and a decoder network, the loss equation comprising a first term associated with reconstruction of an input by the beta-variational autoencoder, a second term associated with regularization of a latent space of the latent layer, and a hyperparameter weight associated with a ratio of the first term and the second term, wherein the latent space has a plurality of dimensions;

determining, with at least one processor, a respective mean of each respective dimension of the plurality of dimensions, wherein determining the respective mean of each respective dimension of the plurality of dimensions comprises determining a respective entropy of the respective mean of each respective dimension of the plurality of dimensions;

determining whether each respective dimension of the plurality of dimensions comprises a useful dimension or a deprecated dimension based on the respective mean for the respective dimension, wherein the respective entropy of the respective mean for each useful dimension is higher than a threshold, and wherein the respective entropy of the respective mean for each deprecated dimension is lower than the threshold;

determining, with at least one processor, a first subset of the plurality of dimensions based on the respective mean for each respective dimension of the plurality of dimensions, wherein the first subset of the plurality of dimensions comprises useful dimensions;

discarding, with at least one processor, a second subset of the plurality of dimensions based on the respective mean of each respective dimension of the plurality of dimensions, wherein the second subset of the plurality of dimensions is different than the first subset of the plurality of dimensions, wherein the second subset of the plurality of dimensions comprises deprecated dimensions, and wherein discarding the second subset of the plurality of dimensions reduces a dimensionality of the latent space;

generating, with at least one processor, the compressed representation of the embedding vector data based on the first subset of dimensions; and

inputting, with at least one processor, the compressed representation of the embedding vector data into at least one machine learning model.

2 . The computer-implemented method of claim 1 , wherein determining the respective mean of each respective dimension of the plurality of dimensions comprises:

determining a respective entropy of the respective mean of each respective dimension of the plurality of dimensions; and

wherein training the beta-variational autoencoder comprises:

iteratively adjusting the hyperparameter weight and repeating the training, the determining of the respective entropy of the respective mean, the determining of the first subset, the discarding of the second subset, and the generating of the compressed representation.

3 . The computer-implemented method of claim 1 , wherein the first subset of the plurality of dimensions comprises each useful dimension, and wherein the second subset of the plurality of dimensions comprises each deprecated dimension.

4 . The computer-implemented method of claim 1 , wherein determining the respective mean of each respective dimension of the plurality of dimensions comprises:

determining, with at least one processor, a respective entropy of the respective mean of each respective dimension of the plurality of dimensions; and

wherein determining the first subset of the plurality of dimensions comprises:

determining, with at least one processor, the first subset of the plurality of dimensions based on the respective entropy of the respective mean for each respective dimension of the plurality of dimensions.

5 . The computer-implemented method of claim 4 , wherein discarding the second subset of the plurality of dimensions comprises:

discarding the second subset of the plurality of dimensions based on the respective entropy of the respective mean of each respective dimension of the plurality of dimensions.

6 . The computer-implemented method of claim 1 , wherein the loss equation is:

ℒ

=

∑

i

=

1

n

(

x

i

-

x

^

i

)

2

+

β

⁢

∑

i

=

1

m

D

KL

(

𝒩

⁡

(

μ

i

,

σ

i

2

)

⁢

𝒩

⁡

(

0

,

1

)

)

;

wherein the first term associated with reconstruction of an input by the beta-variational autoencoder is:

∑

i

=

1

n

(

x

i

-

x

^

i

)

2

;

wherein the second term associated with regularization of the latent space of the latent layer is:

∑

i

=

1

m

D

KL

(

𝒩

⁡

(

μ

i

,

σ

i

2

)

⁢

𝒩

⁡

(

0

,

1

)

)

;

wherein x i is the input and {circumflex over (x)} i is a reconstruction of the input provided by the beta-variational autoencoder;

wherein D KL is a function for a Kullback-Leibler divergence;

wherein is a function for a normal distribution; and

wherein the hyperparameter weight associated with the ratio of the first term and the second term is β.

7 . A system for generating a compressed representation of embedding vector data comprising at least one processor programmed or configured to:

receive the embedding vector data comprising a plurality of embedding vectors;

train a beta-variational autoencoder based on the embedding vector data and a loss equation, the beta-variational autoencoder comprising an encoder network, a latent layer, and a decoder network, the loss equation comprising a first term associated with reconstruction of an input by the beta-variational autoencoder, a second term associated with regularization of a latent space of the latent layer, and a hyperparameter weight associated with a ratio of the first term and the second term, wherein the latent space has a plurality of dimensions;

determine a respective mean of each respective dimension of the plurality of dimensions, wherein, when determining the respective mean of each respective dimension of the plurality of dimensions, the at least one processor is programmed or configured to:

determine a respective entropy of the respective mean of each respective dimension of the plurality of dimensions;

determine whether each respective dimension of the plurality of dimensions comprises a useful dimension or a deprecated dimension based on the respective mean for the respective dimension, wherein the respective entropy of the respective mean for each useful dimension is higher than a threshold, and wherein the respective entropy of the respective mean for each deprecated dimension is lower than the threshold;

determine a first subset of the plurality of dimensions based on the respective mean for each respective dimension of the plurality of dimensions, wherein the first subset of the plurality of dimensions comprises useful dimensions;

discard a second subset of the plurality of dimensions based on the respective mean of each respective dimension of the plurality of dimensions;

generate the compressed representation of the embedding vector data based on the first subset of dimensions; and

input the compressed representation of the embedding vector data into at least one machine learning model.

8 . The system of claim 7 , wherein, when determining the respective mean of each respective dimension of the plurality of dimensions, the at least one processor is programmed or configured to:

determine a respective entropy of the respective mean of each respective dimension of the plurality of dimensions; and

wherein, when training the beta-variational autoencoder, the at least one processor is programmed or configured to:

iteratively adjust the hyperparameter weight and repeat the training, the determining of the respective entropy of the respective mean, the determining of the first subset, the discarding of the second subset, and the generating of the compressed representation.

9 . The system of claim 7 , wherein the first subset of the plurality of dimensions comprises each useful dimension, and wherein the second subset of the plurality of dimensions comprises each deprecated dimension.

10 . The system of claim 7 , wherein, when determining the respective mean of each respective dimension of the plurality of dimensions, the at least one processor is programmed or configured to:

determine a respective entropy of the respective mean of each respective dimension of the plurality of dimensions; and

wherein, when determining the first subset of the plurality of dimensions, the at least one processor is programmed or configured to:

determine the first subset of the plurality of dimensions based on the respective entropy of the respective mean for each respective dimension of the plurality of dimensions.

11 . The system of claim 10 , wherein, when discarding the second subset of the plurality of dimensions, the at least one processor is programmed or configured to:

discard the second subset of the plurality of dimensions based on the respective entropy of the respective mean of each respective dimension of the plurality of dimensions.

12 . The system of claim 7 , wherein the loss equation is:

ℒ

=

∑

i

=

1

n

(

x

i

-

x

^

i

)

2

+

β

⁢

∑

i

=

1

m

D

KL

(

𝒩

⁡

(

μ

i

,

σ

i

2

)

⁢

𝒩

⁡

(

0

,

1

)

)

wherein the first term associated with reconstruction of an input by the beta-variational autoencoder is:

∑

i

=

1

n

(

x

i

-

x

^

i

)

2

;

wherein the second term associated with regularization of the latent space of the latent layer is:

∑

i

=

1

m

D

KL

(

𝒩

⁡

(

μ

i

,

σ

i

2

)

⁢

𝒩

⁡

(

0

,

1

)

)

;

wherein x i is the input and {circumflex over (x)} i is a reconstruction of the input provided by the beta-variational autoencoder;

wherein D KL is a function for a Kullback-Leibler divergence;

wherein is a function for a normal distribution; and

wherein the hyperparameter weight associated with the ratio of the first term and the second term is β.

13 . A computer program product for generating a compressed representation of embedding vector data, the computer program product comprising at least one non-transitory computer readable medium including one or more instructions that, when executed by at least one processor, cause the at least one processor to:

receive the embedding vector data comprising a plurality of embedding vectors;

train a beta-variational autoencoder based on the embedding vector data and a loss equation, the beta-variational autoencoder comprising an encoder network, a latent layer, and a decoder network, the loss equation comprising a first term associated with reconstruction of an input by the beta-variational autoencoder, a second term associated with regularization of a latent space of the latent layer, and a hyperparameter weight associated with a ratio of the first term and the second term, wherein the latent space has a plurality of dimensions;

determine a respective mean of each respective dimension of the plurality of dimensions, wherein, when determining the respective mean of each respective dimension of the plurality of dimensions, the one or more instructions cause the at least one processor to:

determine a respective entropy of the respective mean of each respective dimension of the plurality of dimensions;

determine whether each respective dimension of the plurality of dimensions comprises a useful dimension or a deprecated dimension based on the respective mean for the respective dimension, wherein the respective entropy of the respective mean for each useful dimension is higher than a threshold, and wherein the respective entropy of the respective mean for each deprecated dimension is lower than the threshold;

determine a first subset of the plurality of dimensions based on the respective mean for each respective dimension of the plurality of dimensions, wherein the first subset of the plurality of dimensions comprises useful dimensions;

discard a second subset of the plurality of dimensions based on the respective mean of each respective dimension of the plurality of dimensions, wherein the second subset of the plurality of dimensions is different than the first subset of the plurality of dimensions, wherein the second subset of the plurality of dimensions comprises deprecated dimensions, and wherein discarding the second subset of the plurality of dimensions reduces a dimensionality of the latent space;

generate the compressed representation of the embedding vector data based on the first subset of dimensions; and

input the compressed representation of the embedding vector data into at least one machine learning model.

14 . The computer program product of claim 13 , wherein, when determining the respective mean of each respective dimension of the plurality of dimensions, the one or more instructions cause the at least one processor to:

determine a respective entropy of the respective mean of each respective dimension of the plurality of dimensions; and

wherein, when training the beta-variational autoencoder, the one or more instructions further cause the at least one processor to:

iteratively adjust the hyperparameter weight and repeat the training, the determining of the respective entropy of the respective mean, the determining of the first subset, the discarding of the second subset, and the generating of the compressed representation.

15 . The computer program product of claim 13 , wherein the first subset of the plurality of dimensions comprises each useful dimension, and wherein the second subset of the plurality of dimensions comprises each deprecated dimension.

16 . The computer program product of claim 13 , wherein, when determining the respective mean of each respective dimension of the plurality of dimensions, the one or more instructions cause the at least one processor to:

determine a respective entropy of the respective mean of each respective dimension of the plurality of dimensions; and

wherein, when determining the first subset of the plurality of dimensions, the one or more instructions cause the at least one processor to:

determine the first subset of the plurality of dimensions based on the respective entropy of the respective mean for each respective dimension of the plurality of dimensions.

17 . The computer program product of claim 16 , wherein, when discarding the second subset of the plurality of dimensions, the one or more instructions cause the at least one processor to:

discard the second subset of the plurality of dimensions based on the respective entropy of the respective mean of each respective dimension of the plurality of dimensions.

18 . The computer program product of claim 13 , wherein the loss equation is:

ℒ

=

∑

i

=

1

n

(

x

i

-

x

^

i

)

2

+

β

⁢

∑

i

=

1

m

D

KL

(

𝒩

⁡

(

μ

i

,

σ

i

2

)

⁢

𝒩

⁡

(

0

,

1

)

)

wherein the first term associated with reconstruction of an input by the beta-variational autoencoder is:

∑

i

=

1

n

(

x

i

-

x

^

i

)

2

;

wherein the second term associated with regularization of the latent space of the latent layer is:

∑

i

=

1

m

D

KL

(

𝒩

⁡

(

μ

i

,

σ

i

2

)

⁢

𝒩

⁡

(

0

,

1

)

)

;

wherein x i is the input and {circumflex over (x)} i is a reconstruction of the input provided by the beta-variational autoencoder;

wherein D KL is a function for a Kullback-Leibler divergence;

wherein is a function for a normal distribution; and

wherein the hyperparameter weight associated with the ratio of the first term and the second term is β.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2024
From: LI, HAOYU; WANG, JUNPENG; WANG, LIANG; ZHENG, YAN; ZHANG, WEI
To: VISA INTERNATIONAL SERVICE ASSOCIATION
Reel/Frame 067324/0454 →
Continuity (3)
Continuation 18027028
Provisional Application 63270321 · Oct 21, 2021
Related Publication 20240289613A1 · Aug 29, 2024
References Cited (39)
US 20170357896A1 · Tsatsin et al. · 2017 [cited by applicant]
US 20180165554A1 · Zhang et al. · 2018 [cited by applicant]
US 20200310370A1 · Bogo et al. · 2020 [cited by applicant]
US 20200395117A1 · Schnorr · 2020 [cited by applicant]
US 20210089884A1 · Macready et al. · 2021 [cited by applicant]
US 20210271980A1 · Polykovskiy et al. · 2021 [cited by applicant]
US 20220147818A1 · Zhang · 2022 [cited by examiner]
US 20220277192A1 · Gou · 2022 [cited by examiner]
US 20220385907A1 · Zhang · 2022 [cited by examiner]
US 20230169325A1 · Xie et al. · 2023 [cited by applicant]
US 20230378976A1 · Mizutani · 2023 [cited by examiner]
CN 112668690A · 2021 [cited by applicant]
CN 112771541A · 2021 [cited by applicant]
CN 113052309A · 2021 [cited by applicant]
Eastwood C, Williams CK. A framework for the quantitative evaluation of disentangled representations. In6th International Conference on Learning Representations May 3, 2018. (Year: 2018). [cited by examiner]
Wei R, Garcia C, El-Sayed A, Peterson V, Mahmood A. Variations in variational autoencoders—a comparative evaluation. Ieee Access. Aug. 20, 2020;8:153651-70. (Year: 2020). [cited by examiner]
Tang Y, Chen D, Li X. Dimensionality reduction methods for brain imaging data analysis. ACM Computing Surveys (CSUR). May 3, 2021;54(4):1-36. (Year: 2021). [cited by examiner]
Dahl, Joakim. “Analysis of the effect of latent dimensions on disentanglement in Variational Autoencoders.” (2021). (Year: 2021). [cited by examiner]
Abdella, A statistical comparative study on image reconstruction and clustering with novel VAE cost function, IEEE Access, 2020, pp. 25626-25637. [cited by applicant]
Burgess et al., “Understanding disentangling in B-VAE”, 31st Conference on Neural Information Processing Systems, arXiv:1804.03599v1, 2017, pp. 1-11. [cited by applicant]
Camacho-Collados et al., “SemEval-2017 Task 2: Multilingual and Cross-lingual Semantic Word Similarity”, Proceedings of the 11th International Workshop on Semantic Evaluations (SemEval-2017), 2017, pp. 15-26. [cited by applicant]
Dewangan et al., “Fault Diagnosis of machines using deep convolutional beta-variational autoencoder”, IEEE Transaction on Artificial Intelligence, 2021, pp. 287-296, vol. 3, No. 2. [cited by applicant]
Higgins et al., “B-VAE: Learning Basic Visual Concepts With a Constrained Variational Framework”, 5th International Conference on Learning Representations, ICLR 2017—Conference Track Proceedings, 2017, pp. 1-13. [cited by applicant]
Higgins et al., “Scan: Learning hierarchical compositional visual concepts”, arXiv:1707.03889, 2018, pp. 1-24. [cited by applicant]
Huang et al., “Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding”, arXiv:1804.07875v1, 2018, pp. 1-18. [cited by applicant]
Inselburg et al., “Parallel Coordinates: A Tool for Visualizing Multi-Dimensional Geometry”, Proceedings of the First IEEE Conference on Visualization, 1990, pp. 361-378. [cited by applicant]
Ji et al., “Visual Exploration of Neural Document Embedding in Information Retrieval: Semantics and Feature Selection”, IEEE Transactions on Visualization and Computer Graphics, 2019, pp. 1-12, vol. 25, No. 6. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, arXiv:1312.6114v10, 2014, pp. 1-14. [cited by applicant]
Li et al., “EmbeddingVis: A Visual Analytics Approach to Comparative Network Embedding Inspection”, IEEE Conference on Visual Analytics Science and Technology (VAST), arXiv:1808.09074v1, 2018, pp. 1-12. [cited by applicant]
Liu et al., “Visual Exploration of Semantic Relationships in Neural Word Embeddings”, IEEE Transactions on Graphics and Visualization, 2017, pp. 553-562. [cited by applicant]
Liu et al., “Latent Space Cartography: Visual Analysis of Vector Space Embeddings”, Eurographics Conference on Visualization (EuroVis), 2019, pp. 67-78, vol. 38, No. 3. [cited by applicant]
Mikolov et al., “Exploiting Similarities among Languages for Machine Translation”, arXiv:1309.4168v1, 2013, pp. 1-10. [cited by applicant]
Mohiuddin et al., “LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-linear Mapping in Latent Space”, arXiv:2004.13889v2, 2020, pp. 1-12. [cited by applicant]
Ruder et al., “A Survey of Cross-lingual Word Embedding Models”, Journal of Artificial Intelligence Research, arXiv:1706.04902v4, 2019, pp. 569-631, vol. 65. [cited by applicant]
Sinha et al., “Variational Autoencoder Anomaly-Detection of Avalanche Deposits in Satellite SAR Imagery”, Association for Computing Machinery, 2020, pp. 113-119. [cited by applicant]
Tang et al., “Dimensionality Reduction Methods for Brain Imaging Data Analysis”, ACM Computing Surveys, 2021, pp. 1-36, vol. 54, No. 4. [cited by applicant]
Ulger et al., “Anomaly Detection for Solder Joints Using B-VAE”, IEEE Transactions on Components, Packaging and Manufacturing Technology, 2021, pp. 2214-2221, vol. 11, No. 12. [cited by applicant]
Voigt et al., “The EU General Data Protection Regulation (GDPR): A Practical Guide”, 2017, 1st Ed., Cham: Springer International Publishing, Cham, Switzerland. [cited by applicant]
Wang et al., “SCANViz: Interpreting the Symbol-Concept Association Captured by Deep Neural Networks through Visual Analytics”, IEEE Pacific Visualization Symposium (PacificVis), 2020, pp. 51-60. [cited by applicant]