IP Library › Granted Patent US 12,321,712
Granted Patent B2
US 12,321,712 · App. 18/530,710 · Granted Jun 3, 2025

Method, system, and computer program product for normalizing embeddings for cross-embedding alignment

Inventors: Yan Zheng (Los Gatos, CA); Michael Yeh (Newark, CA); Junpeng Wang (Santa Clara, CA); Wei Zhang (Fremont, CA); Liang Wang (San Jose, CA); Hao Yang (San Jose, CA); Prince Osei Aboagye (Salt Lake City, UT)
Assignee: Visa International Service Association
G06F5/01G06F17/16G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,712
App. No.
18/530,710
Granted
Jun 3, 2025
Kind
B2
Abstract

Provided is a method for normalizing embeddings for cross-embedding alignment. The method may include applying mean centering to the at least one embedding set, applying spectral normalization to the at least one embedding set, and/or applying length normalization to the at least one embedding set. Spectral normalization may include decomposing the at least one embedding set, determining an average singular value of the at least one embedding set, determining a respective substitute singular value for each respective singular value of a diagonal matrix, and/or replacing the at least one embedding set with a product of the at least one embedding set, a right singular vector, and an inverse of the substitute diagonal matrix. The mean centering, spectral normalization, and/or length normalization may be iteratively repeated for a configurable number of iterations. A system and computer program product are also disclosed.

Claims (53)

1. A computer-implemented method, comprising:

receiving, with at least one processor, at least one embedding set;

applying, with at least one processor, mean centering to the at least one embedding set;

applying, with at least one processor, spectral normalization to the at least one embedding set, wherein applying spectral normalization to the at least one embedding set comprises:

decomposing the at least one embedding set to provide a left singular vector, a right singular vector, and a diagonal matrix;

for each respective singular value of the diagonal matrix, if the respective singular value is greater than a configurable multiple of an average singular value of the at least one embedding set, determining a respective substitute singular value based on the respective singular value and the configurable multiple of the average singular value or, if the respective singular value is not greater than the configurable multiple of the average singular value, determining the respective substitute singular value to be 1, wherein a substitute diagonal matrix comprises the respective substitute singular value for each respective singular value of the diagonal matrix; and

replacing the at least one embedding set with at least one replacement embedding set based on the at least one embedding set, the right singular vector, and the substitute diagonal matrix; and

applying, with at least one processor, length normalization to the at least one embedding set.

2. The method of claim 1 , wherein each embedding set of the at least one embedding set comprises a set of embedding vectors, and wherein applying mean centering comprises:

determining, with at least one processor, a mean based on all embedding vectors of the set of embedding vectors; and

subtracting, with at least one processor, the mean from each embedding vector of the set of embedding vectors.

3. The method of claim 1 , wherein decomposing the at least one embedding set comprises performing singular value decomposition on the at least one embedding set.

4. The method of claim 1 , wherein the average singular value is determined based on a square root of an average squared singular value.

5. The method of claim 1 , wherein each embedding set of the at least one embedding set comprises a set of embedding vectors, and wherein applying length normalization comprises:

adjusting, with at least one processor, each embedding vector of the set of embedding vectors to have a 2-norm of 1.

6. The method of claim 1 , further comprising:

iteratively repeating, with at least one processor, applying mean centering, applying spectral normalization, and applying length normalization to the at least one embedding set for a configurable number of iterations.

7. The method of claim 1 , wherein the at least one embedding set comprises a first embedding set and a second embedding set, the method further comprising:

aligning, with at least one processor, the first embedding set with the second embedding set.

8. The method of claim 1 , wherein the at least one embedding set comprises a first language embedding set and a second language embedding set, the first language embedding set comprising a first set of word embedding vectors for a first language, the second language embedding set comprising a second set of word embedding vectors for a second language.

9. The method of claim 1 , wherein the at least one embedding set comprises a first embedding set representing an entity in a first embedding space associated with a first time period and a second embedding set representing the entity in a second embedding space associated with a second time period different than the first time period.

10. The method of claim 9 , wherein the entity comprises at least one of a merchant, a customer, an issuer, an acquirer, or a payment gateway.

11. A system, comprising:

at least one processor configured to:

receive at least one embedding set;

apply mean centering to the at least one embedding set;

apply spectral normalization to the at least one embedding set, wherein applying spectral normalization to the at least one embedding set comprises:

decomposing the at least one embedding set to provide a left singular vector, a right singular vector, and a diagonal matrix;

for each respective singular value of the diagonal matrix, if the respective singular value is greater than a configurable multiple of an average singular value of the at least one embedding set, determining a respective substitute singular value based on the respective singular value and the configurable multiple of the average singular value or, if the respective singular value is not greater than the configurable multiple of the average singular value, determining the respective substitute singular value to be 1, wherein a substitute diagonal matrix comprises the respective substitute singular value for each respective singular value of the diagonal matrix; and

replacing the at least one embedding set with at least one replacement embedding set based on the at least one embedding set, the right singular vector, and the substitute diagonal matrix; and

apply length normalization to the at least one embedding set.

12. The system of claim 11 , wherein each embedding set of the at least one embedding set comprises a set of embedding vectors, and wherein applying mean centering comprises:

determining a mean based on all embedding vectors of the set of embedding vectors; and

subtracting the mean from each embedding vector of the set of embedding vectors.

13. The system of claim 11 , wherein decomposing the at least one embedding set comprises performing singular value decomposition on the at least one embedding set.

14. The system of claim 11 , wherein the average singular value is determined based on a square root of an average squared singular value.

15. The system of claim 11 , wherein each embedding set of the at least one embedding set comprises a set of embedding vectors, and wherein applying length normalization comprises:

adjusting each embedding vector of the set of embedding vectors to have a 2-norm of 1.

16. The system of claim 11 , wherein the instructions, when executed by the at least one processor, further direct the at least one processor to:

iteratively repeat applying mean centering, applying spectral normalization, and applying length normalization to the at least one embedding set for a configurable number of iterations.

17. The system of claim 11 , wherein the at least one embedding set comprises a first embedding set and a second embedding set, wherein the instructions, when executed by the at least one processor, further direct the at least one processor to:

align the first embedding set with the second embedding set.

18. The system of claim 11 , wherein the at least one embedding set comprises a first language embedding set and a second language embedding set, the first language embedding set comprising a first set of word embedding vectors for a first language, the second language embedding set comprising a second set of word embedding vectors for a second language.

19. The system of claim 11 , wherein the at least one embedding set comprises a first embedding set representing an entity in a first embedding space associated with a first time period and a second embedding set representing the entity in a second embedding space associated with a second time period different than the first time period, and

wherein the entity comprises at least one of a merchant, a customer, an issuer, an acquirer, or a payment gateway.

20. At least one non-transitory computer-readable medium including one or more instructions that, when executed by at least one processor, cause the at least one processor to:

receive at least one embedding set;

apply mean centering to the at least one embedding set;

apply spectral normalization to the at least one embedding set, wherein applying spectral normalization to the at least one embedding set comprises:

decomposing the at least one embedding set to provide a left singular vector, a right singular vector, and a diagonal matrix;

for each respective singular value of the diagonal matrix, if the respective singular value is greater than a configurable multiple of an average singular value of the at least one embedding set, determining a respective substitute singular value based on the respective singular value and the configurable multiple of the average singular value or, if the respective singular value is not greater than the configurable multiple of the average singular value, determining the respective substitute singular value to be 1, wherein a substitute diagonal matrix comprises the respective substitute singular value for each respective singular value of the diagonal matrix; and

replacing the at least one embedding set with at least one replacement embedding set based on the at least one embedding set, the right singular vector, and the substitute diagonal matrix; and

apply length normalization to the at least one embedding set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2023
From: ZHENG, YAN; YEH, MICHAEL; WANG, JUNPENG; ZHANG, WEI; WANG, LIANG; YANG, HAO; ABOAGYE, PRINCE OSEI
To: VISA INTERNATIONAL SERVICE ASSOCIATION
Reel/Frame 065806/0428 →
Continuity (3)
Continuation 18006649
Provisional Application 63192779 · May 25, 2021
Related Publication 20240134599A1 · Apr 25, 2024
References Cited (86)
US 10671697B1 · Batruni · 2020 [cited by applicant]
US 20180330729A1 · Golipour et al. · 2018 [cited by applicant]
US 20190355346A1 · Bellegarda · 2019 [cited by applicant]
US 20200134473A1 · Miyato · 2020 [cited by applicant]
US 20210065260A1 · Zheng et al. · 2021 [cited by applicant]
US 20210109951A1 · Yeh et al. · 2021 [cited by applicant]
CN 111047546A · 2020 [cited by applicant]
CN 112732921A · 2021 [cited by applicant]
JP 2020107199A · 2020 [cited by applicant]
Nivre et al., “Universal Dependencies 2.2: LINDAT/CLARIAH-CZ”, Digital Library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. Retrieved from http://hdl… [cited by applicant]
Ormazabal et al., “Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context Anchoring”, Proceedings of the 59th annual meeting of the Association for Computational Linguistics, 2021, pp. 6479-6489,… [cited by applicant]
Overton, “A Quadratically Convergent Method For Minimizing a Sum of Euclidean Norms”, Mathematical Programming, 1983, pp. 34-63, vol. 27(1), North-Holland. [cited by applicant]
Patra et al. “Bliss Lexicon Induction with Semi-supervision in Non-Isometric Embedding Spaces”, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2018, pp. 184-193, Associatio… [cited by applicant]
Pennington et al., “GloVe: Global Vectors for Word Representation”, Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Oct. 2014, pp. 1532-1543, Association for Computational… [cited by applicant]
Perozzi et al., “DeepWalk: Online Learning of Social Representations”, Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, 2014, pp. 701-710, New York, NY. [cited by applicant]
Radinsky et al., “A Word at a Time: Computing Word Relatedness using Temporal Semantic Analysis”, Proceedings of the 20th International Conference on World Wide Web, 2011, pp. 337-346, Hyderabad, India. [cited by applicant]
Rubenstein et al., “Computational Linguistics: Contextual Correlates of Synonymy”, Communications of the ACM, 1965, pp. 627-633, vol. 8(10). [cited by applicant]
Ruder et al., “A Discriminative Latent-Variable Model for Bilingual Lexicon Induction”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, arXiv:1808.09334v2, Oct.-Nov. 2018, pp. 458… [cited by applicant]
Ruder et al., “A Survey of Cross-lingual Word Embedding Models”, Journal of Artificial Intelligence Research, 2019, pp. 569-631, vol. 65. [cited by applicant]
Ruder et al., “Unsupervised Cross-Lingual Representation Learning”, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts, Jul. 2019, pp. 31-38, Association for Comp… [cited by applicant]
Sachidanada et al., “Filtered Inner Product Projection for Crosslingual Embedding Alignment”, ICLR 2021: The 9th International Conference on Learning Representations, arXiv:2006.03652v2, 2021, pp. 1-26. [cited by applicant]
Safaya et al., “Kuisail at SemEval-2020 Task 12: BERT-CNN for Offensive Speech Identification in Social Media”, Proceedings of the 14th International Workshop on Semantic Evaluation, Dec. 2020, pp. 2054-2059, Barcelona,… [cited by applicant]
Smith et al., “Offline Bilingual Word Vectors, Orthogonal Transformations and the Inverted Softmax”, ICLR (Poster), arXiv: 1702.03859v1, 2017, pp. 1-10. [cited by applicant]
Sogaard et al., “On the Limitations of Unsupervised Bilingual Dictionary Induction”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Jul. 2018, pp. 778-788, (vol. 1: Long Papers)… [cited by applicant]
Vardi et al., “The multivariate l1-median and associated data depth”, Proceedings of the National Academy of Sciences, 2000, pp. 1423-1426, vol. 97(4). [cited by applicant]
Vulic et al., “Monolingual and Cross-Lingual Information Retrieval Models Based on (Bilingual) Word Embeddings”, Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retr… [cited by applicant]
Vulic et al., “Do We Really Need Fully Unsupervised Cross-Lingual Embeddings?”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natura… [cited by applicant]
Wang et al., “Constrained Non-Affine Alignment of Embeddings”, Proceedings of the International Conference on Data Mining (ICDM), arXiv:1910.05862v4, 2021, pp. 1-9. [cited by applicant]
Wang et al., “Cross-Lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework”, Published as a conference paper at ICLR, arXiv: 19010.04708v4, 2020. Retrieved from http://arxiv.org/abs/1910… [cited by applicant]
Williams et al., “A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference”, Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human… [cited by applicant]
Wolf et al., “Transformers: State-of-the-Art Natural Language Processing”, arXiv:1910.03771, 2019, pp. 1-8. [cited by applicant]
Xing et al., “Normalized Word Embedding and Orthogonal Transform for Bilingual Word Translation”, Human Language Technologies: The 2015 Annual Conference of the North American Chapter of the ACL, May-Jun. 2015, pp. 1006… [cited by applicant]
Xu et al., “Cross-Lingual BERT Contextual Embedding Space Mapping with Isotropic and Isometric Conditions”, arXiv:2107.09186, 2021, pp. 1-10. [cited by applicant]
Yang et al., “Verb Similarity on the Taxonomy of WordNet”, Proceedings of the Third International WorldNet Conference GWC, 2006, pp. 121-128, South Jeju Island, Korea, Masaryk University. [cited by applicant]
Zhang et al., “Are Girls Neko or Shojo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization”, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2… [cited by applicant]
Agirre et al., “A Study on Similarity and Relatedness Using Distributional and WordNet-based Approaches”, Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Assoc… [cited by applicant]
Ahmad et al., “On Difficulties of Cross-Lingual Transfer with Order Differences: A Case Study on Dependency Parsing”, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational… [cited by applicant]
Artetxe et al., “Learning principled bilingual mappings of word embeddings while preserving monolingual invariance”, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Nov. 2016, pp.… [cited by applicant]
Artetxe et al., “A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, arXiv:1805.06297… [cited by applicant]
Bhattacharya et al., “Deep Speaker Embeddings for Short-Duration Speaker Verification”, Interspeech, 2017, pp. 1517-1521. [cited by applicant]
Bojanowski et al., “Enriching Word Vectors with Subword Information”, Transactions of the Association for Computational Linguistics, 2017, pp. 135-146, vol. 5. doi: 10.1162/tacl a 00051. Retrieved from https://www.aclwe… [cited by applicant]
Bommasani et al., “On the Opportunities and Risks of Foundation Models”, Center for Research on Foundation Models (CRFM), arXiv:2108.07258, 2021, pp. 1-214. [cited by applicant]
Bruni et al., “Distributional Semantics in Technicolor”, Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics, Jul. 2012, pp. 136-145, (vol. 1: Long Papers), Association for Computatio… [cited by applicant]
Cao et al., “Unsupervised Topological Alignment for Single-Cell Multi-Omics Integration”, bioRxiv, 2020, pp. 1-17. doi: 10.1101/2020.02.02.931394. Retrieved from https://www.biorxiv.org/content/early/2020/02/03/2020.02.… [cited by applicant]
Chatelon et al., “A Subgradient Algorithm For Certain Minimax and Minisum Problems”, Mathematical Programming, 1978, pp. 1-39, Research Report No. 77-1. [cited by applicant]
Chen et al., “Enhanced LSTM for Natural Language Inference”, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Jul. 2017, pp. 1657-1668, vol. 1, Association for Computational Lingu… [cited by applicant]
Chen et al., “High-throughput sequencing of the transcriptome and chromatin accessibility in the same cell”, Nature Biotechnology, 2019, pp. 1452-1457, vol. 37(12). [cited by applicant]
Cheow et al., “Single-cell multimodal profiling reveals cellular epigenetic heterogeneity”, Nature Methods, Aug. 2016, pp. 833-836, vol. 13 (10). [cited by applicant]
Conneau et al., “Word Translation Without Parallel Data”, Published as a conference paper at ICLR, arXiv: 1710.04087v3, 2018, pp. 1-14. [cited by applicant]
Conneau et al., “XNLI: Evaluating Cross-lingual Sentence Representations”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct.-Nov. 2018, pp. 2475-2485, Association for Computati… [cited by applicant]
Demetci et al., “Gromov-Wasserstein optimal transport to align single-cell multi-omics data”, bioRxiv, 2020, pp. 1-18. [cited by applicant]
Dev et al., “Absolute Orientation for Word Embedding Alignment”, Knowledge and Information Systems, arXiv: 1806.01330v2, 2021, pp. 1-19. [cited by applicant]
Dorrie, “100 Great Problems of Elementary Mathematics: Their History and Solution”, 1965, pp. 1-402, Dover Publications, Inc., New York, NY. [cited by applicant]
Dubossarsky et al., “The Secret is in the Spectra: Predicting Cross-lingual Task Performance with Spectral Similarity Measures”, EMNLP, arXiv:2001.11136v2, 2020, pp. 1-14. [cited by applicant]
Dyer et al., “A Simple, Fast, and Effective Reparameterization of IBM Model 2”, Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog… [cited by applicant]
Ethayarajh, “How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing an… [cited by applicant]
Eyster et al., “On Solving Multifacility Location Problems using a Hyperboloid Approximation Procedure”, AIIE Transactions, 1973, pp. 1-6, vol. 5(1). [cited by applicant]
Faruqui et al., “Improving Vector Space Word Representations Using Multilingual Correlation”, Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, Apr. 2014, pp. 4… [cited by applicant]
Finkelstein et al., “Placing Search in Context: The Concept Revisited”, ACM Transactions on Information Systems, Jan. 2002, pp. 116-131, vol. 20:1. [cited by applicant]
Frome et al., “DeViSE: A Deep Visual-Semantic Embedding Model”, 2013, pp. 1-9. [cited by applicant]
Glavas et al., “How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some Misconceptions”, Proceedings of the 57th Annual Meeting of the Association for Computational … [cited by applicant]
Gouws et al., “Fast Bilingual Distributed Representations without Word Alignments”, Proceedings of the 32nd International Conference on Machine Learning, Jul. 2015, pp. 748-776, vol. 37, Lille, France. Retrieved from ht… [cited by applicant]
Grave et al., “Unsupervised Alignment of Embeddings with Wasserstein Procrustes”, Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS), 2019, pp. 1880-1890, vol. 89. [cited by applicant]
Grover et al., “node2vec: Scalable Feature Learning for Networks”, Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining, 2016, pp. 855-864. [cited by applicant]
Guo et al., “Cross-Lingual Dependency Parsing Based on Distributed Representations”, Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on … [cited by applicant]
Halawi et al., “Large-Scale Learning of Word Relatedness with Constraints”, Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining, 2012, pp. 1406-1414. [cited by applicant]
Hartmann et al., “Comparing Unsupervised Word Translation Methods Step by Step”, 33rd Conference on Neural Information Processing Systems, 2019, vol. 32, Vancouver, Canada. Retrieved from https://proceedings.neurips.cc/… [cited by applicant]
Hermann et al., “Multilingual Models for Compositional Distributed Semantics”, Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, Jun. 2014, pp. 58-68, (vol. 1: Long Papers), Associ… [cited by applicant]
Heyman et al., “Bilingual Lexicon Induction by Learning to Combine Word-Level and Character-Level Representations”, Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguis… [cited by applicant]
Jenkins et al., “Unsupervised Representation Learning of Spatial Data via Multimodal Embedding”, Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Nov. 2019, pp. 1993-2002, As… [cited by applicant]
Joulin et al., “Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, arXiv: 1804.07745v3, Oct.-Nov. 20… [cited by applicant]
Karan et al., “Classification-Based Self-Learning for Weakly Supervised Bilingual Lexicon Induction”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, pp. 6915-6922, As… [cited by applicant]
Kiela et al., “Learning Image Embeddings using Convolutional Neural Networks for Improved Multi-Modal Semantics” Proceedings of the 2014 Conference on empirical methods in natural language processing (EMNLP), 2014, pp. … [cited by applicant]
Klementiev et al., “Inducing Crosslingual Distributed Representations of Words”, Proceedings of COLING 2012: Technical Papers, Dec. 2012, pp. 1459-1474, COLING 2012, Mumbai, India. Retrieved from https://www.aclweb.org/… [cited by applicant]
Lample et al., “Phrase-Based & Neural Unsupervised Machine Translation”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct.-Nov. 2018, pp. 5039-5049, Association for Computation… [cited by applicant]
Liu et al., “Jointly embedding multiple single-cell omics measurements”, bioRxiv, 2019, pp. 1-13. doi: 10.1101/644310. Retrieved from https://www.biorxiv.org/content/early/2019/05/20/644310. [cited by applicant]
Love et al., “The Nature of Facilities Location Problems”, Facilities Location: Models and Methods, 1988, pp. 7-10, Elsevier Science Publishing Co., New York, NY. [cited by applicant]
Luong et al., “Better Word Representations with Recursive Neural Networks for Morphology”, Proceedings of the seventeenth conference on computational natural language learning, 2013, pp. 104-113. [cited by applicant]
Mikolov et al., “Efficient Estimation of Word Representations in Vector Space”, ICLR (Workshop Poster), arXiv:1301.3781v3, 2013, pp. 1-12. [cited by applicant]
Mikolov et al., “Exploiting Similarities among Languages for Machine Translation”, arXiv:1309.4168, 2013, pp. 1-10. [cited by applicant]
Mikolov et al., “Linguistic Regularities in Continuous Space Word Representations”, Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn… [cited by applicant]
Miller et al., “Contextual correlates of semantic similarity”, Language and Cognitive Processes, 1991, pp. 1-28, vol. 6(1). [cited by applicant]
Miyato et al., “Spectral Normalization for Generative Adversarial Networks”, article accessed at arXiv: 1802.05957, 2018, pp. 1-26. [cited by applicant]
Mu et al., “All-But-The-Top: Simple and Effective Post-Processing for Word Representations”, 6th International Conference on Learning Representations, arXiv:1702.01417v2, 2018, pp. 1-25. [cited by applicant]
Nakashole et al., “Characterizing Departures from Linearity in Word Translation”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018, pp. 221-227, vol. 2, Association for Compu… [cited by applicant]
Gogianu et al., “Spectral Normalisation for Deep Reinforcement Learning: An Optimisation Perspective”, Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 2021, pp. 1-24. [cited by applicant]