IP Library › Granted Patent US 12,254,281
Granted Patent B2
US 12,254,281 · App. 17/842,288 · Granted Mar 18, 2025

Systems and methods for generating improved embeddings while consuming fewer computational resources

Inventor: Anna Darling Goldie (Palo Alto, CA)
Assignee: GOOGLE LLC
G06F40/58G06F18/21375G06F40/216G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,281
App. No.
17/842,288
Granted
Mar 18, 2025
Kind
B2
Abstract

Example aspects of the present disclosure are directed to systems and methods for generation of improved language embeddings (e.g., entity embeddings for natural language tokens) which provide improved model performance. In addition, the proposed techniques require less computational consumption relative to previous approaches.

Claims (48)

1. A computer-implemented method to provide improved embeddings-based model performance with reduced computational consumption, the method comprising:

obtaining, by a computing system comprising one or more computing devices, an input set of entity embeddings, wherein the input set of entity embeddings comprises a plurality of embeddings respectively associated with a plurality of entities included in a vocabulary of entities;

selecting, by the computing system, a subset of entity embeddings based on respective frequencies of appearance of the plurality of embeddings in a corpus, the subset of entity embeddings comprising a subset of the plurality of embeddings respectively associated with a subset of the plurality of entities;

performing, by the computing system, one or more embedding modifications on at least the subset of entity embeddings to produce a modified set of entity embeddings, the one or more embedding modifications comprising one or both of:

subtracting, by the computing system, a mean of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings; and

removing, by the computing system, one or more principal components of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings;

outputting, by the computing system, the modified set of entity embeddings as an output set of entity embeddings; and

using the output set of entity embeddings to predict one or more predicted tokens based on a context.

2. The computer-implemented method of claim 1 , wherein selecting, by the computing system, the subset of entity embeddings based on the frequency of appearance of the plurality of embeddings in the corpus comprises selecting, by the computing system, a percentage of the plurality of embeddings that appear most frequently in the corpus.

3. The computer-implemented method of claim 2 , wherein the percentage is between two to five percent.

4. The computer-implemented method of claim 1 , wherein selecting, by the computing system, the subset of entity embeddings comprises selecting, by the computing system, the subset of entity embeddings to include only entity embeddings that correspond to nouns.

5. The computer-implemented method of claim 1 , wherein selecting, by the computing system, the subset of entity embeddings comprises selecting, by the computing system, the subset of entity embeddings to include only entity embeddings that are included in an expected vocabulary that is different from the vocabulary.

6. The computer-implemented method of claim 1 , wherein the one or more embedding modifications comprise both of:

said subtracting, by the computing system, the mean of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings; and

said removing, by the computing system, the one or more principal components of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings.

7. The computer-implemented method of claim 6 , wherein said subtracting, by the computing system, the mean of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings is performed prior to said removing, by the computing system, the one or more principal components of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings.

8. The computer-implemented method of claim 1 , wherein subtracting, by the computing system, the mean of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings comprises subtracting, by the computing system, the mean of the subset of entity embeddings from each entity embedding included in the input set of entity embeddings.

9. The computer-implemented method of claim 1 , wherein removing, by the computing system, the one or more principal components of the subset of entity embeddings from at least each entity embedding included in the subset of the plurality of embeddings comprises removing, by the computing system, the one or more principal components of the subset of entity embeddings from each entity embedding included in the input set of entity embeddings.

10. The computer-implemented method of claim 1 , wherein the input set of entity embeddings comprise a set of static entity embeddings.

11. The computer-implemented method of claim 1 , wherein the input set of entity embeddings comprise a set of contextual entity embeddings that have been reduced to a set of static entity embeddings.

12. The computer-implemented method of claim 1 , further comprising:

using, by the computing system, the output set of entity embeddings to predict one or more sequences of predicted tokens based on a context.

13. The computer-implemented method of claim 1 , further comprising:

training, by the computing system, a machine-learned language model using the output set of entity embeddings.

14. A computing system configured to provide improved embeddings-based model performance with reduced computational consumption, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing:

a machine-learned language model; and

instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:

obtaining an input set of entity embeddings, wherein the input set of entity embeddings comprises a plurality of embeddings respectively associated with a plurality of entities included in a vocabulary of entities;

selecting a subset of entity embeddings based on respective frequencies of appearance of the plurality of embeddings in a corpus, the subset of entity embeddings comprising a subset of the plurality of embeddings respectively associated with a subset of the plurality of entities;

performing one or more embedding modifications on at least the subset of entity embeddings to produce a modified set of entity embeddings;

outputting the modified set of entity embeddings as the output set of entity embeddings;

training the machine-learned language model using the output set of entity embeddings; and

using the machine-learned language model to predict one or more predicted tokens based on a context.

15. The computing system of claim 14 , wherein selecting the subset of entity embeddings based on the frequency of appearance of the plurality of embeddings in the corpus comprises selecting, by the computing system, a percentage of the plurality of embeddings that appear most frequently in the corpus.

16. The computing system of claim 14 , wherein selecting the subset of entity embeddings comprises selecting the subset of entity embeddings to include only entity embeddings that correspond to nouns.

17. The computing system of claim 14 , wherein selecting the subset of entity embeddings comprises selecting the subset of entity embeddings to include only entity embeddings that are included in an expected vocabulary that is different from the vocabulary.

18. The computing system of claim 14 , wherein the input set of entity embeddings comprise a set of static entity embeddings.

19. The computing system of claim 14 , wherein the operations further comprise using the output set of entity embeddings to predict one or more predicted tokens based on a context.

20. One or more non-transitory, computer-readable media that collectively store:

a machine-learned language model; and

instructions that, when implemented, cause one or more processors to perform operations, the operations comprising:

obtaining an input set of entity embeddings, wherein the input set of entity embeddings comprises a plurality of embeddings respectively associated with a plurality of entities included in a vocabulary of entities;

selecting a subset of entity embeddings based on respective frequencies of appearance of the plurality of embeddings in a corpus, the subset of entity embeddings comprising a subset of the plurality of embeddings respectively associated with a subset of the plurality of entities;

performing one or more embedding modifications on at least the subset of entity embeddings to produce a modified set of entity embeddings;

outputting the modified set of entity embeddings as an output set of entity embeddings; and

using the output set of entity embeddings to predict one or more predicted tokens based on a context.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2022
From: GOLDIE, ANNA DARLING
To: GOOGLE LLC
Reel/Frame 060352/0621 →
Continuity (2)
Provisional Application 63211233 · Jun 16, 2021
Related Publication 20220405493A1 · Dec 22, 2022
References Cited (29)
US 20200410157A1 · van de Kerkhof · 2020 [cited by examiner]
US 20210027157A1 · Chen · 2021 [cited by examiner]
US 20210042800A1 · Chandra · 2021 [cited by examiner]
US 20210174020A1 · Sohn · 2021 [cited by examiner]
US 20210241309A1 · Wolf, Jr. · 2021 [cited by examiner]
US 20220261545A1 · Lauber · 2022 [cited by examiner]
US 20230214177A1 · Zheng · 2023 [cited by examiner]
US 20240232572A1 · Wang · 2024 [cited by examiner]
WO WO2020091829 · 2020 [cited by examiner]
Adiwardana et al, “Towards a Human-Like Open-Domain Chatbot”, arXiv:2001.09977v3, Feb. 27, 2020, 38 pages. [cited by applicant]
Arora et al, “A Simple but Tough-to-Beat Baseline for Sentence Embeddings”, International Conference on Learning Representations, Apr. 24-26, 2017, Toulon, France, 16 pages. [cited by applicant]
Baroni et al, “The WaCky Wide Web: A Collection of Very Large Linguistically Processed Web-Crawled Corpora”, Language Resources and Evaluation, 2009, 22 pages. [cited by applicant]
Brants et al, “Web It 5-gram version 1 ldc2006t13”, Linguistic Data Consortium, https://catalog.ldc.upenn.edu/LDC2006T13, retrieved on Aug. 16, 2022, 2 pages. [cited by applicant]
Brown et al, “Language Models are Few-Shot Learners” arXiv:2005.14165v4, Jul. 22, 2020, 75 pages. [cited by applicant]
Bruni et al, “Multimodal Distributional Semantics”, Journal of Artificial Intelligence Research, vol. 49, 2014, 47 pages. [cited by applicant]
Ethayarajh, “How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings”, arXiv:1909.00512v1, Sep. 2, 2019, 11 pages. [cited by applicant]
Finkelstein et al, “Placing Search in Context: The Concept Revisited”, Transactions on Information Systems, vol. 20, No. 1, pp. 116-131. [cited by applicant]
Gerz et al, “SimVerb-3500: A Large-Scale Evaluation Set of Verb Similarity”, arXiv:1608.00869v4, Sep. 20, 2016, 12 pages. [cited by applicant]
Hill et al, SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation. Computational Linguistics, vol. 41, No. 4, pp. 665-695. [cited by applicant]
Jacob Devlin et al, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv:1810.04805v2, May 24, 2019, 16 pages. [cited by applicant]
Luong et al, “Better Word Representations with Recursive Neural Networks for Morphology”, Conference on Computational Natural Language Learning, 2013, 10 pages. [cited by applicant]
Mikolov et al, “Efficient Estimation of Word Representations in Vector Space”, arXiv:1301.3781v3, Sep. 7, 2013, 12 pages. [cited by applicant]
Mu et al., “All-but-the-Top: Simple and Effective Postprocessing for Word Representations”, arXiv:1702.01417v2, Mar. 19. 2018. 25 pages. [cited by applicant]
Pennington et al., “GloVe: Global Vectors for Word Representation”, Conference on Empirical Methods in Natural Language Processing, Oct. 2014, Doha, Qatar, 12 pages. [cited by applicant]
Radford et al., “Language Models are Unsupervised Multitask Learners”, 2019, 24 pages. [cited by applicant]
Rishi Bommasani et al, “Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings”, Meeting of the Association for Computational Linguistics, Jul. 5-10, 2020, Seattle, Washington, United… [cited by applicant]
Rubenstein et al., “Contextual Correlates of Synonymy”, Communications of the ACM, vol. 8, No. 10, 1965, pp. 627-633. [cited by applicant]
Sahlgren et al, “The Gavagai Living Lexicon”, International Conference on Language Resources and Evaluation, May 2016, Portoroz, Slovenia, pp. 344-350. [cited by applicant]
Wolf et al, “Transformers: State-of-the-Art Natural Language Processing”, arXiv:1910.03771v5, Jul. 14, 2020, 8 pages. [cited by applicant]