IP Library › Granted Patent US 12,737,341
Granted Patent B2
US 12,737,341 · App. 19/029,909 · Granted Sep 15, 2026

Efficient embedding table storage and lookup

Inventor: Gaurav Menghani (Santa Clara, CA)
Assignee: GOOGLE LLC
G06F16/2282G06F16/215G06F16/2255G06F16/3347G06F30/27G06N3/04G06N3/08G06F16/1744G06F2211/007G06F2211/1014H03M7/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,341
App. No.
19/029,909
Filed
Jan 17, 2025
Granted
Sep 15, 2026
Kind
B2
Art Unit
2166
USPC
707/693
Abstract

The present disclosure provides systems, methods, and computer program products for providing efficient embedding table storage and lookup in machine-learning models. A computer-implemented method may include obtaining an embedding table comprising a plurality of embeddings respectively associated with a corresponding index of the embedding table, compressing each particular embedding of the embedding table individually allowing each respective embedding of the embedding table to be decompressed independent of any other embedding in the embedding table, packing the embedding table comprising individually compressed embeddings with a machine-learning model, receiving an input to use for locating an embedding in the embedding table, determining a lookup value based on the input to search indexes of the embedding table, locating the embedding based on searching the indexes of the embedding table for the determined lookup value, and decompressing the located embedding independent of any other embedding in the embedding table.

Claims (42)

1 . A computer-implemented method, comprising:

receiving, by a computing system comprising one or more processors and a memory, an input for processing using a machine-learned model;

accessing, by the computing system, an embedding table associated with the machine-learned model, the embedding table comprising a plurality of embeddings;

locating, by the computing system based on one or more lookup values associated with the input, one or more embeddings in the embedding table that respectively correspond to the one or more lookup values;

loading, by the computing system and into the memory, and based on the locating of the one or more embeddings, the one or more embeddings without loading the entirety of the embedding table into the memory; and

processing, by the computing system, the one or more embeddings in the memory using the machine-learned model.

2 . The computer-implemented method of claim 1 , wherein the embedding table is associated with a plurality of items included in a vocabulary for the machine-learned model.

3 . The computer-implemented method of claim 2 , wherein the plurality of items comprise a plurality of tokens.

4 . The computer-implemented method of claim 3 , wherein the machine-learned model is configured to perform natural language processing.

5 . The computer-implemented method of claim 1 , wherein the one or more embeddings are embeddings compressed using pruning.

6 . The computer-implemented method of claim 1 , wherein the one or more lookup values correspond to one or more index values that index the one or more embeddings in the embedding table.

7 . The computer-implemented method of claim 1 , comprising:

loading, by the computing system and into the memory, the one or more embeddings without loading, into the memory, any other embedding that is not related to the input.

8 . The computer-implemented method of claim 1 , comprising:

accessing the one or more embeddings using an index value associated with a logical storage unit storing at least one of the one or more embeddings.

9 . The computer-implemented method of claim 1 , wherein the one or more embeddings correspond to features of word content that are mapped to vectors of real numbers used by the machine-learned model to generate a prediction output.

10 . A computing system, comprising:

one or more processors; and

a memory comprising one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising:

receiving an input for processing using a machine-learned model;

accessing an embedding table associated with the machine-learned model, the embedding table comprising a plurality of embeddings;

locating, based on one or more lookup values associated with the input, one or more embeddings in the embedding table that respectively correspond to the one or more lookup values;

loading, into the memory, and based on the locating of the one or more embeddings, the one or more embeddings without loading the entirety of the embedding table into the memory; and

processing the one or more embeddings in the memory using the machine-learned model.

11 . The computing system of claim 10 , wherein the embedding table is associated with a plurality of items included in a vocabulary for the machine-learned model.

12 . The computing system of claim 11 , wherein the plurality of items comprise a plurality of tokens.

13 . The computing system of claim 12 , wherein the machine-learned model is configured to perform natural language processing.

14 . The computing system of claim 10 , wherein the one or more lookup values correspond to one or more index values that index the one or more embeddings in the embedding table.

15 . The computing system of claim 10 , the operations comprising:

loading, by the computing system and into the memory, the one or more embeddings without loading, into the memory, any other embedding that is not related to the input.

16 . The computing system of claim 10 , the operations comprising:

accessing the one or more embeddings using an index value associated with a logical storage unit storing at least one of the one or more embeddings.

17 . The computing system of claim 10 , wherein the one or more embeddings correspond to features of word content that are mapped to vectors of real numbers used by the machine-learned model to generate a prediction output.

18 . One or more non-transitory computer-readable media that store instructions that, when executed by one or more processors, cause a computing system to perform operations comprising:

receiving an input for processing using a machine-learned model;

accessing an embedding table associated with the machine-learned model, the embedding table comprising a plurality of embeddings;

locating, based on one or more lookup values associated with the input, one or more embeddings in the embedding table that respectively correspond to the one or more lookup values;

loading, into a memory of the computing system, and based on the locating of the one or more embeddings, the one or more embeddings without loading the entirety of the embedding table into the memory; and

processing the one or more embeddings in the memory using the machine-learned model.

19 . The one or more non-transitory computer-readable media of claim 18 , the operations comprising:

accessing the one or more embeddings using an index value associated with a logical storage unit storing at least one of the one or more embeddings.

20 . The one or more non-transitory computer-readable media of claim 18 , wherein the one or more embeddings correspond to features of word content that are mapped to vectors of real numbers used by the machine-learned model to generate a prediction output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2025
From: MENGHANI, GAURAV
To: GOOGLE LLC
Reel/Frame 071445/0137 →
Continuity (4)
Continuation 18390524 · Dec 20, 2023
Continuation 18161352 · Jan 30, 2023
Division 17147844 · Jan 13, 2021
Related Publication 20250217342A1 · Jul 3, 2025
References Cited (29)
US 8201054B2 · Slyz et al. · 2012 [cited by applicant]
US 9503123B1 · Pinho et al. · 2016 [cited by applicant]
US 9977807B1 · Bowman · 2018 [cited by examiner]
US 10872601B1 · Acharya · 2020 [cited by applicant]
US 20050276496A1 · Molgaard et al. · 2005 [cited by applicant]
US 20180089261A1 · Li · 2018 [cited by examiner]
US 20180349282A1 · Brahm et al. · 2018 [cited by applicant]
US 20190384814A1 · Mcgoldrick · 2019 [cited by examiner]
US 20200004836A1 · Rangarajan Sridhar · 2020 [cited by examiner]
US 20200107072A1 · Lomada et al. · 2020 [cited by applicant]
US 20200201908A1 · Ma et al. · 2020 [cited by applicant]
US 20210034701A1 · Fei et al. · 2021 [cited by applicant]
US 20210158176A1 · Wan et al. · 2021 [cited by applicant]
US 20210240925A1 · Kwon et al. · 2021 [cited by applicant]
US 20210311923A1 · Gohad · 2021 [cited by examiner]
US 20210365772A1 · Zhang et al. · 2021 [cited by applicant]
US 20210365807A1 · Ramsl · 2021 [cited by applicant]
US 20220075843A1 · Kumar et al. · 2022 [cited by applicant]
US 20220222437A1 · Lauber · 2022 [cited by examiner]
CN 108415888 · 2018 [cited by applicant]
WO WO2021061159A1 · 2021 [cited by examiner]
WO WO2021064433 · 2021 [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2021/064355, mailed on Jul. 27, 2023, 9 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2021/064355, mailed Mar. 31, 2022, 15 pages. [cited by applicant]
Kim et al., “Implementing Efficient Data Compression an Encryption in a Persistent Key-Value Store for HPC”, International Journal of High Performance Computing Applications, vol. 33, No. 6, 2019, pp. 1098-1112. [cited by applicant]
Larson et al., “SQL Server Column Store Indexes”, SIGMOD '11, 2011 Association for Computing Machinery SIGMOD International Conference on Management of Data, Athens, Greece, Jun. 12-16, 2011, pp. 1177-1184. [cited by applicant]
Yang et al., “Mixed-Precision Embedding Using a Cache”, arXiv:2010.11305, https://doi.org/10.48550/arXiv.2010.11305, 14 pages. [cited by applicant]
Chinese Search Report Corresponding to Application No. 202180082619 on Nov. 10, 2025. [cited by applicant]
Guan et al., “Post-Training 4-bit Quantization on Embedding Tables”, arXiv:1911.02079v1, Nov. 2019, pp. 1-11. [cited by applicant]