IP Library › Granted Patent US 11,892,998
Granted Patent B2
US 11,892,998 · App. 18/161,352 · Granted Feb 6, 2024

Efficient embedding table storage and lookup

Inventor: Gaurav Menghani (Santa Clara, CA)
Assignee: GOOGLE LLC
G06F16/2282G06F16/215G06F16/2255G06F16/3347G06F30/27G06N3/04G06N3/08G06F16/1744G06F2211/007G06F2211/1014H03M7/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,892,998
App. No.
18/161,352
Granted
Feb 6, 2024
Kind
B2
Abstract

The present disclosure provides systems, methods, and computer program products for providing efficient embedding table storage and lookup in machine-learning models. A computer-implemented method may include obtaining an embedding table comprising a plurality of embeddings respectively associated with a corresponding index of the embedding table, compressing each particular embedding of the embedding table individually allowing each respective embedding of the embedding table to be decompressed independent of any other embedding in the embedding table, packing the embedding table comprising individually compressed embeddings with a machine-learning model, receiving an input to use for locating an embedding in the embedding table, determining a lookup value based on the input to search indexes of the embedding table, locating the embedding based on searching the indexes of the embedding table for the determined lookup value, and decompressing the located embedding independent of any other embedding in the embedding table.

Claims (28)

1. A computer-implemented method for performing efficient embedding table storage and lookup in machine-learning models, comprising:

obtaining, by one or more processors, an embedding table associated with a machine-learning model, the embedding table comprising a plurality of individually compressed embeddings allowing each respective embedding to be decompressed independent of any other embedding in the embedding table;

receiving, by the one or more processors, an input to use for locating an embedding in the embedding table;

determining, by the one or more processors, a lookup value based on the input for searching indexes of the embedding table to locate the embedding;

locating, by the one or more processors, the embedding based on searching the indexes of the embedding table for the determined lookup value; and

decompressing, by the one or more processors, the embedding independent of any other embedding in the embedding table in view of the locating.

2. The computer-implemented method of claim 1 , further comprising:

processing, by the one or more processors, the decompressed embedding in association with running the machine-learning model.

3. The computer-implemented method of claim 1 , wherein the determining of the lookup value is based on a hashing operation that maps the input to one of the indexes in the embedding table.

4. The computer-implemented method of claim 1 , wherein respective embeddings in the embedding table are associated with an item from a plurality of items included in a vocabulary.

5. The computer-implemented method of claim 1 , wherein the decompressing is performed independent from a machine-learning platform.

6. The computer-implemented method of claim 2 , wherein the decompressing is performed using compression unavailable from a machine-learning platform.

7. A computing system for performing efficient embedding table storage and lookup in machine-learning models:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a machine-learned model configured to receive model input and process the model input to generate model output, wherein the machine-learned model comprises an embedding table comprising individually compressed embeddings allowing each respective embedding to be decompressed independent of any other embedding in the embedding table, and wherein the machine-learned model is configured to perform operations comprising:

obtaining the model input for processing in view of the embedding table;

determining a lookup value based on the model input for searching indexes of the embedding table to locate an embedding;

locating the embedding based on searching the indexes of the embedding table for the determined lookup value; and

decompressing, by the one or more processors, the embedding independent of any other embedding in the embedding table in view of the locating.

8. The computing system of claim 7 , further comprising:

processing the decompressed embedding with other respective decompressed embeddings located from the embedding table.

9. The computing system of claim 7 , wherein the operations further comprise:

generating the model output based on processing the decompressed embedding.

10. The computing system of claim 7 , wherein the determining of the lookup value is based on a hashing operation that maps the model input to one of the indexes in the embedding table.

11. The computing system of claim 7 , wherein respective embeddings in the embedding table are associated with an item from a plurality of items included in a vocabulary.

12. The computing system of claim 7 , wherein the decompressing is performed independent from a machine-learning platform.

13. The computing system of claim 7 , wherein the decompressing is performed using compression unavailable from a machine-learning platform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2023
From: MENGHANI, GAURAV
To: GOOGLE LLC
Reel/Frame 062532/0001 →
Continuity (2)
Division 17147844 · Jan 13, 2021
Related Publication 20230169058A1 · Jun 1, 2023