IP Library Granted Patent US 12,468,632
Granted Patent B2
US 12,468,632 · App. 17/835,810 · Granted Nov 11, 2025

Method for embedding rows prefetching in recommendation models

Inventors: Mohamed Assem Abd Elmohsen Ibrahim (Santa Clara, CA); Onur Kayiran (Santa Clara, CA); Shaizeen Dilawarhusen Aga (Santa Clara, CA); Yasuko Eckert (Bellevue, WA)
Assignee: Advanced Micro Devices, Inc.
G06F12/0862G06F2212/602
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,632
App. No.
17/835,810
Granted
Nov 11, 2025
Kind
B2
Abstract

A system and method for efficiently accessing sparse data for a workload are described. In various implementations, a computing system includes an integrated circuit and a memory for storing tasks of a workload that includes sparse accesses of data items stored in one or more tables. The integrated circuit receives a user query, and generates a result based on multiple data items targeted by the user query. To reduce the latency of processing the workload even with sparse lookup operations performed on the one or more tables, a prefetch engine of the integrated circuit stores a subset of data items in prefetch data storage. The prefetch engine also determines which data items to store in the prefetch data storage based on one or more of a frequency of reuse, a distance or latency of access of a corresponding table of the one more tables, or other.

Claims (43)

1 . An integrated circuit comprising:

a prefetch engine comprising circuitry, wherein responsive to receipt of a user query that targets two or more data items, the circuitry is configured to:

retrieve, from a prefetch data storage, a first group of one or more data items of the two or more data items, identified as being stored in the prefetch data storage;

send the first group of data items to a data processing stage comprising processing circuitry, prior to sending remaining data items of the two or more data items;

retrieve, from a data storage different from the prefetch data storage, a second group of one or more data items of the two or more data items, identified as not being stored in the prefetch data storage; and

send the second group of data items to the data processing stage.

2 . The integrated circuit as recited in claim 1 , wherein each of the first group of data items is stored in a non-contiguous manner.

3 . The integrated circuit as recited in claim 1 , wherein the prefetch engine is further configured to reorder a priority of data processing of a list of a plurality of data item identifiers specifying data items targeted by a given user query, wherein reordering is based on access statistics of data items currently stored in the prefetch data storage.

4 . The integrated circuit as recited in claim 1 , wherein the prefetch engine is further configured to replace a first data item currently stored in the prefetch data storage with a second data item not stored in the prefetch data storage based on determining that a priority of the second data item is greater than a priority of the first data item.

5 . The integrated circuit as recited in claim 4 , wherein a priority of a given data item is based on a frequency of reuse of the given data item by a plurality of user queries.

6 . The integrated circuit as recited in claim 4 , wherein a priority of a given data item is based on an access latency of a corresponding one of a data storage area that stores the given data item.

7 . The integrated circuit as recited in claim 1 , wherein:

data storage areas targeted by the user query are embedding tables configured to store embedding vectors used by a data model implementing a neural network; and

each data item is an embedding vector.

8 . A method comprising:

responsive to receiving a user query that targets two or more data items:

retrieving, from a prefetch data storage, a first group of one or more data items of the two or more data items, identified as being stored in the prefetch data storage;

sending the first group of data items to a data processing stage comprising processing circuitry, prior to sending remaining data items of the two or more data items;

retrieving, from a data storage different than the prefetch data storage, a second group of one or more data items of the two or more data items, identified as not being stored in the prefetch data storage; and

sending the second group of data items to the data processing stage.

9 . The method as recited in claim 8 , wherein each of the first group of data items is stored in a non-contiguous manner.

10 . The method as recited in claim 8 , further comprising reordering a priority of data processing of a list of a plurality of data item identifiers specifying data items targeted by a given user query, wherein reordering is based on access statistics of data items currently stored in the prefetch data storage.

11 . The method as recited in claim 8 , further comprising replacing a first data item currently stored in the prefetch data storage with a second data item not stored in the prefetch data storage based on determining that a priority of the second data item is greater than a priority of the first data item.

12 . The method as recited in claim 11 , further comprising assigning a priority of a given data item based on a frequency of reuse of the given data item by a plurality of user queries.

13 . The method as recited in claim 11 , further comprising assigning a priority of a given data item based on an access latency of a data storage area that stores the given data item.

14 . The method as recited in claim 8 , wherein:

data storage areas targeted by the user query are embedding tables configured to store embedding vectors used by a data model implementing a neural network; and

each data item is an embedding vector.

15 . A computing system comprising:

a memory configured to store instructions of one or more tasks and source data to be processed by the one or more tasks;

an integrated circuit configured to execute the instructions using the source data, wherein the integrated circuit comprises:

a prefetch engine comprising circuitry, wherein responsive to receipt of a user query that targets two or more data items, the circuitry is configured to:

retrieve, from a prefetch data storage, a first group of one or more data items of the two or more data items, identified as being stored in the prefetch data storage;

send the first group of data items to a data processing stage comprising processing circuitry, prior to sending remaining data items of the two or more data items;

retrieve, from a data storage different than the prefetch data storage, a second group of one or more data items of the two or more data items, identified as not being stored in the prefetch data storage; and

send the second group of data items to the data processing stage.

16 . The computing system as recited in claim 15 , wherein each of the first group of data items is stored in a non-contiguous manner.

17 . The computing system as recited in claim 15 , wherein the prefetch engine is further configured to reorder a priority of data processing of a list of a plurality of data item identifiers specifying data items targeted by a given user query, wherein reordering is based on access statistics of data items currently stored in the prefetch data storage.

18 . The computing system as recited in claim 15 , wherein the prefetch engine is further configured to replace a first data item currently stored in the prefetch data storage with a second data item not stored in the prefetch data storage based on determining that a priority of the second data item is greater than a priority of the first data item.

19 . The computing system as recited in claim 18 , wherein a priority of a given data item is based on an access latency of a data storage area that stores the given data item.

20 . The computing system as recited in claim 15 , wherein:

data storage areas targeted by the user query are embedding tables configured to store embedding vectors used by a data model implementing a neural network; and

each data item is an embedding vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2022
From: IBRAHIM, MOHAMED ASSEM ABD ELMOHSEN; KAYIRAN, ONUR; AGA, SHAIZEEN DILAWARHUSEN; ECKERT, YASUKO
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 060142/0625 →
Continuity (1)
Related Publication 20230401154A1 · Dec 14, 2023
References Cited (12)
US 10915447B1 · Yau · 2021 [cited by examiner]
US 10922231B1 · Gray · 2021 [cited by examiner]
US 20140289332A1 · Frosst · 2014 [cited by examiner]
US 20190243766A1 · Doerner · 2019 [cited by examiner]
US 20210157500A1 · Gu · 2021 [cited by examiner]
Gupta et al., “The Architectural Implications of Facebook's DNN-based Personalized Recommendation”, International Symposium on High-Performance Computer Architecture, Feb. 15, 2020, 14 pages, https://arxiv.org/pdf/1906.… [cited by applicant]
Hwang et al., “Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations”, Proceedings of the 47th IEEE/ACM International Symposium on Computer Architecture, May 12, 2020, 14 pages, http… [cited by applicant]
Ke et al., “RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing”, Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture, Dec. 30, 2019, 14 pages, https://arxiv… [cited by applicant]
Kwon et al., “TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning”, Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, Aug… [cited by applicant]
Naumov et al., “Deep Learning Recommendation Model for Personalization and Recommendation Systems”, May 31. 2019, 10 pages, https://arxiv.org/pdf/1906.00091.pdf. [Retrieved Jun. 2, 2022]. [cited by applicant]
Park et al., “Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications”, Nov. 29, 2018, 12 pages, https://arxiv.org/pdf/1811.09886.pdf. [Retrieved Jun. 2, 2… [cited by applicant]
Yin et al., “TT-Rec: Tensor Train Compression for Deep Learning Recommendation Models”, Conference on Machine Learning and Systems, Jan. 25, 2021, 15 pages, https://arxiv.org/pdf/2101.11714.pdf. [Retrieved Jun. 2, 2022]. [cited by applicant]