IP Library › Granted Patent US 12,197,601
Granted Patent B2
US 12,197,601 · App. 17/560,193 · Granted Jan 14, 2025

Hardware offload circuitry

Inventors: Ren Wang (Portland, OR); Sameh Gobriel (Dublin, CA); Somnath Paul (Hillsboro, OR); Yipeng Wang (Portland, OR); Priya Autee (Chandler, AZ); Abhirupa Layek (Chandler, AZ); Shaman Narayana (Chandler, AZ); Edwin Verplanke (Queen Creek, AZ); Mrittika Ganguli (Chandler, AZ); Jr-Shian Tsai (Portland, OR); Anton Sorokin (Portland, OR); Suvadeep Banerjee (San Jose, CA); Abhijit Davare (Hillsboro, OR); Desmond Kirkpatrick (Portland, OR); Rajesh M. Sankaran (Portland, OR); Jaykant B. Timbadiya (Amreli Gujrat, IN); Sriram Kabisthalam Muthukumar (Bengaluru, IN); Narayan Ranganathan (Bangalore, IN); Nalini Murari (Hillsboro, OR); Brinda Ganesh (Portland, OR); Nilesh Jain (Portland, OR)
Assignee: Intel Corporation
G06F21/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,601
App. No.
17/560,193
Granted
Jan 14, 2025
Kind
B2
Abstract

Examples described herein relate to offload circuitry comprising one or more compute engines that are configurable to perform a workload offloaded from a process executed by a processor based on a descriptor particular to the workload. In some examples, the offload circuitry is configurable to perform the workload, among multiple different workloads. In some examples, the multiple different workloads include one or more of: data transformation (DT) for data format conversion, Locality Sensitive Hashing (LSH) for neural network (NN), similarity search, sparse general matrix-matrix multiplication (SpGEMM) acceleration of hash based sparse matrix multiplication, data encode, data decode, or embedding lookup.

Claims (38)

1. An apparatus comprising:

offload circuitry comprising one or more compute engines that are configurable to perform a workload offloaded from a process executed by a processor based on a descriptor particular to the workload, wherein the offload circuitry is configurable to perform the workload, among multiple different workloads, and wherein the multiple different workloads comprises:

sparse general matrix-matrix multiplication (SpGEMM) acceleration of hash based sparse matrix multiplication, wherein for a workload comprising SpGEMM acceleration of hash based sparse matrix multiplication, for an operation of A×B=C, a descriptor format comprises one or more of: (1) metadata address of matrix A; (2) metadata address of matrix B; (3) address of hash table array to store a partial result; (4) start row identifier of matrix A to be processed; (5) end row identifier of matrix A to be processed; and/or (6) metadata address of matrix C.

2. The apparatus of claim 1 , wherein for a workload of the multiple different workloads that comprises data transformation (DT) for data format conversion, the descriptor format comprises: a pointer to a serialized string, address translation information, an address for an output, an address for a processing status, and/or a pointer to a transformation map.

3. The apparatus of claim 1 , wherein for a workload of the multiple different workloads that comprises data transformation (DT) for data format conversion, the one or more compute engines perform in parallel:

parse a transformation map to separate fields of data and provide the fields to processing queues for the one or more compute engines to perform data serialization or de-serialization.

4. The apparatus of claim 1 , wherein for a workload of the multiple different workloads comprising Locality Sensitive Hashing (LSH) for neural network (NN) or similarity search, the descriptor format comprises: an address of a queried key, addresses of hash tables, and/or address of a result buffer.

5. The apparatus of claim 1 , wherein for a workload of the multiple different workloads comprising Locality Sensitive Hashing (LSH) for neural network (NN) or similarity search, the one or more compute engines perform in parallel:

determine a subset of candidates and compute one or more distances and

perform candidate pruning to determine an integer K number of nearest neighbors.

6. The apparatus of claim 1 , wherein for the workload comprising SpGEMM acceleration of hash based sparse matrix multiplication, the one or more compute engines perform in parallel:

access a series of hash tables, walk the hash tables to determine an accumulated value, perform multiplication and accumulation on matrices, and provide a resulting output matrix.

7. The apparatus of claim 1 , wherein for a workload of the multiple different workloads comprising embedding lookup, the compute engines perform, in parallel, perform:

embedding lookup operations, floating-point multiplication and addition, with scaling, or multiply-accumulate (MAC) to compute a product of two numbers and add the product to an accumulator.

8. The apparatus of claim 1 , comprising the processor coupled to the offload circuitry, wherein the processor is to execute the process that offloaded the workload to the offload circuitry.

9. The apparatus of claim 8 , comprising a server, wherein the server comprises the offload circuitry and the processor.

10. The apparatus of claim 9 , comprising a datacenter, wherein the datacenter comprises the server and at least one memory pool and wherein the at least one memory pool comprises a second offload circuitry and a second processor.

11. The apparatus of claim 1 , wherein the multiple different workloads also comprise one or more of: data transformation (DT) for data format conversion, Locality Sensitive Hashing (LSH) for neural network (NN), similarity search, data encode, data decode, or embedding lookup.

12. At least one non-transitory computer-readable medium comprising instructions stored thereon, that if executed by one or more processors, cause the one or more processors to:

execute an operating system to provide capability to one or more processes to offload a workload to offload circuitry comprising one or more compute engines that are configurable to perform a workload offloaded from a process executed by a processor based on a descriptor particular to the workload, wherein the offload circuitry is configurable to perform the workload, among multiple different workloads, and wherein the multiple different workloads comprises:

sparse general matrix-matrix multiplication (SpGEMM) acceleration of hash based sparse matrix multiplication, wherein for a workload comprising SpGEMM acceleration of hash based sparse matrix multiplication, for an operation of A×B=C, a descriptor format comprises one or more of: (1) metadata address of matrix A; (2) metadata address of matrix B; (3) address of hash table array to store a partial result; (4) start row identifier of matrix A to be processed; (5) end row identifier of matrix A to be processed; and/or (6) metadata address of matrix C.

13. The computer-readable medium of claim 12 , wherein for a workload of the multiple different workloads that comprises data transformation (DT) for data format conversion, the descriptor format comprises: a pointer to a serialized string, address translation information, an address for an output, an address for a processing status, and/or a pointer to a transformation map.

14. The computer-readable medium of claim 12 , wherein for a workload of the multiple different workloads that comprises data transformation (DT) for data format conversion, the one or more compute engines perform in parallel:

parse a transformation map to separate fields of data and provide the fields to processing queues for the one or more compute engines to perform data serialization or de-serialization.

15. The computer-readable medium of claim 12 , wherein for a workload of the multiple different workloads comprising Locality Sensitive Hashing (LSH) for neural network (NN) or similarity search, the descriptor format comprises: an address of a queried key, addresses of hash tables, and/or address of a result buffer.

16. The computer-readable medium of claim 12 , wherein for a workload of the multiple different workloads comprising Locality Sensitive Hashing (LSH) for neural network (NN) or similarity search, the one or more compute engines perform in parallel:

determine a subset of candidates and compute one or more distances and

perform candidate pruning to determine an integer K number of nearest neighbors.

17. The computer-readable medium of claim 12 , wherein for the workload comprising SpGEMM acceleration of hash based sparse matrix multiplication, the one or more compute engines perform in parallel:

access a series of hash tables, walk the hash tables to determine an accumulated value, perform multiplication and accumulation on matrices, and provide a resulting output matrix.

18. The computer-readable medium of claim 12 , wherein for a workload of the multiple different workloads comprising embedding lookup, the compute engines perform, in parallel, perform:

embedding lookup operations, floating-point multiplication and addition, with scaling, or multiply-accumulate (MAC) to compute a product of two numbers and add the product to an accumulator.

19. The computer-readable medium of claim 12 , wherein the multiple different workloads also comprise one or more of: data transformation (DT) for data format conversion, Locality Sensitive Hashing (LSH) for neural network (NN), similarity search, data encode, data decode, or embedding lookup.

20. A method comprising:

offloading a workload to offload circuitry comprising one or more compute engines that are configurable to perform a workload offloaded from a process executed by a processor based on a descriptor particular to the workload, wherein the offload circuitry is configurable to perform the workload, among multiple different workloads, and wherein the multiple different workloads comprises:

sparse general matrix-matrix multiplication (SpGEMM) acceleration of hash based sparse matrix multiplication, wherein for a workload comprising SpGEMM acceleration of hash based sparse matrix multiplication, for an operation of A×B=C, a descriptor format comprises one or more of: (1) metadata address of matrix A; (2) metadata address of matrix B; (3) address of hash table array to store a partial result; (4) start row identifier of matrix A to be processed; (5) end row identifier of matrix A to be processed; and/or (6) metadata address of matrix C.

21. The method of claim 20 , wherein a descriptor associated with the workload configures the offload circuitry to perform the workload.

22. The method of claim 20 , wherein the multiple different workloads also comprise one or more of: data transformation (DT) for data format conversion, Locality Sensitive Hashing (LSH) for neural network (NN), similarity search, data encode, data decode, or embedding lookup.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2023
From: MURARI, NALINI; GANESH, BRINDA; JAIN, NILESH
To: INTEL CORPORATION
Reel/Frame 063304/0590 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2022
From: WANG, REN; GOBRIEL, SAMEH; PAUL, SOMNATH; WANG, YIPENG; AUTEE, PRIYA; LAYEK, ABHIRUPA; NARAYANA, SHAMAN; VERPLANKE, EDWIN; GANGULI, MRITTIKA; TSAI, JR-SHIAN; SOROKIN, ANTON; BANERJEE, SUVADEEP; DAVARE, ABHIJIT; KIRKPATRICK, DESMOND; SANKARAN, RAJESH M.; TIMBADIYA, JAYKANT B.; KABISTHALAM MUTHUKUMAR, SRIRAM; RANGANATHAN, NARAYAN
To: INTEL CORPORATION
Reel/Frame 059824/0511 →
Continuity (2)
Provisional Application 63130663 · Dec 26, 2020
Related Publication 20220114270A1 · Apr 14, 2022
References Cited (17)
US 20180322390A1 · Das · 2018 [cited by examiner]
US 20190317802A1 · Bachmutsky · 2019 [cited by examiner]
US 20210374162A1 · Kirdey · 2021 [cited by examiner]
Bentley, Jon Louis, “Multidimensional divide-and-conquer”, Communications of the ACM, vol. 23, Issue 4, https://doi.org/10.1145/358841.358850, Apr. 1980, pp. 214-229. [cited by applicant]
Gupta, Udit et al., “The Architectural Implications of Facebook's DNN-based Personalized Recommendation”, 2019, 14 pages. [cited by applicant]
HPCA Submission, “QEI: Query Acceleration Can be Generic and Efficient in the Cloud”, 2021, 14 pages. [cited by applicant]
Jegou, Herve et al., “Faiss: A library for efficient similarity search”, https://engineering.fb.com/2017/03/29/data-infrastructure/faiss-a-library-for-efficient-similarity-search/, Mar. 29, 2017, 7 pages. [cited by applicant]
Ke, Liu et al., “RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing”, Dec. 2019, 14 pages. [cited by applicant]
Pourhabibi, Arash et al., “Optimus Prime: Accelerating Data Transformation in Servers”, ASPLOS '20, Mar. 16-20, 2020, Lausanne, Switzerland, 14 pages. [cited by applicant]
Sun, Philip, “Announcing ScaNN: Efficient Vector Similarity Search”, https://ai.googleblog.com/2020/07/announcing-scann-efficient-vector.html, Jul. 28, 2020, 7 pages. [cited by applicant]
University of Bristol, “What is a Voronoi diagram?”, https://www.bristol.ac.uk/maths/fry-building/public-art-strategy/what-is-a-voronoi-diagram/, May 2020, 2 pages. [cited by applicant]
Vaidya, Pravin M., “AnO(n logn) algorithm for the all-nearest-neighbors Problem”, http://rd.springer.com/article/10.1007/BF02187718, Mar. 1989, 15 pages. [cited by applicant]
Yuan, Yifan et al. “QEI: Query Acceleration Can be Generic and Efficient in the Cloud,” 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), Seoul, Korea (South), 2021, 14 pages. [cited by applicant]
Yuan, Yifan et al., “HALO: Accelerating Flow Classification for Scalable Packet Processing in NFV”, 2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA), Phoenix, AZ, USA, Jun. 2019, 14 page… [cited by applicant]
Extended European Search Report for Patent Application No. 21217626.7, Mailed May 10, 2022, 8 pages. [cited by applicant]
Dong, Wei et al., “Efficient K-Nearest Neighbor Graph Construction for Generic Similarity Measures”, WWW'11: Proceedings of the 20th international conference on World wide web, Mar. 2011, 10 pages. [cited by applicant]
Kaul, Himanshu et al., “14.4 A 21.5M-query-vectors/s 3.37nJ/vector reconfigurable k-nearest-neighbor accelerator with adaptive precision in 14nm tri-gate CMOS,” 2016 IEEE International Solid-State Circuits Conference (I… [cited by applicant]