IP Library › Granted Patent US 12,405,890
Granted Patent B2
US 12,405,890 · App. 17/561,571 · Granted Sep 2, 2025

Method and apparatus for leveraging simultaneous multithreading for bulk compute operations

Inventors: Anant Nori (Bangalore, IN); Rahul Bera (Zürich, CH); Shankar Balachandran (Bangalore, IN); Joydeep Rakshit (Bengaluru, IN); Om Ji Omer (Bangalore, IN); Sreenivas Subramoney (Bangalore, IN); Avishaii Abuhatzera (Amir, IL); Belliappa Kuttanna (Austin, TX)
Assignee: Intel Corporation
G06F12/0811G06F9/3012G06F9/3851G06F9/3888G06F9/4881G06N3/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,890
App. No.
17/561,571
Granted
Sep 2, 2025
Kind
B2
Abstract

Apparatus and method for leveraging simultaneous multithreading for bulk compute operations. For example, one embodiment of a processor comprises: a plurality of cores including a first core to simultaneously process instructions of a plurality of threads; a cache hierarchy coupled to the first core and the memory, the cache hierarchy comprising a Level 1 (L1) cache, a Level 2 (L2) cache, and a Level 3 (L3) cache; and a plurality of compute units coupled to the first core including a first compute unit associated with the L1 cache, a second compute unit associated with the L2 cache, and a third compute unit associated with the L3 cache, wherein the first core is to offload instructions for execution by the compute units, the first core to offload instructions from a first thread to the first compute unit, instructions from a second thread to the second compute unit, and instructions from a third thread to the third compute unit.

Claims (51)

1. A processor comprising:

a plurality of cores including a first core to simultaneously process instructions of a plurality of threads;

a cache hierarchy coupled to the first core and the memory, the cache hierarchy comprising a Level 1 (L1) cache, a Level 2 (L2) cache, and a Level 3 (L3) cache; and

a plurality of compute units coupled to the first core including a first compute unit associated with the L1 cache, a second compute unit associated with the L2 cache, and a third compute unit associated with the L3 cache,

wherein the first core is to offload instructions for execution by the compute units, the first core to offload instructions from a first thread to the first compute unit, instructions from a second thread to the second compute unit, and instructions from a third thread to the third compute unit.

2. The processor of claim 1 wherein the instructions offloaded by the core comprise instructions associated with structured, repetitive loops with a plurality of iterations.

3. The processor of claim 2 wherein the instructions offloaded by the core comprise instructions to perform matrix operations and matrix-vector operations.

4. The processor of claim 2 wherein the instructions offloaded by the core comprise instructions of a Deep Neural Network kernel.

5. The processor of claim 1 wherein the first compute unit is bound exclusively to the first thread, the second compute unit is bound exclusively to the second thread, and the third compute unit is bound exclusively to the third thread, and wherein the first core is to process instructions from each of the first thread, second thread, and third thread.

6. The processor of claim 5 wherein a compute unit of the first compute unit, second compute unit, and third compute unit comprises:

a code register file comprising a first plurality of entries to store instructions to be executed by the compute unit;

a scheduler to schedule the instructions for execution;

a data register file comprising a second plurality of entries to store data associated with the instructions; and

compute circuitry to execute the instructions and store results in the data register file.

7. The processor of claim 6 wherein the compute unit further comprises:

an address generation unit (AGU) to determine virtual addresses required by the compute circuitry to execute the instructions; and

a translation cache (TC) to store translations from the virtual addresses to physical addresses of system memory accessible by the first core.

8. The processor of claim 7 further comprising:

an in-order compute issue queue coupled between the scheduler and the compute circuitry and configured to store compute instructions scheduled by the scheduler for execution by the compute circuitry.

9. The processor of claim 7 further comprising:

an in-order load/store issue queue coupled between the scheduler and the AGU and configured to store load/store instructions scheduled by the scheduler for processing by the AGU.

10. The processor of claim 1 wherein the first compute unit is to directly access the L1 cache and not directly access the L2 cache or the L3 cache; the second compute unit is to directly access the L2 cache and not directly access the L1 cache or L3 cache; and the third compute unit is to directly access the L3 cache and not directly access the L1 cache or L2 cache.

11. The processor of claim 10 further comprising:

cache allocation hardware logic to allocate a specified portion of the L3 cache to the third compute unit.

12. The processor of claim 10 wherein the L2 cache is to store a first plurality of cache lines, each cache line to include a bit to indicate whether the cache line is owned by the L2 cache or the L1 cache.

13. A system comprising:

a memory to store instructions of a plurality of threads;

a plurality of cores including a first core to simultaneously process instructions of a plurality of threads;

a cache hierarchy coupled to the first core and the memory, the cache hierarchy comprising a Level 1 (L1) cache, a Level 2 (L2) cache, and a Level 3 (L3) cache; and

a plurality of compute units coupled to the first core including a first compute unit associated with the L1 cache, a second compute unit associated with the L2 cache, and a third compute unit associated with the L3 cache,

wherein the first core is to offload instructions for execution by the compute units, the first core to offload instructions from a first thread to the first compute unit, instructions from a second thread to the second compute unit, and instructions from a third thread to the third compute unit.

14. The system of claim 13 wherein the instructions offloaded by the core comprise instructions associated with structured, repetitive loops with a plurality of iterations.

15. The system of claim 14 wherein the instructions offloaded by the core comprise instructions to perform matrix operations and matrix-vector operations.

16. The system of claim 14 wherein the instructions offloaded by the core comprise instructions of a Deep Neural Network kernel.

17. The system of claim 13 wherein the first compute unit is bound exclusively to the first thread, the second compute unit is bound exclusively to the second thread, and the third compute unit is bound exclusively to the third thread, and wherein the first core is to process instructions from each of the first thread, second thread, and third thread.

18. The system of claim 17 wherein a compute unit of the first compute unit, second compute unit, and third compute unit comprises:

a code register file comprising a first plurality of entries to store instructions to be executed by the compute unit;

a scheduler to schedule the instructions for execution;

a data register file comprising a second plurality of entries to store data associated with the instructions; and

compute circuitry to execute the instructions and store results in the data register file.

19. The system of claim 18 wherein the compute unit further comprises:

an address generation unit (AGU) to determine virtual addresses required by the compute circuitry to execute the instructions; and

a translation cache (TC) to store translations from the virtual addresses to physical addresses of system memory accessible by the first core.

20. The system of claim 19 further comprising:

an in-order compute issue queue coupled between the scheduler and the compute circuitry and configured to store compute instructions scheduled by the scheduler for execution by the compute circuitry.

21. The system of claim 19 further comprising:

an in-order load/store issue queue coupled between the scheduler and the AGU and configured to store load/store instructions scheduled by the scheduler for processing by the AGU.

22. The system of claim 13 wherein the first compute unit is to directly access the L1 cache and not directly access the L2 cache or the L3 cache; the second compute unit is to directly access the L2 cache and not directly access the L1 cache or L3 cache; and the third compute unit is to directly access the L3 cache and not directly access the L1 cache or L2 cache.

23. The system of claim 22 further comprising:

cache allocation hardware logic to allocate a specified portion of the L3 cache to the third compute unit.

24. The system of claim 22 wherein the L2 cache is to store a first plurality of cache lines, each cache line to include a bit to indicate whether the cache line is owned by the L2 cache or the L1 cache.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2022
From: NORI, ANANT; BERA, RAHUL; BALACHANDRAN, SHANKAR; RAKSHIT, JOYDEEP; OMER, OM JI; SUBRAMONEY, SREENIVAS; ABUHATZERA, AVISHAII; KUTTANNA, BELLIAPPA
To: INTEL CORPORATION
Reel/Frame 058987/0937 →
Continuity (1)
Related Publication 20230205692A1 · Jun 29, 2023
References Cited (20)
US 20170097891A1 · Shaikh · 2017 [cited by examiner]
US 20190004810A1 · Jayasimha et al. · 2019 [cited by applicant]
US 20200301733A1 · Zhang · 2020 [cited by examiner]
US 20210303357A1 · Varma · 2021 [cited by examiner]
US 20220100514A1 · Nori et al. · 2022 [cited by applicant]
Aga et al., “Compute Caches”, IEEE International Symposium on High Performance Computer Architecture, 2017, pp. 481-492. [cited by applicant]
Carlson et al., “The Sniper User Manual”, Sniper, Nov. 13, 2013, 39 pages. [cited by applicant]
Eckert et al., “Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks”, ACM/IEEE 45th Annual International Symposium on Computer Architecture, 2018, pp. 383-396. [cited by applicant]
Feichtenhofer et al., “Convolutional Two-Stream Network Fusion for Video Action Recognition”, IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, 9 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, pp. 770-778. [cited by applicant]
Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, CoRR, Apr. 17, 2017, pp. 1-10. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks”, IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
IDF AC9494, “Loop Support Extensions—Late unrolling of loops on CPUs for data parallel algorithms like DNN inference”, filed Sep. 2020, 9 pages. [cited by applicant]
Intel, “Improving Real-Time Performance by Utilizing Cache Allocation Technology Enhancing Performance via Allocation of the Processor's Cache”, White Paper, Apr. 2015, Document No. 331843-001US, pp. 1-16. [cited by applicant]
LLVM/OpenMP, “LLVM/OpenMP 21.0.ogit documentation”, available online at <https://openmp.llvm.org/index.html>, 2025, 2 pages. [cited by applicant]
Lockerman et al., “Livia: Data-Centric Computing Throughout the Memory Hierarchy”, ASPLOS '20, Mar. 16-20, 2020, 18 pages. [cited by applicant]
Mutlu et al., “Processing Data Where It Makes Sense: Enabling In-Memory Computation”, Mar. 10, 2019, 21 pages. [cited by applicant]
Rodriguez et al., “Lower numerical precision deep learning inference and training,” Intel White Paper, 2018, pp. 1-19. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, Dec. 4-9, 2017, 11 pages. [cited by applicant]
Xie et al., “Aggregated Residual Transformations for Deep Neural Networks”, IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Jul. 21-26, 2017, pp. 5987-5995. [cited by applicant]