IP Library Granted Patent US 12,306,771
Granted Patent B2
US 12,306,771 · App. 18/671,095 · Granted May 20, 2025

Efficient data sharing for graphics data processing operations

Inventors: Joydeep Ray (Folsom, CA); Altug Koker (El Dorado Hills, CA); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Michael Macpherson (Portland, OR); Aravindh V. Anantaraman (Folsom, CA); Vasanth Ranganathan (El Dorado Hills, CA); Lakshminarayanan Striramassarma (Folsom, CA); Varghese George (Folsom, CA); Abhishek Appu (El Dorado Hills, CA); Prasoonkumar Surti (Folsom, CA)
Assignee: INTEL CORPORATION
G06F13/1605G06F9/3004G06F9/3887G06F9/3888G06F9/38885G06F9/5016G06T1/20G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,306,771
App. No.
18/671,095
Granted
May 20, 2025
Kind
B2
Abstract

An apparatus to facilitate efficient data sharing for graphics data processing operations is disclosed. The apparatus includes a processing resource to generate a stream of instructions, an L1 cache communicably coupled to the processing resource and comprising an on-page detector circuit to determine that a set of memory requests in the stream of instructions access a same memory page; and set a marker in a first request of the set of memory requests; and arbitration circuitry communicably coupled to the L1 cache, the arbitration circuitry to route the set of memory requests to memory comprising the memory page and to, in response to receiving the first request with the marker set, remain with the processing resource to process the set of memory requests.

Claims (34)

1. An apparatus comprising:

a processing resource of a chiplet; and

an L1 cache communicably coupled to the processing resource and comprising synchronization hardware circuit to:

allocate a first thread group as a member of a super thread group (STG) comprising a collection of thread groups running on the chiplet;

receive from a second thread group within the STG, a request to communicate with the first thread group identified by a thread group identifier (ID) within the STG;

access a routing table to determine a location of the first thread group based on the thread group ID; and

route the request to the determined location of the first thread group using communication links between L1 caches of the chiplet.

2. The apparatus of claim 1 , wherein the processing resource is an execution unit in a graphics processing unit (GPU).

3. The apparatus of claim 1 , wherein the synchronization hardware circuit comprises one or more of a router, a receiver/transmitter (Rx/Tx), or the routing table.

4. The apparatus of claim 3 , wherein the router is to implement an application programming interface (API) defined to allow communication between sub-slices of the chiplet.

5. The apparatus of claim 3 , wherein the router is to communicate with the Rx/Tx to cause communications from the L1 cache to be sent or received over buses of a point-to-point communication network.

6. The apparatus of claim 1 , wherein a kernel is to utilize query functions to inquire an STG ID number, to inquire a number of thread groups in the STG, or to send and receive data from/to thread in a same STG as the kernel using a logical thread number.

7. The apparatus of claim 1 , wherein the chiplet further comprises a chiplet-level cache for sub-slice communication within the chiplet and to enable communication between threads on different sub-slices of the chiplet that are within the STG of the chiplet.

8. The apparatus of claim 7 , wherein the chiplet-level cache comprises an L2 cache that is in communication with a global cache of a base die of the chiplet.

9. The apparatus of claim 1 , wherein the apparatus is at least one of a single instruction multiple data (SIMD) machine or a single instruction multiple thread (SIMT) machine.

10. A method comprising:

allocating, by a synchronization hardware circuit of an L1 cache of a graphics processor, a first thread group as a member of a super thread group (STG) comprising a collection of thread groups running on the graphics processor;

receiving, by the synchronization hardware circuit, from a second thread group within the STG, a request to communicate with the first thread group identified by a thread group identifier (ID) within the STG;

accessing, by the synchronization hardware circuit, a routing table to determine a location of the first thread group based on the thread group ID; and

routing, by the synchronization hardware circuit, the request to the determined location of the first thread group using communication links between L1 caches of the graphics processor.

11. The method of claim 10 , wherein the synchronization hardware circuit comprises one or more of a router, a receiver/transmitter (Rx/Tx), or the routing table.

12. The method of claim 11 , wherein the router is to implement an application programming interface (API) defined to allow communication between sub-slices of the graphics processor.

13. The method of claim 11 , wherein the router is to communicate with the Rx/Tx to cause communicate from the L1 cache to be sent or received over buses of a point-to-point communication network.

14. The method of claim 10 , wherein a kernel is to utilize query functions to inquire an STG ID number, to inquire a number of thread groups in the STG, or to send and receive data from/to thread in a same STG as the kernel using a logical thread number.

15. The method of claim 10 , wherein the graphics processor further comprises a chiplet-level cache for sub-slice communication within the graphics processor and to enable communication between threads on different sub-slices of the graphics processor that are within the STG of the graphics processor.

16. A non-transitory computer-readable medium having instructions stored thereon, which when executed by one or more processors, cause the one or more processors to:

allocate, by a synchronization hardware circuit of an L1 cache of the one or more processors, a first thread group as a member of a super thread group (STG) comprising a collection of thread groups running on the one or more processors;

receive, by the synchronization hardware circuit, from a second thread group within the STG, a request to communicate with the first thread group identified by a thread group identifier (ID) within the STG;

access, by the synchronization hardware circuit, a routing table to determine a location of the first thread group based on the thread group ID; and

route, by the synchronization hardware circuit, the request to the determined location of the first thread group using communication links between L1 caches of the one or more processors.

17. The non-transitory computer-readable medium of claim 16 , wherein the synchronization hardware circuit comprises one or more of a router, a receiver/transmitter (Rx/Tx), or the routing table.

18. The non-transitory computer-readable medium of claim 17 , wherein the router is to implement an application programming interface (API) defined to allow communication between sub-slices of the one or more processors.

19. The non-transitory computer-readable medium of claim 16 , wherein a kernel is to utilize query functions to inquire an STG ID number, to inquire a number of thread groups in the STG, or to send and receive data from/to thread in a same STG as the kernel using a logical thread number.

20. The non-transitory computer-readable medium of claim 16 , wherein the one or more processors further comprise a chiplet-level cache for sub-slice communication within the one or more processors and to enable communication between threads on different sub-slices of the one or more processors that are within the STG of the one or more processors.

Continuity (4)
Continuation 18358550 · Jul 25, 2023
Continuation 17212503 · Mar 25, 2021
Provisional Application 63000784 · Mar 27, 2020
Related Publication 20240385975A1 · Nov 21, 2024
References Cited (24)
US 6282232B1 · Fleming, III · 2001 [cited by examiner]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 11755501B2 · Ray et al. · 2023 [cited by applicant]
US 12032496B2 · Ray et al. · 2024 [cited by applicant]
US 20030108050A1 · Black · 2003 [cited by examiner]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20210303481A1 · Ray et al. · 2021 [cited by applicant]
US 20220091986A1 · Kotra et al. · 2022 [cited by applicant]
US 20240012767A1 · Ray et al. · 2024 [cited by applicant]
CN 101159747B · 2010 [cited by examiner]
CN 105786447A · 2016 [cited by examiner]
English translation of CN-105786447-A (Year: 2016). [cited by examiner]
English translation of CN-101159747-B (Year: 2010). [cited by examiner]
U.S. Appl. No. 17/212,503 “Notice of Allowance” mailed May 1, 2023, 11 pages. [cited by applicant]
U.S. Appl. No. 18/358,550 “Notice of Allowance” mailed Feb. 29, 2024, 10 pages. [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]