IP Library Granted Patent US 12,645,490
Granted Patent B2
US 12,645,490 · App. 17/557,927 · Granted Jun 2, 2026

Variable dispatch walk

Inventors: Saurabh Sharma (Santa Clara, CA); Jeremy Lukacs (Santa Clara, CA); Hashem Hashemi (Roseville, CA); Gianpaolo Tommasi (Santa Clara, CA); Guennadi Riguer (Markham, CA); Mark Fowler (Boxborough, MA); Randy Ramsey (Orlando, FL)
Assignees: ADVANCED MICRO DEVICES, INC.; ATI TECHNOLOGIES ULC
G06F9/4831
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,490
App. No.
17/557,927
Granted
Jun 2, 2026
Kind
B2
Abstract

A processing unit performs a dispatch walk of a set of thread groups based on a programmable access pattern. The access pattern is stored at a table that is programmed with the access pattern based upon a specified command. By using the command to program the table with different access patterns, the dispatch order of the set of thread groups is adapted to better suit the processing of different data sets, thereby reducing power consumption at the processing unit, and improving overall processing efficiency.

Claims (39)

1 . A non-transitory computer readable medium embodying a set of executable instructions, the set of executable instructions to manipulate at least one processor to:

in response to a first dispatch command, dispatch a first plurality of thread groups to a set of processing circuits having access to a cache in a first dispatch order based on shared data to be used by some thread groups but not all thread groups of the first plurality of thread groups, wherein the first dispatch order comprises dispatching a first thread group of the first plurality of thread groups immediately after a second thread group of the first plurality of thread groups based on an amount of data stored at the cache that is shared by the first thread group and the second thread group; and

in response to a second dispatch command, dispatch a second plurality of thread groups to the set of processing circuits in a second dispatch order, the second dispatch order different from the first dispatch order.

2 . The non-transitory computer readable medium of claim 1 , wherein the set of executable instructions are to manipulate at least one processor to:

store an access pattern that indicates the first dispatch order at a programmable table.

3 . The non-transitory computer readable medium of claim 2 , wherein the set of executable instructions are to manipulate at least one processor to:

store the access pattern at a programmable table in response to a pattern command.

4 . The non-transitory computer readable medium of claim 2 , wherein:

the second dispatch order is based on a second access pattern that is different from the access pattern.

5 . The non-transitory computer readable medium of claim 1 , wherein:

the first plurality of thread groups is organized according to at least two dimensions.

6 . The non-transitory computer readable medium of claim 1 , wherein the first dispatch order comprises one of:

a Hilbert curve, Morton curve pattern, or a z-order walking pattern.

7 . The non-transitory computer readable medium of claim 1 , wherein:

the set of processing circuits comprise a shader of a graphics processing unit.

8 . A method, comprising:

in response to receiving a first command, dispatching a first plurality of thread groups to a set of processing circuits having access to a cache in a first dispatch order based on shared data to be used by some thread groups but not all thread groups of the first plurality of thread groups, wherein the first dispatch order comprises dispatching a first thread group of the first plurality of thread groups immediately after a second thread group of the first plurality of thread groups based on an amount of data stored at the cache that is shared by the first thread group and the second thread group; and

in response to receiving a second command, dispatching a second plurality of thread groups in a second dispatch order to the set of processing circuits, the second dispatch order different from the first dispatch order.

9 . The method of claim 8 , wherein:

the first plurality of thread groups is organized according to at least two dimensions.

10 . The method of claim 8 , wherein:

an access pattern that indicates the first dispatch order is stored at a programmable table.

11 . The method of claim 8 , wherein:

the first dispatch order is determined based on data to be used by the first plurality of thread groups as part of one or more data swaps.

12 . The method of claim 8 , wherein the first dispatch order comprises one of:

a Hilbert curve, Morton curve pattern, or a z-order walking pattern.

13 . A processor comprising:

a set of processing circuits having access to a cache; and

a dispatch unit configured to:

in response to a first dispatch command, dispatch a first plurality of thread groups to a set of processing circuits in a first dispatch order based on shared data to be used by some thread groups but not all thread groups of the first plurality of thread groups, wherein the first dispatch order comprises dispatching a first thread group of the first plurality of thread groups immediately after a second thread group of the first plurality of thread groups based on an amount of data stored at the cache that is shared by the first thread group and the second thread group; and

in response to a second dispatch command, dispatch a second plurality of thread groups to the set of processing circuits in a second dispatch order, the second dispatch order different from the first dispatch order.

14 . The processor of claim 13 , further comprising:

a programmable table to store a plurality of access patterns including an access pattern that indicates the first dispatch order.

15 . The processor of claim 13 , wherein:

the first plurality of thread groups is organized according to at least three dimensions.

16 . The processor of claim 13 , wherein the first dispatch order comprises one of:

a Hilbert curve, Morton curve pattern, or a z-order walking pattern.

17 . The processor of claim 13 , wherein:

the processor comprises a graphics processing unit and the set of processing circuits comprise a shader.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2022
From: SHARMA, SAURABH; HASHEMI, HASHEM; TOMMASI, GIANPAOLO; LUKACS, JEREMY; RAMSEY, RANDY; FOWLER, MARK
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 059016/0739 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2022
From: RIGUER, GUENNADI
To: ATI TECHNOLOGIES ULC
Reel/Frame 059016/0905 →
Continuity (1)
Related Publication 20230195509A1 · Jun 22, 2023
References Cited (41)
US 7681014B2 · Jensen · 2010 [cited by examiner]
US 9514506B2 · Lee · 2016 [cited by examiner]
US 9615104B2 · Wu · 2017 [cited by examiner]
US 10037149B2 · Nazarov et al. · 2018 [cited by applicant]
US 10261903B2 · Sakthivel · 2019 [cited by examiner]
US 10402224B2 · Veernapu · 2019 [cited by examiner]
US 10521875B2 · Koker · 2019 [cited by examiner]
US 10733012B2 · Nugteren · 2020 [cited by examiner]
US 10796397B2 · Valerio · 2020 [cited by examiner]
US 11822956B2 · Luo · 2023 [cited by examiner]
US 11954062B2 · Ray · 2024 [cited by examiner]
US 11995737B2 · Koker · 2024 [cited by examiner]
US 20030200396A1 · Musumeci · 2003 [cited by applicant]
US 20060184747A1 · Guthrie et al. · 2006 [cited by applicant]
US 20090303245A1 · Soupikov et al. · 2009 [cited by applicant]
US 20110213937A1 · Barry et al. · 2011 [cited by applicant]
US 20120278558A1 · Dufter et al. · 2012 [cited by applicant]
US 20120320069A1 · Lee · 2012 [cited by examiner]
US 20130265318A1 · Schneider · 2013 [cited by applicant]
US 20150160970A1 · Nugteren et al. · 2015 [cited by applicant]
US 20160239441A1 · Chun · 2016 [cited by examiner]
US 20160364828A1 · Valerio · 2016 [cited by examiner]
US 20170124742A1 · Hasselgren et al. · 2017 [cited by applicant]
US 20170178386A1 · Redshaw et al. · 2017 [cited by applicant]
US 20180329712A1 · Palani et al. · 2018 [cited by applicant]
US 20190324757A1 · Valerio · 2019 [cited by examiner]
US 20200327060A1 · Desai · 2020 [cited by applicant]
US 20210141649A1 · Xu et al. · 2021 [cited by applicant]
US 20220066931A1 · Ray · 2022 [cited by examiner]
US 20220156875A1 · Koker · 2022 [cited by examiner]
US 20230104199A1 · Volkov · 2023 [cited by examiner]
KR 101799978B1 · 2017 [cited by applicant]
Final Office Action issued in U.S. Appl. No. 17/558,008, mailed Dec. 18, 2023, 16 pages. [cited by applicant]
Office Action issued in U.S. Appl. No. 17/558,008, mailed Jun. 28, 2023, 22 pages. [cited by applicant]
Office Action issued in U.S. Appl. No. 17/558,008, mailed May 28, 2024, 18 pages. [cited by applicant]
International Preliminary Report on Patentability issued in Application No. PCT/US2022/053381, mailed Jul. 4, 2024, 8 pages. [cited by applicant]
Non-Final Office Action issued in U.S. Appl. No. 17/558,008, mailed Dec. 13, 2024, 24 pages. [cited by applicant]
Final Office Action issued in U.S. Appl. No. 17/558,008, mailed Jun. 30, 2025, 21 pages. [cited by applicant]
International Search Report and Written Opinion issued in Application No. PCT/US2022/053381, mailed Apr. 21, 2023, 11 pages. [cited by applicant]
Extended European Search Report mailed Nov. 27, 2025 for European Application No. 22912350.0, 10 pages. [cited by applicant]
Notice of Allowance mailed Oct. 1, 2025 for in U.S. Appl. No. 17/558,008, 6 pages. [cited by applicant]