IP Library › Granted Patent US 12,487,930
Granted Patent B2
US 12,487,930 · App. 18/537,716 · Granted Dec 2, 2025

Dynamic accounting of cache for GPU scratch memory usage

Inventors: Nicolai Haehnle (Munich, DE); Michael John Bedy (Groton, MA)
Assignee: Advanced Micro Devices, Inc.
G06F12/0811G06F9/3851G06F9/4843G06F12/0871G06F12/0897
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,487,930
App. No.
18/537,716
Granted
Dec 2, 2025
Kind
B2
Abstract

An apparatus and method for efficiently scheduling wavefronts for execution on an integrated circuit. In various implementations, a computing system includes a parallel data processing circuit with multiple, replicated compute circuits. Each compute circuit executes one or more wavefronts. Each compute circuit includes a cache configured to store temporary data that cannot fit in the vector general-purpose register file of the compute circuit. Each wavefront requests a corresponding amount of storage space in the cache for storing the temporary data. When the available data storage space in the cache is less than a data size requested by a wavefront waiting to be dispatched, a control circuit of the compute circuit reduces a dispatch rate of wavefronts. The control circuit also reduces an issue rate of instructions of one or more dispatched wavefronts to assigned execution circuits of the compute circuit.

Claims (31)

1 . An apparatus comprising:

circuitry configured to:

receive, from a plurality of wavefronts waiting to be dispatched, requests to reserve storage space in a cache; and

reduce a dispatch rate of a first set of wavefronts of the plurality of wavefronts in response to an available amount of data storage space in the cache being less than a reservation data size of a first wavefront of the plurality of wavefronts.

2 . The apparatus as recited in claim 1 , wherein in further response to the available amount of data storage space in the cache being less than the reservation data size of the first wavefront, the circuitry is further configured to reduce an instruction issue rate of a second set of wavefronts that have already been dispatched.

3 . The apparatus as recited in claim 2 , wherein to reduce the dispatch rate and the instruction issue rate, the circuitry is further configured to pause dispatching the first set of wavefronts and pause issuing instructions of the second set of wavefronts.

4 . The apparatus as recited in claim 2 , wherein the circuitry is further configured to select the first set of wavefronts and the second set of wavefronts, in response to each of the first set and the second set of wavefronts being assigned to a same workgroup.

5 . The apparatus as recited in claim 2 , wherein the circuitry is further configured to select the first set of wavefronts and the second set of wavefronts, in response to each of the first set and the second set of wavefronts having a priority level below a priority level threshold.

6 . The apparatus as recited in claim 2 , wherein the circuitry is further configured to increase one or more of the dispatch rate of the first set of wavefronts and the instruction issue rate of the second set of wavefronts, in response to the available amount of data storage space in the cache being greater than the reservation data size of the first wavefront.

7 . The apparatus as recited in claim 1 , wherein the reservation data size of the first wavefront of the plurality of wavefronts is different from a reservation data size of a second wavefront of the plurality of wavefronts.

8 . A method, comprising:

receiving, from a plurality of wavefronts waiting to be dispatched, requests to reserve storage space in a cache; and

reducing, by a control circuit, a dispatch rate of a first set of wavefronts of the plurality of wavefronts in response to an available amount of data storage space in the cache being less than a reservation data size of a first wavefront of the plurality of wavefronts.

9 . The method as recited in claim 8 , wherein in further response to the available amount of data storage space in the cache being less than the reservation data size of the first wavefront, the method comprises reducing, by the control circuit, an instruction issue rate of a second set of wavefronts that have already been dispatched.

10 . The method as recited in claim 9 , wherein to reduce the dispatch rate and the instruction issue rate, the method comprises pausing, by the control circuit, dispatching the first set of wavefronts and pausing issuing instructions of the second set of wavefronts.

11 . The method as recited in claim 9 , further comprising selecting, by the control circuit, the first set of wavefronts and the second set of wavefronts, in response to each of the first set and the second set of wavefronts being assigned to a same workgroup.

12 . The method as recited in claim 9 , further comprising selecting, by the control circuit, the first set of wavefronts and the second set of wavefronts, in response to each of the first set and the second set of wavefronts having a priority level below a priority level threshold.

13 . The method as recited in claim 9 , further comprising increasing, by the control circuit, one or more of the dispatch rate of the first set of wavefronts and the instruction issue rate of the second set of wavefronts, in response to the available amount of data storage space in the cache being greater than the reservation data size of the first wavefront.

14 . The method as recited in claim 8 , wherein the reservation data size of the first wavefront of the plurality of wavefronts is different from a reservation data size of a second wavefront of the plurality of wavefronts.

15 . A computing system comprising:

a plurality of chiplets, each comprising:

a cache;

processing circuitry configured to process tasks of one or more wavefronts of a plurality of wavefronts that store temporary data in the cache; and

a control circuit configured to:

receive, from a plurality of wavefronts waiting to be dispatched, requests to reserve storage space in a cache; and

reduce a dispatch rate of a first set of wavefronts of the plurality of wavefronts in response to an available amount of data storage space in the cache being less than a reservation data size of a first wavefront of the plurality of wavefronts.

16 . The computing system as recited in claim 15 , wherein in further response to the available amount of data storage space in the cache being less than the reservation data size of the first wavefront the control circuit is further configured to reduce an instruction issue rate of a second set of wavefronts that have already been dispatched.

17 . The computing system as recited in claim 16 , wherein to reduce the dispatch rate and the instruction issue rate, the control circuit is further configured to pause dispatching of the first set of wavefronts and pause issuance of instructions of the second set of wavefronts.

18 . The computing system as recited in claim 16 , wherein the control circuit is further configured to select the first set of wavefronts and the second set of wavefronts, in response to each of the first set and the second set of wavefronts being assigned to a same workgroup.

19 . The computing system as recited in claim 16 , wherein the control circuit is further configured to select the first set of wavefronts and the second set of wavefronts, in response to each of the first set and the second set of wavefronts having a priority level below a priority level threshold.

20 . The computing system as recited in claim 16 , wherein the control circuit is further configured to increase one or more of the dispatch rate of the first set of wavefronts and the instruction issue rate of the second set of wavefronts, in response to the available amount of data storage space in the cache having increased to being greater than the reservation data size of the first wavefront.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2024
From: HAEHNLE, NICOLAI; BEDY, MICHAEL JOHN
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 066592/0673 →
Continuity (1)
Related Publication 20250190347A1 · Jun 12, 2025
References Cited (36)
US 11875197B2 · Beckmann et al. · 2024 [cited by applicant]
US 20070150880A1 · Mitran et al. · 2007 [cited by applicant]
US 20070198813A1 · Favor et al. · 2007 [cited by applicant]
US 20070288909A1 · Cheung et al. · 2007 [cited by applicant]
US 20090100249A1 · Eichenberger et al. · 2009 [cited by applicant]
US 20130205123A1 · Vorbach · 2013 [cited by applicant]
US 20140122841A1 · Abernathy et al. · 2014 [cited by applicant]
US 20160026507A1 · Muckle et al. · 2016 [cited by applicant]
US 20170053374A1 · Howes et al. · 2017 [cited by applicant]
US 20170220346A1 · Beckmann et al. · 2017 [cited by applicant]
US 20180082470A1 · Nijasure · 2018 [cited by examiner]
US 20180239604A1 · Cain, III et al. · 2018 [cited by applicant]
US 20190206110A1 · Gierach et al. · 2019 [cited by applicant]
US 20190266092A1 · Robinson et al. · 2019 [cited by applicant]
US 20190278605A1 · Emberling et al. · 2019 [cited by applicant]
US 20190324757A1 · Valerio et al. · 2019 [cited by applicant]
US 20190332420A1 · Ukidave et al. · 2019 [cited by applicant]
US 20190370059A1 · Puthoor · 2019 [cited by examiner]
US 20200065073A1 · Pan et al. · 2020 [cited by applicant]
US 20200183697A1 · Targowski et al. · 2020 [cited by applicant]
US 20200293369A1 · Maiyuran · 2020 [cited by examiner]
US 20210149673A1 · Li et al. · 2021 [cited by applicant]
US 20210173796A1 · Puthoor · 2021 [cited by examiner]
US 20220100680A1 · Chrysos et al. · 2022 [cited by applicant]
US 20220179787A1 · Koker · 2022 [cited by examiner]
US 20220206876A1 · Beckmann et al. · 2022 [cited by applicant]
US 20220308884A1 · Øygard · 2022 [cited by examiner]
US 20230069890A1 · Kazakov · 2023 [cited by examiner]
US 20230102843A1 · Vishnuswaroop Ramesh · 2023 [cited by examiner]
US 20230289215A1 · Palmer · 2023 [cited by examiner]
International Search Report and Written Opinion for International Application No. PCT/US2024/033367, mailed Oct. 10, 2024, 11 pp. [cited by applicant]
Non-Final Office Action in U.S. Appl. No. 17/136,725, mailed Feb. 2, 2022, 8 pages. [cited by applicant]
Beckmann et al., U.S. Appl. No. 17/136,725, entitled “Dynamic Graphical Processing Unit Register Allocation”, filed Dec. 29, 2020, 51 pages. [cited by applicant]
Jeon et al., “GPU Register File Virtualization”, MICR0-48: Proceedings of the 48th International Symposium on Microarchitecture, Dec. 2015, pp. 420-432. [cited by applicant]
Kloosterman et al., “Regless: Just-In-Time Operand Staging for GPUs” MICR0-50 '17: Proceedings of the 5oth I\nnual IEEE/ACM International Symposium on Microarchitecture, Oct. 2017, pp. 151-164. [cited by applicant]
Vijaykumar et al., “Zorua: A Holistic Approach to Resource Virtualization in GPUs”, 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), Oct. 2016, 14 pages. [cited by applicant]