IP Library Granted Patent US 12,646,249
Granted Patent B2
US 12,646,249 · App. 18/762,389 · Granted Jun 2, 2026

System and method for executing a task

Inventors: Brian Emberling (Santa Clara, CA); Michael Y. Chow (Santa Clara, CA)
Assignee: Advanced Micro Devices, Inc.
G06T15/80G06F9/5016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,249
App. No.
18/762,389
Granted
Jun 2, 2026
Kind
B2
Abstract

A method, system, and computer-readable medium for executing a task is disclosed. The method includes receiving input data and computing instructions, launching a workgroup including wavefronts to execute the task, wherein the launching causes the wavefronts to process the input data by sharing intermediate results and resources, and adjusting the operation based on characteristics of the wavefronts. The characteristics include data dependencies, computational load, memory usage, and execution timing requirements. The wavefronts execute the task in stages, where each stage processes portions of input data and data generated by other wavefronts.

Claims (32)

1 . A method for executing a task comprising:

receiving, by a system, input data and computing instructions associated with the task;

launching, by the system, a workgroup including wavefronts to execute the task in a workgroup processor (WGP), wherein the launching of the workgroup causes the wavefronts to process the input data by sharing intermediate results and resources; and

adjusting, by the system, an operation of the WGP based on characteristics of the wavefronts.

2 . The method of claim 1 , wherein the characteristics of the wavefronts include any one or a combination of: data dependencies, computational load, memory usage, and execution timing requirements.

3 . The method of claim 1 , wherein the wavefronts are executed in stages of operation and each wavefront uses the computing instructions on a portion of the input data and data generated by other wavefronts.

4 . The method of claim 3 , wherein in a first stage of operation, each wavefront processes a portion of the input data using the computing instructions to generate data, and the operation of the WGP is adjusted based on the characteristics of the wavefronts.

5 . The method of claim 4 , wherein in a second stage of operation, each wavefront processes the generated data in the first stage of operation using computing instructions associated with the task, and the operation of the WGP is further adjusted based on the characteristics of the wavefronts.

6 . The method of claim 1 , wherein a wavefront of the workgroup has access to memory resources that are associated with at least one other wavefront of the workgroup.

7 . The method of claim 1 , wherein memory resources of the WGP enable symmetric memory access for wavefronts of the workgroup.

8 . The method of claim 1 , further comprising activating, by the system, a cache management policy for the WGP based on a workload pattern of the task.

9 . The method of claim 1 , wherein the WGP comprises one or more single instruction multiple data (SIMD) units where each unit is used to execute a subset of the wavefronts.

10 . A system comprising:

a processor;

a work group processor (WGP); and

a memory configured to store instructions that, when executed by the processor, cause the system to:

receive input data and computing instructions associated with a task;

launch a workgroup including wavefronts to execute the task in the WGP, wherein the launching of the workgroup causes the wavefronts to process the input data by sharing intermediate results and resources; and

adjust an operation of the WGP based on characteristics of the wavefronts.

11 . The system of claim 10 , wherein the characteristics of the wavefronts include any one or a combination of: data dependencies, computational load, memory usage, and execution timing requirements.

12 . The system of claim 10 , wherein the wavefronts are executed in stages of operation and each wavefront uses the computing instructions on a portion of the input data and data generated by other wavefronts.

13 . The system of claim 12 , wherein in a first stage of operation, each wavefront processes a portion of the input data using the computing instructions to generate data, and the operation of the WGP is adjusted based on the characteristics of the wavefronts.

14 . The system of claim 13 , wherein in a second stage of operation, each wavefront processes the generated data in the first stage of operation using computing instructions associated with the task, and the operation of the WGP is further adjusted based on the characteristics of the wavefronts.

15 . The system of claim 10 , wherein a wavefront of the workgroup has access to memory resources that are associated with at least one other wavefront of the workgroup.

16 . The system of claim 10 , wherein memory resources of the WGP enable symmetric memory access for wavefronts of the workgroup.

17 . The system of claim 10 , wherein the WGP comprises any one or a combination of the following resources: vector general purpose registers (VGPRs); and local data share (LDS) memory.

18 . The system of claim 10 , wherein the WGP comprises one or more single instruction multiple data (SIMD) units where each unit is used to execute a subset of the wavefronts.

19 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform a method for executing a task comprising:

receiving, by a system, input data and computing instructions associated with the task;

launching, by the system, a workgroup including wavefronts to execute the task in a workgroup processor (WGP), wherein the launching of the workgroup causes the wavefronts to process the input data by sharing intermediate results and resources; and

adjusting, by the system, an operation of the WGP based on characteristics of the wavefronts.

20 . The non-transitory computer readable medium of claim 19 , wherein the characteristics of the wavefronts include any one or a combination of: data dependencies, computational load, memory usage, and execution timing requirements.

Continuity (2)
Continuation 17489724 · Sep 29, 2021
Related Publication 20240355044A1 · Oct 24, 2024
References Cited (11)
US 10831490B2 · Jin · 2020 [cited by applicant]
US 12033275B2 · Emberling · 2024 [cited by examiner]
US 20130147816A1 · Hartog · 2013 [cited by examiner]
US 20130155077A1 · Hartog et al. · 2013 [cited by applicant]
US 20190034151A1 · Dutu et al. · 2019 [cited by applicant]
US 20190155604A1 · Emberling et al. · 2019 [cited by applicant]
US 20190332420A1 · Ukidave et al. · 2019 [cited by applicant]
US 20210373899A1 · Vembu et al. · 2021 [cited by applicant]
KR 1020180128075A · 2018 [cited by applicant]
“Radeon RX 5700: Navi and the RDNA architecture, WikiChip Fuse”, David Schor, Feb. 23, 2020. [Retrieved on Dec. 19, 2022] Retrieved from <URL: https://fuse.wikichip.org/news/3331/radeon-rx -5700-navi-and-the-rdna-archit… [cited by applicant]
Anonymous: ““RDNA 2” Instruction Set Architecture”, Nov. 30, 2020, pp. 1-291, XP093291077, Retrieved from the Internet: URL:https://www.amd.com/content/dam/amd/en/documents/radeon-tech-docs/instructions-set-architecture… [cited by applicant]