IP Library Granted Patent US 12,333,311
Granted Patent B2
US 12,333,311 · App. 17/691,621 · Granted Jun 17, 2025

Cooperative group arrays

Inventors: Greg Palmer (Cedar Park, TX); Gentaro Hirota (San Jose, CA); Ronny Krashinsky (Portola Valley, CA); Ze Long (San Jose, CA); Brian Pharris (Cary, NC); Rajballav Dash (San Jose, CA); Jeff Tuckey (Saratoga, CA); Jerome F. Duluk, Jr. (Palo Alto, CA); Lacky Shah (Los Altos Hills, CA); Luke Durant (San Jose, CA); Jack Choquette (Palo Alto, CA); Eric Werness (San Jose, CA); Naman Govil (Sunnyvale, CA); Manan Patel (San Jose, CA); Shayani Deb (Seattle, WA); Sandeep Navada (San Jose, CA); John Edmondson (Arlington, MA); Prakash Bangalore Prabhakar (San Jose, CA); Wish Gandhi (Sunnyvale, CA); Ravi Manyam (San Ramon, CA); Apoorv Parle (San Jose, CA); Olivier Giroux (Santa Clara, CA); Shirish Gadre (Fremont, CA); Steve Heinrich (Madison, AL)
Assignee: NVIDIA Corporation
G06F9/3888G06F9/3009G06F9/3851G06F9/4881G06F9/544
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,311
App. No.
17/691,621
Granted
Jun 17, 2025
Kind
B2
Abstract

A new level(s) of hierarchy—Cooperate Group Arrays (CGAs)—and an associated new hardware-based work distribution/execution model is described. A CGA is a grid of thread blocks (also referred to as cooperative thread arrays (CTAs)). CGAs provide co-scheduling, e.g., control over where CTAs are placed/executed in a processor (such as a GPU), relative to the memory required by an application and relative to each other. Hardware support for such CGAs guarantees concurrency and enables applications to see more data locality, reduced latency, and better synchronization between all the threads in tightly cooperating collections of CTAs programmably distributed across different (e.g., hierarchical) hardware domains or partitions.

Claims (27)

1. A processing system comprising:

a work distributor hardware circuit configured to launch a collection of thread groups on a set of plural processors while providing a hardware-based guarantee that all thread groups of the collection can be launched at the same time;

the work distributor hardware circuit being further configured to speculatively launch the thread groups in the collection to confirm that the thread groups are able to launch and/or run concurrently on the set of plural processors before launching any of the thread groups in the collection.

2. The processing system of claim 1 wherein the work distributor hardware circuit comprises a multilevel hardware circuit configured to distribute the collection of thread groups to processors in a predefined hardware cluster.

3. The processing system of claim 1 wherein the set of processors comprise a predefined hardware domain and the work distributor hardware circuit is configured to launch the thread groups on any processor(s) within the predefined hardware domain.

4. The processing system of claim 3 wherein the predefined hardware domain comprises a GPU, a μGPU, a GPC or a TPC.

5. The processing system of claim 3 wherein the predefined hardware domain comprises a nested hierarchy of processors, and the work distributor hardware circuit is configured to schedule the thread groups to execute concurrently on processors in different levels of the nested hierarchy of processors.

6. The processing system of claim 1 wherein the collection of thread groups comprises a cooperative group array representable as a multidimensional grid.

7. The processing system of claim 1 wherein the work distributor hardware circuit is further configured to broadcast a grid launch packet to the plural processors.

8. A method of executing instructions on at least one processing system comprising:

determining a cooperative group array of plural thread blocks, each thread block comprising plural threads;

speculatively launching the cooperative group array of thread blocks to determine whether the plural thread blocks will be able to execute concurrently on plural parallel processors; and

when the speculative launching reveals the cooperative group array of thread blocks will be able to execute concurrently on plural parallel processors, launching the cooperative group array of thread blocks on the plural parallel processors.

9. The method of claim 8 wherein the launching throttles, at a hardware level, shared memory usage of the concurrently launched cooperative group array of thread blocks.

10. The method of claim 8 wherein the plural parallel processors comprise streaming multiprocessors.

11. The method of claim 8 wherein the cooperative group array is determined based on a grid.

12. The method of claim 8 wherein the respective plural processors are all within a same hardware domain associated with the cooperative group array.

13. The method of claim 12 wherein the hardware domain comprises a GPC, a μGPU, a GPC, or a TPC.

14. A processing system comprising:

a memory storing a cooperative group array (CGA) comprising plural cooperative thread arrays; and

a work distributor hardware circuit configured to provide a hardware-based guarantee that the plural cooperative thread arrays of the CGA can be launched concurrently, the work distributor hardware circuit being configured to speculatively launch the cooperative group array across multiple processing cores to determine whether the cooperative group array can launch concurrently, the work distributor hardware circuit actually launching the cooperative group array across the multiple processing cores only if the speculative launch determines that all cooperative thread arrays can launch concurrently.

15. The processing system of claim 14 wherein the processing cores comprise streaming multiprocessors.

16. The processing system of claim 14 wherein the work distributor hardware circuit comprises registers, combinatorial logic and a hardware state machine.

17. The processing system of claim 14 wherein the work distributor hardware circuit comprises a multi-level work distribution architecture to provide CGA launch on associated hardware affinity/domain and support nesting of multiple levels of CGAs.

18. The processing system of claim 14 wherein the work distributor hardware circuit comprises a load balancer, resource trackers, a TPC enable table, a local memory (LMEM) block index table, credit counters, a task table, and a priority-sorted task table.

19. The processing system of claim 14 wherein the work distributor hardware circuit is configured to receive a launch command specifying a CGA grid, including an enumeration of various dimensions of composite thread blocks and CGAs within the specified CGA grid.

20. The processing system of claim 14 wherein the work distributor hardware circuit is configured to query and launch CTAs from multiple CGAs but works on one CGA at a time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2022
From: PALMER, GREG; HIROTA, GENTARO; KRASHINSKY, RONNY; LONG, ZE; PHARRIS, BRIAN; DASH, RAJBALLAV; TUCKEY, JEFF; DULUK JR., JEROME F.; SHAH, LACKY; DURANT, LUKE; CHOQUETTE, JACK; WERNESS, ERIC; GOVIL, NAMAN; PATEL, MANAN; DEB, SHAYANI; NAVADA, SANDEEP; EDMONDSON, JOHN; BANGALORE PRABHAKAR, PRAKASH; GANDHI, WISH; MANYAM, RAVI; PARLE, APOORV; GIROUX, OLIVIER; GADRE, SHIRISH; HEINRICH, STEVE
To: NVIDIA CORPORATION
Reel/Frame 060364/0185 →
Continuity (1)
Related Publication 20230289215A1 · Sep 14, 2023
References Cited (37)
US 7447873B1 · Nordquist · 2008 [cited by applicant]
US 7506134B1 · Juffa et al. · 2009 [cited by applicant]
US 7788468B1 · Nickolls et al. · 2010 [cited by applicant]
US 7836118B1 · Juffa et al. · 2010 [cited by applicant]
US 7937567B1 · Nickolls et al. · 2011 [cited by applicant]
US 8112614B2 · Nickolls et al. · 2012 [cited by applicant]
US 8643656B2 · Li · 2014 [cited by examiner]
US 8997103B2 · Gadre et al. · 2015 [cited by applicant]
US 9223578B2 · Nickolls et al. · 2015 [cited by applicant]
US 9324175B2 · Bolz et al. · 2016 [cited by applicant]
US 9513975B2 · Jones et al. · 2016 [cited by applicant]
US 9928109B2 · Durant · 2018 [cited by applicant]
US 10217183B2 · Palmer et al. · 2019 [cited by applicant]
US 10817338B2 · Duluk, Jr. et al. · 2020 [cited by applicant]
US 10909033B1 · Fang et al. · 2021 [cited by applicant]
US 10977037B2 · Tirumala et al. · 2021 [cited by applicant]
US 20140122809A1 · Robertson et al. · 2014 [cited by applicant]
US 20150178879A1 · Palmer et al. · 2015 [cited by applicant]
US 20200043123A1 · Dash et al. · 2020 [cited by applicant]
US 20210124582A1 · Kerr et al. · 2021 [cited by applicant]
“Cuda C ++ Best Practices Guide” Release 12.6, https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html, Sep. 24, 2024. [cited by applicant]
“Scaling tests” HPCWIKI, https://hpc-wiki.info/hpc/Scaling_tests, lasted edited on Jul. 19, 2024. [cited by applicant]
Lindholm et al, “NVIDIA Tesla: A Unified Graphics and Computing Architecture,” IEEE Micro (2008). [cited by applicant]
“PTX ISA Release 8.5”, CUDA, https://docs.nvidia.com/cuda/parallel-thread-execution/index.html, Sep. 24, 2024. [cited by applicant]
Choquette et al, “Volta: Performance and Programmability”, IEEE Micro (vol. 38, Issue: 2, DOI: 10.1109/MM.2018.022071134, Mar./Apr. 2018. [cited by applicant]
Harris et al., “Cooperative Groups: Flexible CUDA Thread Programming”, https://developer.nvidia.com/blog/cooperative-groups/, Oct. 4, 2017. [cited by applicant]
Harris, “CUDA 9 Features Revealed: Volta, Cooperative Groups and More”, https://developer.nvidia.com/blog/cuda-9-features-revealed/, May 11, 2017. [cited by applicant]
Bob Crovella et al, “Cooperative Groups” https://vimeo.com/461821629, (Sep. 17, 2020). [cited by applicant]
Zhang et al, A Study of Single and Multi-device Synchronization Methods in NVIDIA GPUs, (arXiv:2004.05371v1 [cs. DC] Apr. 11, 2020). [cited by applicant]
Lustig et al, “A Formal Analysis of the NVIDIA PTX Memory Consistency Model”, Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, pp. 257-2… [cited by applicant]
Weber et al, “Toward a Multi-GPU Implementation of the Modular Integer GCD Algorithm Extended Abstract” ICPP August 13-16, 2018, Eugene, OR USA (ACM 2018). [cited by applicant]
Jog et al., “OWL: Cooperative Thread Array Aware Scheduling Techniques for Improving GPGPU Performance” (ASPLOS'13, Houston, Texas, USA) Mar. 16-20, 2013. [cited by applicant]
Parallel Thread Execution ISA: Application Guide (NVIDIA v5.0 Jun. 2017. [cited by applicant]
Lal et al, “A Quantitative Study of Locality in GPU Caches”, in: Orailoglu et al (eds), Embedded Computer Systems: Architectures, Modeling, and Simulation, Lecture Notes in Computer Science, vol. 12471. Springer, Cham. … [cited by applicant]
Adinets, “CUDA Dynamic Parallelism API and Principles”, https://developer.nvidia.com/blog/cuda-dynamic-parallelism-api-principles/, May 20, 2014. [cited by applicant]
Breshears, The Art of Concurrency: A Thread Monkey's Guide to Writing Parallel Applications (O'Reilly 2009). [cited by applicant]
Rajwar et al, “Speculative lock elision: enabling highly concurrent multithreaded execution,” Proceedings of the 34th ACM/IEEE International Symposium on Microarchitecture MICRO-34 (Dec. 1-5, 2001). [cited by applicant]
Cited By (1)
US 12,657,023