IP Library › Granted Patent US 12,248,788
Granted Patent B2
US 12,248,788 · App. 17/691,690 · Granted Mar 11, 2025

Distributed shared memory

Inventors: Prakash Bangalore Prabhakar (San Jose, CA); Gentaro Hirota (San Jose, CA); Ronny Krashinsky (Portola Valley, CA); Ze Long (San Jose, CA); Brian Pharris (Cary, NC); Rajballav Dash (San Jose, CA); Jeff Tuckey (Saratoga, CA); Jerome F. Duluk, Jr. (Palo Alto, CA); Lacky Shah (Los Altos Hills, CA); Luke Durant (San Jose, CA); Jack Choquette (Palo Alto, CA); Eric Werness (San Jose, CA); Naman Govil (Sunnyvale, CA); Manan Patel (San Jose, CA); Shayani Deb (Seattle, WA); Sandeep Navada (San Jose, CA); John Edmondson (Arlington, MA); Greg Palmer (Cedar Park, TX); Wish Gandhi (Sunnyvale, CA); Ravi Manyam (San Ramon, CA); Apoorv Parle (San Jose, CA); Olivier Giroux (Santa Clara, CA); Shirish Gadre (Fremont, CA); Steve Heinrich (Madison, AL)
Assignee: NVIDIA Corporation
G06F9/3851G06F9/522G06F9/544G06F13/1663G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,788
App. No.
17/691,690
Granted
Mar 11, 2025
Kind
B2
Abstract

Distributed shared memory (DSMEM) comprises blocks of memory that are distributed or scattered across a processor (such as a GPU). Threads executing on a processing core local to one memory block are able to access a memory block local to a different processing core. In one embodiment, shared access to these DSMEM allocations distributed across a collection of processing cores is implemented by communications between the processing cores. Such distributed shared memory provides very low latency memory access for processing cores located in proximity to the memory blocks, and also provides a way for more distant processing cores to also access the memory blocks in a manner and using interconnects that do not interfere with the processing cores' access to main or global memory such as hacked by an L2 cache. Such distributed shared memory supports cooperative parallelism and strong scaling across multiple processing cores by permitting data sharing and communications previously possible only within the same processing core.

Claims (38)

1. A processing system comprising:

a first semiconductor memory block co-located with at least one associated first processing core executing a first cooperative thread array (CTA) of a cooperative group array (CGA);

a second semiconductor memory block co-located with at least one associated second processing core executing a second CTA of the CGA; and

an interconnect operatively connecting the first processing core with the second processing core,

wherein the first CTA is enabled to access the second semiconductor memory block via the second processing core.

2. The processing system of claim 1 wherein the processing system conditions access of the second semiconductor memory block by the first CTA on allocation of the second semiconductor memory block to the second CTA.

3. The processing system of claim 1 wherein the processing system is configured to guarantee that the first and second CTAs run concurrently.

4. The processing system of claim 1 wherein the first and second processing cores each comprise streaming multiprocessors, the first semiconductor memory block is part of a first streaming multiprocessor, the second semiconductor memory block is part of a second streaming multiprocessor, and the interconnect provides intercommunications between the first streaming multiprocessor and the second streaming multiprocessor.

5. The processing system of claim 1 wherein the processing system is configured to enable the first CTA to issue read, write, and atomic commands to access the second semiconductor memory block.

6. The processing system of claim 1 wherein the interconnect comprises a switch or a network.

7. The processing system of claim 1 wherein the interconnect enables direct communication between processing cores to all shared memory allocated to CTAs that are part of the same CGA.

8. The processing system of claim 1 wherein the processing system is configured to provide segmented addressing of distributed shared memory blocks based on a unique identifier per executing CGA.

9. The processing system of claim 1 wherein the processing system is configured to support hardware resident CGA identifier allocation and recycling protocols.

10. The processing system of claim 1 wherein the first and second processing cores each store an associating structure to enable distributed shared memory addressing from processing cores remote thereto.

11. The processing system of claim 1 wherein the interconnect comprises a low-latency processor-to-processor communication network.

12. The processing system of claim 1 wherein the processing system coalesces write acknowledgements for communications across the interconnect.

13. The processing system of claim 1 wherein the processing system includes hardware barriers to synchronize usage of the first semiconductor memory block and the second semiconductor memory block at CGA creation and during CGA context switching.

14. The processing system of claim 13 wherein the hardware barriers ensure the first semiconductor memory block and the second semiconductor memory block are available in all CTAs of a CGA before any CTA references the memory blocks over the interconnect.

15. The processing system of claim 13 wherein the hardware barriers ensure all CTAs in a CGA have completed all accesses before state save, and all distributed shared memory has been restored during state restore.

16. The processing system of claim 1 wherein the processing system is further configured to provide CGA/CTA exit and error handling protocols with distributed shared memory.

17. A method of operating a processing system comprising:

launching a first thread block on a first processing core;

allocating to the first thread block a first semiconductor memory block local to the first processing core;

launching a second thread block on a second processing core;

allocating to the second thread block a second semiconductor memory block local to the second processing core;

launching a third thread block on a third processing core;

allocating to the third thread block a third semiconductor memory block local to the third processing core; and

enabling the first thread block to access each of the allocated second and third semiconductor memory blocks via the second and third processing cores, respectively; enabling the second thread block to access each of the allocated first and third semiconductor memory blocks via the first and third processing cores, respectively; and enabling the third thread block to access each of the first and second semiconductor memory blocks via the first and second processing cores, respectively.

18. The method of claim 17 wherein the processing system is disposed on a substrate, and the first, second and third memory blocks are disposed on different portions of the substrate and are separated from one another.

19. The method of claim 17 further including using a hardware barrier to ensure that the first, second and third semiconductor memory blocks are allocated before permitting access by any of the first, second and third thread blocks.

20. The method of claim 17 wherein the enabling comprises providing communications between the first, second and third processing cores.

21. A processing system comprising:

a hardware work distributor circuit that launches a collection of thread groups concurrently on plural processors; and

a communications interface between the plural processors that enables each thread group to access memory allocated to any other thread group in the collection,

wherein the communications interface comprises communication links between the plural processors and inter-processor messaging functions performed by the plural processors, and each thread group is enabled to access memory allocated to other thread groups via respective processor(s) executing the other thread groups to which memory to be accessed is allocated.

22. The processing system of claim 21 wherein the communications interface prohibits thread groups not in the collection from accessing memory portions allocated to any other thread group in the collection.

23. The processing system of claim 21 wherein the plural processors comprise streaming multiprocessors.

24. The processing system of claim 21 wherein the plural processors each comprise a group of processing cores and memory local to and directly addressable by each of the processing cores in the group.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2022
From: BANGALORE PRABHAKAR, PRAKASH; HIROTA, GENTARO; KRASHINSKY, RONNY; LONG, ZE; PHARRIS, BRIAN; DASH, RAJBALLAV; TUCKEY, JEFF; DULUK JR., JEROME F.; SHAH, LACKY; DURANT, LUKE; CHOQUETTE, JACK; WERNESS, ERIC; GOVIL, NAMAN; PATEL, MANAN; DEB, SHAYANI; NAVADA, SANDEEP; EDMONDSON, JOHN; PALMER, GREG; GANDHI, WISH; MANYAM, RAVI; PARLE, APOORV; GIROUX, OLIVIER; GADRE, SHIRISH; HEINRICH, STEVE
To: NVIDIA CORPORATION
Reel/Frame 060364/0226 →
Continuity (1)
Related Publication 20230289189A1 · Sep 14, 2023
References Cited (45)
US 7447873B1 · Nordquist · 2008 [cited by applicant]
US 7506134B1 · Juffa et al. · 2009 [cited by applicant]
US 7761688B1 · Park · 2010 [cited by examiner]
US 7788468B1 · Nickolls et al. · 2010 [cited by applicant]
US 7836118B1 · Juffa et al. · 2010 [cited by applicant]
US 7937567B1 · Nickolls et al. · 2011 [cited by applicant]
US 8112614B2 · Nickolls et al. · 2012 [cited by applicant]
US 8997103B2 · Gadre et al. · 2015 [cited by applicant]
US 9223578B2 · Nickolls et al. · 2015 [cited by applicant]
US 9324175B2 · Bolz et al. · 2016 [cited by applicant]
US 9513975B2 · Jones et al. · 2016 [cited by applicant]
US 9928109B2 · Durant · 2018 [cited by applicant]
US 10217183B2 · Palmer et al. · 2019 [cited by applicant]
US 10817338B2 · Duluk, Jr. et al. · 2020 [cited by applicant]
US 10909033B1 · Fang et al. · 2021 [cited by applicant]
US 10977037B2 · Tirumala et al. · 2021 [cited by applicant]
US 20120131309A1 · Johnson · 2012 [cited by examiner]
US 20140122809A1 · Robertson et al. · 2014 [cited by applicant]
US 20150178879A1 · Palmer et al. · 2015 [cited by applicant]
US 20150339126A1 · Bonanno · 2015 [cited by examiner]
US 20200043123A1 · Dash et al. · 2020 [cited by applicant]
US 20200293367A1 · Andrei · 2020 [cited by examiner]
US 20210124582A1 · Kerr et al. · 2021 [cited by applicant]
US 20220058053A1 · Andrei · 2022 [cited by examiner]
US 20230289190A1 · Parle et al. · 2023 [cited by applicant]
US 20230289211A1 · Hirota et al. · 2023 [cited by applicant]
US 20230289242A1 · Guo et al. · 2023 [cited by applicant]
US 20230315655A1 · Choquette et al. · 2023 [cited by applicant]
“CUDA C++ Best Practices Guide”, Release 12.6, https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html (Jul. 23, 2024). [cited by applicant]
https://hpc-wiki.info/hpc/Scaling_tests (last edited on Jul. 19, 2024). [cited by applicant]
Lindholm et al, “NVIDIA Tesla: A Unified Graphics and Computing Architecture,” IEEE Micro (2008). [cited by applicant]
“PTX ISA” CUDA, Release 8.5, https://docs.nvidia.com/cuda/parallel-thread-execution/index.html (Jul. 23, 2024). [cited by applicant]
Choquette et al, “Volta: Performance and Programmability”, IEEE Micro (vol. 38, Issue: 2, Mar./Apr. 2018), DOI: 10.1109/MM.2018.022071134. [cited by applicant]
“Cooperative Groups: Flexible CUDA Thread Programming”, https://developer.nvidia.com/blog/cooperative-groups/ (Oct. 4, 2017). [cited by applicant]
“CUDA 9 Features Revealed: Volta, Cooperative Groups and More”, https://developer.nvidia.com/blog/cuda-9-features-revealed/ (May 11, 2017). [cited by applicant]
Bob Crovella et al., “Cooperative Groups” (Sep. 17, 2020), https://vimeo.com/461821629. [cited by applicant]
Zhang et al, A Study of Single and Multi-device Synchronization Methods in NVIDIA GPUs, (arXiv:2004.05371v1 [cs.DC] Apr. 11, 2020). [cited by applicant]
Lustig et al, “A Formal Analysis of the NVIDIA PTX Memory Consistency Model”, Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, pp. 257-2… [cited by applicant]
Weber et al, “Toward a Multi-GPU Implementation of the Modular Integer GCD Algorithm Extended Abstract” ICPP Aug. 13-16, 2018, Eugene, OR USA (ACM 2018). [cited by applicant]
Jog et al, “OWL: Cooperative Thread Array Aware Scheduling Techniques for Improving GPGPU Performance” (ASPLOS'13, Mar. 16-20, 2013, Houston, Texas, USA). [cited by applicant]
Parallel Thread Execution ISA: Application Guide (NVidia v5.0 Jun. 2017). [cited by applicant]
Lal et al, “A Quantitative Study of Locality in GPU Caches”, in: Orailoglu et al (eds), Embedded Computer Systems: Architectures, Modeling, and Simulation, (SAMOS 2020), Lecture Notes in Computer Science, vol. 12471. Sp… [cited by applicant]
Adinets, “CUDA Dynamic Parallelism API and Principles” https://developer.nvidia.com/blog/cuda-dynamic-parallelism-api-principles/ (May 20, 2014). [cited by applicant]
Breshears, The Art of Concurrency: A Thread Monkey's Guide to Writing Parallel Applications (O'Reilly 2009). [cited by applicant]
Rajwar et al, “Speculative lock elision: enabling highly concurrent multithreaded execution,” Proceedings of the 34th ACM/IEEE International Symposium on Microarchitecture MICRO-34 (Dec. 1-5, 2001). [cited by applicant]
Cited By (1)
US 12,657,023