IP Library › Granted Patent US 12,190,405
Granted Patent B2
US 12,190,405 · App. 17/853,711 · Granted Jan 7, 2025

Direct memory writes by network interface of a graphics processing unit

Inventors: Todd Rimmer (Exton, PA); Mark Debbage (Santa Clara, CA); Bruce G. Warren (Poulsbo, WA); Sayantan Sur (Portland, OR); Nayan Amrutlal Suthar (Pune, IN); Ajaya Durg (Austin, TX)
Assignee: Intel Corporation
G06T1/20G06T1/60G06F2213/3808
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,405
App. No.
17/853,711
Granted
Jan 7, 2025
Kind
B2
Abstract

Examples described herein relate to a first graphics processing unit (GPU) with at least one integrated communications system, wherein the at least one integrated communications system is to apply a reliability protocol to communicate with a second at least one integrated communications system associated with a second GPU to copy data from a first memory region to a second memory region and wherein the first memory region is associated with the first GPU and the second memory region is associated with the second GPU.

Claims (43)

1. An apparatus comprising:

a first graphics processing unit (GPU) with at least one integrated communications system, wherein the at least one integrated communications system is to apply a reliability protocol to communicate with a second at least one integrated communications system associated with a second GPU to copy data from a first memory region to a second memory region and wherein the first memory region is associated with the first GPU and the second memory region is associated with the second GPU.

2. The apparatus of claim 1 , wherein

the at least one integrated communications system comprises a communications system integrated into a same integrated circuit or system on chip (SoC) as that of the first GPU and

the second at least one integrated communications system comprises a communications system integrated into a same integrated circuit or SoC as that of the first GPU and the second GPU.

3. The apparatus of claim 1 , wherein the at least one integrated communications system comprises:

direct memory access (DMA) circuitry;

reliable transport circuitry; and

a network interface controller.

4. The apparatus of claim 1 , wherein the at least one integrated communications system is to perform topology discovery to discover the second at least one integrated communications system.

5. The apparatus of claim 1 , comprising a device interface to communicatively couple the at least one integrated communications system to one or more execution units of the first GPU.

6. The apparatus of claim 1 , comprising a first memory associated with the first GPU and a second memory associated with the second GPU, wherein the first memory comprises a source of data and the second memory comprises a destination of data.

7. The apparatus of claim 1 , comprising:

a network interface device to receive and apply control configurations for the at least one integrated communications system and the second at least one integrated communications system.

8. The apparatus of claim 1 , wherein the at least one integrated communications system and the second at least one integrated communications system are to communicate with at least one GPU of another system.

9. The apparatus of claim 1 , comprising:

a network interface device to provide communications among the first GPU, the second GPU, and at least one GPU of another system.

10. The apparatus of claim 1 , comprising:

at least one central processing unit communicatively coupled to the first GPU and the second GPU, wherein the at least one central processing unit is to communicate with at least one GPU of another system using the at least one integrated communications system and the second at least one integrated communications system.

11. The apparatus of claim 1 , comprising:

a GPU-to-GPU connection to provide communication between the at least one integrated communications system and the second at least one integrated communications system.

12. The apparatus of claim 1 , comprising:

fabric to provide communication among the first and second GPUs and at least one GPU of another system.

13. At least one non-transitory computer-readable medium comprising instructions stored thereon, that if executed, cause one or more processors to:

configure at least one network interface device of a first graphics processing unit (GPU) to configure data planes of one or more other network interface devices, wherein the at least one network interface device is to use a reliability protocol to communicate with another at least one network interface device associated with a second GPU to copy data from a first memory to a second memory and wherein the first memory is associated with the first GPU and the second memory is associated with the second GPU.

14. The computer-readable medium of claim 13 , comprising instructions stored thereon, that if executed, cause one or more processors to:

perform topology discovery to discover at least one network interface device.

15. The computer-readable medium of claim 13 , wherein:

the at least one network interface device of a first GPU comprises a communications system integrated into a same integrated circuit or system on chip (SoC) as that of the first GPU and

the another at least one network interface device associated with a second GPU comprises a communications system integrated into a same integrated circuit or SoC as that of the second GPU.

16. A method comprising:

in an integrated circuit with multiple graphics processing units (GPUs):

providing communications among the multiple GPUs by communications systems integrated into the multiple GPUs and a GPU-to-GPU connection,

providing communications among the multiple GPUs and at least one GPU of another integrated circuitry by a communications systems integrated into the multiple GPUs and a switching network,

configuring the communications systems integrated into the multiple GPUs by a network interface device coupled to a control network.

17. The method of claim 16 , wherein:

a first communications systems of the communications systems integrated into the multiple GPUs is integrated into a same integrated circuit or system on chip (SoC) as that of a first GPU of the multiple GPUs and

a second communications systems of the communications systems integrated into the multiple GPUs is integrated into a same integrated circuit or SoC as that of a second GPU of the multiple GPUs.

18. The method of claim 17 , wherein the communications among the multiple GPUs utilize reliable transport.

19. The method of claim 17 , comprising:

one or more switches, in a scale out network, performing topology discovery to discover at least one GPU of the multiple GPUs.

20. The method of claim 17 , comprising:

providing communications among central processing units, accelerators, and GPUs by selection among at least one of the communications systems integrated into the multiple GPUs or a network interface device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2023
From: RIMMER, TODD; DEBBAGE, MARK; WARREN, BRUCE G.; SUR, SAYANTAN; SUTHAR, NAYAN AMRUTLAL; DURG, AJAYA
To: INTEL CORPORATION
Reel/Frame 062261/0861 →
Continuity (2)
Provisional Application 63218832 · Jul 6, 2021
Related Publication 20220351326A1 · Nov 3, 2022
References Cited (41)
US 6393483B1 · Latif et al. · 2002 [cited by applicant]
US 9602437B1 · Bernath · 2017 [cited by applicant]
US 10997106B1 · Bandaru et al. · 2021 [cited by applicant]
US 11150963B2 · Nainar et al. · 2021 [cited by applicant]
US 11531752B2 · Abodunrin et al. · 2022 [cited by applicant]
US 11645104B2 · Rajagopal · 2023 [cited by applicant]
US 11863376B2 · Ang et al. · 2024 [cited by applicant]
US 11995024B2 · Ang et al. · 2024 [cited by applicant]
US 20070139423A1 · Kong et al. · 2007 [cited by applicant]
US 20070273699A1 · Sasaki et al. · 2007 [cited by applicant]
US 20100013839A1 · Rawson · 2010 [cited by applicant]
US 20140223131A1 · Agarwal et al. · 2014 [cited by applicant]
US 20160337193A1 · Rao · 2016 [cited by applicant]
US 20180314638A1 · LeBeane et al. · 2018 [cited by applicant]
US 20180349062A1 · Sines · 2018 [cited by applicant]
US 20190042741A1 · Abodunrin et al. · 2019 [cited by applicant]
US 20200226096A1 · Koker et al. · 2020 [cited by applicant]
US 20200278892A1 · Nainar et al. · 2020 [cited by applicant]
US 20210203704A1 · Pohl · 2021 [cited by applicant]
US 20220085916A1 · Debbage et al. · 2022 [cited by applicant]
US 20220100491A1 · Voltz et al. · 2022 [cited by applicant]
US 20220197681A1 · Rajagopal · 2022 [cited by applicant]
US 20230041806A1 · Singh · 2023 [cited by applicant]
US 20230195675A1 · Ang et al. · 2023 [cited by applicant]
US 20230198833A1 · Ang et al. · 2023 [cited by applicant]
WO 2020172692A2 · 2020 [cited by applicant]
WO 2020190798A1 · 2020 [cited by applicant]
WO WO2020190801A1 · 2020 [cited by examiner]
AWS, “Amazon EC2 P4 Instances”, https://aws.amazon.com/ec2/instance-types/p4/, Nov. 2, 2020, 8 pages. [cited by applicant]
Bouffler, Brendan, “In the search for performance, there's more than one way to build a network”, AWS HPC Blog, https://aws.amazon.com/blogs/hpc/in-the-search-for-performance-theres-more-than-one-way-to-build-a-network/… [cited by applicant]
Gebara, Nadeen et al., “In-Network Aggregation for Shared Machine Learning Clusters”, Proceedings of the 4th MLSys Conference, San Jose, CA, USA, 2021, 16 pages. [cited by applicant]
Nvidia, “DGX Station A100 User Guide”, DU-10189-001 v5.0.2, https://docs.nvidia.com/dgx/dgx-station-a100-user-guide/index.html, Nov. 2022, 74 pages. [cited by applicant]
Singhvi, Arjun et al., “1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters”, SIGCOMM '20, Aug. 10-14, 2020, Virtual Event, NY, USA, 14 pages. [cited by applicant]
‘Enhancing Networks with SmartNICs’ by Mellanox, copyright 2018. (Year: 2018). [cited by applicant]
First Office Action for U.S. Appl. No. 17/853,793, Mailed Jul. 26, 2024, 26 pages. [cited by applicant]
‘HyperV: A High Performance Hypervisor for Virtualization of the Programmable Data Plane’ by Cheng Zhang et al., 2017 26th International Conference on Computer Communication and Networks (ICCCN), Jul. 2017. (Year: 2017). [cited by applicant]
‘HyperVDP: High-Performance Virtualization of the Programmable Data Plane’ by Cheng Zhang et al., IEEE Journal on Selected Areas in Communications, vol. 37, No. 3, Mar. 2019. (Year: 2019). [cited by applicant]
‘P4NFV: P4 Enabled NFV Systems with SmartNICs’ by Ali Mohammadkhan et al., 2019 IEEE Conference on Network Function Virtualization and Software Defined Networks (NFV-SDN). (Year: 2019). [cited by applicant]
‘Uno: Unifying Host and Smart NIC Offload for Flexible Packet Processing’ by Yanfang Le et al., SoCC '17, Sep. 24-27, 2017. (Year: 2017). [cited by applicant]
International Search Report and Written Opinion for PCT Patent Application No. PCT/US22/36397, Mailed Oct. 24, 2022, 12 pages. [cited by applicant]
U.S. Appl. No. 17/853,793, filed Jun. 29, 2022, 60 pages. [cited by applicant]
Cited By (1)
US 12,277,080