IP Library › Granted Patent US 12,641,025
Granted Patent B2
US 12,641,025 · App. 18/356,475 · Granted May 26, 2026

Data transmission system and method, and related device

Inventor: Qihang Duan (Hangzhou, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
H04L47/10H04L41/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,641,025
App. No.
18/356,475
Granted
May 26, 2026
Kind
B2
Abstract

A system includes a plurality of nodes, and a plurality of accelerators in each node are connected to each other through a first communication link. A plurality of communication planes is constructed between accelerators in the plurality of nodes. Each communication plane includes one accelerator in each node. A first accelerator in a first node obtains first data sent by another accelerator in the first node, and the first data includes data to be sent by the other accelerator in the first node to a second accelerator in a second node. Then, the first accelerator sends the first data to the second accelerator through the second communication link.

Claims (60)

1 . A system, comprising:

a first node comprising:

a first accelerator located on a first communication plane of the system and configured to send first data; and

a second accelerator located on a second communication plane of the system, interconnected to the first accelerator through a first communication link, and configured to:

receive, from the first accelerator and through the first communication link, the first data; and

send the first data; and

a second node comprising:

a third accelerator located on the second communication plane of the system and configured to receive, from the second accelerator and through a second communication link, the first data, wherein a first data transmission speed of the first communication link is higher than a second data transmission speed of the second communication link; and

a fourth accelerator located on the first communication plane of the system, wherein the second accelerator is further configured to:

aggregate, from all accelerators on the second node, second data to be sent to the fourth accelerator; and

send, to the first accelerator, based on the first accelerator being located on the same first communication plane of the system as the fourth accelerator, and after all intra-node data exchange is completed on the second node, the second data to be sent to the fourth accelerator, and

wherein the first accelerator is further configured to:

receive, from the second accelerator, the second data; and

send, to the fourth accelerator, the second data.

2 . The system of claim 1 , wherein the first node further comprises a fifth accelerator, wherein the fifth accelerator is configured to send, to the second accelerator and through the first communication link, third data to be sent to the third accelerator, and wherein the second accelerator is further configured to send, to the third accelerator and through the second communication link, a first data set comprising the first data and the third data.

3 . The system of claim 1 , wherein the first node further comprises a fifth accelerator, wherein the fifth accelerator is configured to send, to the first accelerator and through the first communication link, third data to be sent to the fourth accelerator, wherein the first accelerator is further configured to send, to the fourth accelerator and through the second communication link, a first data set, and wherein the first data set comprises the second data and the third data.

4 . The system of claim 1 , wherein the first node and the second node are configured to perform neural network model training in a model parallelism manner.

5 . The system of claim 1 , wherein the first node and the second node are located in different computing devices.

6 . The system of claim 1 , wherein the first accelerator, the second accelerator, the third accelerator, and the fourth accelerator comprise a graphics processing unit (GPU), a neural-network processing unit (NPU), or a Tensor Processing Unit (TPU).

7 . The system of claim 1 , wherein the second accelerator is further configured to receive, from a processor, group information indicating which accelerators are on each communication plane in the system.

8 . The system of claim 7 , wherein the second accelerator is further configured to identify, based on the group information, that the first accelerator is located on the same first communication plane of the system as the fourth accelerator.

9 . The system of claim 8 , wherein the group information comprises addresses of the accelerators, and wherein the second accelerator is configured to identify that the first accelerator is located on the same first communication plane of the system as the fourth accelerator based on the addresses of the accelerators.

10 . A system, comprising:

a first computing device comprising a first node, wherein the first node comprises:

a first accelerator located on a first communication plane of the system and configured to send first data; and

a second accelerator located on a second communication plane of the system, interconnected to the first accelerator through a first communication link, and configured to:

receive, from the first accelerator and through the first communication link, the first data; and

send the first data; and

a second computing device comprising a second node, wherein the second node comprises:

a third accelerator located on the second communication plane of the system and configured to receive, from the second accelerator and through a second communication link, the first data, wherein a first data transmission speed of the first communication link is higher than a second data transmission speed of the second communication link; and

a fourth accelerator located on the first communication plane of the system,

wherein the second accelerator is further configured to:

aggregate, from all accelerators on the second node, second data to be sent to the fourth accelerator; and

send, to the first accelerator, based on the first accelerator being located on the same first communication plane of the system as the fourth accelerator, and after all intra-node data exchange is complete on the second node, the second data to be sent to the fourth accelerator, and

wherein the first accelerator is further configured to:

receive, from the second accelerator, the second data; and

send, to the fourth accelerator, the second data.

11 . The system of claim 10 , wherein the first node further comprises a fifth accelerator, wherein the fifth accelerator is configured to send, to the second accelerator and through the first communication link, third data to be sent to the third accelerator, and wherein the second accelerator is further configured to send, to the third accelerator and through the second communication link, a first data set comprising the first data and the third data.

12 . The system of claim 10 , wherein the first node further comprises a fifth accelerator, wherein the fifth accelerator is configured to send, to the first accelerator and through the first communication link, third data to be sent to the fourth accelerator, wherein the first accelerator is further configured to send, to the fourth accelerator and through the second communication link, a first data set, and wherein the first data set comprises the second data and the third data.

13 . The system of claim 10 , wherein the first node and the second node perform neural network model training in a model parallelism manner.

14 . The system of claim 10 , further comprising a network interface card, wherein the second accelerator corresponds to the network interface card, and wherein the second accelerator is further configured to send, to the network interface card, the first data for the network interface card to send the first data to the third accelerator.

15 . The system of claim 14 , wherein the network interface card is configured to send, to the third accelerator and by using a switch between the first computing device and the second computing device, the first data.

16 . The system of claim 10 , wherein the first computing device further comprises a central processing unit (CPU), and wherein the CPU is configured to manage the first node.

17 . The system of claim 10 , wherein the second accelerator is further configured to receive, from a processor, group information indicating which accelerators are on each communication plane in the system.

18 . The system of claim 17 , wherein the second accelerator is further configured to identify, based on the group information, that the first accelerator is located on the same first communication plane of the system as the fourth accelerator.

19 . The system of claim 18 , wherein the group information comprises addresses of the accelerators, and wherein the second accelerator is configured to identify that the first accelerator is located on the same first communication plane of the system as the fourth accelerator based on the addresses of the accelerators.

20 . A method, comprising:

sending, by a first accelerator in a first node, first data to a second accelerator in the first node through a first communication link;

sending, by the second accelerator, the first data to a third accelerator in a second node through a second communication link, wherein a first data transmission speed of the first communication link is higher than a second data transmission speed of the second communication link;

aggregating, from all accelerators on the second node, second data to be sent to a fourth accelerator;

sending, by the second accelerator, to the first accelerator, based on the first accelerator being located on a same first communication plane as the fourth accelerator in the second node, and after all intra-node data exchange is completed on the second node, the second data to be sent to the fourth accelerator;

receiving, by the first accelerator and from the second accelerator, the second data; and

sending, by the first accelerator and to the fourth accelerator, the second data.

21 . The method of claim 20 , further comprising sending, by a fifth accelerator in the first node, to the second accelerator, and through the first communication link, third data to be sent to the third accelerator, wherein sending the first data comprises sending, to the third accelerator and through the second communication link, a first data set comprising the first data and the third data.

22 . The method of claim 20 , further comprising sending, by a fifth accelerator in the first node, to the first accelerator, and through the first communication link, third data to be sent to the third accelerator, wherein sending the second data comprises sending, by the first accelerator, to the fourth accelerator, and through the second communication link, a first data set, and wherein the first data set comprises the second data and the third data.

23 . The method of claim 20 , further comprising performing, by the first node and the second node, neural network model training in a model parallelism manner.

24 . The method of claim 20 , wherein sending the first data comprises sending, by the second accelerator, the first data to a network interface card corresponding to the second accelerator for the network interface card to send the first data to the third accelerator.

25 . The method of claim 20 , further comprising receiving, by the second accelerator and from a processor, group information indicating which accelerators are on each communication plane in a system.

26 . The method of claim 25 , further comprising identifying, by the second accelerator and based on the group information, that the first accelerator is located on the same first communication plane of the system as the fourth accelerator.

27 . The method of claim 26 , wherein the group information comprises addresses of the accelerators, and wherein the method further comprises identifying, by the second accelerator, that the first accelerator is located on the same first communication plane of the system as the fourth accelerator based on the addresses of the accelerators.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2023
From: DUAN, QIHANG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 065209/0471 →
Priority Claims (1)
CN 202210073931.9 · Jan 21, 2022 · national
Continuity (2)
Continuation PCTCN2022106309 · Jul 18, 2022
Related Publication 20230403232A1 · Dec 14, 2023
References Cited (11)
US 20180287964A1 · Gray · 2018 [cited by examiner]
US 20190205146A1 · Patel et al. · 2019 [cited by applicant]
US 20190258251A1 · Ditty · 2019 [cited by examiner]
US 20200166977A1 · Jayaraman et al. · 2020 [cited by applicant]
US 20200410323A1 · Vinod · 2020 [cited by examiner]
US 20210103544A1 · Guim Bernat et al. · 2021 [cited by applicant]
US 20220116138A1 · Das Sharma · 2022 [cited by examiner]
US 20220121903A1 · Zhang · 2022 [cited by examiner]
US 20220351326A1 · Rimmer · 2022 [cited by examiner]
US 20230004433A1 · Kan · 2023 [cited by examiner]
US 20230109990A1 · Pappu · 2023 [cited by examiner]