IP Library Granted Patent US 12,124,889
Granted Patent B2
US 12,124,889 · App. 18/059,368 · Granted Oct 22, 2024

Efficient and more advanced implementation of ring-allreduce algorithm for distributed parallel deep learning

Inventors: Liang Han (San Mateo, CA); Yang Jiao (San Mateo, CA)
Assignee: T-Head (Shanghai) Semiconductor Co., Ltd
G06F9/52G06F9/4881G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,124,889
App. No.
18/059,368
Granted
Oct 22, 2024
Kind
B2
Abstract

The present disclosure provides a method for syncing data of a computing task across a plurality of groups of computing nodes, each group comprising a set of computing nodes A-D, a set of intra-group interconnects that communicatively couple computing node A with computing nodes B and C and computing node D with computing nodes B and C, and a set of inter-group interconnects that communicatively couple a computing node A of a first group of the plurality of groups with a computing node A of a second group neighboring the first group, a computing node B of the first group with a computing node B of the second group, a computing node C of the first group with the computing node C of the second group, and a computing node D of the first group with a computing node D of the second group, the method comprising: syncing across a first dimension of computing nodes using a first set of ring connections, wherein the first set of ring connections are formed using inter-group and intra-group interconnects that communicatively couple the computing nodes along the first dimension; and broadcasting synced data across a second dimension of computing nodes using a second ring connection.

Claims (42)

1. A method for syncing data of a computing task across a plurality of groups of computing nodes, each group comprising a set of computing nodes A-D, a set of intra-group interconnects that communicatively couple computing node A with computing nodes B and C and computing node D with computing nodes B and C, and a set of inter-group interconnects comprising: a first inter-group interconnect configured to communicatively and directly couple a computing node A of a first group of the plurality of groups with a computing node A of a second group neighboring the first group, a second inter-group interconnect configured to communicatively and directly couple a computing node B of the first group with a computing node B of the second group, a third inter-group interconnect configured to communicatively and directly couple a computing node C of the first group with a computing node C of the second group, and a forth inter-group interconnect configured to communicatively and directly couple a computing node D of the first group with a computing node D of the second group, wherein the second group is aligned with the first group in a first dimension, the method comprising:

syncing data across the first dimension of computing nodes of the first group and the second group using a first set of ring connections, wherein the first set of ring connections are formed using inter-group and intra-group interconnects that communicatively couple the computing nodes of the first group and the second group along the first dimension, and the syncing data across the first dimension comprises transferring, in a unit time, sub-data from a computing node along the first dimension to another computing node via a connection on the first set of ring connections; and

broadcasting synced data across the first dimension of computing nodes using the first set of ring connections.

2. The method of claim 1 , wherein the data to be synced comprises a plurality of sub-data, and each computing node comprises a different version of each sub-data.

3. The method of claim 2 , wherein:

syncing data across the first dimension of computing nodes using a first set of ring connections further comprises:

in a clock cycle, receiving a version of sub-data by each computing node in a row, wherein the version of the sub-data is transferred from another computing node in the row via the connection on the first set of ring connections; and

continuing data transferring until each computing node along the first dimension receives all versions of a sub-data from all computing nodes in the row.

4. The method of claim 1 , wherein:

the plurality of groups of computing nodes comprise a third group of computing nodes aligned with the first group in a second dimension that is different from the first dimension.

5. The method of claim 3 , wherein broadcasting synced data across the first dimension of computing nodes using the first set of ring connections further comprises:

in a clock cycle, receiving sub-data by each computing node across the first dimension, wherein the sub-data transferred from another computing node across the first dimension via the connection on the first set of ring connections; and

continuing data transferring until all computing nodes along the first dimension receive sub-data from all computing nodes.

6. The method of claim 1 , wherein the set of intra-group interconnects and the set of inter-group interconnects comprise inter-chip interconnects.

7. The method of claim 1 , wherein the computing nodes are artificial intelligence (“AI”) training processors, AI training chips, neural processing units (“NPU”), or graphic processing units (“GPU”).

8. The method of claim 7 , wherein the computing task is an AI computing task involving an allreduce algorithm.

9. A system for syncing data of a computing task across a plurality of groups of computing nodes, each group comprising a set of computing nodes A-D, a set of intra-group interconnects that communicatively couple computing node A with computing nodes B and C and computing node D with computing nodes B and C, and a set of inter-group interconnects comprising: a first inter-group interconnect configured to communicatively and directly couple a computing node A of a first group of the plurality of groups with a computing node A of a second group neighboring the first group, a second inter-group interconnect configured to communicatively and directly couple a computing node B of the first group with a computing node B of the second group, a third inter-group interconnect configured to communicatively and directly couple a computing node C of the first group with a computing node C of the second group, and a forth inter-group interconnect configured to communicatively and directly couple a computing node D of the first group with a computing node D of the second group, wherein the second group is aligned with the first group in a first dimension, the system comprising:

a memory storing a set of instructions; and

one or more processors configured to execute the set of instructions to cause the system to:

sync data across the first dimension of computing nodes of the first group and the second group using a first set of ring connections, wherein the first set of ring connections are formed using inter-group and intra-group interconnects that communicatively couple the computing nodes of the first group and the second group along the first dimension, and the syncing data across the first dimension comprises transferring, in a unit time, sub-data from a computing node along the first dimension to another computing node via a connection on the first set of ring connections; and

broadcast synced data across the first dimension of computing nodes using the first set of ring connections.

10. The system of claim 9 , wherein the data to be synced comprises a plurality of sub-data, and each computing node comprises a different version of each sub-data.

11. The system of claim 10 , wherein the one or more processors are further configured to execute the set of instructions to cause the system to:

in a clock cycle, receive a version of sub-data by each computing node in a row, wherein the version of the sub-data is transferred from another computing node in the row via the connection on the first set of ring connections; and

continue data transferring until each computing node along the first dimension receives all versions of a sub-data from all computing nodes in the row.

12. The system of claim 11 , wherein the plurality of groups of computing nodes comprise a third group of computing nodes aligned with the first group in a second dimension that is different from the first dimension.

13. The system of claim 10 , wherein the one or more processors are further configured to execute the set of instructions to cause the system to:

in a clock cycle, receive sub-data by each computing node across the first dimension, wherein the sub-data transferred from another computing node across the first dimension via the connection on the first set of ring connections; and

continue data transferring until all computing nodes along the first dimension receive sub-data from all computing nodes.

14. The system of claim 9 , wherein the set of intra-group interconnects and the set of inter-group interconnects comprise inter-chip interconnects.

15. A non-transitory computer readable medium that stores a set of instructions that is executable by one or more processors of an apparatus to cause the apparatus to initiate a method for syncing data of a computing task across a plurality of groups of computing nodes, each group comprising a set of computing nodes A-D, a set of intra-group interconnects that communicatively couple computing node A with computing nodes B and C and computing node D with computing nodes B and C, and a set of inter-group interconnects comprising: a first inter-group interconnect configured to communicatively and directly couple a computing node A of a first group of the plurality of groups with a computing node A of a second group neighboring the first group, a second inter-group interconnect configured to communicatively and directly couple a computing node B of the first group with a computing node B of the second group, a third inter-group interconnect configured to communicatively and directly couple a computing node C of the first group with a computing node C of the second group, and a forth inter-group interconnect configured to communicatively and directly couple a computing node D of the first group with a computing node D of the second group, wherein the second group is aligned with the first group in a first dimension, the method comprising:

syncing data across the first dimension of computing nodes of the first group and the second group using a first set of ring connections, wherein the first set of ring connections are formed using inter-group and intra-group interconnects that communicatively couple the computing nodes of the first group and the second group along the first dimension, and the syncing data across the first dimension comprises transferring, in a unit time, sub-data from a computing node along the first dimension to another computing node via a connection on a first set of ring connections; and

broadcasting synced data across the first dimension of computing nodes using the first set of ring connections.

16. The non-transitory computer readable medium of claim 15 , wherein the data to be synced comprises a plurality of sub-data, and each computing node comprises a different version of each sub-data.

17. The non-transitory computer readable medium of claim 16 , wherein the method further comprises:

in a clock cycle, receiving a version of sub-data by each computing node in a row, wherein the version of the sub-data is transferred from another computing node in the row via the connection on the first set of ring connections; and

continuing data transferring until each computing node along the first dimension receives all versions of a sub-data from all computing nodes in the row.

18. The non-transitory computer readable medium of claim 17 , wherein the plurality of groups of computing nodes comprise a third group of computing nodes aligned with the first group in a second dimension that is different from the first dimension.

19. The non-transitory computer readable medium of claim 17 , wherein the method further comprises:

in a clock cycle, receiving sub-data by each computing node across the first dimension, wherein the sub-data transferred from another computing node across the first dimension via the connection on the first set of ring connections; and

continuing data transferring until all computing nodes along the first dimension receives sub-data from all computing nodes.

20. The non-transitory computer readable medium of claim 15 , wherein the set of intra-group interconnects and the set of inter-group interconnects comprise inter-chip interconnects.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2024
From: ALIBABA GROUP HOLDING LIMITED
To: T-HEAD (SHANGHAI) SEMICONDUCTOR CO., LTD.
Reel/Frame 066348/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2022
From: HAN, LIANG; JIAO, YANG
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 061896/0425 →
Continuity (2)
Continuation 16777711 · Jan 30, 2020
Related Publication 20230088237A1 · Mar 23, 2023