IP Library › Granted Patent US 12,632,409
Granted Patent B1
US 12,632,409 · App. 18/756,633 · Granted May 19, 2026

Skip-hop collective compute data transfer

Inventors: Yongseok Koh (San Jose, CA); Se Wang Oh (Campbell, CA); Zhaoqi Zhu (San Jose, CA); Ron Diamant (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G06F15/17375G06F13/28G06F15/17306
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,409
App. No.
18/756,633
Granted
May 19, 2026
Kind
B1
Abstract

Techniques for performing collective compute operations are described. A collective compute operation can be performed in a logical ring of processing ranks formed by a set of rank groups that each contain a number of processing ranks including a primary rank and one or more secondary ranks. At each rank group, a primary rank receives an incoming data slice via an intranode interconnect from a previous primary rank at multiple hops away on the logical ring. A data transfer is performed between the primary rank and each secondary rank of the rank group. An outgoing data slice is then transferred from the primary rank of the rank group to the next primary rank at multiple hops away on the logical ring.

Claims (38)

1 . A system of processing ranks in a logical ring formed by a set of rank groups each containing m number of processing ranks, m being an integer greater than 1, wherein each rank group includes:

a primary rank; and

m−1 number of one or more secondary ranks coupled to the primary rank via an intranode interconnect,

wherein each rank group is operable to perform operations for a collective computation, the operations including:

receiving, at a channel buffer of the primary rank of the rank group via the intranode interconnect of the rank group, an incoming data slice from a previous primary rank at m hops away on the logical ring;

reducing, by a direct memory access (DMA) engine of the primary rank, a primary rank data slice from an input buffer of the primary rank, a secondary rank data slice via the intranode interconnect from an input buffer of each secondary rank, and the incoming data slice from the channel buffer of the primary rank to generate an outgoing data slice; and

transferring, by the DMA engine of the primary rank, the outgoing data slice from the primary rank of the rank group to a next primary rank at m hops away on the logical ring.

2 . The system of claim 1 , wherein the operations further include:

at each rank group:

receiving, at an output buffer of the primary rank of the rank group via an intranode interconnect of the rank group, another incoming data slice from the previous primary rank on the logical ring;

transferring, by a DMA engine of each secondary rank, the other incoming data slice from the output buffer of the primary rank to an output buffer of a corresponding secondary rank; and

transferring, by the DMA engine of the primary rank, the other incoming data slice as another outgoing data slice from the primary rank of the rank group to a next primary rank on the logical ring.

3 . The system of claim 1 , wherein the collective computation is an all-reduce operation.

4 . The system of claim 1 , wherein the system of processing ranks is implemented using a plurality of system-on-chips that each include multiple rank groups.

5 . A method for performing a collective compute operation on tensor data in a logical ring of processing ranks formed by a set of rank groups each containing m number of processing ranks including a primary rank and one or more secondary ranks, m being an integer greater than 1, the method comprising:

at each rank group:

receiving, at the primary rank of the rank group via an intranode interconnect of the rank group, an incoming data slice from a previous primary rank at m hops away on the logical ring;

performing data transfer between the primary rank and each secondary rank of the rank group; and

transferring an outgoing data slice from the primary rank of the rank group to a next primary rank at m hops away on the logical ring.

6 . The method of claim 5 , wherein the incoming data slice is received at a channel buffer of the primary rank.

7 . The method of claim 6 , wherein performing the data transfer between the primary rank and each secondary rank of the rank group includes obtaining, by a direct memory access (DMA) engine of the primary rank, a primary rank data slice from an input buffer of the primary rank, a secondary rank data slice from an input buffer of each secondary rank, and the incoming data slice from the channel buffer of the primary rank.

8 . The method of claim 7 , wherein the outgoing data slice is generated by reducing, by the DMA engine, the incoming data slice, the primary rank data slice, and each secondary rank data slice.

9 . The method of claim 8 , wherein the outgoing data slice is transferred to a channel buffer of the next primary rank on the logical ring.

10 . The method of claim 9 , wherein the collective compute operation is a reduce-scatter operation.

11 . The method of claim 5 , wherein the incoming data slice is received at an output buffer of the primary rank.

12 . The method of claim 11 , wherein performing the data transfer between the primary rank and each secondary rank of the rank group includes obtaining, by a DMA engine of each secondary rank, the incoming data slice from the output buffer of the primary rank, and transferring the incoming data slice to an output buffer of each secondary rank.

13 . The method of claim 12 , wherein the incoming data slice from the output buffer of the primary rank is provided as the outgoing data slice.

14 . The method of claim 13 , wherein the outgoing data slice is transferred to an output buffer of the next primary rank on the logical ring.

15 . The method of claim 14 , wherein the collective compute operation is an all-gather operation.

16 . The method of claim 5 , wherein each rank group is implemented as an integrated circuit die containing m processing ranks.

17 . The method of claim 5 , wherein the rank group is part of a system-on-chip containing multiple rank groups.

18 . A non-transitory computer readable medium having stored therein instructions that, when executed by one or more processors, cause the one or more processors to perform operations for a collective computation on tensor data in a logical ring of processing ranks formed by a set of rank groups each containing m number of processing ranks including a primary rank and one or more secondary ranks, m being an integer greater than 1, the operations including:

at each rank group:

receiving, at the primary rank of the rank group via an intranode interconnect of the rank group, an incoming data slice from a previous primary rank at m hops away on the logical ring;

performing data transfer between the primary rank and each secondary rank of the rank group; and

transferring an outgoing data slice from the primary rank of the rank group to a next primary rank at m hops away on the logical ring.

19 . The non-transitory computer readable medium of claim 18 , wherein the operations include generating the outgoing data slice by reducing the incoming data slice, a primary rank data slice from the primary rank, and a secondary rank data slice obtained via the intranode interconnect from each secondary rank of the rank group.

20 . The non-transitory computer readable medium of claim 18 , wherein the data transfer between the primary rank and each secondary rank of the rank group includes transferring the incoming data slice from the primary rank to each secondary rank via the intranode interconnect of the rank group.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2024
From: KOH, YONGSEOK; OH, SE WANG; ZHU, ZHAOQI; DIAMANT, RON
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 067863/0356 →
References Cited (16)
US 5386532A · Sodos · 1995 [cited by examiner]
US 8738860B1 · Griffin · 2014 [cited by examiner]
US 20090040946A1 · Archer · 2009 [cited by examiner]
US 20090240915A1 · Faraj · 2009 [cited by examiner]
US 20090292905A1 · Faraj · 2009 [cited by examiner]
US 20090307467A1 · Faraj · 2009 [cited by examiner]
US 20130151713A1 · Faraj · 2013 [cited by examiner]
US 20160094435A1 · Goss · 2016 [cited by examiner]
US 20190312772A1 · Zhao · 2019 [cited by examiner]
US 20200380344A1 · Lie · 2020 [cited by examiner]
US 20210117130A1 · Davis · 2021 [cited by examiner]
US 20210248453A1 · Lauterbach · 2021 [cited by examiner]
US 20210286752A1 · Modukuri · 2021 [cited by examiner]
US 20220057937A1 · Askar · 2022 [cited by examiner]
US 20220374288A1 · Kibardin · 2022 [cited by examiner]
US 20220413759A1 · Shen · 2022 [cited by examiner]