IP Library Granted Patent US 12,332,823
Granted Patent B2
US 12,332,823 · App. 17/824,824 · Granted Jun 17, 2025

Parallel dataflow routing scheme systems and methods

Inventors: Liang Han (Campbell, CA); ChengYuan Wu (Fremont, CA); Guoyu Zhu (San Jose, CA); Yang Jiao (San Jose, CA); Rong Zhong (Fremont, CA); Yunxiao Zou (Shanghai, CN)
Assignee: T-Head (Shanghai) Semiconductor Co., Ltd.
G06F13/4068G06F15/17312
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,332,823
App. No.
17/824,824
Granted
Jun 17, 2025
Kind
B2
Abstract

The presented systems enable efficient and effective network communications. In one embodiment, a system comprises a parallel processing unit (PPU) included in a chip and a plurality of interconnects in an inter-chip network (ICN) configured to communicatively couple a plurality of PPUs that communicate with one another via the ICN. Corresponding communications are configured in accordance with routing tables. The routing tables can be stored and reside in registers of an ICN subsystem included in the PPU and include indications of minimum links available to forward a communication from the PPU and another PPU. Respective ones of the routing tables include indications of a correlation between the minimum links and respective ones of a plurality of egress ports that are available for communication coupling to the respective ones of PPUs that are possible destination PPUs. The routing tables can include indications of a single path between two PPUs per respective information communication flow.

Claims (49)

1. A system comprising:

a plurality of parallel processing units (PPUs) included in a first chip, each of the PPUs includes:

a plurality of processing cores;

a plurality of memories, wherein a first set of the memories couple to a first set of the plurality of processing cores; and

a plurality of interconnects in an inter-chip network (ICN) configured to communicatively couple the plurality of PPUs,

wherein each of the PPUs is configured to communicate over the ICN in accordance with respective routing tables that are stored and reside in registers included in the respective PPUs, and wherein the respective routing tables include indications of minimum links available to forward a communication from the PPU to another PPU.

2. The system of claim 1 , wherein the PPUs include respective plurality of egress ports configured to communicatively couple to the ICN, wherein the respective routing tables include indications of a correlation between the minimum links and the respective plurality of egress ports.

3. The system of claim 1 , wherein the respective routing tables include indications of a single path between two PPUs per respective information communication flow, wherein the single path is based upon the correlation between the minimum links and respective ones of a plurality of egress ports that are available for communication coupling to the respective ones of PPUs that are possible destination PPUs.

4. The system of claim 1 , wherein the respective routing tables are loaded and stored in registers of an ICN subsystem included in the PPUs, wherein the respective routing tables are loaded as part of configuration and setup operations of the communication capabilities of the ICN and the plurality of PPUs before running normal processing operations.

5. The system of claim 1 , wherein the respective routing tables are:

static and predetermined in between execution of the configuration and setup operations of the communication capabilities of the ICN and the plurality of PPUs; and

re-configurable as part of ICN configuration and setup operations of the communication capabilities of the ICN and the plurality of PPUs, wherein the setup operations include loading and storing the respective routing tables in registers of an ICN subsystem included in the PPU.

6. The system of claim 5 , wherein the respective ones of the plurality of parallel processing units include respective ones of the routing tables.

7. The system of claim 5 , a respective one of a plurality of parallel processing units is considered a source PPU when a communication originates at the respective one of a plurality of parallel processing units, a relay PPU when a communication passes through the respective one of a plurality of parallel processing units, or a destination PPU when a communication ends at the respective one of a plurality of parallel processing units.

8. The system of claim 7 , wherein respective ones of the plurality of interconnects are configured for single-flow balancing, wherein the source PPU supports parallel communication flows up to a narrowest number of links between the source PPU and the destination PPU.

9. The system of claim 1 , wherein the respective routing tables are compatible with a basic dimension ordered X-Y routing scheme, in which schematically the PPUs are organized in a two-dimensional array in which Y corresponds to one dimension in the two-dimensional array and X corresponds to the other dimension in the two-dimensional array and the routing scheme restricts routing direction turns from the Y dimension to the X dimension.

10. The system of claim 1 , wherein the PPU is included in a first set of the plurality of PPUs and the first set of the plurality of PPUs is included in a first compute node, and a second set of the plurality of PPUs is included in a second node of the plurality of PPUs.

11. The system of claim 10 , wherein respective ones of the plurality of interconnects are configured for many-flow balancing, wherein a relay PPU runs routing to balance flows among egresses.

12. The system of claim 1 , wherein communications are processed in accordance with a routing scheme that provides balanced workloads, guaranteed dependency, and guaranteed access orders.

13. A communication method comprising:

performing a setup operation including creation of static pre-determined routing tables;

forwarding a communication packet from a source parallel processing unit (PPU), wherein the communication packet is formed and forwarded in accordance with the static pre-determined routing tables;

receiving the communication packet at a destination parallel processing unit (PPU); and

wherein the source PPU and destination PPU are included in respective ones of a plurality of parallel processing units (PPUs) included in a network,

wherein a first set of the plurality processing cores are included in a first chip and a second set of the plurality processing cores are included in a second chip, and wherein the plurality of processing units communicate over a plurality of interconnects and corresponding communications are configured in accordance with the static pre-determined routing tables

wherein a routing scheme at the source PPU is determined by:

creating a flow ID associated with a unique communication path through the interconnects;

utilizing a corresponding one of the routing tables to ascertain a minimum number of links path to a destination; and

establishing a routing selection based upon the flow ID and the minimum number of links path.

14. The communication method of claim 13 , wherein the flow ID is established by hashing a selected number of bits in a physical address.

15. The communication method of claim 13 , wherein a routing scheme at the relay PPU includes selecting an egress port, wherein selection of the egress port includes:

creating a flow ID associated with a unique communication path through the interconnects, wherein the flow ID is established by hashing a selected number of bits in a physical address;

mapping a source PPU ID and the flow ID;

determining the number of possible egress ports available based upon the mapping;

utilizing a corresponding one of the routing tables to ascertain a minimum links path to a destination; and

establishing a routing selection based upon the flow ID, the number of possible egress ports, and the minimum links path.

16. The communication method of claim 13 , further comprising balancing the forwarding of the communication packet, including distributing the communication packet via physical address based interleaving.

17. A system, comprising:

a first set of parallel processing units (PPUs) included in a first compute node, wherein respective PPUs included in the first set of PPUs are included in separate respective chips;

a second set of parallel processing units (PPUs) included in a second compute node, wherein respective PPUs included in the second set of PPUs are included in separate respective chips; and

a plurality of interconnects in an inter-chip network (ICN) configured to communicatively couple the first set of PPUs and the second set of PPUs,

wherein PPUs included in the first set of PPUs and the second set of PPUs communicate over the plurality of interconnects and corresponding communications are configured in accordance with routing tables that reside in storage features of respective ones of the PPUs included in the first set of PPUs and the second set of PPUs, and wherein the routing tables include indications of minimum links available to forward communications between a source and destination included in the respective ones of the PPUs.

18. The system of claim 17 wherein respective ones of the PPUs included in the first set of PPUs and the second set of PPUs comprises:

a plurality of processing cores, wherein respective sets of processing cores are included in respective ones of the PPUs; and

a plurality of memories, wherein respective sets of memories are communicatively coupled to the respective sets of sets of processing cores and included in the respective ones of the PPUs.

19. The system of claim 17 , wherein the plurality of processing units communicates over the plurality of interconnects and corresponding communications are configured in accordance with routing tables, wherein the routing tables are static and predetermined, wherein the routing tables are loaded in registers associated with the plurality of processing units as part of a setup of the plurality of processors before running normal processing operations.

20. The system of claim 17 wherein balanced workloads are provided by multi-flow through minimal links and interleaving with flow ID and source parallel processing unit (PPU) ID.

21. The system of claim 17 wherein guaranteed dependency is provided by utilization of physical address (PA) interleaving with flow ID and hashing at the source parallel processing unit (PPU).

22. The system of claim 17 wherein guaranteed access orders are provided by utilization of flow ID along routs and splittable remote-fence.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2024
From: ALIBABA (CHINA) CO., LTD.
To: T-HEAD (SHANGHAI) SEMICONDUCTOR CO., LTD.
Reel/Frame 066348/0611 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: HAN, LIANG; ZHU, GUOYU; JIAO, YANG; ZHONG, RONG; ZOU, YUNXIAO
To: ALIBABA (CHINA) CO., LTD.
Reel/Frame 065151/0241 →
Continuity (1)
Related Publication 20230244626A1 · Aug 3, 2023
References Cited (43)
US 4204251A · Brudevold · 1980 [cited by applicant]
US 4816993A · Takahashi · 1989 [cited by applicant]
US 5504918A · Collette · 1996 [cited by applicant]
US 5734872A · Kelly · 1998 [cited by applicant]
US 7149876B2 · Kirsch · 2006 [cited by applicant]
US 7483428B2 · Goodman · 2009 [cited by applicant]
US 7840914B1 · Agarwal et al. · 2010 [cited by applicant]
US 8131975B1 · Cismas · 2012 [cited by applicant]
US 8145880B1 · Cismas · 2012 [cited by applicant]
US 8291400B1 · Lee · 2012 [cited by applicant]
US 8327114B1 · Cismas · 2012 [cited by applicant]
US 8375395B2 · Prasanna · 2013 [cited by applicant]
US 8510535B2 · Deng · 2013 [cited by applicant]
US 8583896B2 · Cadambi · 2013 [cited by applicant]
US 8990460B2 · Chang · 2015 [cited by applicant]
US 10580190B2 · S. · 2020 [cited by applicant]
US 10635631B2 · Hutton et al. · 2020 [cited by applicant]
US 10915328B2 · Pearce et al. · 2021 [cited by applicant]
US 11093277B2 · Sankaran et al. · 2021 [cited by applicant]
US 11211334B2 · Lin et al. · 2021 [cited by applicant]
US 11354135B2 · Qin et al. · 2022 [cited by applicant]
US 11487541B2 · Corbal et al. · 2022 [cited by applicant]
US 11693691B2 · Sankaran et al. · 2023 [cited by applicant]
US 11824754B2 · Magnezi · 2023 [cited by examiner]
US 11960437B2 · Han et al. · 2024 [cited by applicant]
US 20020199017A1 · Russell · 2002 [cited by applicant]
US 20080285562A1 · Scott et al. · 2008 [cited by applicant]
US 20100325257A1 · Goel · 2010 [cited by examiner]
US 20120331268A1 · Konig · 2012 [cited by applicant]
US 20150006776A1 · Liu · 2015 [cited by applicant]
US 20170286169A1 · Ravindran et al. · 2017 [cited by applicant]
US 20190122415A1 · S. · 2019 [cited by applicant]
US 20190138890A1 · Liang et al. · 2019 [cited by applicant]
US 20190258796A1 · Paczkowski et al. · 2019 [cited by applicant]
US 20190294575A1 · Dennison et al. · 2019 [cited by applicant]
US 20200192676A1 · Pearce et al. · 2020 [cited by applicant]
US 20200249957A1 · Qin et al. · 2020 [cited by applicant]
US 20200401440A1 · Sankaran et al. · 2020 [cited by applicant]
US 20210081198A1 · Corbal et al. · 2021 [cited by applicant]
US 20220164218A1 · Sankaran et al. · 2022 [cited by applicant]
US 20230083705A1 · Corbal et al. · 2023 [cited by applicant]
US 20230089800A1 · Chadalavada · 2023 [cited by examiner]
US 20230090604A1 · Han · 2023 [cited by examiner]