IP Library Granted Patent US 12,040,949
Granted Patent B2
US 12,040,949 · App. 18/070,040 · Granted Jul 16, 2024

Connecting processors using twisted torus configurations

Inventor: Brian Patrick Towles (Chapel Hill, NC)
Assignee: Google LLC
H04L41/12H04L49/10H04L49/15H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,040,949
App. No.
18/070,040
Granted
Jul 16, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer-storage media, for connecting processors using twisted torus configurations. In some implementations, a cluster of processing nodes is coupled using a reconfigurable interconnect fabric. The system determines a number of processing nodes to allocate as a network within the cluster and a topology for the network. The system selects an interconnection scheme for the network, where the interconnection scheme is selected from a group that includes at least a torus interconnection scheme and a twisted torus interconnection scheme. The system allocates the determined number of processing nodes of the cluster in the determined topology, sets the reconfigurable interconnect fabric to provide the selected interconnection scheme for the processing nodes in the network, and provides access to the network for performing a computing task.

Claims (56)

1. A method performed by one or more computers, the method comprising:

providing an interface for requesting access to a cluster of processing nodes over a communication network, wherein the processing nodes are coupled using a reconfigurable interconnect fabric;

receiving, over the communication network and through the interface, a request from each of multiple remote devices;

allocating different subsets of the processing nodes in the cluster for separate networks based on the requests, wherein each of the networks includes a different subset of the processing nodes in the cluster, wherein allocating the different subsets of the processing nodes in the cluster for separate networks comprises, for each of the requests:

determining, based on the request, (i) a number of processing nodes within the cluster to allocate in a network that is formed in response to the request and (ii) a topology to be created in the network that is formed in response to the request;

selecting an interconnection scheme for the network, wherein the interconnection scheme is selected from a group that includes at least a torus interconnection scheme and a twisted torus interconnection scheme;

allocating a subset of the processing nodes comprising the determined number of processing nodes, wherein the network provides the allocated subset of processing nodes in the determined topology; and

altering connections in the reconfigurable interconnect fabric to establish data links among the processing nodes that provide the selected interconnection scheme for the processing nodes in the network; and

providing, to each of the remote devices, access to one of the allocated networks that was allocated in response to the request from the remote device, wherein the allocated networks are configured to separately and concurrently perform computing tasks requested by the corresponding remote devices.

2. The method of claim 1 , wherein, for a first request of the requests, selecting the interconnection scheme comprises selecting between the torus interconnection scheme and the twisted torus interconnection scheme based on the number of processing nodes determined for the first request.

3. The method of claim 2 , wherein, for the first request, selecting the interconnection scheme comprises selecting the torus interconnection scheme based on determining that, for a network of a number of dimensions that the reconfigurable interconnect fabric supports, the number of processing nodes allows the network to have equal size in each of the dimensions.

4. The method of claim 2 , wherein, for the first request, the determined topology has a first size in a first dimension and a second size in a second dimension; and

wherein, for the first request, selecting the interconnection scheme comprises selecting the twisted torus interconnection scheme based on determining that the first size is a multiple of the second size.

5. The method of claim 1 , wherein, for a first request of the requests, the determined topology for the network comprises an arrangement of nodes that extends along multiple dimensions and includes multiple nodes along each of the multiple dimensions, wherein the determined topology has different amounts of nodes along at least two of the multiple dimensions; and

wherein, for the first request, selecting the interconnection scheme comprises selecting the twisted torus interconnection scheme such that the network is symmetric.

6. The method of claim 5 , wherein the twisted torus interconnection scheme includes wraparound connections made using switching elements of the reconfigurable interconnect fabric; and

wherein wraparound connections connect nodes or edges of the network that face opposite directions along a same dimension, wherein the wraparound connections for a first dimension in which the network is longest do not include any offsets in other dimensions, and wherein the wraparound connections for a second dimension in which the network is shorter than the first dimension has an offset in the first dimension.

7. The method of claim 6 , wherein the wraparound connections for the second dimension are each determined by connecting a starting node with an ending node that has:

(i) a position in the second dimension that is the same as the starting node, and

(ii) a position in the first dimension that is equal to a result of a modulo operation involving (a) a sum of a position of the starting node in the first dimension and a predetermined twist increment determined based on the determined topology and (b) a length of the longest dimension.

8. The method of claim 1 , wherein, for a first request of the requests, selecting the interconnection scheme comprises selecting the twisted torus interconnection scheme;

wherein, for the network allocated based on the first request, (i) an amount of twist in the twisted torus interconnection scheme is based on lengths of the topology based on the first request and (ii) dimensions in which to apply an offset for the twist is based on lengths of the topology determined based on the first request.

9. The method of claim 1 , wherein, for at least one of the requests, the determined topology is a two-dimensional topology.

10. The method of claim 1 , wherein, for at least one of the requests, the determined topology is a three-dimensional topology.

11. The method of claim 1 , wherein the cluster of processing nodes comprises multiple segments having a predetermined size and arrangement of multiple processing nodes, the segments having mesh connections between the nodes in each segment; and

wherein the reconfigurable interconnect fabric comprises switching elements to permit dynamic, programmable reconfiguration of connections for external-facing data ports of processing nodes in each segment.

12. The method of claim 11 , wherein the segments are each a 4×4×4 group of processing nodes.

13. The method of claim 1 , wherein the processing nodes are each separate application specific integrated circuits (ASICs).

14. The method of claim 1 , wherein the reconfigurable interconnect fabric comprises switches for data-carrying optical signals.

15. The method of claim 1 , wherein the computing tasks comprise training a machine learning model.

16. The method of claim 1 , comprising:

storing multiple configuration profiles specifying different configurations of the reconfigurable interconnect fabric to connect subsets of the processing nodes in the cluster; and

selecting a configuration profile from among the multiple configuration profiles;

wherein switching elements of the reconfigurable interconnect fabric are set according to the selected configuration profile.

17. The method of claim 16 , comprising initializing routing tables for processing nodes in the network based on stored routing information corresponding to the selected configuration profile.

18. The method of claim 1 , wherein at least some of the networks allocated based on the requests have different interconnection schemes.

19. A system comprising:

one or more computers; and

one or more computer-readable media storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

providing an interface for requesting access to a cluster of processing nodes over a communication network, wherein the processing nodes are coupled using a reconfigurable interconnect fabric;

receiving, over the communication network and through the interface, a request from each of multiple remote devices;

allocating different subsets of the processing nodes in the cluster for separate networks based on the requests, wherein each of the networks includes a different subset of the processing nodes in the cluster, wherein allocating the different subsets of the processing nodes in the cluster for separate networks comprises, for each of the requests:

determining, based on the request, (i) a number of processing nodes within the cluster to allocate in a network that is formed in response to the request and (ii) a topology to be created in the network that is formed in response to the request;

selecting an interconnection scheme for the network, wherein the interconnection scheme is selected from a group that includes at least a torus interconnection scheme and a twisted torus interconnection scheme;

allocating a subset of the processing nodes comprising the determined number of processing nodes, wherein the network provides the allocated subset of processing nodes in the determined topology; and

altering connections in the reconfigurable interconnect fabric to establish data links among the processing nodes that provide the selected interconnection scheme for the processing nodes in the network; and

providing, to each of the remote devices, access to one of the allocated networks that was allocated in response to the request from the remote device, wherein the allocated networks are configured to separately and concurrently perform computing tasks requested by the corresponding remote devices.

20. One or more non-transitory computer-readable media storing instructions that are operable, when executed by one or more computers, to cause the one or more computers to perform operations comprising:

providing an interface for requesting access to a cluster of processing nodes over a communication network, wherein the processing nodes are coupled using a reconfigurable interconnect fabric;

receiving, over the communication network and through the interface, a request from each of multiple remote devices;

allocating different subsets of the processing nodes in the cluster for separate networks based on the requests, wherein each of the networks includes a different subset of the processing nodes in the cluster, wherein allocating the different subsets of the processing nodes in the cluster for separate networks comprises, for each of the requests:

determining, based on the request, (i) a number of processing nodes within the cluster to allocate in a network that is formed in response to the request and (ii) a topology to be created in the network that is formed in response to the request;

selecting an interconnection scheme for the network, wherein the interconnection scheme is selected from a group that includes at least a torus interconnection scheme and a twisted torus interconnection scheme;

allocating a subset of the processing nodes comprising the determined number of processing nodes, wherein the network provides the allocated subset of processing nodes in the determined topology; and

altering connections in the reconfigurable interconnect fabric to establish data links among the processing nodes that provide the selected interconnection scheme for the processing nodes in the network; and

providing, to each of the remote devices, access to one of the allocated networks that was allocated in response to the request from the remote device, wherein the allocated networks are configured to separately and concurrently perform computing tasks requested by the corresponding remote devices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2023
From: TOWLES, BRIAN PATRICK
To: GOOGLE LLC
Reel/Frame 062552/0116 →
Continuity (3)
Continuation 17120051 · Dec 11, 2020
Provisional Application 63119329 · Nov 30, 2020
Related Publication 20230094933A1 · Mar 30, 2023