System for allocation of network resources for executing large language model (LLM) tasks
Systems, computer program products, and methods are described herein for allocation of network resources for executing large language model (LLM) tasks. An example system receives an LLM task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a parallelism parameter and a communication pattern; determines a plurality of hosts based on at least the parallelism parameter and the communication pattern; determines a plurality of switches based on the plurality of hosts; operatively couples the plurality of hosts to the plurality of switches to configure a network point of delivery (POD); and triggers execution of the LLM task using the network POD.
1 . A method, the method comprising:
receiving a large language model (LLM) task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a data parallelism parameter, a pipeline parallelism parameter, and a communication pattern;
determining a plurality of hosts based on at least the parallelism parameter and the communication pattern;
segmenting the execution of the LLM task into a plurality of pipelines based on at least the data parallelism parameter;
segmenting each pipeline into a plurality of pipeline stages based on the pipeline parallelism parameter;
allocating the plurality of pipelines and the plurality of pipeline stages among the plurality of hosts, wherein the plurality of hosts is interconnected for data portion communication and pipeline communication;
determining a plurality of switches based on the plurality of hosts;
determining, based on the data parallelism parameter, a first set of optical circuit connections for each switch to facilitate the data portion communication between the plurality of hosts across the plurality of switches;
determining, based on the pipeline parallelism parameter, a second set of optical circuit connections for each switch to facilitate the pipeline communication between the plurality of hosts across the plurality of switches;
operatively coupling the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections;
dynamically configuring a network point of delivery (POD) by operatively coupling the plurality of hosts to the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections; and
triggering execution of the LLM task using the network POD,
wherein a count of the first set of optical circuit connections and a count of the second set of optical circuit connections is determined to satisfy a full-bisection bandwidth requirement.
2 . The method of claim 1 , wherein the data parallelism parameter indicates a number of pipelines for executing the LLM task, wherein each pipeline represents a data partition, wherein the pipeline parallelism parameter indicates a number of pipeline stages for each pipeline, wherein each pipeline stage represents a portion of the corresponding data partition.
3 . The method of claim 1 , wherein the data portion communication is based on the communication pattern associated with the data parallelism parameter and the pipeline communication is based on the communication pattern associated with the pipeline parallelism parameter.
4 . The method of claim 1 , wherein,
the count of the first set of optical circuit connections is greater than or equal to 2*ps*k, wherein ps is number of pipeline stages allocated to a subset of the plurality of hosts that are operatively coupled to each switch, wherein k is a fractional bandwidth requirement for each data portion communication in each direction relative to a total bandwidth of an optical circuit connection in the first set of optical circuit connections, and
the count of the second set of optical circuit connections is greater than or equal to 2*p*m, wherein p is the number of pipelines allocated to the subset of the plurality of hosts that are operatively coupled to each switch, and wherein m is a fractional bandwidth requirement for each pipeline communication in each direction relative to a total bandwidth of an optical circuit connection in the second set of optical circuit connections.
5 . The method of claim 1 , wherein the plurality of hosts is operatively coupled to a same switch.
6 . The method of claim 1 , wherein the network POD is configured based on a closed loop topology to allow the plurality of hosts to communicate with one another via the plurality of switches, wherein the closed loop topology comprises at least one of a ring topology or a torus topology.
7 . The method of claim 1 , wherein the network POD is configured based on an in-network collective, wherein the in-network collective comprises at least a scalable hierarchical aggregation and reduction protocol (SHARP) model in which the network POD is configured by allocating a plurality of circuits from each switch such that an aggregate count of the plurality of switches is equal to an aggregate count of distinct reductions associated with the switch, thereby ensuring full bandwidth utilization, and constructing a network topology that includes designated root switches for facilitating the reductions.
8 . The method of claim 1 , further comprising:
configuring a network structure with a plurality of network PODs;
determining a plurality of spine switches based on at least the plurality of network PODs;
interconnecting the plurality of network PODs via the plurality of spine switches; and
triggering the execution of the LLM task using the network structure.
9 . The method of claim 2 , wherein the communication pattern associated with the pipeline parallelism parameter comprises at least a point-to-point communication, and wherein the communication pattern associated with the data parallelism parameter comprises at least a reduction operation.
10 . A system, the system comprising:
a processing device;
a non-transitory storage device containing instructions that, when executed by the processing device, cause the processing device to:
receive a large language model (LLM) task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a data parallelism parameter, a pipeline parallelism parameter, and a communication pattern;
determine a plurality of hosts based on at least the parallelism parameter and the communication pattern;
segment the execution of the LLM task into a plurality of pipelines based on at least the data parallelism parameter;
segment each pipeline into a plurality of pipeline stages based on the pipeline parallelism parameter;
allocate the plurality of pipelines and the plurality of pipeline stages among the plurality of hosts, wherein the plurality of hosts is interconnected for data portion communication and pipeline communication;
determine a plurality of switches based on the plurality of hosts;
determine, based on the data parallelism parameter, a first set of optical circuit connections for each switch to facilitate the data portion communication between the plurality of hosts across the plurality of switches;
determine, based on the pipeline parallelism parameter, a second set of optical circuit connections for each switch to facilitate the pipeline communication between the plurality of hosts across the plurality of switches;
operatively couple the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections;
dynamically configure a network point of delivery (POD) by operatively coupling the plurality of hosts to the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections; and
trigger execution of the LLM task using the network POD,
wherein a count of the first set of optical circuit connections and a count of the second set of optical circuit connections is determined to satisfy a full-bisection bandwidth requirement.
11 . The system of claim 10 , wherein the data parallelism parameter indicates a number of pipelines for executing the LLM task, wherein each pipeline represents a data partition, wherein the pipeline parallelism parameter indicates a number of pipeline stages for each pipeline, wherein each pipeline stage represents a portion of the corresponding data partition.
12 . The system of claim 10 , wherein the plurality of hosts is interconnected for data portion communication and pipeline communication, wherein the data portion communication is based on the communication pattern associated with the data parallelism parameter and pipeline communication is based on the communication pattern associated with the pipeline parallelism parameter.
13 . The system of claim 10 , wherein,
the count of the first set of optical circuit connections is greater than or equal to 2*ps*k, wherein ps is number of pipeline stages allocated to a subset of the plurality of hosts that are operatively coupled to each switch, wherein k is a fractional bandwidth requirement for each data portion communication in each direction relative to a total bandwidth of an optical circuit connection in the first set of optical circuit connections, and
the count of the second set of optical circuit connections is greater than or equal to 2*p*m, wherein p is the number of pipelines allocated to the subset of the plurality of hosts that are operatively coupled to each switch, and wherein m is a fractional bandwidth requirement for each pipeline communication in each direction relative to a total bandwidth of an optical circuit connection in the second set of optical circuit connections.
14 . The system of claim 10 , wherein the plurality of hosts is operatively coupled to a same switch.
15 . The system of claim 10 , wherein the network POD is configured based on a closed loop topology to allow the plurality of hosts to communicate with one another via the plurality of switches, wherein the closed loop topology comprises at least one of a ring topology or a torus topology.
16 . The system of claim 10 , wherein the network POD is configured based on an in-network collective, wherein the in-network collective comprises at least a scalable hierarchical aggregation and reduction protocol (SHARP) model in which the network POD is configured by allocating a plurality of circuits from each switch such that an aggregate count of the plurality of switches is equal to an aggregate count of distinct reductions associated with the switch, thereby ensuring full bandwidth utilization, and constructing a network topology that includes designated root switches for facilitating the reductions.
17 . The system of claim 10 , wherein the instructions, when executed, cause the processing device to:
configure a network structure with a plurality of network PODs;
determine a plurality of spine switches based on at least the plurality of network PODs;
interconnect the plurality of network PODs via the plurality of spine switches; and
trigger the execution of the LLM task using the network structure.
18 . A computer program product, the computer program product comprising a non-transitory computer-readable medium comprising code configured to cause an apparatus to:
receive a large language model (LLM) task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a data parallelism parameter, a pipeline parallelism parameter, and a communication pattern;
determine a plurality of hosts based on at least the parallelism parameter and the communication pattern;
segment the execution of the LLM task into a plurality of pipelines based on at least the data parallelism parameter;
segment each pipeline into a plurality of pipeline stages based on the pipeline parallelism parameter;
allocate the plurality of pipelines and the plurality of pipeline stages among the plurality of hosts, wherein the plurality of hosts is interconnected for data portion communication and pipeline communication;
determine a plurality of switches based on the plurality of hosts;
determine, based on the data parallelism parameter, a first set of optical circuit connections for each switch to facilitate the data portion communication between the plurality of hosts across the plurality of switches;
determine, based on the pipeline parallelism parameter, a second set of optical circuit connections for each switch to facilitate the pipeline communication between the plurality of hosts across the plurality of switches;
operatively couple the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections;
dynamically configure a network point of delivery (POD) by operatively coupling the plurality of hosts to the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections; and
trigger execution of the LLM task using the network POD,
wherein a count of the first set of optical circuit connections and a count of the second set of optical circuit connections is determined to satisfy a full-bisection bandwidth requirement.
19 . The computer program product of claim 18 , wherein the data parallelism parameter indicates a number of pipelines for executing the LLM task, wherein each pipeline represents a data partition, wherein the pipeline parallelism parameter indicates a number of pipeline stages for each pipeline, wherein each pipeline stage represents a portion of the corresponding data partition.
20 . The computer program product of claim 18 , wherein the plurality of hosts is interconnected for data portion communication and pipeline communication, wherein the data portion communication is based on the communication pattern associated with the data parallelism parameter and pipeline communication is based on the communication pattern associated with the pipeline parallelism parameter.
21 . The computer program product of claim 18 , wherein,
the count of the first set of optical circuit connections is greater than or equal to 2*ps*k, wherein ps is number of pipeline stages allocated to a subset of the plurality of hosts that are operatively coupled to each switch, wherein k is a fractional bandwidth requirement for each data portion communication in each direction relative to a total bandwidth of an optical circuit connection in the first set of optical circuit connections, and
the count of the second set of optical circuit connections is greater than or equal to 2*p*m, wherein p is the number of pipelines allocated to the subset of the plurality of hosts that are operatively coupled to each switch, and wherein m is a fractional bandwidth requirement for each pipeline communication in each direction relative to a total bandwidth of an optical circuit connection in the second set of optical circuit connections.
22 . The computer program product of claim 18 , wherein the code further causes the apparatus to:
configure a network structure with a plurality of network PODs;
determine a plurality of spine switches based on at least the plurality of network PODs;
interconnect the plurality of network PODs via the plurality of spine switches; and
trigger the execution of the LLM task using the network structure.