IP Library › Granted Patent US 12,726,448
Granted Patent B2
US 12,726,448 · App. 18/416,028 · Granted Sep 1, 2026

System for allocation of network resources for executing large language model (LLM) tasks

Inventors: Ioannis (Giannis) Patronas (Piraeus, GR); Nikolaos Terzenidis (Kilkis, GR); Eitan Zahavi (Zichron Yaakov, IL); Paraskevas Bakopoulos (Ilion, GR); Zsolt-Alon Wertheimer (Kiryat Ata, IL); Dimitrios Syrivelis (Volos, GR); Prethvi Ramesh Kashinkunti (Chicago, IL); Louis Bennie Capps, Jr. (Georgetown, TX); Julie Irene Marcelle Bernauer (Palo Alto, CA); Elad Mentovich (Tel Aviv, IL)
Assignee: Mellanox Technologies, Ltd.
H04L47/808H04L47/781H04L47/827
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,726,448
App. No.
18/416,028
Granted
Sep 1, 2026
Kind
B2
Abstract

Systems, computer program products, and methods are described herein for allocation of network resources for executing large language model (LLM) tasks. An example system receives an LLM task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a parallelism parameter and a communication pattern; determines a plurality of hosts based on at least the parallelism parameter and the communication pattern; determines a plurality of switches based on the plurality of hosts; operatively couples the plurality of hosts to the plurality of switches to configure a network point of delivery (POD); and triggers execution of the LLM task using the network POD.

Claims (78)

1 . A method, the method comprising:

receiving a large language model (LLM) task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a data parallelism parameter, a pipeline parallelism parameter, and a communication pattern;

determining a plurality of hosts based on at least the parallelism parameter and the communication pattern;

segmenting the execution of the LLM task into a plurality of pipelines based on at least the data parallelism parameter;

segmenting each pipeline into a plurality of pipeline stages based on the pipeline parallelism parameter;

allocating the plurality of pipelines and the plurality of pipeline stages among the plurality of hosts, wherein the plurality of hosts is interconnected for data portion communication and pipeline communication;

determining a plurality of switches based on the plurality of hosts;

determining, based on the data parallelism parameter, a first set of optical circuit connections for each switch to facilitate the data portion communication between the plurality of hosts across the plurality of switches;

determining, based on the pipeline parallelism parameter, a second set of optical circuit connections for each switch to facilitate the pipeline communication between the plurality of hosts across the plurality of switches;

operatively coupling the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections;

dynamically configuring a network point of delivery (POD) by operatively coupling the plurality of hosts to the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections; and

triggering execution of the LLM task using the network POD,

wherein a count of the first set of optical circuit connections and a count of the second set of optical circuit connections is determined to satisfy a full-bisection bandwidth requirement.

2 . The method of claim 1 , wherein the data parallelism parameter indicates a number of pipelines for executing the LLM task, wherein each pipeline represents a data partition, wherein the pipeline parallelism parameter indicates a number of pipeline stages for each pipeline, wherein each pipeline stage represents a portion of the corresponding data partition.

3 . The method of claim 1 , wherein the data portion communication is based on the communication pattern associated with the data parallelism parameter and the pipeline communication is based on the communication pattern associated with the pipeline parallelism parameter.

4 . The method of claim 1 , wherein,

the count of the first set of optical circuit connections is greater than or equal to 2*ps*k, wherein ps is number of pipeline stages allocated to a subset of the plurality of hosts that are operatively coupled to each switch, wherein k is a fractional bandwidth requirement for each data portion communication in each direction relative to a total bandwidth of an optical circuit connection in the first set of optical circuit connections, and

the count of the second set of optical circuit connections is greater than or equal to 2*p*m, wherein p is the number of pipelines allocated to the subset of the plurality of hosts that are operatively coupled to each switch, and wherein m is a fractional bandwidth requirement for each pipeline communication in each direction relative to a total bandwidth of an optical circuit connection in the second set of optical circuit connections.

5 . The method of claim 1 , wherein the plurality of hosts is operatively coupled to a same switch.

6 . The method of claim 1 , wherein the network POD is configured based on a closed loop topology to allow the plurality of hosts to communicate with one another via the plurality of switches, wherein the closed loop topology comprises at least one of a ring topology or a torus topology.

7 . The method of claim 1 , wherein the network POD is configured based on an in-network collective, wherein the in-network collective comprises at least a scalable hierarchical aggregation and reduction protocol (SHARP) model in which the network POD is configured by allocating a plurality of circuits from each switch such that an aggregate count of the plurality of switches is equal to an aggregate count of distinct reductions associated with the switch, thereby ensuring full bandwidth utilization, and constructing a network topology that includes designated root switches for facilitating the reductions.

8 . The method of claim 1 , further comprising:

configuring a network structure with a plurality of network PODs;

determining a plurality of spine switches based on at least the plurality of network PODs;

interconnecting the plurality of network PODs via the plurality of spine switches; and

triggering the execution of the LLM task using the network structure.

9 . The method of claim 2 , wherein the communication pattern associated with the pipeline parallelism parameter comprises at least a point-to-point communication, and wherein the communication pattern associated with the data parallelism parameter comprises at least a reduction operation.

10 . A system, the system comprising:

a processing device;

a non-transitory storage device containing instructions that, when executed by the processing device, cause the processing device to:

receive a large language model (LLM) task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a data parallelism parameter, a pipeline parallelism parameter, and a communication pattern;

determine a plurality of hosts based on at least the parallelism parameter and the communication pattern;

segment the execution of the LLM task into a plurality of pipelines based on at least the data parallelism parameter;

segment each pipeline into a plurality of pipeline stages based on the pipeline parallelism parameter;

allocate the plurality of pipelines and the plurality of pipeline stages among the plurality of hosts, wherein the plurality of hosts is interconnected for data portion communication and pipeline communication;

determine a plurality of switches based on the plurality of hosts;

determine, based on the data parallelism parameter, a first set of optical circuit connections for each switch to facilitate the data portion communication between the plurality of hosts across the plurality of switches;

determine, based on the pipeline parallelism parameter, a second set of optical circuit connections for each switch to facilitate the pipeline communication between the plurality of hosts across the plurality of switches;

operatively couple the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections;

dynamically configure a network point of delivery (POD) by operatively coupling the plurality of hosts to the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections; and

trigger execution of the LLM task using the network POD,

wherein a count of the first set of optical circuit connections and a count of the second set of optical circuit connections is determined to satisfy a full-bisection bandwidth requirement.

11 . The system of claim 10 , wherein the data parallelism parameter indicates a number of pipelines for executing the LLM task, wherein each pipeline represents a data partition, wherein the pipeline parallelism parameter indicates a number of pipeline stages for each pipeline, wherein each pipeline stage represents a portion of the corresponding data partition.

12 . The system of claim 10 , wherein the plurality of hosts is interconnected for data portion communication and pipeline communication, wherein the data portion communication is based on the communication pattern associated with the data parallelism parameter and pipeline communication is based on the communication pattern associated with the pipeline parallelism parameter.

13 . The system of claim 10 , wherein,

the count of the first set of optical circuit connections is greater than or equal to 2*ps*k, wherein ps is number of pipeline stages allocated to a subset of the plurality of hosts that are operatively coupled to each switch, wherein k is a fractional bandwidth requirement for each data portion communication in each direction relative to a total bandwidth of an optical circuit connection in the first set of optical circuit connections, and

the count of the second set of optical circuit connections is greater than or equal to 2*p*m, wherein p is the number of pipelines allocated to the subset of the plurality of hosts that are operatively coupled to each switch, and wherein m is a fractional bandwidth requirement for each pipeline communication in each direction relative to a total bandwidth of an optical circuit connection in the second set of optical circuit connections.

14 . The system of claim 10 , wherein the plurality of hosts is operatively coupled to a same switch.

15 . The system of claim 10 , wherein the network POD is configured based on a closed loop topology to allow the plurality of hosts to communicate with one another via the plurality of switches, wherein the closed loop topology comprises at least one of a ring topology or a torus topology.

16 . The system of claim 10 , wherein the network POD is configured based on an in-network collective, wherein the in-network collective comprises at least a scalable hierarchical aggregation and reduction protocol (SHARP) model in which the network POD is configured by allocating a plurality of circuits from each switch such that an aggregate count of the plurality of switches is equal to an aggregate count of distinct reductions associated with the switch, thereby ensuring full bandwidth utilization, and constructing a network topology that includes designated root switches for facilitating the reductions.

17 . The system of claim 10 , wherein the instructions, when executed, cause the processing device to:

configure a network structure with a plurality of network PODs;

determine a plurality of spine switches based on at least the plurality of network PODs;

interconnect the plurality of network PODs via the plurality of spine switches; and

trigger the execution of the LLM task using the network structure.

18 . A computer program product, the computer program product comprising a non-transitory computer-readable medium comprising code configured to cause an apparatus to:

receive a large language model (LLM) task and an input specifying information associated with execution of the LLM task, wherein the input comprises at least a data parallelism parameter, a pipeline parallelism parameter, and a communication pattern;

determine a plurality of hosts based on at least the parallelism parameter and the communication pattern;

segment the execution of the LLM task into a plurality of pipelines based on at least the data parallelism parameter;

segment each pipeline into a plurality of pipeline stages based on the pipeline parallelism parameter;

allocate the plurality of pipelines and the plurality of pipeline stages among the plurality of hosts, wherein the plurality of hosts is interconnected for data portion communication and pipeline communication;

determine a plurality of switches based on the plurality of hosts;

determine, based on the data parallelism parameter, a first set of optical circuit connections for each switch to facilitate the data portion communication between the plurality of hosts across the plurality of switches;

determine, based on the pipeline parallelism parameter, a second set of optical circuit connections for each switch to facilitate the pipeline communication between the plurality of hosts across the plurality of switches;

operatively couple the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections;

dynamically configure a network point of delivery (POD) by operatively coupling the plurality of hosts to the plurality of switches using the first set of optical circuit connections and the second set of optical circuit connections; and

trigger execution of the LLM task using the network POD,

wherein a count of the first set of optical circuit connections and a count of the second set of optical circuit connections is determined to satisfy a full-bisection bandwidth requirement.

19 . The computer program product of claim 18 , wherein the data parallelism parameter indicates a number of pipelines for executing the LLM task, wherein each pipeline represents a data partition, wherein the pipeline parallelism parameter indicates a number of pipeline stages for each pipeline, wherein each pipeline stage represents a portion of the corresponding data partition.

20 . The computer program product of claim 18 , wherein the plurality of hosts is interconnected for data portion communication and pipeline communication, wherein the data portion communication is based on the communication pattern associated with the data parallelism parameter and pipeline communication is based on the communication pattern associated with the pipeline parallelism parameter.

21 . The computer program product of claim 18 , wherein,

the count of the first set of optical circuit connections is greater than or equal to 2*ps*k, wherein ps is number of pipeline stages allocated to a subset of the plurality of hosts that are operatively coupled to each switch, wherein k is a fractional bandwidth requirement for each data portion communication in each direction relative to a total bandwidth of an optical circuit connection in the first set of optical circuit connections, and

the count of the second set of optical circuit connections is greater than or equal to 2*p*m, wherein p is the number of pipelines allocated to the subset of the plurality of hosts that are operatively coupled to each switch, and wherein m is a fractional bandwidth requirement for each pipeline communication in each direction relative to a total bandwidth of an optical circuit connection in the second set of optical circuit connections.

22 . The computer program product of claim 18 , wherein the code further causes the apparatus to:

configure a network structure with a plurality of network PODs;

determine a plurality of spine switches based on at least the plurality of network PODs;

interconnect the plurality of network PODs via the plurality of spine switches; and

trigger the execution of the LLM task using the network structure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2024
From: PATRONAS, IOANNIS (GIANNIS); TERZENIDIS, NIKOLAOS; ZAHAVI, EITAN; BAKOPOULOS, PARASKEVAS; WERTHEIMER, ZSOLT-ALON; SYRIVELIS, DIMITRIOS; KASHINKUNTI, PRETHVI RAMESH; CAPPS, LOUIS BENNIE, JR.; BERNAUER, JULIE IRENE MARCELLE; MENTOVICH, ELAD
To: MELLANOX TECHNOLOGIES, LTD.
Reel/Frame 066166/0086 →
Priority Claims (1)
GR 20230101059 · Dec 20, 2023 · national
Continuity (1)
Related Publication 20250211548A1 · Jun 26, 2025
References Cited (96)
US 9419902B1 · Sites · 2016 [cited by applicant]
US 9755948B1 · Viljoen · 2017 [cited by applicant]
US 9929933B1 · Viljoen · 2018 [cited by applicant]
US 10129135B1 · Viljoen · 2018 [cited by applicant]
US 10637685B2 · Goel et al. · 2020 [cited by applicant]
US 10664438B2 · Sity et al. · 2020 [cited by applicant]
US 10686729B2 · Sindhu et al. · 2020 [cited by applicant]
US 11057301B2 · Thubert et al. · 2021 [cited by applicant]
US 11171882B2 · Levy et al. · 2021 [cited by applicant]
US 11350189B2 · Bakopoulos et al. · 2022 [cited by applicant]
US 11615285B2 · Reimann et al. · 2023 [cited by applicant]
US 11797353B2 · Abdulaal et al. · 2023 [cited by applicant]
US 12086080B2 · Chrysos et al. · 2024 [cited by applicant]
US 12175375B2 · Kapoor et al. · 2024 [cited by applicant]
US 12176945B2 · Bakopoulos et al. · 2024 [cited by applicant]
US 12177133B2 · Morrison et al. · 2024 [cited by applicant]
US 12189570B2 · Du et al. · 2025 [cited by applicant]
US 12190084B2 · Prabhakar et al. · 2025 [cited by applicant]
US 20040098104A1 · Sirhan et al. · 2004 [cited by applicant]
US 20060165070A1 · Hall et al. · 2006 [cited by applicant]
US 20080138067A1 · Beshai · 2008 [cited by applicant]
US 20110270987A1 · Schlansker et al. · 2011 [cited by applicant]
US 20140211622A1 · Kumer et al. · 2014 [cited by applicant]
US 20150043905A1 · Graves et al. · 2015 [cited by applicant]
US 20150074246A1 · Premji et al. · 2015 [cited by applicant]
US 20150127797A1 · Attar et al. · 2015 [cited by applicant]
US 20160044393A1 · Graves · 2016 [cited by examiner]
US 20160241491A1 · Tripathi et al. · 2016 [cited by applicant]
US 20160316005A1 · Thirumurthi et al. · 2016 [cited by applicant]
US 20160352824A1 · Miwa et al. · 2016 [cited by applicant]
US 20160380886A1 · Blair et al. · 2016 [cited by applicant]
US 20170063613A1 · Bloch · 2017 [cited by examiner]
US 20170187607A1 · Shaikh et al. · 2017 [cited by applicant]
US 20170212918A1 · Johnsen et al. · 2017 [cited by applicant]
US 20170272355A1 · Shimizu et al. · 2017 [cited by applicant]
US 20170279690A1 · Tripathi et al. · 2017 [cited by applicant]
US 20170294961A1 · Anand et al. · 2017 [cited by applicant]
US 20180070157A1 · Menard et al. · 2018 [cited by applicant]
US 20180227169A1 · Nakashima · 2018 [cited by applicant]
US 20180322387A1 · Sridharan · 2018 [cited by examiner]
US 20190080239A1 · Yang · 2019 [cited by applicant]
US 20190123961A1 · Fedyk · 2019 [cited by applicant]
US 20190165997A1 · Shaikh et al. · 2019 [cited by applicant]
US 20190166013A1 · Shaikh et al. · 2019 [cited by applicant]
US 20190245751A1 · Wong · 2019 [cited by applicant]
US 20190246187A1 · Wong · 2019 [cited by applicant]
US 20190260658A1 · Gell et al. · 2019 [cited by applicant]
US 20190294972A1 · Keller et al. · 2019 [cited by applicant]
US 20190312772A1 · Zhao et al. · 2019 [cited by applicant]
US 20190324811A1 · Ganguli et al. · 2019 [cited by applicant]
US 20190379477A1 · Tien et al. · 2019 [cited by applicant]
US 20200028774A1 · Pathikonda et al. · 2020 [cited by applicant]
US 20200259740A1 · Wetterwald et al. · 2020 [cited by applicant]
US 20200259746A1 · Thubert et al. · 2020 [cited by applicant]
US 20200304406A1 · Thubert et al. · 2020 [cited by applicant]
US 20200344155A1 · Chhibber et al. · 2020 [cited by applicant]
US 20200403923A1 · Attar et al. · 2020 [cited by applicant]
US 20210076112A1 · Hand · 2021 [cited by applicant]
US 20210097378A1 · Rodrigues et al. · 2021 [cited by applicant]
US 20210149918A1 · Finkler · 2021 [cited by applicant]
US 20210152494A1 · Johnsen et al. · 2021 [cited by applicant]
US 20210234753A1 · Ben-Moshe · 2021 [cited by examiner]
US 20220004897A1 · Jadon et al. · 2022 [cited by applicant]
US 20220014828A1 · Bakopoulos et al. · 2022 [cited by applicant]
US 20220103480A1 · Chiesa et al. · 2022 [cited by applicant]
US 20220124548A1 · Srivastava et al. · 2022 [cited by applicant]
US 20220164297A1 · Sity et al. · 2022 [cited by applicant]
US 20220172072A1 · Keller et al. · 2022 [cited by applicant]
US 20220210058A1 · Bataineh · 2022 [cited by examiner]
US 20220284294A1 · Keller et al. · 2022 [cited by applicant]
US 20220393974A1 · Wen et al. · 2022 [cited by applicant]
US 20230208915A1 · Duan · 2023 [cited by examiner]
US 20230334292A1 · Zhang et al. · 2023 [cited by applicant]
US 20230336414A1 · Miriyala et al. · 2023 [cited by applicant]
US 20230412265A1 · Bakopoulos et al. · 2023 [cited by applicant]
US 20240020265A1 · Fu et al. · 2024 [cited by applicant]
US 20240037063A1 · Suh et al. · 2024 [cited by applicant]
US 20240098088A1 · Adogla et al. · 2024 [cited by applicant]
US 20240098104A1 · Patronas et al. · 2024 [cited by applicant]
US 20240113943A1 · Patronas et al. · 2024 [cited by applicant]
US 20240119267A1 · Kierat et al. · 2024 [cited by applicant]
US 20240121158A1 · Bakopoulos et al. · 2024 [cited by applicant]
US 20240129161A1 · Miriyala et al. · 2024 [cited by applicant]
US 20240152396A1 · Brar et al. · 2024 [cited by applicant]
US 20240152409A1 · Brar et al. · 2024 [cited by applicant]
US 20240160495A1 · Brar et al. · 2024 [cited by applicant]
US 20240160496A1 · Brar et al. · 2024 [cited by applicant]
US 20240160691A1 · Kim · 2024 [cited by examiner]
US 20240185048A1 · Sebastian et al. · 2024 [cited by applicant]
US 20240214294A1 · Miriyala et al. · 2024 [cited by applicant]
US 20240314197A1 · Qu · 2024 [cited by applicant]
US 20240406072A1 · Bhardwaj · 2024 [cited by examiner]
US 20250181358A1 · Bilal · 2025 [cited by examiner]
Urman et al., Pending U.S. Appl. No. 18/242,637, filed Sep. 6, 2023. [cited by applicant]
Patronas et al., Pending U.S. Appl. No. 18/427,046, filed Jan. 30, 2024. [cited by applicant]
Patronas et al., Pending U.S. Appl. No. 18/415,878, filed Jan. 18, 2024. [cited by applicant]