IP Library › Granted Patent US 12,585,929
Granted Patent B2
US 12,585,929 · App. 17/668,200 · Granted Mar 24, 2026

Layered gradient accumulation and modular pipeline parallelism for improved training of machine learning models

Inventor: Joel Lamy-Poirier (Montreal, CA)
Assignee: ServiceNow, Inc.
G06N3/063G06F18/2155G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,929
App. No.
17/668,200
Filed
Feb 9, 2022
Granted
Mar 24, 2026
Kind
B2
Art Unit
2143
USPC
706/33
Abstract

A method is provided including: (i) assigning sequentially-ordered layers of a machine learning model to a plurality of compute nodes, each of the layers being assigned to exactly one of the nodes; (ii) dividing training data into micro-batches; (iii) forward-propagating the micro-batches through the model, each node operating in parallel to generate respective activation states for the micro-batches with their assigned layers, and with the activation states being communicated between the nodes according to the layers' sequential ordering; and (iv) backward-propagating the micro-batches through the model, each node operating in parallel to generate respective error states for the micro-batches with their assigned layers, with the error states being communicated between the nodes according to the layers' reverse sequential ordering, wherein each of the nodes completes the backward-propagation of all micro-batches through a given layer prior to performing backward-propagation through any layer that precedes the given layer in the sequential ordering.

Claims (69)

1 . A computer-implemented method comprising:

operating M compute nodes to train a machine learning model based on a batch of N training examples, wherein the batch of N training examples is divided into n micro-batches, wherein the machine learning model comprises L layers, each layer defined by a respective plurality of parameters of the machine learning model, wherein operating the M compute nodes to train the machine learning model based on the batch of N training examples comprises updating the parameters of the machine learning model by:

sequentially applying, by a first compute node of the M compute nodes, each of the micro-batches to a first layer of the L layers to generate respective first-layer activation states;

transmitting, from the first compute node to a second compute node of the M compute nodes, the first-layer activation states;

sequentially applying, by the second compute node, each of the first-layer activation states to a second layer of the L layers to generate respective second-layer activation states for each of the micro-batches, wherein the second compute node applying a particular one of the first layer activation states to the second layer at least partially overlaps in time with the first compute node transmitting a subsequent one of the first layer activation states to the second compute node;

transmitting, from an M th compute node of the M compute nodes to the first compute node, M th -layer activation states for each of the micro-batches;

sequentially applying, by the first compute node subsequent to generating the first-layer activation states, each of the M th -layer activation states to an (M+1) th layer of the L layers to generate respective (M+1) th -layer activation states for each of the micro-batches;

transmitting, from an (M−1) th compute node of the M compute nodes to the M th compute node, (L−1) th -layer activation states for each of the micro-batches;

sequentially applying, by the M th compute node, each of the (L−1) th -layer activation states to an L th layer of the L layers to generate respective L th -layer activation states for each of the micro-batches;

based on the L th -layer activation states, generating respective L th -layer error states for each of the micro-batches;

sequentially applying, by the M th compute node, each of the L th -layer error states to the L th layer to generate respective (L−1) th -layer error states for each of the micro-batches and respective L th -layer parameter update information for each of the micro-batches;

transmitting, from the M th compute node to the (M−1) th compute node, the (L−1) th -layer error states;

sequentially applying, by the (M−1) th compute node, each of the (L−1) th -layer error states to the (L−1) th -layer to generate respective (L−2) th -layer error states for each of the micro-batches and respective (L−1) th -layer parameter update information for each of the micro-batches, wherein the (M−1) th compute node applying a particular one of the (L−1) th -layer error states to the (L−1) th -layer at least partially overlaps in time with the M th compute node transmitting a subsequent one of the (L−1) th -layer error states to the (M−1) th compute node;

transmitting, from the first compute node to the M th compute node, (L−M) th -layer error states for each of the micro-batches; and

sequentially applying, by the M th compute node subsequent to generating the (L−1) th -layer error states, each of the (L−M) th -layer error states to an (L−M) th layer of the L layers to generate respective (L−M−1) th -layer activation states for each of the micro-batches and respective (L−M) th -layer parameter update information for each of the micro-batches.

2 . The computer-implemented method of claim 1 , wherein the machine learning model comprises a transformer.

3 . The computer-implemented method of claim 2 , wherein each layer of the L layers includes a discrete number of layers of the transformer.

4 . The computer-implemented method of claim 1 , further comprising:

transmitting, to the first compute node, a portion of the plurality of parameters of the machine learning model that define the (M+1) th -layer prior to the first compute node applying each of the M th -layer activation states to the (M+1) th layer to generate respective (M+1) th -layer activation states for each of the micro-batches,

wherein the first compute node applying at least one of the micro-batches to the first layer of the L layers to generate at least one respective first-layer activation states at least partially overlaps in time with transmitting the portion of the plurality of parameters of the machine learning model that define the (M+1) th -layer to the first compute node.

5 . The computer-implemented method of claim 1 , wherein the number of micro-batches n equals the number of compute nodes M.

6 . The computer-implemented method of claim 1 , wherein each micro-batch includes a single training example from the batch of N training examples.

7 . The computer-implemented method of claim 1 , further comprising:

based on the L th -layer parameter update information for each of the micro-batches, determining parameter updates for the plurality of parameters of the machine learning model that define the L th -layer.

8 . The computer-implemented method of claim 7 , wherein determining parameter updates for the plurality of parameters of the machine learning model that define the L th -layer commences prior to the (M−1) th compute node generating at least one of the (L−2) th -layer error states.

9 . The computer-implemented method of claim 7 , wherein determining parameter updates for the plurality of parameters of the machine learning model that define the L th -layer comprises:

transmitting, from the M th compute node to an additional compute node that is not one of the M compute nodes, the L th -layer parameter update information for each of the micro-batches; and

transmitting, from the additional compute node to the M th compute node, L th -layer parameter update information for one or more additional micro-batches that are not part of the n micro-batches.

10 . The computer-implemented method of claim 7 , further comprising:

updating the plurality of parameters of the machine learning model that define the L th -layer based on the parameter updates.

11 . An article of manufacture including a non-transitory computer-readable medium, having stored thereon program instructions that, upon execution by a computing device, cause the computing device to perform operations comprising:

operating M compute nodes to train a machine learning model based on a batch of N training examples, wherein the batch of N training examples is divided into n micro-batches, wherein the machine learning model comprises L layers, each layer defined by a respective plurality of parameters of the machine learning model, wherein operating the M compute nodes to train the machine learning model based on the batch of N training examples comprises updating the parameters of the machine learning model by:

sequentially applying, by a first compute node of the M compute nodes, each of the micro-batches to a first layer of the L layers to generate respective first-layer activation states;

transmitting, from the first compute node to a second compute node of the M compute nodes, the first-layer activation states;

sequentially applying, by the second compute node, each of the first-layer activation states to a second layer of the L layers to generate respective second-layer activation states for each of the micro-batches, wherein the second compute node applying a particular one of the first layer activation states to the second layer at least partially overlaps in time with the first compute node transmitting a subsequent one of the first layer activation states to the second compute node;

transmitting, from an M th compute node of the M compute nodes to the first compute node, M th -layer activation states for each of the micro-batches;

sequentially applying, by the first compute node subsequent to generating the first-layer activation states, each of the M th -layer activation states to an (M+1) th layer of the L layers to generate respective (M+1) th -layer activation states for each of the micro-batches;

transmitting, from an (M−1) th compute node of the M compute nodes to the M th compute node, (L−1) th -layer activation states for each of the micro-batches;

sequentially applying, by the M th compute node, each of the (L−1) th -layer activation states to an L th layer of the L layers to generate respective L th -layer activation states for each of the micro-batches;

based on the L th -layer activation states, generating respective L th -layer error states for each of the micro-batches;

sequentially applying, by the M th compute node, each of the L th -layer error states to the L th layer to generate respective (L−1) th -layer error states for each of the micro-batches and respective L th -layer parameter update information for each of the micro-batches;

transmitting, from the M th compute node to the (M−1) th compute node, the (L−1) th -layer error states;

sequentially applying, by the (M−1) th compute node, each of the (L−1) th -layer error states to the (L−1) th -layer to generate respective (L−2) th -layer error states for each of the micro-batches and respective (L−1) th -layer parameter update information for each of the micro-batches, wherein the (M−1) th compute node applying a particular one of the (L−1) th -layer error states to the (L−1) th layer at least partially overlaps in time with the M th compute node transmitting a subsequent one of the (L−1) th -layer error states to the (M−1) th compute node;

transmitting, from the first compute node to the M th compute node, (L−M) th -layer error states for each of the micro-batches; and

sequentially applying, by the M th compute node subsequent to generating the (L−1) th -layer error states, each of the (L−M) th -layer error states to an (L−M) th layer of the L layers to generate respective (L−M−1) th -layer activation states for each of the micro-batches and respective (L−M) th -layer parameter update information for each of the micro-batches.

12 . The article of manufacture of claim 11 , wherein the operations further comprise:

transmitting, to the first compute node, a portion of the plurality of parameters of the machine learning model that define the (M+1) th -layer prior to the first compute node applying each of the M th -layer activation states to the (M+1) th layer to generate respective (M+1) th -layer activation states for each of the micro-batches,

wherein the first compute node applying at least one of the micro-batches to the first layer of the L layers to generate at least one respective first-layer activation states at least partially overlaps in time with transmitting the portion of the plurality of parameters of the machine learning model that define the (M+1) th -layer to the first compute node.

13 . The article of manufacture of claim 11 , wherein the operations further comprise:

based on the L th -layer parameter update information for each of the micro-batches, determining parameter updates for the plurality of parameters of the machine learning model that define the L th -layer.

14 . The article of manufacture of claim 13 , wherein determining parameter updates for the plurality of parameters of the machine learning model that define the L th -layer comprises:

transmitting, from the M th compute node to an additional compute node that is not one of the M compute nodes, the L th -layer parameter update information for each of the micro-batches; and

transmitting, from the additional compute node to the M th compute node, L th -layer parameter update information for one or more additional micro-batches that are not part of the n micro-batches.

15 . A computer-implemented method comprising:

operating M compute nodes to train a machine learning model based on a batch of N training examples, wherein the batch of N training examples is divided into n micro-batches, wherein the machine learning model comprises L layers, each layer defined by a respective plurality of parameters of the machine learning model, wherein operating the M compute nodes to train the machine learning model based on the batch of N training examples comprises updating the parameters of the machine learning model by:

sequentially applying, by a first compute node of the M compute nodes, each of the micro-batches to a first layer of the L layers to generate respective first-layer activation states;

transmitting, from the first compute node to a second compute node of the M compute nodes, the first-layer activation states;

sequentially applying, by the second compute node, each of the first-layer activation states to a second layer of the L layers to generate respective second-layer activation states for each of the micro-batches, wherein the second compute node applying a particular one of the first layer activation states to the second layer at least partially overlaps in time with the first compute node transmitting a subsequent one of the first layer activation states to the second compute node;

transmitting, from an Mth compute node of the M compute nodes to the first compute node, Mth-layer activation states for each of the micro-batches;

sequentially applying, by the first compute node subsequent to generating the first-layer activation states, each of the Mth-layer activation states to an (M+1) layer of the L layers to generate respective (M+1)th-layer activation states for each of the micro-batches;

sequentially applying, by the Mth compute node, (L−1)th-layer activation states received from an (M−1)th compute node of the M compute nodes to an L th layer of the L layers to generate respective L th -layer activation states for each of the micro-batches; and

based on the L th -layer activation states, generating parameter update information for each of the micro-batches and for each of the L layers of the machine learning model.

16 . The computer-implemented method of claim 15 , further comprising:

transmitting, to the first compute node, a portion of the plurality of parameters of the machine learning model that define the (M+1) th -layer prior to the first compute node applying each of the M th -layer activation states to the (M+1) th layer to generate respective (M+1) th -layer activation states for each of the micro-batches,

wherein the first compute node applying at least one of the micro-batches to the first layer of the L layers to generate at least one respective first-layer activation states at least partially overlaps in time with transmitting the portion of the plurality of parameters of the machine learning model that define the (M+1) th -layer to the first compute node.

17 . The computer-implemented method of claim 15 , wherein the number of micro-batches n equals the number of compute nodes M.

18 . The computer-implemented method of claim 15 , wherein each micro-batch includes a single training example from the batch of N training examples.

19 . The computer-implemented method of claim 15 , wherein generating parameter update information for each of the micro-batches and for each of the L layers of the machine learning model comprises:

commencing, by the M th compute node, determination of parameter updates for the plurality of parameters of the machine learning model that define the L th -layer prior to the (M−1) th compute node completing the generation of (L−2) th -layer error states by applying (L−1) th -layer error states received from the M th compute node to the (L−1) th -layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2026
From: LAMY POIRIER, JOEL
To: SERVICENOW, INC.
Reel/Frame 074610/0148 →
Continuity (2)
Provisional Application 63194389 · May 28, 2021
Related Publication 20220383084A1 · Dec 1, 2022
References Cited (185)
US 4941084A · Terada et al. · 1990 [cited by applicant]
US 5185860A · Wu · 1993 [cited by applicant]
US 5237518A · Sztipanovits et al. · 1993 [cited by applicant]
US 5261097A · Saxon · 1993 [cited by applicant]
US 5265252A · Rawson, III et al. · 1993 [cited by applicant]
US 5367685A · Gosling · 1994 [cited by applicant]
US 5390297A · Barber et al. · 1995 [cited by applicant]
US 5442791A · Wrabetz et al. · 1995 [cited by applicant]
US 5452415A · Hotka · 1995 [cited by applicant]
US 5522042A · Fee et al. · 1996 [cited by applicant]
US 5533116A · Vesterinen · 1996 [cited by applicant]
US 5655081A · Bonnell et al. · 1997 [cited by applicant]
US 5659736A · Hasegawa et al. · 1997 [cited by applicant]
US 5671412A · Christiano · 1997 [cited by applicant]
US 5696701A · Burgess et al. · 1997 [cited by applicant]
US 5715463A · Merkin · 1998 [cited by applicant]
US 5745879A · Wyman · 1998 [cited by applicant]
US 5761502A · Jacobs · 1998 [cited by applicant]
US 5764913A · Jancke et al. · 1998 [cited by applicant]
US 5887139A · Madison, Jr. et al. · 1999 [cited by applicant]
US 5909217A · Bereiter · 1999 [cited by applicant]
US 5937165A · Schwaller et al. · 1999 [cited by applicant]
US 5949976A · Chappelle · 1999 [cited by applicant]
US 5978594A · Bonnell et al. · 1999 [cited by applicant]
US 6021437A · Chen et al. · 2000 [cited by applicant]
US 6041347A · Harsham et al. · 2000 [cited by applicant]
US 6088717A · Reed et al. · 2000 [cited by applicant]
US 6101500A · Lau · 2000 [cited by applicant]
US 6128016A · Coelho et al. · 2000 [cited by applicant]
US 6131118A · Stupek, Jr. et al. · 2000 [cited by applicant]
US 6134581A · Ismael et al. · 2000 [cited by applicant]
US 6138122A · Smith et al. · 2000 [cited by applicant]
US 6148335A · Haggard et al. · 2000 [cited by applicant]
US 6166732A · Mitchell et al. · 2000 [cited by applicant]
US 6167448A · Hemphill et al. · 2000 [cited by applicant]
US 6175866B1 · Holloway et al. · 2001 [cited by applicant]
US 6175878B1 · Seaman et al. · 2001 [cited by applicant]
US 6260050B1 · Yost et al. · 2001 [cited by applicant]
US 6263457B1 · Anderson et al. · 2001 [cited by applicant]
US 6272150B1 · Hrastar et al. · 2001 [cited by applicant]
US 6336138B1 · Caswell et al. · 2002 [cited by applicant]
US 6363421B2 · Barker et al. · 2002 [cited by applicant]
US 6393386B1 · Zager et al. · 2002 [cited by applicant]
US 6397245B1 · Johnson, II et al. · 2002 [cited by applicant]
US 6434626B1 · Prakash et al. · 2002 [cited by applicant]
US 6438592B1 · Killian · 2002 [cited by applicant]
US 6456306B1 · Chin et al. · 2002 [cited by applicant]
US 6466932B1 · Dennis et al. · 2002 [cited by applicant]
US 6487590B1 · Foley et al. · 2002 [cited by applicant]
US 6505248B1 · Casper et al. · 2003 [cited by applicant]
US 6526442B1 · Stupek, Jr. et al. · 2003 [cited by applicant]
US 6621823B1 · Mellquist et al. · 2003 [cited by applicant]
US 6707795B1 · Noorhosseini et al. · 2004 [cited by applicant]
US 6742015B1 · Bowman-Amuah · 2004 [cited by applicant]
US 6763380B1 · Mayton et al. · 2004 [cited by applicant]
US 6816898B1 · Scarpelli et al. · 2004 [cited by applicant]
US 6895586B1 · Brasher et al. · 2005 [cited by applicant]
US 6948175B1 · Fong et al. · 2005 [cited by applicant]
US 6985901B1 · Sachse et al. · 2006 [cited by applicant]
US 7003564B2 · Greuel et al. · 2006 [cited by applicant]
US 7028228B1 · Lovy et al. · 2006 [cited by applicant]
US 7043537B1 · Pratt · 2006 [cited by applicant]
US 7043661B2 · Valadarsky et al. · 2006 [cited by applicant]
US 7062683B2 · Warpenburg et al. · 2006 [cited by applicant]
US 7096459B2 · Keller et al. · 2006 [cited by applicant]
US 7146574B2 · Goldthwaite et al. · 2006 [cited by applicant]
US 7197466B1 · Peterson et al. · 2007 [cited by applicant]
US 7215360B2 · Gupta · 2007 [cited by applicant]
US 7216304B1 · Gourdol et al. · 2007 [cited by applicant]
US 7222147B1 · Black et al. · 2007 [cited by applicant]
US 7281170B2 · Taylor et al. · 2007 [cited by applicant]
US 7412502B2 · Fearn et al. · 2008 [cited by applicant]
US 7505872B2 · Keller et al. · 2009 [cited by applicant]
US 7593013B2 · Agutter et al. · 2009 [cited by applicant]
US 7596716B2 · Frost et al. · 2009 [cited by applicant]
US 7617073B2 · Trinon et al. · 2009 [cited by applicant]
US 7660731B2 · Chaddha et al. · 2010 [cited by applicant]
US 7676294B2 · Baier et al. · 2010 [cited by applicant]
US 7676437B2 · Satkunanathan et al. · 2010 [cited by applicant]
US 7840490B1 · Sellers et al. · 2010 [cited by applicant]
US 7877783B1 · Cline et al. · 2011 [cited by applicant]
US 7890869B1 · Mayer et al. · 2011 [cited by applicant]
US 7966398B2 · Wiles, Jr. · 2011 [cited by applicant]
US 8060396B1 · Bessler et al. · 2011 [cited by applicant]
US 8196210B2 · Sterin · 2012 [cited by applicant]
US 8321948B2 · Robinson et al. · 2012 [cited by applicant]
US 8407669B2 · Yee et al. · 2013 [cited by applicant]
US 8554750B2 · Rangarajan et al. · 2013 [cited by applicant]
US 8595647B2 · Sabin et al. · 2013 [cited by applicant]
US 8620818B2 · Hughes et al. · 2013 [cited by applicant]
US 8646093B2 · Myers et al. · 2014 [cited by applicant]
US 8674992B2 · Poston et al. · 2014 [cited by applicant]
US 8725647B2 · Disciascio et al. · 2014 [cited by applicant]
US 9053460B2 · Gilbert et al. · 2015 [cited by applicant]
US 10673963B1 · Feiguine et al. · 2020 [cited by applicant]
US 10749943B1 · Feiguine et al. · 2020 [cited by applicant]
US 10771344B2 · Bitterfeld et al. · 2020 [cited by applicant]
US 10824650B2 · Bar Oz et al. · 2020 [cited by applicant]
US 10944654B2 · Rimar et al. · 2021 [cited by applicant]
US 11089115B2 · Garty et al. · 2021 [cited by applicant]
US 11095506B1 · Erblat et al. · 2021 [cited by applicant]
US 11520592B2 · Pudipeddi · 2022 [cited by examiner]
US 20020116340A1 · Hellberg et al. · 2002 [cited by applicant]
US 20020133584A1 · Greuel et al. · 2002 [cited by applicant]
US 20020158969A1 · Gupta · 2002 [cited by applicant]
US 20030118087A1 · Goldthwaite et al. · 2003 [cited by applicant]
US 20030200293A1 · Fearn et al. · 2003 [cited by applicant]
US 20050015217A1 · Weidl et al. · 2005 [cited by applicant]
US 20050091356A1 · Izzo · 2005 [cited by applicant]
US 20060026453A1 · Frost et al. · 2006 [cited by applicant]
US 20060095461A1 · Raymond · 2006 [cited by applicant]
US 20060179058A1 · Bram et al. · 2006 [cited by applicant]
US 20060293942A1 · Chaddha et al. · 2006 [cited by applicant]
US 20070033279A1 · Battat et al. · 2007 [cited by applicant]
US 20070188494A1 · Agutter et al. · 2007 [cited by applicant]
US 20070288389A1 · Vaughan et al. · 2007 [cited by applicant]
US 20080133289A1 · Armour et al. · 2008 [cited by applicant]
US 20080148253A1 · Badwe et al. · 2008 [cited by applicant]
US 20080319779A1 · Hughes et al. · 2008 [cited by applicant]
US 20090088875A1 · Baier et al. · 2009 [cited by applicant]
US 20090228984A1 · Sterin · 2009 [cited by applicant]
US 20100110932A1 · Doran et al. · 2010 [cited by applicant]
US 20180123940A1 · Rimar et al. · 2018 [cited by applicant]
US 20190104398A1 · Owen et al. · 2019 [cited by applicant]
US 20190130292A1 · N · 2019 [cited by examiner]
US 20200050689A1 · Tal et al. · 2020 [cited by applicant]
US 20200204443A1 · Bar Oz et al. · 2020 [cited by applicant]
US 20200311536A1 · Venkataramani · 2020 [cited by examiner]
US 20210019151A1 · Pudipeddi · 2021 [cited by examiner]
US 20210019152A1 · Pudipeddi · 2021 [cited by examiner]
US 20210019634A1 · Pudipeddi · 2021 [cited by examiner]
US 20210042620A1 · Chen · 2021 [cited by examiner]
US 20210194764A1 · Badyan et al. · 2021 [cited by applicant]
EP 0433979 · 1991 [cited by applicant]
EP 1607824 · 2005 [cited by applicant]
WO WO9934285 · 1999 [cited by applicant]
WO WO0052559 · 2000 [cited by applicant]
WO WO0179970 · 2001 [cited by applicant]
Hu et al. (“PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices”, 2022 25th Euromicro Conference on Digital System Design (DSD), pp. 298-307). [cited by examiner]
Brown et al., “Language Models are Few-Shot Learners”, Jul. 22, 2020. [cited by applicant]
Chen et al., “Training Deep Nets with Sublinear Memory Cost”, Apr. 22, 2016. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, May 24, 2019. [cited by applicant]
Duchi et al., “Asynchronous stochastic convex optimization”, Aug. 4, 2015. [cited by applicant]
Fedus et al., “Switch Transformers: Scaling To Trillion Parameter Models With Simple And Efficient Sparsity”, Jan. 11, 2021. [cited by applicant]
Golmant et al., “On The Computational Inefficiency Of Large Batch Sizes For Stochastic Gradient Descent”, Nov. 30, 2018. [cited by applicant]
Hannah et al., “On Unbounded Delays in Asychronous Parallel Fixed-Point Algorithms”, Aug. 18, 2017. [cited by applicant]
Harlap et al., “PipeDream: Fast and Efficient Pipeline Parallel DNN Training”, Jun. 8, 2018. [cited by applicant]
Hendrycks et al., “Gaussian Error Linear Units (GELUs)”, Jul. 8, 2020. [cited by applicant]
Huang et al., “GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism”, Jul. 25, 2019. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, Mar. 2, 2015. [cited by applicant]
Kaplan et al., “Scaling Laws for Neural Language Models”, Jan. 23, 2020. [cited by applicant]
Kwon et al., “Beyond the Memory Wall: A Case for Memory-centric HPC System for Deep Learning”, Feb. 18, 2019. [cited by applicant]
Lepikhin et al., “GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding”, Jun. 30, 2020. [cited by applicant]
Lewis et al., “BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension”, Oct. 29, 2019. [cited by applicant]
Li et al., “Evaluating Modern GPU Interconnect: PCle, NVLink, NV-SLI, NVSwitch and GPUDirect”, Mar. 11, 2019. [cited by applicant]
McCandlish et al., “An Empirical Model of Large-Batch Training”, Dec. 14, 2018. [cited by applicant]
Narang et al., “Mixed Precision Training”, Feb. 15, 2018. [cited by applicant]
Narayanan et al., “Memory-Efficient Pipeline-Parallel DNN Training”, Jul. 22, 2021. [cited by applicant]
Narayanan et al., “Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM”, Aug. 23, 2021. [cited by applicant]
Niu et al., “HOGWILD !: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent”, Nov. 11, 2011. [cited by applicant]
Radford et al., “Improving Language Understanding by Generative Pre-Training”, 12pages. [cited by applicant]
Radford et al., “Language Models are Unsupervised Multitask Learners”, 24 pages. [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, Journal of Machine earning Research 21, Jul. 28, 2020. [cited by applicant]
Rajbhandari et al., “ZeRO: Memory Optimizations Toward Training Trillion Parameter Models”, May 13, 2020. [cited by applicant]
Rajbhandari et al., “ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning”, Apr. 16, 2021. [cited by applicant]
Ren et al., “ZeRO-Offload: Democratizing Billion-Scale Model Training”, Jan. 18, 2021. [cited by applicant]
Shallue et al., “Measuring the Effects of Data Parallelism on Neural Network Training”, Journal of Machine Learning Research 20, Jul. 19, 2019. [cited by applicant]
Shoeybi et al., “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”, Mar. 13, 2020. [cited by applicant]
Smith et al., “Don't Decay The Learning Rate, Increase The Batch Size”, Feb. 24, 2018. [cited by applicant]
Stich et al., “Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates”, Mar. 3, 2021. [cited by applicant]
Stich et al., “Sparsified SGD with Memory”, Nov. 28, 2018. [cited by applicant]
Tang et al., “1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed”, Jun. 29, 2021. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, Dec. 6, 2017. [cited by applicant]
Yang et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding”, Jan. 2, 2020. [cited by applicant]
“Solutions for the Data Center”, NVIDIA, 5 pages. [cited by applicant]
“Visual intuition on ring-Allreduce for distributed Deep Learning”, Towards Data Science, Aug. 1, 2019, 8 pages. [cited by applicant]
NVIDIA / Megatron-LM, Github, https://github.com/NVIDIA/Megatron-LM, downloaded Apr. 27, 2022. [cited by applicant]
Beaumont et al., “Optimal GPU-CPU Offloading Strategies for Deep Neural Network Training”, HAL open science, Oct. 21, 2019, 28 pages. [cited by applicant]
Xue et al., “mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer”, Mar. 11, 2021. [cited by applicant]
Tay et al., “Efficient Transformers: A Survey”, Mar. 14, 2022. [cited by applicant]
Gale et al., “Sparse GPU Kernels for Deep Learning”, Aug. 31, 2020. [cited by applicant]
Zaheer et al., “Big Bird: Transformers for Longer Sequences”, Jan. 8, 2021. [cited by applicant]
Beltagy et al., “Longformer: The Long-Document Transformer”, Dec. 2, 2020. [cited by applicant]
Child et al., “Generating Long Sequences with Sparse Transformers”, Apr. 23, 2019. [cited by applicant]
Li et al., “Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers”, Jun. 23, 2020. [cited by applicant]