IP Library Granted Patent US 12,511,543
Granted Patent B2
US 12,511,543 · App. 16/675,069 · Granted Dec 30, 2025

Distributed weight update for backpropagation of a neural network

Inventors: Sharan Chetlur (San Jose, CA); Natalia Gimelshein (San Jose, CA); Thor Mikal Johnsen (Houston, TX); Simon Layton (West Hartford, CT)
Assignee: NVIDIA Corporation
G06N3/084G01C21/3608G06N3/04G06N3/063G10L15/16G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,543
App. No.
16/675,069
Filed
Nov 5, 2019
Granted
Dec 30, 2025
Kind
B2
Art Unit
2655
USPC
706/25
Abstract

Speed of training a neural network is improved by updating the weights of the neural network in parallel. In at least one embodiment, after back propagation, gradients are distributed to a plurality of processors, each of which calculate a portion of the updated weights of the neural network.

Claims (66)

1 . One or more processors, A processor comprising:

circuitry to:

cause different processing cores to perform updates to respective weights of different portions of one or more neural networks in parallel;

cause at least one of the different processing cores to receive the updates to the respective weights of the different portions of the one or more neural networks from others of the different processing cores; and

cause the at least one of the different processing cores to generate an updated copy of the one or more neural networks based, at least in part, on the received updates.

2 . The one or more processors of claim 1 , wherein the different processing cores apply one or more gradients to different sets of nodes of the one or more neural networks.

3 . The one or more processors of claim 1 , where in the one or more neural networks are is trained, at least in part, by generating the updates to the respective weights by combining the updates to the respective weights produced in parallel by the different processing cores.

4 . The one or more processors of claim 1 , wherein the one or more processors further divide a weight update operation into a plurality of partial weight update operations and distributes individual partial weight update operations to the different processing cores.

5 . The one or more processors of claim 4 , wherein the partial weight update operations are produced by dividing an initial weight and gradient update into a plurality of distinct portions.

6 . The one or more processors of claim 4 , wherein each partial weight update is executed using a different thread.

7 . The one or more processors of claim 4 , wherein the one or more processors further gather the partial weight updates to generate the updated copy.

8 . A system, comprising:

one or more processors to perform software to:

cause different processing cores to perform updates to respective weights of different portions of one or more neural networks in parallel;

cause at least one of the different processing cores to receive the updates to the respective weights of the different portions of the one or more neural networks from others of the different processing cores; and

cause the at least one of the different processing cores to generate an updated copy of the one or more neural networks based, at least in part, on the received updates; and

one or more memories to store the one or more neural networks.

9 . The system of claim 8 where in the one or more neural networks are is trained, at least in part, by:

generating the updates to the respective weights in parallel using the different processing cores one or more second processors; and

combining the updates to the respective weights to produce a weight gradient update.

10 . The system of claim 9 , wherein the updates to the respective weights are produced in parallel using a plurality of worker threads.

11 . The system of claim 9 , wherein the one or more neural networks are is trained at least in part by:

forward propagating an input through the one or more neural networks to produce an output;

determining an error based at least in part on a difference between the output and an expected value; and

backpropagating the error to determine a gradient, the updates to the respective weights based at least in part on the gradient.

12 . The system of claim 9 , wherein the updates to the respective weights are produced at least by:

identifying a plurality of subsets of network nodes of the one or more neural networks; and

producing the updates to the respective weights for each subset in the plurality of subsets.

13 . The system of claim 12 , wherein the plurality of subsets are non-overlapping subsets of weights of the one or more neural networks.

14 . The system of claim 12 , wherein:

an individual subset of the plurality of subsets includes a quantity of node weights; and

the quantity of node weights is determined based at least in part on an amount of processing power available to a worker assigned to process the individual subset relative to other workers assigned to process other subsets.

15 . The system of claim 8 , wherein:

the system determines a set of gradients for each input of a set of input values; and

the set of gradients is distributed to each of the different processing cores.

16 . A method, comprising training one or more neural networks by, at least in part:

causing different processing cores to perform updates to respective weights of different portions of one or more neural networks in parallel;

causing at least one of the different processing cores to receive the updates to the respective weights of the different portions of the one or more neural networks from others of the different processing cores; and

causing the at least one of the different processing cores to generate an updated copy of the one or more neural networks based, at least in part, on the received updates.

17 . The method of claim 16 wherein:

the one or more neural networks are trained at least in part by distributing gradient information to plurality of workers;

the plurality of workers calculate the updates to the respective weights in parallel; and

the updates to the respective weights are aggregated to produce new weight values of the one or more neural networks.

18 . The method of claim 17 , wherein the one or more neural networks are is trained at least in part by:

forward propagating an input through the one or more neural networks to produce an output;

determining an error based at least in part on the output; and

determining a gradient by at least backpropagating the error, the updates to the respective weights based at least in part on the gradient.

19 . The method of claim 18 , wherein:

the gradient is distributed to each of the plurality of workers; and

the plurality of workers calculate the updates to the respective weights.

20 . The method of claim 16 , wherein each worker of a plurality of workers executes on a different one of the different processor cores.

21 . The method of claim 16 , wherein each worker of a plurality of workers executes in parallel using the different processing cores on a graphical processing unit.

22 . The method of claim 17 , wherein a number of the updates to the respective weights matches a number of available different processing cores of a computer system.

23 . The method of claim 17 , wherein the updates to the respective weights are is divided into substantially equal nonoverlapping groups of node weights to produce the different portions.

24 . A speech processing system comprising one or more neural networks that take a digital representation of sound as input and identifies elements of human speech, wherein respective weights of different portions of the one or more neural networks are trained to recognize human speech separately in parallel by different processing cores, caused by:

performing, by the different processing cores, updates to the respective weights;

receiving, by at least one of the different processing cores, the updates to the respective weights of the different portions of the one or more neural networks from others of the different processing cores; and

generating, by the at least one of the different processing cores, an updated copy of the one or more neural networks based, at least in part, on the received updates.

25 . The speech processing system of claim 24 wherein the one or more neural networks:

generates updates to the respective weights in parallel; and

combines the updates to the respective wights into the updated copy.

26 . The speech processing system of claim 24 , wherein the speech processing system further comprises one or more processors and memory to store executable instructions that, as a result of being executed by the different processing cores, cause the speech processing system to at least:

obtain data representing audio from a microphone;

process the data using the one or more neural networks to identify a spoken word represented in the data; and

perform an action based at least in part on the identity of the spoken word.

27 . The speech processing system of claim 26 , wherein the action is a navigation request to be processed by a navigation system of a vehicle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2020
From: CHETLUR, SHARAN; GIMELSHEIN, NATALIA; JOHNSEN, THOR MIKAL; LAYTON, SIMON
To: NVIDIA CORPORATION
Reel/Frame 054242/0105 →
Continuity (1)
Related Publication 20210133583A1 · May 6, 2021
References Cited (35)
US 10210860B1 · Ward · 2019 [cited by examiner]
US 10325200B2 · Yu · 2019 [cited by examiner]
US 10776684B1 · Agarwal · 2020 [cited by examiner]
US 12387082B2 · Datta · 2025 [cited by examiner]
US 20150371132A1 · Gemello et al. · 2015 [cited by applicant]
US 20160379111A1 · Bittner, Jr. · 2016 [cited by examiner]
US 20180293492A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180322606A1 · Das et al. · 2018 [cited by applicant]
US 20190138934A1 · Prakash et al. · 2019 [cited by applicant]
US 20190188569A1 · Naumov · 2019 [cited by applicant]
US 20190205745A1 · Sridharan et al. · 2019 [cited by applicant]
US 20200097822A1 · Vishnu · 2020 [cited by examiner]
US 20200104718A1 · Taba · 2020 [cited by examiner]
US 20200327884A1 · Bui · 2020 [cited by examiner]
US 20210125040A1 · Cassidy · 2021 [cited by examiner]
US 20210132688A1 · Kim · 2021 [cited by examiner]
CN 108876702A · 2018 [cited by applicant]
CN 109034385A · 2018 [cited by applicant]
CN 109993277A · 2019 [cited by applicant]
WO 2019172878A1 · 2019 [cited by applicant]
IEEE Computer Society, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” The Institute of Electrical and Electronics Engineers, Inc., Aug. 29, 2008, 70 pages. [cited by applicant]
International Electrotechnical Commission, “Functional safety of electrical/electronic/programmable electronic safety-related systems,” IEC Standard 61508-1, Apr. 2014, 23 pages. [cited by applicant]
International Organization for Standardization, “Road vehicles—Functional safety,” ISO Standard 26262, https://www.iso.org/obp/ui/#iso:std:iso:26262:-1:ed-1:v1:en, Nov. 11, 2011, 35 pages. [cited by applicant]
Rajbhandari et al., “ZeRO: Memory Optimizations Toward Training Trillion Parameter Models,” Oct. 4, 2019, retreived Oct. 28, 2020, from https://arxiv.org/abs/1910.02054, 24 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
International Search Report and Written Opinion mailed Mar. 16, 2021, for Application No. PCT/US2020/058424, filed Oct. 30, 2020, 18 pages. [cited by applicant]
Kumar et al., “Scale MLPerf-0.6 Models on Google TPU-v3 Pods,” Oct. 2, 2019, 8 pages. [cited by applicant]
Office Action for Chinese Application No. 202080076627.9, mailed Aug. 14, 2024, 26 pages. [cited by applicant]
Office Action for Japanese Application No. 2022-523930, mailed Sep. 25, 2024, 6 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2205672.5, mailed Apr. 27, 2023, 2 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2205672.5, mailed Jun. 10, 2024, 4 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2205672.5, mailed May 10, 2024, 6 pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 202080076627.9, mailed May 29, 2025, 16 pages. [cited by applicant]
Office Action for Chinese Application No. 202080076627.9, mailed Feb. 28, 2025, 16 pages. [cited by applicant]