IP Library › Granted Patent US 12,481,861
Granted Patent B2
US 12,481,861 · App. 16/033,926 · Granted Nov 25, 2025

Hierarchical parallelism in a network of distributed neural network cores

Inventors: John V. Arthur (Mountain View, CA); Andrew S. Cassidy (San Jose, CA); Myron D. Flickner (San Jose, CA); Pallab Datta (San Jose, CA); Hartmut Penner (San Jose, CA); Rathinakumar Appuswamy (San Jose, CA); Jun Sawada (Austin, TX); Dharmendra S. Modha (San Jose, CA); Steven K. Esser (San Jose, CA); Brian Taba (Cupertino, CA); Jennifer Klamo (San Jose, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/04G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,861
App. No.
16/033,926
Granted
Nov 25, 2025
Kind
B2
Abstract

Networks of distributed neural cores are provided with hierarchical parallelism. In various embodiments, a plurality of neural cores is provided. Each of the plurality of neural cores comprises a plurality of vector compute units configured to operate in parallel. Each of the plurality of neural cores is configured to compute in parallel output activations by applying its plurality of vector compute units to input activations. Each of the plurality of neural cores is assigned a subset of output activations of a layer of a neural network for computation. Upon receipt of a subset of input activations of the layer of the neural network, each of the plurality of neural cores computes a partial sum for each of its assigned output activations, and computes its assigned output activations from at least the computed partial sums.

Claims (51)

1 . A neural inference chip, comprising:

a plurality of neural cores comprising a first neural core and a second neural core, each neural core of the plurality of neural cores comprising a plurality of vector compute units and a parallel sum and nonlinearity unit such that the first neural core comprises a first plurality of vector compute units and a first parallel sum and nonlinearity unit and the second plurality of vector compute units comprises a second plurality of vector compute units and a second parallel sum and nonlinearity unit, wherein:

each of the plurality of neural cores is configured to compute in parallel output activations by applying its plurality of vector compute units to input activations and applying a nonlinear function by its parallel sum and nonlinearity unit such that the first neural core is configured to compute output activations by applying the first plurality of vector compute units to input activations and applying a nonlinear function by the first parallel sum and nonlinearity unit;

each output activation of a layer of an artificial neural network is assigned to at least one neural core of the plurality of neural cores such that the first neural core is assigned to a subset of output activations of the layer for computation;

the first and second pluralities of vector compute units each comprise multiplication and addition units;

the second neural core of the plurality of neural cores computes zero and non-zero partial sums;

the second neural core directly communicates its non-zero partial sums with the first parallel sum and nonlinearity unit of via a wire without communicating its zero partial sums.

2 . The neural inference chip of claim 1 , wherein:

upon receipt of a subset of input activations of the layer of the artificial neural network, each of the plurality of neural cores

computes a partial sum for each of its assigned output activations, and

computes its assigned output activations from at least the computed partial sums.

3 . The neural inference chip of claim 2 , wherein upon receipt of a subset of input activations of the layer of the artificial neural network, each of the plurality of neural cores

receives partial sums for at least one of its assigned output activations from another of the plurality of neural cores, and

computes its assigned output activations from the computed partial sums and the received partial sums.

4 . The neural inference chip of claim 2 , wherein the plurality of neural cores perform said partial sum computation in parallel.

5 . The neural inference chip of claim 2 , wherein the plurality of neural cores perform said output activation computation in parallel.

6 . The neural inference chip of claim 2 , wherein computing the partial sum comprises applying at least one of the plurality of vector compute units to multiply the input activations and synaptic weights.

7 . The neural inference chip of claim 2 , wherein computing the assigned output activations comprises applying a plurality of addition units.

8 . The neural inference chip of claim 2 , wherein the plurality of vector compute units are configured to:

perform a plurality of multiply operations in parallel;

perform a plurality of additions in parallel; and

accumulate the partial sum.

9 . The neural inference chip of claim 2 , wherein the plurality of vector compute units are configured to compute partial sums in parallel.

10 . The neural inference chip of claim 1 , wherein the plurality of vector compute units comprise accumulation units.

11 . The neural inference chip of claim 1 , wherein said computation of the subset of output activations of a layer of the artificial neural network by each of the plurality of neural cores is pipelined.

12 . The neural inference chip of claim 11 , wherein each of the plurality of neural cores is configured to concurrently perform each stage of a plurality of stages of said computation of the subset of output activations of a layer of the artificial neural network.

13 . The neural inference chip of claim 12 , wherein said computation of the subset of output activations of a layer of the artificial neural network maintains parallelism.

14 . The neural inference chip of claim 1 , wherein each neural core of the plurality of neural cores computes its zero and non-zero partial sums simultaneously and separately computes all of its assigned subset of output activations simultaneously.

15 . A method comprising:

receiving, at each neural core of a plurality of neural cores, a subset of input activations of a layer of an artificial neural network such that a first subset of input activations of a first layer is received at a first neural core of the plurality, wherein:

each neural core of the plurality of neural cores comprises a plurality of vector compute units and a parallel sum and nonlinearity unit such that the first neural core comprises a first plurality of vector compute units and a first parallel sum and nonlinearity unit and the second plurality of vector compute units comprises a second plurality of vector compute units and a second parallel sum and nonlinearity unit, and

each of the plurality of neural cores is configured to compute in parallel output activations by applying its plurality of vector multipliers to its subset of input activations and applying a nonlinear function by its parallel sum and nonlinearity unit such that the first neural core is configured to compute output activations by applying the first plurality of vector compute units to the first subset of input activations of the first layer and applying a nonlinear function by the first parallel and nonlinearity unit;

assigning each output activation of the first layer to a neural core of the plurality of neural cores such that the first neural core is assigned to a first subset of output activations of the first layer for computation; and

upon receipt of the first subset of input activations of the layer of the artificial neural network, the first neural core

computing a partial sum for each output activation of the first subset of output activations, and

computing the first subset of output activations from at least the computed partial sums, wherein

 the first and second plurality of vector compute units each comprise multiplication and addition units,

 the second neural core of the plurality of neural cores computes zero and non-zero partial sums, and

 the second neural core directly communicates its non-zero partial sums with the first parallel sum and nonlinearity unit of via a wire without communicating its zero partial sums.

16 . The method of claim 15 , wherein upon receipt of a subset of input activations of the layer of the artificial neural network, each of the plurality of neural cores

receives non-zero partial sums for at least one of its assigned output activations from another of the plurality of neural cores, and

computes its assigned output activations from the computed partial sums and the received non-zero partial sums.

17 . The method of claim 15 , wherein the plurality of vector compute units comprise accumulation units.

18 . The method of claim 15 , wherein the plurality of neural cores perform said partial sum computation in parallel.

19 . The method of claim 15 , wherein the plurality of neural cores perform said output activation computation in parallel.

20 . The method of claim 15 , wherein computing the partial sum comprises applying at least one of the plurality of vector compute units to multiply the input activations and synaptic weights.

21 . The method of claim 15 , wherein computing the assigned output activations comprises applying a plurality of addition units.

22 . The method of claim 15 , further comprising:

performing a plurality of multiply operations in parallel;

performing a plurality of additions in parallel; and

accumulating the partial sum.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2018
From: ARTHUR, JOHN V.; CASSIDY, ANDREW S.; FLICKNER, MYRON D.; DATTA, PALLAB; PENNER, HARTMUT; APPUSWAMY, RATHINAKUMAR; SAWADA, JUN; MODHA, DHARMENDRA S.; ESSER, STEVEN K.; TABA, BRIAN; KLAMO, JENNIFER
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046572/0364 →
Continuity (1)
Related Publication 20200019836A1 · Jan 16, 2020
References Cited (30)
US 6654730B1 · Kato · 2003 [cited by examiner]
US 9710265B1 · Temam · 2017 [cited by examiner]
US 10049322B2 · Ross · 2018 [cited by applicant]
US 20160098629A1 · Lipasti et al. · 2016 [cited by applicant]
US 20170154259A1 · Burr et al. · 2017 [cited by applicant]
US 20180046906A1 · Dally · 2018 [cited by examiner]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180121377A1 · Woo · 2018 [cited by examiner]
US 20180121786A1 · Narayanaswami · 2018 [cited by examiner]
US 20180189651A1 · Henry · 2018 [cited by examiner]
CN 107454966A · 2017 [cited by applicant]
CN 107622303A · 2018 [cited by applicant]
CN 112384935A · 2021 [cited by applicant]
EP 3821376A1 · 2021 [cited by applicant]
JP H04021153A · 1992 [cited by applicant]
JP H05346914A · 1993 [cited by applicant]
JP 2001188767A · 2001 [cited by applicant]
JP 2021532451A · 2021 [cited by applicant]
WO 2016174113A1 · 2016 [cited by applicant]
WO 2016186810A1 · 2016 [cited by applicant]
WO 2020011936A1 · 2020 [cited by applicant]
Li, W. et al.; “Parallel Multiclass Support Vector Machine for Remote Sensing Data Classification on Multicore and Many-Core Architectures”; IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensi… [cited by applicant]
Hegde, V. et al.; “Parallel and Distributed Deep Learning”; Stanford University, [email protected] <mailto:[email protected]>. 2016. [cited by applicant]
“An FPGA Architecture for Accelerating Convolutional Neural Network in Speech Recognition”; <http://ip.com/IPCOM/000247151D>; Aug. 11, 2016. [cited by applicant]
“Reconfigurable Communication Network with Distributed Control”; <http://ip.com/IPCOM/000240011D>; Dec. 22, 2014. [cited by applicant]
“Egps: Exploratory Graph Processing System”;< http://ip.com/IPCOM/000239478D>; Nov. 11, 2014. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/EP2019/068713 dated Oct. 21, 2019. [cited by applicant]
Notice of Reasons for Refusal (English Translation), JP Patent Application No. 2021-500263, issued Dec. 20, 2022, 4 pages. [cited by applicant]
Notice of Reasons for Refusal (English Translation), JP Patent Application No. 2021-500263, issued July, 5, 2023. [cited by applicant]
Office Action for EP Application No. 19748724.2 dated Feb. 14, 2024. [cited by applicant]