IP Library › Granted Patent US 12,387,082
Granted Patent B2
US 12,387,082 · App. 16/051,034 · Granted Aug 12, 2025

Scheduler for mapping neural networks onto an array of neural cores in an inference processing unit

Inventors: Pallab Datta (San Jose, CA); Andrew S. Cassidy (San Jose, CA); Myron D. Flickner (San Jose, CA); Hartmut Penner (San Jose, CA); Rathinakumar Appuswamy (San Jose, CA); Jun Sawada (Austin, TX); John V. Arthur (Mountain View, CA); Dharmendra S. Modha (San Jose, CA); Steven K. Esser (San Jose, CA); Brian Taba (Cupertino, CA); Jennifer Klamo (San Jose, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/04G06F9/30076G06F9/4881G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,082
App. No.
16/051,034
Granted
Aug 12, 2025
Kind
B2
Abstract

Mapping of neural network layers to physical neural cores is provided. In various embodiments, a neural network description describing a plurality of neural network layers is read. Each of the plurality of neural network layers has an associated weight tensor, input tensor, and output tensor. A plurality of precedence relationships among the plurality of neural network layers is determined. The weight tensor, input tensor, and output tensor of each of the plurality of neural network layers are mapped onto an array of neural cores.

Claims (51)

1. A method comprising:

reading a neural network description describing a plurality of neural network layers, each of the plurality of neural network layers having an associated parameter tensor, weight tensor, input tensor, and output tensor;

determining a plurality of precedence relationships among the plurality of neural network layers and a plurality of logical cores, wherein each logical core corresponds to a neural network layer, wherein the precedence relationships are based on connections among the plurality of neural network layers;

based on the plurality of precedence relationships, generating a sequence of the plurality of neural network layers and a sequence of the plurality of logical cores;

identifying a set of identical logical cores from the plurality of logical cores, wherein the identical logical cores correspond to identical neural network layers, wherein the identical neural network layers are associated with identical parameter tensors and identical weight tensors; and

mapping the plurality of logical cores onto an array of neural cores such that the identical logical cores are mapped onto a same neural core, wherein the mapping of the plurality of logical cores on the array of neural cores comprises:

determining a plurality of sequence numbers, each sequence number of the plurality of sequence numbers corresponding to one neural core of the array of neural cores,

mapping the associated parameter tensor, weight tensor, input tensor, and output tensor of each of the plurality of neural network layers onto the array of neural cores, and

determining a schedule for delivery of the weight tensor, the input tensor, and the output tensor of each of the plurality of neural network layers in accordance with the plurality of sequence numbers, the data delivery being at each of the plurality of neural cores for computation of each neural network layer of the plurality of neural network layers.

2. The method of claim 1 , wherein the mapping of the weight tensor, the input tensor, and the output tensor comprises computation and communication operations.

3. The method of claim 2 , further comprising:

generating microcode, the microcode executable by a chip to compute the plurality of neural network layers, the chip comprising the array of neural cores.

4. The method of claim 3 , further comprising:

generating core microcode, the core microcode executable by the array of neural cores to generate partial sums.

5. The method of claim 3 , further comprising:

generating chip microcode, the chip microcode executable by at least one chip microengine to distribute weights, parameters, instructions, and/or activation data to the array of neural cores.

6. The method of claim 1 , wherein the mapping of the weight tensor, input tensor, and output tensor comprises determining a memory allocation for the weight tensor, input tensor, and output tensor of each of the plurality of neural network layers.

7. The method of claim 6 , wherein the mapping of the weight tensor, input tensor, and output tensor comprises: determining a memory allocation for plurality of partial sums of each of the plurality of neural network layers.

8. The method of claim 6 , wherein the mapping of the weight tensor, input tensor, and output tensor comprises: determining a memory allocation for the weight tensor in a global memory.

9. The method of claim 8 , wherein the mapping of the weight tensor, input tensor, and output tensor comprises: determining a plurality of mapping functions of the weight tensor to local memories of the neural cores.

10. The method of claim 6 , wherein the mapping of the weight tensor, input tensor, and output tensor comprises: determining a memory allocation for the input tensor and output tensor in local memories of the neural cores.

11. The method of claim 10 , wherein the memory allocation minimizes a transfer of activations among cores.

12. The method of claim 10 , wherein determining the memory allocation comprises sequencing the plurality of neural network layers to conform to a local memory limit.

13. The method of claim 10 , wherein the schedule for the data delivery is determined such that the memory allocation is applied when the execution requires the input tensor and output tensor.

14. The method of claim 13 , wherein the local memories are reallocated repeatedly according to the schedule for the data delivery.

15. The method of claim 1 , wherein each layer of the neural network has an associated tensor operation, the method further comprising:

selecting a parametrized scheme from a library of parametrized schemes, the selected scheme corresponding to the tensor operation, wherein:

the mapping comprises instantiating the selected scheme with parameters corresponding to the tensor operation.

16. The method of claim 15 , wherein the mapping comprises determining an execution schedule comprising a plurality of operations for the array of neural cores, the execution schedule guaranteeing data delivery at each of the plurality of neural cores for the computation of each neural network layer.

17. The method of claim 16 , wherein the execution schedule comprises computation and communication operations.

18. The method of claim 17 , wherein the execution schedule further comprises no operations (NOPs).

19. The method of claim 1 , wherein the mapping comprises determining a network schedule, wherein the network schedule determines a timing of communications on one or more networks.

20. The method of claim 19 , wherein the one or more networks comprises a weight network, an instruction network, an activation network, or a partial sum network.

21. The method of claim 20 , wherein the network schedule precludes simultaneous transmission on each network.

22. The method of claim 19 , wherein the network schedule guarantees data delivery at each of the plurality of neural cores for each computation using said data.

23. The method of claim 1 , wherein the mapping comprises determining a batch size.

24. The method of claim 23 , wherein the batch size is unique for at least one neural network layer.

25. A system comprising:

a chip comprising a controller and an array of neural cores;

a computing node comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor of the computing node to cause the processor to perform a method comprising:

reading a neural network description describing a plurality of neural network layers, each of the plurality of neural network layers having an associated parameter tensor, weight tensor, input tensor, and output tensor;

determining a plurality of precedence relationships among the plurality of neural network layers and a plurality of logical cores, wherein each logical core corresponds to a neural network layer, wherein the precedence relationships are based on connections among the plurality of neural network layers;

based on the plurality of precedence relationships, generating a sequence of the plurality of neural network layers and a sequence of the plurality of logical cores;

identifying a set of identical logical cores from the plurality of logical cores, wherein the identical logical cores correspond to identical layers, wherein the identical layers are associated with identical parameter tensors and identical weight tensors;

mapping the plurality of logical cores onto an array of neural cores such that the identical logical cores are mapped onto a same neural core, wherein the mapping of the plurality of logical cores on the array of neural cores comprises:

determining a plurality of sequence numbers, each sequence number of the plurality of sequence numbers corresponding to one neural core of the array of neural cores,

mapping the associated parameter tensor, weight tensor, input tensor, and output tensor of each of the plurality of neural network layers onto the array of neural cores, and

determining a schedule for delivery of the weight tensor, the input tensor, and the output tensor of each of the plurality of neural network layers in accordance with the plurality of sequence numbers, the data delivery being at each of the plurality of neural cores for computation of each neural network layer of the plurality of neural network layers;

generating microcode; and

distributing the microcode to the chip, wherein:

the chip is adapted to execute the microcode to compute the plurality of neural network layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2018
From: DATTA, PALLAB; CASSIDY, ANDREW S.; FLICKNER, MYRON D.; PENNER, HARTMUT; APPUSWAMY, RATHINAKUMAR; SAWADA, JUN; ARTHUR, JOHN V.; MODHA, DHARMENDRA S.; ESSER, STEVEN K.; TABA, BRIAN; KLAMO, JENNIFER
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046517/0008 →
Continuity (1)
Related Publication 20200042856A1 · Feb 6, 2020
References Cited (62)
US 5714768A · Ovshinsky et al. · 1998 [cited by applicant]
US 6389404B1 · Carson et al. · 2002 [cited by applicant]
US 9015096B2 · Hunzinger · 2015 [cited by applicant]
US 9165243B2 · Yu et al. · 2015 [cited by applicant]
US 9245222B2 · Modha · 2016 [cited by applicant]
US 9424284B2 · Alvarez-Icaza Rivera et al. · 2016 [cited by applicant]
US 9852006B2 · Akopyan et al. · 2017 [cited by applicant]
US 20050149936A1 · Pilkington · 2005 [cited by applicant]
US 20080208372A1 · Pannese · 2008 [cited by applicant]
US 20080235700A1 · Iguchi · 2008 [cited by applicant]
US 20100268912A1 · Conte et al. · 2010 [cited by applicant]
US 20120053948A1 · Mustiere · 2012 [cited by examiner]
US 20130073497A1 · Akopyan et al. · 2013 [cited by applicant]
US 20150106314A1 · Birdwell et al. · 2015 [cited by applicant]
US 20150324684A1 · Alvarez-Icaza Rivera et al. · 2015 [cited by applicant]
US 20150324690A1 · Chilimbi · 2015 [cited by examiner]
US 20160098629A1 · Lipasti et al. · 2016 [cited by applicant]
US 20160239074A1 · Lee et al. · 2016 [cited by applicant]
US 20160247062A1 · Amir et al. · 2016 [cited by applicant]
US 20170103311A1 · Henry et al. · 2017 [cited by applicant]
US 20170169326A1 · Diamos et al. · 2017 [cited by applicant]
US 20170236053A1 · Lavigueur et al. · 2017 [cited by applicant]
US 20170316312A1 · Goyal · 2017 [cited by examiner]
US 20170323198A1 · Tan et al. · 2017 [cited by applicant]
US 20170344882A1 · Ambrose et al. · 2017 [cited by applicant]
US 20170357935A1 · Fabjanski et al. · 2017 [cited by applicant]
US 20180107766A1 · Dumitrescu et al. · 2018 [cited by applicant]
US 20180174041A1 · Imam et al. · 2018 [cited by applicant]
US 20180189648A1 · Sengupta et al. · 2018 [cited by applicant]
US 20180197075A1 · Modha · 2018 [cited by examiner]
US 20180285727A1 · Baum · 2018 [cited by examiner]
US 20180307980A1 · Barik · 2018 [cited by examiner]
US 20190235866A1 · Das Sarma · 2019 [cited by examiner]
US 20190325295A1 · Modha et al. · 2019 [cited by applicant]
US 20200026992A1 · Zhang · 2020 [cited by examiner]
US 20200201697A1 · Torng · 2020 [cited by examiner]
JP 2010092483A · 2010 [cited by applicant]
JP 2015517712A · 2015 [cited by applicant]
JP 2016001417A · 2016 [cited by applicant]
WO 1993014459A1 · 1993 [cited by applicant]
US 9,846,837 B1, 12/2017, Narayanaswami et al. (withdrawn) [cited by applicant]
Daulani, Hitesh P., Radhakrishna Naik, and Pavan S. Wankhade. “Precedence Constraint Task Scheduling for Multicore Multikernel Architecture.” IOSR Journal of Computer Engineering (IOSR-JCE) 16.4 (2014): 43-53. (Year: 20… [cited by examiner]
Sawada, Jun, et al. “Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications.” SC'16: Proceedings of the International Conference for High Performance Computing, Networking, Storag… [cited by examiner]
Campos, Pedro, et al. “XL-STaGe: A cross-layer scalable tool for graph generation, evaluation and implementation.” 2016 International Conference on Embedded Computer Systems: Architectures, Modeling and Simulation (SAMO… [cited by examiner]
Liu, Shaoli, et al. “Cambricon: An instruction set architecture for neural networks.” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA). IEEE, 2016. (Year: 2016). [cited by examiner]
Peemen, Maurice, et al. “Memory-centric accelerator design for convolutional neural networks.” 2013 IEEE 31st International Conference on Computer Design (ICCD). IEEE, 2013. (Year: 2013). [cited by examiner]
Xin, Chen, et al. “COSY: An energy-efficient hardware architecture for deep convolutional neural networks based on systolic array.” 2017 IEEE 23rd International Conference on Parallel and Distributed Systems (ICPADS). I… [cited by examiner]
Park, Jongse, et al. “Scale-out acceleration for machine learning.” Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture. 2017. (Year: 2017). [cited by examiner]
Santara, Anirban, et al. “Faster learning of deep stacked autoencoders on multi-core systems using synchronized layer-wise pre-training.” arXiv preprint arXiv:1603.02836 (2016). (Year: 2016). [cited by examiner]
Chillet et al.; “Real-Time Scheduling on Heterogeneous System-On-Chip ArchitecturesUsing an Optimized Artificial Neural Network”, Journal of Systems Architecture 57 (2011) 340-353. [cited by applicant]
Ghosh et al.; “Mapping Neural Networks Onto Message-Passing Multicomputers”, Journal of Parallel and Distributed Computing 6,29 1-330 (1989). [cited by applicant]
Bengio et al.; “Scheduled Sampling for Sequence Prediction With Recurrent Neural Networks”, (Sep. 2015) Google Research. [cited by applicant]
Hasan et al.; “High Throughput Neural Network Based Embedded Streaming Multicore Processors.”. [cited by applicant]
Akopyan et al., “TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip,” IEEE Transactions on Computer-Aided Design of Integrated Circuit and Systems, 34(10): 1537-1557 (2015). [cited by applicant]
Amir et al: “Cognitive computing programming paradigm: A Corelet Languagefor composing networks of neurosynaptic cores,” The 2013 International Joint Conference on Neural Networks (IJCNN), IEEE, Aug. 4, 2013. [cited by applicant]
Braga et al., “VANNGen: A Flexible CAD Tool for Hardware Implementation of Artificial Neural Networks,” Proc of the 2005 Intl Conf on Reconfigurable Computing and FPGAs (ReConFig 2005), IEEE Computer Society, 2005. [cited by applicant]
EPO Oral Proceedings Written Submissions Letter filed for EP Application No. 17825508.9 submitted Feb. 1, 2022. [cited by applicant]
Filipp et al: “TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip”, IEEE Transactions on Computer Aided Design of Integrated Circuits and Systems, IEEE Service Center, Piscataway… [cited by applicant]
International Search Report and Written Opinion for PCT/EP2017/083881 mailed Mar. 28, 2018. [cited by applicant]
JP Notice of Reasons for Refusal for JP Application No. 2019-529248 dated Jul. 26, 2021. [cited by applicant]
Preissl et al: “Compass: A scalable simulator for an architecture for cognitive computing”, High Performance Computing, Networking, Storage and Analysis (SC), 2012 International Conference for, IEEE, Nov. 10, 2012. [cited by applicant]
Saifullah et al: “Parallel Real-Time Scheduling of DAGs”, IEEE Transactions on Parallel and Distributed Systems, IEEE Service Center, Los Alamitos, CA, US, vol. 25, No. 12, Dec. 1, 2014. [cited by applicant]
Cited By (1)
US 12,511,543