IP Library Granted Patent US 12,254,399
Granted Patent B2
US 12,254,399 · App. 17/159,312 · Granted Mar 18, 2025

Hierarchical hybrid network on chip architecture for compute-in-memory probabilistic machine learning accelerator

Inventors: Deepak Dasalukunte (Beaverton, OR); Richard Dorrance (Hillsboro, OR); Hechen Wang (Hillsboro, OR)
Assignee: Intel Corporation
G06N3/065G06N3/047H04L45/586H04L49/109
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,399
App. No.
17/159,312
Granted
Mar 18, 2025
Kind
B2
Abstract

Systems, methods, apparatuses, and computer-readable media. An analog router of a first supertile of a plurality of supertiles of a network on a chip (NoC) may receive a first analog output from a first compute-in-memory tile of a plurality of compute-in-memory tiles of the first supertile. The analog router may determine, based on a configuration of a neural network executing on the NoC, that a destination of the first analog output includes a second supertile of the plurality of supertiles. An analog-to-digital converter (ADC) of the analog router may convert the first analog output to a first digital output and transmit the first digital output to the second supertile via a communications bus of the NoC.

Claims (63)

1. An apparatus for a network on a chip (NoC), the apparatus comprising:

a plurality of supertiles, each supertile including an analog router and a plurality of compute-in-memory tiles; and

logic circuitry, at least a portion of which is implemented in a first analog router of a first supertile of the plurality of supertiles, the logic circuitry to:

receive a first analog output from a first compute-in-memory tile of a first plurality of compute-in-memory tiles of the first supertile, the first analog output representing a weight of a neural network to be executed on the NoC;

determine, based on a configuration of the neural network, that a destination of the first analog output identifies a second supertile of the plurality of supertiles;

generate, by a Gaussian distribution circuit of the analog router, a Gaussian distribution for the weight:

convert, by an analog-to-digital converter (ADC) of the analog router, the first analog output to a first digital output; and

transmit, by the analog router, the first digital output including the Gaussian distribution to the second supertile via at least one digital router of a communications bus of the NoC, the communications bus including a plurality of digital routers.

2. The apparatus of claim 1 , wherein the logic circuitry is to perform, by the Gaussian distribution circuit, a stochastic rounding operation on the weight.

3. The apparatus of claim 1 , wherein the destination is a first destination, and the logic circuitry is to:

receive, by the analog router, a second analog output of the first compute-in-memory tile;

determine, by the analog router, that a second destination of the second analog output identifies a second compute-in-memory tile of the first plurality of compute-in-memory tiles of the first supertile; and

transmit, by the analog router, the second analog output to the second compute-in-memory tile.

4. The apparatus of claim 3 , wherein the weight is a first weight, the second analog output represents a second weight of the neural network, the Gaussian distribution is a first Gaussian distribution, and the logic circuitry is to:

generate, by the Gaussian distribution circuit of the analog router, a second Gaussian distribution for the second weight; and

transmit, by the analog router to the second compute-in-memory tile, the second Gaussian distribution and the second analog output.

5. An apparatus for a network on a chip (NoC), the apparatus comprising:

a plurality of supertiles, each supertile including an analog router and a plurality of compute-in-memory tiles; and

logic circuitry, at least a portion of which is implemented in a first analog router of a first supertile of the plurality of supertiles, the logic circuitry to:

receive a first analog output from a first compute-in-memory tile of a first plurality of compute-in-memory tiles of the first supertile;

determine, based on a configuration of a neural network to be executed on the NoC, that a destination of the first analog output identifies a second supertile of the plurality of supertiles;

convert, by an analog-to-digital converter (ADC) of the analog router, the first analog output to a first digital output;

transmit, by the analog router, the first digital output to the second supertile via a communications bus of the NoC;

receive, by the analog router via a digital router of the communications bus, a digital packet from the second supertile;

convert, by a digital-to-analog converter (DAC) of the analog router, the digital packet to an analog signal; and

provide, by the analog router, the analog signal to one of the first plurality of compute-in-memory tiles of the first supertile.

6. The apparatus of claim 1 , wherein each of the plurality of compute-in-memory tiles includes processor circuitry and a memory, and the logic circuitry is to packetize the first digital output into one or more packets prior to transmission of the first digital output to the second supertile.

7. A non-transitory computer-readable storage medium comprising instructions that cause a network on a chip (NoC) to:

receive, by an analog router of a first supertile of a plurality of supertiles of the NoC, a first analog output from a first compute-in-memory tile of a plurality of compute-in-memory tiles of the first supertile, the first analog output including a weight of a neural network to be executed on the NoC;

determine, by the analog router and based on a configuration of the neural network, that a destination of the first analog output identifies a second supertile of the plurality of supertiles;

convert, by an analog-to-digital converter (ADC) of the analog router, the first analog output to a first digital output;

transmit, by the analog router, the first digital output to the second supertile via a communications bus of the NoC;

receive, by the analog router via a digital router of the communications bus, a digital packet from the second supertile;

convert, by a digital-to-analog converter (DAC) of the analog router, the digital packet to an analog signal; and

provide, by the analog router, the analog signal to one of the plurality of compute-in- memory tiles of the first supertile.

8. The non-transitory computer-readable storage medium of claim 7 , wherein the instructions cause the NoC to:

generate, by a Gaussian distribution circuit of the analog router, a Gaussian distribution for the weight; and

transmit, by the analog router, the first digital output to the second supertile via at least one digital router of a plurality of digital routers of the communications bus, the first digital output including the Gaussian distribution.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the instructions cause the NoC to perform, by the Gaussian distribution circuit, a stochastic rounding operation on the weight.

10. The non-transitory computer-readable storage medium of claim 7 , wherein the destination is a first destination, and the instructions cause the NoC to:

receive, by the analog router, a second analog output of the first compute-in-memory tile;

determine, by the analog router, that a second destination of the second analog output identifies a second compute-in-memory tile of the plurality of compute-in-memory tiles of the first supertile; and

transmit, by the analog router, the second analog output to the second compute-in-memory tile.

11. The non-transitory computer-readable storage medium of claim 10 , wherein the instructions cause the NoC to:

generate, by a Gaussian distribution circuit of the analog router, a Gaussian distribution for the weight; and

transmit, by the analog router to the second compute-in-memory tile, the Gaussian distribution and the second analog output via at least one digital router of a plurality of digital routers of the NoC.

12. The non-transitory computer-readable storage medium of claim 7 , wherein each of the plurality of compute-in-memory tiles includes processor circuitry and a memory, and the instructions cause the NoC to packetize the first digital output into one or more packets prior to transmission of the first digital output to the second supertile.

13. An apparatus comprising:

means for routing to:

receive a first analog output from a first compute-in-memory tile of a plurality of compute-in-memory tiles of a first supertile of a plurality of supertiles of a network on a chip (NoC), the first analog output to represent a weight of a neural network to be executed on the NoC;

determine, based on a configuration of the neural network, that a destination of the first analog output identifies a second supertile of the plurality of supertiles;

transmit a first digital output to the second supertile via a communications bus of the NoC:

receive, from a digital router of the communications bus, a digital packet from the second supertile; and

provide an analog signal to one of the plurality of compute-in-memory tiles of the first supertile;

means for converting the first analog output to the first digital output; and

means for converting the digital packet to the analog signal.

14. The apparatus of claim 13 , wherein the apparatus further includes including means for generating a Gaussian distribution for the weight, and the means for routing is to transmit the first digital output to the second supertile via at least one digital router of a plurality of digital routers of the communications bus, the first digital output including the Gaussian distribution.

15. The apparatus of claim 14 , further including means for performing a stochastic rounding operation on the weight.

16. The apparatus of claim 13 , wherein the destination is a first destination, and the means for routing is to:

receive a second analog output of the first compute-in-memory tile;

determine that a second destination of the second analog output identifies a second compute-in-memory tile of the plurality of compute-in-memory tiles of the first supertile; and

transmit the second analog output to the second compute-in-memory tile.

17. The apparatus of claim 16 , wherein the apparatus further includes means for generating a Gaussian distribution for the weight, and the means for routing is to transmit the Gaussian distribution and the second analog output to the second compute-in-memory tile.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2021
From: DASALUKUNTE, DEEPAK; DORRANCE, RICHARD; WANG, HECHEN
To: INTEL CORPORATION
Reel/Frame 055127/0816 →
Continuity (1)
Related Publication 20210150328A1 · May 20, 2021
References Cited (20)
US 11100193B2 · Gu · 2021 [cited by examiner]
US 11847560B2 · Kolter · 2023 [cited by examiner]
US 20180095930A1 · Lu · 2018 [cited by examiner]
Carrillo et al., “Scalable Hierarchical Network-on-Chip Architecture for Spiking Neural Network Hardware Implementation”, IEEE Transactions on Parallel and Distributed Systems, vol. 24, No. 12, Dec. 2013 (Year: 2013). [cited by examiner]
A. D. Patil, H. Hua, S. Gonugondla, M. Kang and N. R. Shanbhag, “An MRAM-Based Deep In-Memory Architecture for Deep Neural Networks,” 2019 IEEE International Symposium on Circuits and Systems (ISCAS), 2019, pp. 1-5, doi… [cited by applicant]
Sebastian et al., “Tutorial: Brain-inspired computing using phase-change memory devices” Journal of Applied Physics 124, 111101 (2018); https://doi.org/10.1063/1.5042413. [cited by applicant]
D. Miyashita, S. Kousai, T. Suzuki and J. Deguchi, “Time-domain neural network: A 48.5 TSOp/s/W neuromorphic chip optimized for deep learning and CMOS technology,” 2016 IEEE Asian Solid-State Circuits Conference (A-SSCC… [cited by applicant]
S. K. Gonugondla, M. Kang and N. Shanbhag, “A 42pJ/decision 3.12TOPS/W robust in-memory machine learning classifier with on-chip training,” 2018 IEEE International Solid-State Circuits Conference—(ISSCC), 2018, pp. 490-… [cited by applicant]
J. Song, K. Ragab, X. Tang and N. Sun, “A 10-b 800-MS/s Time-Interleaved SAR ADC With Fast Variance-Based Timing-Skew Calibration,” in IEEE Journal of Solid-State Circuits, vol. 52, No. 10, pp. 2563-2575, Oct. 2017, doi… [cited by applicant]
H. Valavi, P. J. Ramadge, E. Nestler and N. Verma, “A 64-Tile 2.4-Mb In-Memory-Computing CNN Accelerator Employing Charge-Domain Compute,” in IEEE Journal of Solid-State Circuits, vol. 54, No. 6, pp. 1789-1799, Jun. 201… [cited by applicant]
Hopkins Michael, Mikaitis Mantas, Lester Dave R. and Furber Steve 2020Stochastic rounding and reduced-precision fixed-point arithmetic for solving neural ordinary differential equationsPhil. Trans. R. Soc. A.37820190052… [cited by applicant]
Gupta et al., “Deep Learning with Limited Numerical Precision.” In Proceedings of the 32nd International Conference on Machine Learning—vol. 37 (ICML'15). JMLR.org, 1737-1746. [cited by applicant]
Author Unknown, “Network on a chip” Wikipedia—accessed Aug. 10, 2020. [cited by applicant]
Wang, J. et al. “A 28-nm Compute SRAM With Bit-Serial Logic/Arithmetic Operations for Programmable In-Memory Vector Computing.” IEEE Journal of Solid-State Circuits 55 (2020): 76-86. [cited by applicant]
X. Liu et al., “RENO: A high-efficient reconfigurable neuromorphic computing accelerator design,” 2015 52nd ACM/EDAC/IEEE Design Automation Conference (DAC), 2015, pp. 1-6, doi: 10.1145/2744769.2744900. [cited by applicant]
Kang, Mingu et al. “A Multi-Functional In-Memory Inference Processor Using a Standard 6T SRAM Array.” IEEE Journal of Solid-State Circuits 53 (2018): 642-655. [cited by applicant]
Dong, Qing et al. “15.3 A 351TOPS/W and 372.4GOPS Compute-in-Memory SRAM Macro in 7nm FinFET CMOS for Machine-Learning Applications.” 2020 IEEE International Solid-State Circuits Conference—(ISSCC) (2020): 242-244. [cited by applicant]
“Goya™” Habana—An Intel Company—accessed Aug. 10, 2020. [cited by applicant]
“Mythic” https://www.mythic-ai.com/technology/—accessed Aug. 10, 2020. [cited by applicant]
“MNIST database”—Wikipedia—accessed Aug. 10, 2020. [cited by applicant]