IP Library › Granted Patent US 12,314,837
Granted Patent B2
US 12,314,837 · App. 18/339,954 · Granted May 27, 2025

Multi-memory on-chip computational network

Inventors: Randy Huang (Morgan Hill, CA); Ron Diamant (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G06N3/045G06F3/061G06F3/065G06F3/0683G06F13/28G06F13/4068G06F15/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,837
App. No.
18/339,954
Granted
May 27, 2025
Kind
B2
Abstract

Provided are systems, methods, and integrated circuits for neural network processing. In various implementations, an integrated circuit for neural network processing can include a plurality of memory banks storing weight values for a neural network. The memory banks can be on the same chip as an array of processing engines. Upon receiving input data, the circuit can be configured to use the set of weight values to perform a task defined for the neural network. Performing the task can include reading weight values from the memory banks, inputting the weight values into the array of processing engines, and computing a result using the array of processing engines, where the result corresponds to an outcome of performing the task.

Claims (44)

1. An integrated circuit comprising:

a first array of processing engines; and

a plurality of memory banks for storing a set of weight values for a neural network, wherein each memory bank from the plurality of memory banks is independently accessible, wherein the plurality of memory banks and the first array of processing engines are on a same die, and wherein the plurality of memory banks supports simultaneous accesses including reading a first value from a first memory bank, while a second value is being read from or written to a second memory bank;

wherein, upon receiving input data, the integrated circuit is configured to use the set of weight values to perform a task defined for the neural network, and wherein performing the task includes:

reading weight values from the plurality of memory banks;

inputting the weight values and the input data into the first array of processing engines; and

computing a result using the first array of processing engines, wherein the result corresponds to an outcome of performing the task.

2. The integrated circuit of claim 1 , wherein each of the first and second values includes one of a weight value, an input value, or an intermediate result.

3. The integrated circuit of claim 1 , further comprising:

a second array of processing engines, wherein a first set of memory banks from the plurality of memory banks is initially configured for use by the first array of processing engines, wherein a second set of memory banks from the plurality of memory banks is initially configured for use by the second array of processing engines, and wherein the first set of memory banks and the second set of memory banks each include a portion of the set of weight values.

4. The integrated circuit of claim 3 , wherein performing the task further includes:

computing, by the first array of processing engines, an intermediate result, wherein the first array of processing engines computes the intermediate result using weight values from the first set of memory banks; and

reading, by the first array of processing engines, additional weight values from the second set of memory banks, wherein the first array of processing engines uses the intermediate result and the additional weight values to compute the result.

5. The integrated circuit of claim 3 , wherein the set of weight values occupies less than all of the second set of memory banks, and wherein the second array of processing engines performs computations using a part of the second set of memory banks that is not occupied by the set of weight values.

6. The integrated circuit of claim 1 , wherein performing the task further includes:

determining that an amount of memory needed for storing an intermediate result has decreased;

reading an additional set of weight values from another memory; and

storing the additional set of weight values in a portion of the plurality of memory banks that is no longer needed for the intermediate result.

7. The integrated circuit of claim 1 , wherein the plurality of memory banks includes at least as many memory banks as a number of rows in the first array of processing engines.

8. The integrated circuit of claim 1 , wherein the plurality of memory banks includes at least as many memory banks as a number of columns in the first array of processing engines.

9. The integrated circuit of claim 1 , further comprising a pooling circuit to perform a pooling function on an output of the first array of processing engines to generate an intermediate result before writing the intermediate result to a portion of the plurality of memory banks.

10. The integrated circuit of claim 1 , further comprising an activation circuit to perform an activation function on an output of the first array of processing engines to generate an intermediate result before writing the intermediate result to a portion of the plurality of memory banks.

11. A method comprising:

storing a set of weight values in a plurality of memory banks of a neural network processing circuit, wherein the neural network processing circuit includes an array of processing engines on a same die as the plurality of memory banks, wherein the plurality of memory banks supports simultaneous accesses including reading a first value from a first memory bank, while a second value is being read from or written to a second memory bank, and wherein the set of weight values is stored prior to receiving input data;

receiving input data; and

using the set of weight values to perform a task defined for a neural network, wherein performing the task includes:

reading weight values from the plurality of memory banks;

inputting the weight values and the input data into the array of processing engines; and

computing a result using the array of processing engines, wherein the result corresponds to an outcome of performing the task.

12. The method of claim 11 , wherein each of the first and second values includes one of a weight value, an input value, or an intermediate result.

13. The method of claim 11 , wherein performing the task further includes:

computing, by the array of processing engines, an intermediate result, wherein the array of processing engines computes the intermediate result using weight values from a first set of memory banks in the plurality of memory banks.

14. The method of claim 13 , wherein performing the task further includes:

writing the intermediate result to a first portion of the plurality of memory banks, and moving the intermediate result to a second portion of the plurality of memory banks from which the intermediate result is read.

15. The method of claim 13 , wherein performing the task further includes:

reading, by the array of processing engines, additional weight values from a second set of memory banks in the plurality of memory banks, wherein the array of processing engines uses the intermediate result and the additional weight values to compute the result.

16. The method of claim 15 , wherein a second array of processing engines performs computations using a part of the second set of memory banks that is not occupied by the set of weight values.

17. The method of claim 11 , wherein performing the task further includes:

determining that an amount of memory needed for storing an intermediate result has decreased;

reading an additional set of weight values from another memory; and

storing the additional set of weight values in a portion of the plurality of memory banks that is no longer needed for the intermediate result.

18. The method of claim 11 , wherein each row of the array of processing engines reads from a respective memory bank of the plurality of memory banks.

19. The method of claim 11 , wherein each column of the array of processing engines writes to a respective memory bank of the plurality of memory banks.

20. The method of claim 11 , further performing an activation function or a pooling function on an output of the array of processing engines to generate an intermediate result before writing the result to a portion of the plurality of memory banks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: HUANG, RANDY; DIAMANT, RON
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064035/0441 →
Continuity (3)
Division 17033573 · Sep 25, 2020
Continuation 15839301 · Dec 12, 2017
Related Publication 20230334294A1 · Oct 19, 2023
References Cited (61)
US 5422983A · Castelaz et al. · 1995 [cited by applicant]
US 9971540B2 · Herrero et al. · 2018 [cited by applicant]
US 10096134B2 · Yan et al. · 2018 [cited by applicant]
US 10332028B2 · Talathi et al. · 2019 [cited by applicant]
US 10504022B2 · Temam et al. · 2019 [cited by applicant]
US 10803379B2 · Huang et al. · 2020 [cited by applicant]
US 10846621B2 · Huang et al. · 2020 [cited by applicant]
US 11741345B2 · Huang et al. · 2023 [cited by applicant]
US 20070094481A1 · Snook et al. · 2007 [cited by applicant]
US 20140310217A1 · Sarah et al. · 2014 [cited by applicant]
US 20150206050A1 · Talathi et al. · 2015 [cited by applicant]
US 20150242741A1 · Campos et al. · 2015 [cited by applicant]
US 20150331832A1 · Minoya et al. · 2015 [cited by applicant]
US 20160071005A1 · Wang et al. · 2016 [cited by applicant]
US 20160179434A1 · Herrero Abellanas et al. · 2016 [cited by applicant]
US 20160342891A1 · Ross et al. · 2016 [cited by applicant]
US 20170061326A1 · Talathi et al. · 2017 [cited by applicant]
US 20170286825A1 · Akopyan et al. · 2017 [cited by applicant]
US 20190180183A1 · Diamant et al. · 2019 [cited by applicant]
US 20190243764A1 · Sakthivel et al. · 2019 [cited by applicant]
CN 106462803A · 2017 [cited by applicant]
CN 107454965A · 2017 [cited by applicant]
CN 107454966A · 2017 [cited by applicant]
JP H04293151A · 1992 [cited by applicant]
JP 2002049405A · 2002 [cited by applicant]
JP 2002117389A · 2002 [cited by applicant]
JP 2013529342A · 2013 [cited by applicant]
JP 2013178294A · 2013 [cited by applicant]
JP 2015215837A · 2015 [cited by applicant]
WO WO2011146147A1 · 2011 [cited by applicant]
WO WO2016186810A1 · 2016 [cited by applicant]
CN Notice of Allowance dated Aug. 29, 2023 in Application No. CN201880080107.8 with English translation. [cited by applicant]
CN Office Action dated Feb. 15, 2023 in Application No. CN201880080107.8 with English translation. [cited by applicant]
Communication under Rules 161(1) and 162 EPC dated Jul. 21, 2020 in Application No. EP 18836949.0. [cited by applicant]
EP Office Action dated Jul. 5, 2022, in Application No. EP18836949.0. [cited by applicant]
Final Office Action dated Apr. 1, 2022 in Application No. JP 2020-531932. [cited by applicant]
Galindo Sanchez, F., “Energy proportional streaming spiking neural network in a reconfigurable system,” Accepted Author Manuscript (AAM), University of Bristol, 2017, 20 pages. URL: https://research-information.bris.ac.… [cited by applicant]
Gill, “Everything You Always Wanted to Know About SDRAM (Memory): But Were Afraid to Ask”, AnandTech, Available Online at: https://www.anandtech.com/show/3851/everything-you-always-wanted-to-know-about-sdram-memory-but-… [cited by applicant]
International Search Report and Written Opinion, mailed Apr. 1, 2019 in PCT/US2018/064777. [cited by applicant]
Jouppi, N. P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” To appear at the 44th International Symposium on Computer Architecture (ISCA), Toronto, Canada, Jun. 26, 2017, 17 pages. [cited by applicant]
Liao et al., “Machine Learning-Based Prefetch Optimization for Data Center Applications”, IEEE Xplore, Nov. 14-20, 2009, 10 pages. [cited by applicant]
Machine translation of Ye, Y., et al., “Design of Non-Volatile Synapse Array and Neuron Circuits,” Microelectronics & Computer, Nov. 2017, vol. 34(11), 5 pages. [cited by applicant]
Notice of Allowance dated Oct. 21, 2022, in Application No. JP 2020-531932. [cited by applicant]
Nurvitadhi et al., “Can FPGAs Beat GP US in Accelerating Next-Generation Deep Neural Networks?”, in FPGA '17, ACM Digital Library, Available Online at: https://dl.acm.org/doi/pdf/10.1145/3020078.3021740, Feb. 22-24, 201… [cited by applicant]
Office Action dated Aug. 4, 2021; JP 2020531932. [cited by applicant]
Peemen et al., “Memory-Centric Accelerator Design For Convolutional Neural Networks”, IEEE Xplore, Available Online at: https://ieeexplore.ieee.org/document/6657019, 2013, 7 pages. [cited by applicant]
Reineke et al., “Timing Predictability of Cache Replacement Policies”, Real Time Systems, vol. 37, No. 2, Nov. 2007, pp. 1-24. [cited by applicant]
Sze, V., et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages. [cited by applicant]
U.S. Restriction Requirement dated Sep. 7, 2022 in U.S. Appl. No. 17/033,573. [cited by applicant]
U.S. Non-Final office Action dated Dec. 30, 2022 in U.S. Appl. No. 17/033,573. [cited by applicant]
U.S. Non-Final Office Action dated Jun. 24, 2020 in U.S. Appl. No. 15/839,017. [cited by applicant]
U.S. Non-Final office Action dated May 22, 2020, in U.S. Appl. No. 15/839,157. [cited by applicant]
U.S. Notice of Allowance dated Apr. 6, 2023 in U.S. Appl. No. 17/033,573. [cited by applicant]
U.S. Corrected Notice of Allowability dated Apr. 24, 2023 in U.S. Appl. No. 17/033,573. [cited by applicant]
U.S. Notice of Allowance dated Jul. 17, 2020, in U.S. Appl. No. 15/839,157. [cited by applicant]
U.S. Notice of Allowance dated Jun. 17, 2020, in U.S. Appl. No. 15/839,301. [cited by applicant]
Ye, Y., “Design of Non-Volatile Synapse Array and Neuron Circuits,” Microelectronics & Computer, Nov. 2017, vol. 34 (11), pp. 1-5. [cited by applicant]
JP Office Action dated Oct. 6, 2023 in JP Application No. 2022-123015 with English Translation. [cited by applicant]
EP Office Action dated Feb. 8, 2024 in EP Application No. 18836949.0. [cited by applicant]
EP Notice of Allowance dated Aug. 27, 2024 in EP Application No. 18836949.0. [cited by applicant]
JP Notice of Allowance dated Feb. 9, 2024 in JP Application No. 2022-123015 with English Translation. [cited by applicant]