IP Library Granted Patent US 11,586,910
Granted Patent B1
US 11,586,910 · App. 16/537,478 · Granted Feb 21, 2023

Write cache for neural network inference circuit

Inventors: Kenneth Duong (San Jose, CA); Jung Ko (San Jose, CA); Steven L. Teig (Menlo Park, CA)
Assignee: PERCEIVE CORPORATION
G06N3/08G06F12/0891G06F12/121G06N5/04G06F2212/604
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,586,910
App. No.
16/537,478
Granted
Feb 21, 2023
Kind
B1
Abstract

Some embodiments provide a neural network inference circuit (NNIC) for executing a neural network that includes computation nodes at multiple layers. The NNIC includes multiple value computation circuits for computing output values of computation nodes. The NNIC includes a set of memories for storing the output values of computation nodes for use as input values to computation nodes in subsequent layers of the neural network. The NNIC includes a set of write control circuits for writing the computed output values to the set of memories. Upon receiving a set of computed output values, a write control circuit (i) temporarily stores the set of computed output values in a cache when adding the set of computed output values to the cache does not cause the cache to fill up and (ii) writes data in the cache to the set of memories when the cache fills up.

Claims (24)

1. A neural network inference circuit for executing a neural network that comprises a plurality of computation nodes at a plurality of layers, the neural network inference circuit comprising:

a plurality of value computation circuits for computing output values of computation nodes;

a set of memories for storing the output values of computation nodes for use as input values to computation nodes in subsequent layers of the neural network; and

a set of write control circuits for writing the computed output values to the set of memories, a particular write control circuit comprising temporary storages for storing (i) a cache of computed output values that stores one memory location worth of data and (ii) data indicating a location in the cache up to which the cache stores valid data,

wherein the particular write control circuit stores computed output values in the cache in a circular manner,

wherein upon receiving a set of computed output values, the particular write control circuit uses a number of output values in the set and the data indicating the location in the cache up to which the cache stores valid data to determine whether the received set of computed output values fill up the cache and (i) temporarily stores the set of computed output values in the cache when adding the set of computed output values to the cache does not cause the cache to fill up and (ii) writes data in the cache to the set of memories when the cache fills up.

2. The neural network inference circuit of claim 1 comprising a plurality of core circuits, wherein each respective core circuit comprises a respective subset of the memories and a respective write control circuit for writing the computed output values to the respective subset of memories of the respective core circuit.

3. The neural network inference circuit of claim 2 , wherein the value computation circuits comprise partial computation circuits in each of the core circuits.

4. The neural network inference circuit of claim 3 , wherein the value computation circuits further comprise post-processing units that generate the output values for the write control circuits to write to the set of memories based on data generated by the partial computation circuits.

5. The neural network inference circuit of claim 4 further comprising an output bus for transporting the output values generated by the post-processing units to the write control circuits in the cores.

6. The neural network inference circuit of claim 1 , wherein the temporary storages are registers.

7. The neural network inference circuit of claim 1 , wherein the number of output values in the set is received by the particular write control circuit as configuration data.

8. The neural network inference circuit of claim 1 , wherein when the particular write control circuit determines that the received set of output values will not completely fill up the cache, the particular write control circuit (i) stores the received set of output values in the cache without writing any data to the set of memories and (ii) updates the data indicating the cache location up to which the cache stores valid data based on the number of output values in the received set of computed output values.

9. The neural network inference circuit of claim 1 , wherein when the particular write control circuit determines that the received set of output values will completely fill up the cache, the particular write control circuit writes to a specified memory location (i) the set of computed output values stored in the cache and (ii) a subset of the received set of output values.

10. The neural network inference circuit of claim 9 , wherein the subset of the received set of output values is an amount of output values equal to the remaining room in the cache prior to receiving the set of output values.

11. The neural network inference circuit of claim 9 , wherein any computed output values of the set of output values that are not in the subset of output values written to the specified memory location are stored in the beginning of the cache to be written to a different memory location at a later time.

12. The neural network inference circuit of claim 9 , wherein when the particular write control circuit determines that the received set of output values will completely fill up the cache and writes the set of output values in the cache and the subset of the received set of output values to the specified memory location, the particular write control circuit further modifies data indicating a next memory location to which to write subsequent output values.

13. The neural network inference circuit of claim 1 , wherein the set of computed output values is a first set of output values, wherein in a particular clock cycle of the neural network inference circuit, the particular write control circuit receives (i) a second set of computed output values and (ii) an instruction to write any output values stored in the cache and the received set of output values to memory irrespective of whether the received set of output values causes the cache to fill up.

14. The neural network inference circuit of claim 13 , wherein when the set of output values received in the particular clock cycle does not cause the cache to fill up, the particular write control circuit writes, to a specified memory location, (i) the set of output values stored in the cache, (ii) the set of output values received in the particular clock cycle, and (iii) zeros for any remaining bits in the specified memory location.

15. The neural network inference circuit of claim 13 , wherein when the set of output values received in the particular clock cycle is larger than remaining space in the cache, the particular write control circuit:

in the particular clock cycle, writes to a first memory location (i) the set of output values stored in the cache and (ii) a first subset of the set of output values received in the particular clock cycle; and

in a subsequent clock cycle, writes to a second memory location (i) a second subset of the set of output values received in the particular clock cycle and (ii) zeros for any remaining bits in the second memory location.

16. The neural network inference circuit of claim 15 , wherein the particular write control circuit stores, in the particular clock cycle, a state value indicating to perform a write operation in the subsequent clock cycle without receiving a write request from a controller circuit.

17. The neural network inference circuit of claim 13 , wherein the particular write control circuit resets the data indicating the location in the cache based on the instruction.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2019
From: DUONG, KENNETH; KO, JUNG; TEIG, STEVEN L.
To: PERCEIVE CORPORATION
Reel/Frame 050253/0110 →
Continuity (11)
Continuation In Part 16120387 · Sep 3, 2018
Provisional Application 62873804 · Jul 12, 2019
Provisional Application 62853128 · May 27, 2019
Provisional Application 62797910 · Jan 28, 2019
Provisional Application 62792123 · Jan 14, 2019
Provisional Application 62773164 · Nov 29, 2018
Provisional Application 62773162 · Nov 29, 2018
Provisional Application 62753878 · Oct 31, 2018
Provisional Application 62742802 · Oct 8, 2018
Provisional Application 62724589 · Aug 29, 2018
Provisional Application 62660914 · Apr 20, 2018
Cited By (6)
US 12,190,892 US 12,217,160 US 12,579,416 US 12,639,234 US 12,639,557 US 12,675,678