IP Library Granted Patent US 11,886,979
Granted Patent B1
US 11,886,979 · App. 16/355,648 · Granted Jan 30, 2024

Shifting input values within input buffer of neural network inference circuit

Inventors: Kenneth Duong (San Jose, CA); Jung Ko (San Jose, CA); Steven L. Teig (Menlo Park, CA)
Assignee: PERCEIVE CORPORATION
G06N3/063G06F9/30134G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,886,979
App. No.
16/355,648
Granted
Jan 30, 2024
Kind
B1
Abstract

Some embodiments provide a method for a neural network inference circuit that executes a neural network. The method loads a first set of inputs into an input buffer and computes a first dot product between the first set of inputs and a set of weights. The method shifts the first set of inputs in the buffer while loading a second set of inputs into the buffer such that a first subset of the first set of inputs is removed from the buffer, a second subset of the first set of inputs is moved to new locations in the buffer, and a second set of inputs are loaded into locations in the buffer vacated by the shifting. The method computes a second dot product between (i) the second set of inputs and the second subset of the first set of inputs and (ii) the set of weights.

Claims (39)

1. For a neural network inference circuit that executes a neural network comprising a plurality of computation nodes, each of a set of the computation nodes comprising a dot product of input values and weight values, a method comprising:

over a first plurality of clock cycles of the neural network inference circuit, loading a first set of input values into an input buffer;

in a first single clock cycle of the neural network inference circuit after the first plurality of clock cycles, simultaneously computing a first plurality of dot products for a first plurality of computation nodes of the neural network, wherein each respective dot product of the first plurality of dot products is between (i) the first set of input values and (ii) a different respective set of a plurality of sets of weight values;

over a second plurality of clock cycles of the neural network inference circuit after the first single clock cycle, shifting the first set of input values in the input buffer while loading a second set of input values into the input buffer such that (i) a first subset of the first set of input values is removed from the input buffer, (ii) a second subset of the first set of input values is moved to new locations in the input buffer, and (iii) the second set of input values are loaded into locations in the input buffer vacated by the shifting of the first set of input values; and

in a second single clock cycle of the neural network inference circuit after the second plurality of clock cycles, simultaneously computing a second plurality of dot products for a second plurality of computation nodes of the neural network, wherein each respective dot product of the second plurality of dot products is between (i) the second set of input values and the second subset of the first set of input values and (ii) a different respective set of the plurality of sets of weight values, wherein the respective sets of weight values used for the first plurality of dot products are the same as the respective sets of weight values used for the second plurality of dot products.

2. The method of claim 1 , wherein during a particular clock cycle of the second plurality of clock cycles, a first group of the first subset of the first set of input values is removed from the input buffer while a first group of the second set of input values are loaded into the input buffer.

3. The method of claim 1 , wherein the number of clock cycles in the second plurality of clock cycles depends on a kernel size associated with the plurality of sets of weight values.

4. The method of claim 1 , wherein loading the first set of input values comprises loading input values from at least one of (i) a memory location of a set of memories of the neural network inference circuit and (ii) a cache of the neural network inference circuit for storing input values previously read from memory locations of the set of memories.

5. The method of claim 1 , wherein the input buffer comprises a plurality of register cells, each register cell for storing an input value.

6. The method of claim 5 , wherein the plurality of register cells are configurably grouped into a set of shift registers, wherein shifting the first set of input values in the input buffer while loading the second set of input values into the input buffer comprises:

loading a first input value of the second set of input values into a first register cell of a particular one of the shift registers; and

shifting each of the input values of the first set of input values that are stored in the register cells of the particular shift register by one register cell within the particular shift register, wherein a first input value of the first set of input values that is initially stored in a last register cell of the shift register is shifted out of the input buffer.

7. The method of claim 6 , wherein the loading of the first input value and the shifting of each of the input values of the first set of input values that are stored in the register cells of the particular shift register occurs in a same third single clock cycle of the second plurality of clock cycles.

8. The method of claim 7 , wherein shifting the first set of input values in the input buffer while loading the second set of input values into the input buffer further comprises, in a fourth single clock cycle that is subsequent to the third clock cycle and part of the second plurality of clock cycles:

loading a second input value of the second set of input values into the first register cell of the particular shift register; and

shifting the first input value of the second set of input values and each of the input values of the first set of input values that are still stored in the register cells of the particular shift register by one register cell within the particular shift register, wherein a second input value of the first set of input values is shifted out of the input buffer.

9. The method of claim 7 , wherein different input values of the second set of input values are loaded into each of the shift registers of the set of shift registers in the same third single clock cycle and the input values of the first set of input values that are stored in each of the shift registers are shifted such that different input values of the first set of input values are shifted out of the input buffer in the same third single clock cycle.

10. The method of claim 6 , wherein the configurable grouping of the plurality of register cells into the set of shift registers depends on a kernel size for a current layer of the neural network.

11. The method of claim 6 , wherein each of the shift registers has a same length.

12. The method of claim 1 , wherein each different set of the plurality of sets of weight values is a different filter with a same set of filter dimensions such that each simultaneously computed dot product of the first plurality of dot products is between (i) the same first set of input values and (ii) a different filter.

13. A neural network inference circuit that executes a neural network comprising a plurality of computation nodes, each computation node of a set of the computation nodes comprising a dot product of input values and weight values, the neural network inference circuit comprising:

an input buffer to store a set of input values for a dot product computation;

a set of dot product computation circuits to compute dot products between the set of input values stored in the input buffer and a plurality of sets of weight values; and

an input buffer configuration circuit to (i) load a first set of input values into the input buffer over a first plurality of clock cycles of the neural network inference circuit and (ii) after a first plurality of dot products for a first plurality of computation nodes of the neural network are simultaneously computed in a first single clock cycle of the neural network inference circuit after the first plurality of clock cycles, wherein each respective dot product of the first plurality of dot products is between the first set of input values and a different respective set of the plurality of sets of weight values, shift the first set of input values in the input buffer while loading a second set of input values into the input buffer over a second plurality of clock cycles of the neural network inference circuit after the first single clock cycle such that (1) a first subset of the first set of input values is removed from the input buffer, (2) a second subset of the first set of input values is moved to new locations in the input buffer, and (3) the second set of input values are loaded into locations in the input buffer vacated by the shifting of the first set of input values,

wherein the set of dot product computation circuits (i) simultaneously computes the first plurality of dot products for the first plurality of computation nodes in the first single clock cycle and (ii) simultaneously computes a second plurality of dot products for a second plurality of computation nodes of the neural network in a second single clock cycle of the neural network inference circuit after the second plurality of clock cycles, wherein each respective dot product of the second plurality of dot products is between (i) the second set of input values and the second subset of the first set of input values and (ii) a different respective set of the plurality of sets of weight values, wherein the respective sets of weight values used for the first plurality of dot products are the same as the respective sets of weight values used for the second plurality of dot products.

14. The neural network inference circuit of claim 13 , wherein during a particular clock cycle of the second plurality of clock cycles, the input buffer configuration circuit removes a first group of the first subset of the first set of input values from the input buffer while a first group of the second set of input values are loaded into the input buffer.

15. The neural network inference circuit of claim 13 further comprising:

a set of memories to store input values for a plurality of computation nodes; and

a cache to store input values previously read from memory locations of the set of memories,

wherein the first set of input values are loaded from at least one of (i) a memory location of the set of memories and (ii) the cache.

16. The neural network inference circuit of claim 13 , wherein the input buffer comprises a plurality of register cells, each register cell for storing an input value.

17. The neural network inference circuit of claim 16 , wherein the plurality of register cells are configurably grouped into a set of shift registers, wherein the input buffer configuration circuit shifts the first set of input values in the input buffer while loading the second set of input values into the input buffer by:

loading a first input value of the second set of input values into a first register cell of a particular one of the shift registers; and

shifting each of the input values of the first set of input values that are stored in the register cells of the particular shift register by one register cell within the particular shift register, wherein a first input value of the first set of input values that is initially stored in a last register cell of the shift register is shifted out of the input buffer.

18. The neural network inference circuit of claim 17 , wherein the loading of the first input value and the shifting of each of the input values of the first set of input values that are stored in the register cells of the particular shift register occurs in a same third single clock cycle of the neural network inference circuit.

19. The neural network inference circuit of claim 17 , wherein the configurable grouping of the plurality of register cells into the set of shift registers depends on a kernel size for a current layer of the neural network.

20. The neural network inference circuit of claim 13 , wherein each different set of the plurality of sets of weight values is a different filter with a same set of filter dimensions such that each simultaneously computed dot product of the first plurality of dot products is between (i) the same first set of input values and (ii) a different filter.

21. The neural network inference circuit of claim 15 , wherein a first group of the first set of input values are loaded from at least one memory location of the set of memories and a second group of the first set of input values are loaded from the cache.

22. The method of claim 1 , wherein loading the first set of input values comprises (i) loading a first group of the first set of input values from a set of memories of the neural network inference circuit and (ii) loading a second group of the first set of input values from a cache of the neural network inference circuit that stores input values previously read from the set of memories.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2019
From: DUONG, KENNETH; KO, JUNG; TEIG, STEVEN L.
To: PERCEIVE CORPORATION
Reel/Frame 048617/0334 →
Continuity (8)
Provisional Application 62797910 · Jan 28, 2019
Provisional Application 62792123 · Jan 14, 2019
Provisional Application 62773164 · Nov 29, 2018
Provisional Application 62773162 · Nov 29, 2018
Provisional Application 62753878 · Oct 31, 2018
Provisional Application 62742802 · Oct 8, 2018
Provisional Application 62724589 · Aug 29, 2018
Provisional Application 62660914 · Apr 20, 2018
Cited By (2)
US 12,282,853 US 12,693,990