IP Library Granted Patent US 11,210,586
Granted Patent B1
US 11,210,586 · App. 16/457,757 · Granted Dec 28, 2021

Weight value decoder of neural network inference circuit

Inventors: Kenneth Duong (San Jose, CA); Jung Ko (San Jose, CA); Steven L. Teig (Menlo Park, CA)
Assignee: PERCEIVE CORPORATION
G06N3/08G06N3/04G06N3/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,210,586
App. No.
16/457,757
Granted
Dec 28, 2021
Kind
B1
Abstract

Some embodiments provide a method for a neural network inference circuit that executes a neural network including multiple computation nodes at multiple layers. Each computation node of a set of the computation nodes includes a dot product of input values and weight values. The method reads a set of encoded weight data for a set of weight values from a memory of the neural network inference circuit. The method decodes the encoded weight data to generate decoded weight data for the set of weight values. The method stores the decoded weight data in a buffer. The method uses the decoded weight data to execute a set of computation nodes. Each computation node of the set of computation nodes includes a dot product between the set of weight values and a different set of input values.

Claims (37)

1. A method for a neural network inference circuit that executes a neural network comprising a plurality of computation nodes at a plurality of layers, each computation node of a set of the computation nodes comprising a dot product of input values and weight values, the method comprising:

reading a set of encoded weight data for a set of weight values from a memory of the neural network inference circuit into a cache;

decoding the encoded weight data to generate decoded weight data for the set of weight values, said decoding the encoded weight data comprising (i) retrieving a first portion of the encoded weight data from the cache in a first clock cycle and (ii) retrieving a second portion of the encoded weight data from the cache in a second clock cycle;

storing the decoded weight data in a buffer; and

using the decoded weight data to execute a set of computation nodes, each computation node of the set of computation nodes comprising a dot product between the set of weight values and a different set of input values.

2. The method of claim 1 , wherein decoding the encoded weight data further comprises:

expanding the encoded weight data; and

aligning data from a first set of the encoded weight data with data from a second set of the encoded weight data.

3. The method of claim 1 , wherein the encoded weight data comprises a fixed width portion and a variable width portion.

4. The method of claim 1 , wherein the neural network inference circuit comprises a plurality of buffers for storing weight data, wherein the encoded weight data comprises an address indicating which buffer of the plurality of buffers stores the decoded weight data during the execution of the set of computation nodes.

5. The method of claim 1 , wherein the encoded weight data is decoded over two clock cycles.

6. The method of claim 1 , wherein the first portion is a fixed amount of data, wherein a size of the second portion depends on data in the first portion.

7. The method of claim 1 , wherein the second portion includes only data from the first portion.

8. The method of claim 1 , wherein the second portion includes data from the first portion and additional data.

9. The method of claim 1 , wherein the second portion includes only data not in the first portion.

10. The method of claim 1 , wherein the encoded weight data comprises (i) a bit for each weight value of the set of weight values indicating whether the weight value is non-zero and (ii) additional data for each of the non-zero weight values.

11. The method of claim 10 , wherein the additional data for each of the non-zero weight values comprises (i) a bit indicating whether the non-zero weight value is positive and (ii) a set of multiplexer select bits that indicates, for each computation node of the set of computation nodes, by which of a subset of the set of input values the weight value is multiplied.

12. The method of claim 10 , wherein retrieving the first portion of the encoded weight data comprises retrieving the bits for each of the weight values of the set of weight values and a fixed amount of additional data, wherein decoding the encoded weight data further comprises, in the first clock cycle:

aligning the additional data for each of the non-zero weight values in a first half of the weight values with the bit indicating that the weight value is non-zero; and

filling in zeros as the additional weight data for the weight values equal to zero in the first half of the weight values for which additional weight data is not stored in memory.

13. The method of claim 12 , wherein decoding the encoded weight data further comprises:

identifying a number of the weight values equal to zero in the second half of the weight values; and

using the identified number to determine the second portion of the encoded weight data to retrieve.

14. The method of claim 13 , wherein decoding the encoded weight data further comprises, in the second clock cycle:

aligning the additional data from the second portion of the encoded weight data for each of the non-zero weight values in the second half of the weight values with the bit indicating that the weight value is non-zero; and

filling in zeros as the additional weight data for the weight values equal to zero in the second half of the weight values for which additional weight data is not stored in memory.

15. A neural network inference circuit that executes a neural network comprising a plurality of computation nodes at a plurality of layers, each computation node of a set of the computation nodes comprising a dot product of input values and weight values, the neural network inference circuit comprising:

a set of memory control circuits to read a set of encoded weight data for a set of weight values from a memory of the neural network inference circuit into a set of caches;

a set of weight decoder circuits to generate decoded weight data for the set of weight values from the encoded weight values, said generation of decoded weight data comprising (i) retrieving a first portion of the encoded weight data from the cache in a first clock cycle and (ii) retrieving a second portion of the encoded weight data from the cache in a second clock cycle;

a set of buffers to store the decoded weight data; and

a set of computation circuits to use the decoded weight data to execute a set of computation nodes, each computation node of the set of computation nodes comprising a dot product between the set of weight values and a different set of input values.

16. The neural network inference circuit of claim 15 , wherein the set of weight decoder circuits decodes the encoded weight data by expanding the encoded weight data and aligning data from a first set of the encoded weight data with data from a second set of the encoded weight data.

17. The neural network inference circuit of claim 15 , wherein the encoded weight data comprises a fixed width portion and a variable width portion.

18. The neural network inference circuit of claim 15 , wherein the set of buffers comprises a plurality of buffers for storing weight data, wherein the encoded weight data comprises an address indicating which buffer of the plurality of buffers stores the decoded weight data during the execution of the set of computation nodes.

19. The neural network inference circuit of claim 15 , wherein:

the encoded weight data comprises (i) a bit for each weight value of the set of weight values indicating whether the weight value is non-zero and (ii) additional data for each of the non-zero weight values; and

the additional data for each of the non-zero weight values comprises (i) a bit indicating whether the non-zero weight value is positive and (ii) a set of multiplexer select bits that indicates, for each computation node of the set of computation nodes, by which of a subset of the set of input values the weight value is multiplied.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2019
From: DUONG, KENNETH; KO, JUNG; TEIG, STEVEN L.
To: PERCEIVE CORPORATION
Reel/Frame 049838/0602 →
Continuity (10)
Continuation In Part 16120387 · Sep 3, 2018
Provisional Application 62853128 · May 27, 2019
Provisional Application 62797910 · Jan 28, 2019
Provisional Application 62792123 · Jan 14, 2019
Provisional Application 62773162 · Nov 29, 2018
Provisional Application 62773164 · Nov 29, 2018
Provisional Application 62753878 · Oct 31, 2018
Provisional Application 62742802 · Oct 8, 2018
Provisional Application 62724589 · Aug 29, 2018
Provisional Application 62660914 · Apr 20, 2018
Cited By (10)
US 12,190,892 US 12,260,337 US 12,288,152 US 12,334,956 US 12,474,890 US 12,591,430 US 12,626,121 US 12,675,268 US 12,711,393 US 12,718,062