IP Library Granted Patent US 11,238,334
Granted Patent B2
US 11,238,334 · App. 16/569,307 · Granted Feb 1, 2022

System and method of input alignment for efficient vector operations in an artificial neural network

Inventors: Avi Baum (Givat Shmuel, IL); Or Danon (Kiryat Ono, IL); Daniel Ciubotariu (Ashdod, IL)
G06N3/063G06F5/06G06F5/10G06F7/78G06F7/785G06F9/30032G06F9/30134G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,238,334
App. No.
16/569,307
Granted
Feb 1, 2022
Kind
B2
Abstract

A novel and useful system and method of input alignment for streamlining vector operations that reduce the required memory read bandwidth. The input aligner as deployed in the NN processor, functions to facilitate the reuse of data read from memory and to avoid having to re-read that data in the context of neural network calculations. The input aligner functions to distribute input data (or weights) to the appropriate compute elements while consuming input data in a single cycle. Thus, the input aligner is operative to lower the required read bandwidth of layer input in an ANN. This reflects the fact that normally in practice, a vector multiplication is performed every time instance. This considers the fact that in many native calculations that take place in an ANN, the same data point is involved in multiple calculations.

Claims (39)

1. An input aligner apparatus for use in an artificial neural network, comprising:

a plurality of multibit registers arranged in linear fashion to function as a shift register;

a first circuit operative to load said plurality of multibit registers in parallel and/or serially with data from a memory;

a plurality of multiplexer circuits coupled to said multibit registers, said plurality of multiplexer circuits configured to:

shift contents of said plurality of multibit registers in accordance with a stride value whereby zero or more multibit registers are skipped;

cycle contents of said plurality of multibit registers in accordance with a stride value whereby zero or more multibit registers are skipped; and

a distribution circuit operative to output the contents of said plurality of multibit registers to appropriate processing elements in parallel simultaneously, said contents effectively re-used multiple times thereby avoiding multiple data reads from memory and reducing required memory bandwidth.

2. The apparatus according to claim 1 , wherein the output of said input aligner is input to processing elements in one or more subclusters.

3. The apparatus according to claim 1 , wherein a shift of the contents of the input aligner is performed when weight calculations with current contents are complete in accordance with a particular kernel.

4. The apparatus according to claim 1 , wherein a parallel load of input data to the input aligner is performed when calculations with current contents are exhausted and new input data is required.

5. The apparatus according to claim 1 , wherein the output of said input aligner is input to processing elements whereby multiple processing elements receive the same input from the input aligner and different weights from weight memory.

6. The apparatus according to claim 1 , further comprising a third circuit operative to expand contents of said plurality of multibit registers in accordance with an expand value whereby output data is expanded and filled in with an expand value.

7. The apparatus according to claim 6 , wherein said third circuit is operative to calculate said expand value dynamically in accordance with input data stream.

8. The apparatus according to claim 1 , wherein data loaded into said input aligner is stored therein while multiple weight data is shifted into and out of memory.

9. An input aligner apparatus for use in an artificial neural network, comprising:

a plurality of multibit registers arranged in linear fashion to function as a shift register;

a first circuit operative to load said plurality of multibit registers in parallel and/or serially with data from a memory;

a plurality of multiplexer circuits coupled to said multibit registers, said plurality of multiplexer circuits configured to expand contents of said plurality of multibit registers in accordance with a non-zero expand value whereby output data is from said expand value rather than said multibit registers; and

a distribution circuit operative to output the contents of said plurality of multibit registers to appropriate processing elements in parallel simultaneously, said contents effectively re-used multiple times thereby avoiding multiple data reads from memory and reducing required memory bandwidth.

10. The apparatus according to claim 9 , wherein the output of said input aligner is input to processing elements in one or more subclusters.

11. The apparatus according to claim 9 , wherein a shift of the contents of the input aligner is performed when weight calculations with current contents are complete.

12. The apparatus according to claim 9 , wherein a parallel load of input data to the input aligner is performed when calculations with current contents are exhausted in accordance with a particular kernel and new input data is required.

13. The apparatus according to claim 9 , wherein the output of said input aligner is input to processing elements whereby multiple processing elements receive the same input from the input aligner and different weights from weight memory.

14. The apparatus according to claim 9 , further comprising a third circuit operative to shift contents of said plurality of multibit registers in accordance with a stride value whereby zero or more multibit registers are skipped.

15. The apparatus according to claim 9 , wherein the number of active multibit registers in said input aligner is configurable in accordance with a length command.

16. The apparatus according to claim 9 , wherein the configuration of processing elements coupled to said input aligner is configurable in accordance with a configuration command.

17. An input alignment method for use in an artificial neural network, the method comprising:

providing a plurality of multibit registers arranged in linear fashion to function as a shift register;

providing a plurality of multiplexer circuits coupled to said multibit registers;

loading said plurality of multibit registers in parallel and/or serially with data from a memory utilizing said multiplexer circuits;

shifting and/or cycling contents of said plurality of multibit registers via said multiplexer circuits in accordance with a shift command;

skipping zero or more multibit registers in accordance with a stride value; and

distributing the contents of said plurality of multibit registers to appropriate processing elements in parallel simultaneously, said contents effectively re-used multiple times thereby avoiding multiple data reads from memory and reducing required memory bandwidth.

18. The method according to claim 17 , further comprising outputting the content of said plurality of multibit registers to processing elements in one or more subclusters.

19. The method according to claim 17 , further comprising shifting the content of said plurality of multibit registers when weight calculations with current contents are complete in accordance with a particular kernel.

20. The method according to claim 17 , further comprising loading in parallel input data to said plurality of multibit registers when calculations with current contents are exhausted and new input data is required in accordance with particular kernel.

21. The method according to claim 17 , further comprising outputting the content of said plurality of multibit registers to processing elements whereby multiple processing elements receive the same input and different weights from a weight memory.

22. The method according to claim 17 , further comprising expanding contents of said plurality of multibit registers in accordance with an expand value whereby output data is expanded and filled in with an expand value.

23. The method according to claim 22 , wherein said expand value is calculated dynamically in accordance with an input data stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2019
From: BAUM, AVI; DANON, OR; CIUBOTARIU, DANIEL
To: HAILO TECHNOLOGIES LTD.
Reel/Frame 050362/0185 →
Continuity (4)
Continuation In Part 15943800 · Apr 3, 2018
Provisional Application 62481492 · Apr 4, 2017
Provisional Application 62531372 · Jul 12, 2017
Related Publication 20200005127A1 · Jan 2, 2020
Cited By (4)
US 12,353,987 US 12,430,902 US 12,602,576 US 12,737,603