IP Library Granted Patent US 11,741,346
Granted Patent B2
US 11,741,346 · App. 15/981,679 · Granted Aug 29, 2023

Systolic neural network engine with crossover connection optimization

Inventor: Luiz M. Franca-Neto (Sunnyvale, CA)
Assignee: Western Digital Technologies, Inc.
G06N3/063G06F15/8046G06N3/04G06N3/048G06N3/08G06N3/084G06N5/046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,741,346
App. No.
15/981,679
Granted
Aug 29, 2023
Kind
B2
Abstract

Devices and methods for systolically processing data according to a neural network. In one aspect, a first arrangement of processing units includes at least first, second, third, and fourth processing units. The first and second processing units are connected to systolically pulse data to one another, and the third and fourth processing units are connected to systolically pulse data to one another. A second arrangement of processing units includes at least fifth, sixth, seventh, and eighth processing units. The fifth and sixth processing units are connected to systolically pulse data to one another, and the seventh and eighth processing units are connected to systolically pulse data to one another. The second processing unit is configured to systolically pulse data to the seventh processing unit along a first interconnect and the third processing unit is configured to systolically pulse data to the sixth processing unit along a second interconnect.

Claims (42)

1. A device for systolically processing data according to a neural network, the device comprising:

a first interconnect between the first and second arrangements that connects the second and seventh processing units, wherein the second processing unit is configured to systolically pulse data to the seventh processing unit along the first interconnect; and

a second interconnect between the first and second arrangements that connects the third and sixth processing units, wherein the third processing unit is configured to systolically pulse data to the sixth processing unit along the second interconnect;

wherein the first and second interconnects form a first pair of interconnects of multiple pairs of interconnects that connect the first arrangement to the second arrangement;

wherein each of the processing units in the first and second arrangements includes a number of convolution engines based on the number of multiple pairs of interconnects connecting the first arrangement to the second arrangement; and

wherein the convolution engines in each processing unit of the first arrangement and the second arrangement includes an equal number of clockwise convolution engines and counterclockwise convolution engines, wherein each of the clockwise convolution engines is configured to perform clockwise convolution operations and each of the counterclockwise convolution engines is configured to perform counterclockwise convolution operations.

2. The device of claim 1 , wherein the number of convolution engines in each processing unit of the first and second arrangements equals the number of multiple pairs of interconnects connecting the first arrangement to the second arrangement.

3. The device of claim 1 , wherein the device further comprises a second pair of interconnects, the second pair of interconnects including a third interconnect between an uppermost processing unit in the first arrangement and an uppermost processing unit in the second arrangement and a fourth interconnect between a lowermost processing unit in the first arrangement and a lowermost processing unit in the second arrangement.

4. The device of claim 1 , wherein, during a single systolic pulse, the sixth processing unit is configured to receive a first piece of data from the third processing unit and a second piece of data from the fifth processing unit.

5. The device of claim 1 , wherein the device further includes a systolic processor chip, and wherein the first and second arrangements of first and second processing units comprise circuitry embedded in the systolic processor chip.

6. The device of claim 1 , wherein the second processing unit includes an output systolic element configured to tag an activation output generated by the second processing unit with an identifier, wherein the identifier indicates an address for the second processing unit.

7. The device of claim 6 , wherein the activation output including the tag is systolically pulsed to an input systolic element of the seventh processing unit.

8. The device of claim 7 , wherein the seventh processing unit is configured to:

receive the activation output and perform processing to generate an additional activation output; and

use the identifier to identify a weight to use for processing the activation output.

9. The device of claim 8 , wherein the weight is stored locally at the seventh processing unit.

10. The device of claim 8 , wherein the weight is retrieved from a memory external to the seventh processing unit.

11. The device of claim 1 , wherein at least a subset of the first arrangement of processing units are assigned to perform computations of a first layer of the neural network, and wherein at least a subset of the second arrangement of processing units are assigned to perform computations of a second layer of the neural network.

12. The device of claim 1 , wherein the first processing unit includes an input systolic element configured to receive data, a first processing circuit configured to perform processing of the received data to generate a first activation output, a first output systolic element, and a data tagger configured to tag the first activation output with an address of the first processing unit.

13. A method for systolically processing data according to a neural network comprising at least a first layer and a second layer, the method comprising:

during a first systolic clock cycle, performing a first set of systolic pulses of data through at least first, second, third, and fourth processing units arranged along a first arrangement of processing units and at least fifth, sixth, seventh, and eighth processing units arranged along a second arrangement of processing units, the first set of systolic pulses including:

systolically pulsing data from the first processing unit of the first arrangement to the second processing unit of the first arrangement;

systolically pulsing data from the seventh processing unit of the second arrangement to the eighth processing unit of the second arrangement; and

systolically pulsing data from the second processing unit of the first arrangement to the seventh processing unit of the second arrangement;

wherein the second processing unit is configured to systolically pulse data to the seventh processing unit along a first interconnect between the first and second arrangements, and wherein the third processing unit is configured to systolically pulse data to the sixth processing unit along a second interconnect between the first and second arrangements;

wherein the first and second interconnects form a first pair of interconnects, and wherein multiple pairs of interconnects connect the first arrangement to the second arrangement of processing units; and

wherein each of the processing units in the first and second arrangements includes an equal number of clockwise and counterclockwise convolution engines based on the number of multiple pairs of interconnects connecting the first arrangement to the second arrangement, wherein each of the clockwise convolution engines is configured to perform clockwise convolution operations and each of the counterclockwise convolution engines is configured to perform counterclockwise convolution operations.

14. The method of claim 13 , further comprising, during the first systolic clock cycle, performing a second set of systolic pulses including:

systolically pulsing data from the second processing unit of the first arrangement to the first processing unit of the first arrangement;

systolically pulsing data from the third processing unit of the first arrangement to the sixth processing unit of the second arrangement;

systolically pulsing data from the fourth processing unit of the first arrangement to the third processing unit of the first arrangement;

systolically pulsing data from the sixth processing unit of the second arrangement to the fifth processing unit of the second arrangement; and

systolically pulsing data from the eighth processing unit of the second arrangement to the seventh processing unit of the second arrangement.

15. The method of claim 14 , wherein the first set of systolic pulses travel in a first direction through the first and second arrangements, and wherein the second set of systolic pulses travel in a second direction through the first and second arrangements, wherein the first direction is opposite to the second direction.

16. The method of claim 13 , further comprising, during a second systolic clock cycle, performing a second set of systolic pulses including:

systolically pulsing, from the second processing unit of the first arrangement to the seventh processing unit of the second arrangement, the data received from the first processing unit during the first systolic clock cycle; and

systolically pulsing, from the third processing unit of the first arrangement to the sixth processing unit of the second arrangement, data received from the fourth processing unit during the first systolic clock cycle.

17. The method of claim 16 , further comprising, via the seventh processing unit during the second systolic clock cycle, processing the data received from the second processing unit during the first systolic clock cycle, the processing performed according to computations of a node of the second layer of the neural network.

18. The method of claim 17 , further comprising, via the seventh processing unit during a third systolic clock cycle, processing the data received from the second processing unit during the second systolic clock cycle, the processing performed according to computations of the node of the second layer of the neural network.

19. The method of claim 17 , further comprising using a tag of the data received from the second processing unit to identify a weight to use for processing the data received from the second processing unit, the tag identifying that the data originated at the second processing unit.

20. The method of claim 13 , wherein the number of convolution engines in each processing unit of the first arrangement and the second arrangement of processing units equals a number of multiple pairs of interconnects connecting the first arrangement to the second arrangement.

21. The device of claim 1 , wherein the clockwise and counterclockwise convolution engines in each of the processing units of the first and second arrangements are configured to concurrently process inputs received by the processing unit.

Assignments (10)
PARTIAL RELEASE OF SECURITY INTERESTS Recorded Apr 25, 2025
From: JPMORGAN CHASE BANK, N.A., AS AGENT
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 071382/0001 →
SECURITY AGREEMENT Recorded Apr 25, 2025
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 071050/0001 →
PATENT COLLATERAL AGREEMENT Recorded Aug 23, 2024
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS THE AGENT
Reel/Frame 068762/0494 →
CHANGE OF NAME Recorded Jun 27, 2024
From: SANDISK TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067982/0032 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067567/0682 →
PATENT COLLATERAL AGREEMENT - DDTL LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 067045/0156 →
PATENT COLLATERAL AGREEMENT - A&R LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064715/0001 →
RELEASE OF SECURITY INTEREST AT REEL 052915 FRAME 0566 Recorded Feb 8, 2022
From: JPMORGAN CHASE BANK, N.A.
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 059127/0001 →
SECURITY INTEREST Recorded Feb 6, 2020
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS AGENT
Reel/Frame 052915/0566 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2018
From: FRANCA-NETO, LUIZ M.
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 047025/0027 →
Continuity (3)
Provisional Application 62628076 · Feb 8, 2018
Provisional Application 62627957 · Feb 8, 2018
Related Publication 20190244082A1 · Aug 8, 2019
Cited By (1)
US 12,693,990