IP Library Granted Patent US 11,157,801
Granted Patent B2
US 11,157,801 · App. 15/637,664 · Granted Oct 26, 2021

Neural network processing with the neural network model pinned to on-chip memories of hardware nodes

Inventors: Eric S. Chung (Woodinville, WA); Douglas C. Burger (Bellevue, WA); Jeremy Fowers (Seattle, WA); Kalin Ovtcharov (Issaquah, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/063G06F9/3867G06N3/04G06F17/16G06N3/0481
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,157,801
App. No.
15/637,664
Granted
Oct 26, 2021
Kind
B2
Abstract

Systems and methods for neural network processing are provided. A method in a system comprising a plurality of nodes interconnected via a network, where each node includes a plurality of on-chip memory blocks and a plurality of compute units, is provided. The method includes upon service activation receiving an N by M matrix of coefficients corresponding to the neural network model. The method includes loading the coefficients corresponding to the neural network model into the plurality of the on-chip memory blocks for processing by the plurality of compute units. The method includes regardless of a utilization of the plurality of the on-chip memory blocks as part of an evaluation of the neural network model, maintaining the coefficients corresponding to the neural network model in the plurality of the on-chip memory blocks until the service is interrupted or the neural network model is modified or replaced.

Claims (29)

1. A method for evaluating a neural network model corresponding to a service in a system comprising a plurality of nodes interconnected via a network, wherein each node comprises a plurality of on-chip memory blocks and a plurality of compute units, the method comprising:

upon service activation receiving an N by M matrix of coefficients corresponding to the neural network model, wherein N is an integer equal to or greater than 8 and M is an integer equal to or greater than 8;

loading the N by M matrix of coefficients corresponding to the neural network model into the plurality of the on-chip memory blocks for processing by the plurality of compute units; and

regardless of a utilization of the plurality of the on-chip memory blocks as part of an evaluation of the neural network model, maintaining the N by M matrix of coefficients corresponding to the neural network model in the plurality of the on-chip memory blocks until the service is interrupted or the neural network model is modified or replaced.

2. The method of claim 1 , wherein the node comprises a field programmable gate array (FPGA) and wherein each of the plurality of the on-chip memory blocks comprises a static random access memory block.

3. The method of claim 2 , wherein each of the plurality of compute units comprises a set of pre-configured resources on the FPGA.

4. The method of claim 1 , wherein the plurality of the on-chip memory blocks is arranged in rows and wherein each of the plurality of compute units is configured to process at least a subset of at least one of the rows per clock cycle.

5. The method of claim 1 , wherein the loading the N by M matrix of coefficients corresponding to the neural network model into the plurality of the on-chip memory blocks comprises streaming data corresponding to the N by M matrix of coefficients corresponding to the neural network model via a broadcast block into the plurality of the on-chip memory blocks.

6. The method of claim 5 , wherein the streaming does not comprise loading any additional data corresponding to the neural network model from an off-chip memory in response to any operation associated with the N by M matrix of coefficients corresponding to the neural network model.

7. The method of claim 1 wherein the N by M matrix of coefficients comprises a Long Short Term Memory (LSTM) weights matrix.

8. A method for evaluating a neural network model corresponding to a service in a system comprising a plurality of nodes interconnected via a network, wherein each node comprises a plurality of on-chip memory blocks and a plurality of compute units, the method comprising:

upon service activation, partitioning the neural network model into separate layers, wherein each layer comprises an N by M matrix of coefficients corresponding to the neural network model, wherein N is an integer equal to or greater than 8 and M is an integer equal to or greater than 8;

loading the N by M matrix of coefficients corresponding to the neural network model into the plurality of the on-chip memory blocks for processing by the plurality of compute units; and

regardless of a utilization of the plurality of the on-chip memory blocks as part of an evaluation of the neural network model, maintaining the N by M matrix of coefficients corresponding to the neural network model in the plurality of the on-chip memory blocks until the service is interrupted or the neural network model is modified or replaced.

9. The method of claim 8 , wherein the node comprises a field programmable gate array (FPGA) and wherein each of the plurality of the on-chip memory blocks comprises a static random access memory block.

10. The method of claim 9 , wherein each of the plurality of compute units comprises a set of pre-configured resources on the FPGA.

11. The method of claim 8 , wherein the plurality of the on-chip memory blocks is arranged in rows and wherein each of the plurality of compute units is configured to process at least a subset of at least one of the rows per dock cycle.

12. The method of claim 8 , wherein the loading the N by M matrix of coefficients corresponding to the neural network model into the plurality of the on-chip memory blocks comprises streaming data corresponding to the N by M matrix of coefficients corresponding to the neural network model via a broadcast block into the plurality of the on-chip memory blocks.

13. The method of claim 12 , wherein, the streaming does not comprise loading any additional data corresponding to the neural network model from an off-chip memory in response to any operation associated with the N by M matrix of coefficients corresponding to the neural network model.

14. The method of claim 8 , wherein the N by M matrix of coefficients comprises a Long Short Term Memory (LSTM) weights matrix.

15. A system comprising a plurality of nodes interconnected via a network for evaluating a neural network model corresponding to a service, wherein each node comprises a plurality of on-chip memory blocks and a plurality of compute units, wherein each node is configured to:

upon service activation, receive an N by M matrix of coefficients corresponding to the neural network model, wherein N is an integer equal to or greater than 8 and M is an integer equal to or greater than 8;

load the N by M matrix of coefficients corresponding to the neural network model into the plurality of the on-chip memory blocks for processing by the plurality of compute units; and

regardless of a utilization of the plurality of the on-chip memory blocks as part of an evaluation of the neural network model, maintain the N by M matrix of coefficients corresponding to the neural network model in the plurality of the on-chip memory blocks until the service is interrupted or the neural network model is modified or replaced.

16. The system of claim 15 , wherein the node comprises a field programmable gate array (FPGA) and wherein each of the plurality of the on-chip memory blocks comprises a static random access memory block.

17. The system of claim 16 , wherein each of the plurality of compute units comprises a set of pre-configured resources on the FPGA.

18. The system of claim 15 , wherein the plurality of the on-chip memory blocks is arranged in rows and wherein each of the plurality of compute units is configured to process at least a subset of at least one of the rows per clock cycle.

19. The system of claim 15 , wherein the system is further configured to stream data corresponding to the N by M matrix of coefficients corresponding to the neural network model via a broadcast block into the plurality of the on-chip memory blocks.

20. The system of claim 15 , wherein, after the service activation, the system is further configured to not load any additional data corresponding to the neural network model from an off-chip memory in response to any operation associated with the N by M matrix of coefficients corresponding to the neural network model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2017
From: BURGER, DOUGLAS C.; CHUNG, ERIC S.; FOWERS, JEREMY; OVTCHAROV, KALIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 042867/0513 →
Continuity (2)
Provisional Application 62465063 · Feb 28, 2017
Related Publication 20180247190A1 · Aug 30, 2018