IP Library › Granted Patent US 11,100,391
Granted Patent B2
US 11,100,391 · App. 15/951,106 · Granted Aug 24, 2021

Power-efficient deep neural network module configured for executing a layer descriptor list

Inventors: Amol Ashok Ambardekar (Redmond, WA); Kent D. Cedola (Bellevue, WA); Larry Marvin Wall (Seattle, WA); Boris Bobrov (Kirkland, WA); George Petre (Redmond, WA); Chad Balling McBride (North Bend, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/063G06F1/324G06F1/3275G06F3/0604G06F3/067G06F3/0631G06F9/30087G06F9/3836G06F9/3887G06F9/46G06F12/0207G06F12/08G06F12/0862G06F12/10G06F13/1673G06F13/1689G06F13/28G06F15/8007G06F17/15G06N3/04G06N3/049G06N3/0454G06N3/06G06N3/0635G06N3/08G06N3/10H03M7/6005H03M7/6011H03M7/70H04L45/04H04L67/02H04L67/1002G06F2209/484G06F2209/485G06F2212/657H03M7/46H04L45/50Y02D10/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,100,391
App. No.
15/951,106
Granted
Aug 24, 2021
Kind
B2
Abstract

A deep neural network (DNN) processor is configured to execute descriptors in layer descriptor lists. The descriptors define instructions for performing a pass of a DNN by the DNN processor. Several types of descriptors can be utilized: memory-to-memory move (M2M) descriptors; operation descriptors; host communication descriptors; configuration descriptors; branch descriptors; and synchronization descriptors. A DMA engine uses M2M descriptors to perform multi-dimensional strided DMA operations. Operation descriptors define the type of operation to be performed by neurons in the DNN processor and the activation function to be used by the neurons. M2M descriptors are buffered separately from operation descriptors and can be executed at soon as possible, subject to explicitly set dependencies. As a result, latency can be reduced and, consequently, the neurons can complete their processing faster. The DNN module can then be powered down earlier than it otherwise would have, thereby saving power.

Claims (65)

1. A neural network processor, comprising:

one or more neurons;

a first memory device storing a layer descriptor list defining configurations for layers of the neural network processor, the layer descriptor list comprising

at least one memory-to-memory (M2M) descriptor, and

at least one operation descriptor;

a second memory device for storing data to be operated on by the one or more neurons; and

a controller configured to

retrieve the layer descriptor list and configuring at least one of the one or more neurons based on the layer descriptor list;

execute the at least one M2M descriptor to perform a M2M operation to transfer the data to be operated on by the one or more neurons from a memory of a host computing device to the second memory device, and

execute the at least one operation descriptor stored in the first memory device to cause the one or more neurons to perform an operation on the data in the second memory device.

2. The neural network processor of claim 1 , wherein the at least one operation descriptor comprises a field specifying the operation to be performed by the one or more neurons, and wherein the operation comprises

an additive combining operation,

a scalar multiply and add operation,

a convolution operation,

a deconvolution operation,

a max pooling operation, or

a fully connected layer operation.

3. The neural network processor of claim 1 , wherein the at least one operation descriptor comprises a field specifying a type of activation function to be used by the one or more neurons during the operation.

4. The neural network processor of claim 1 , wherein the layer descriptor list further comprises a branch descriptor which, when executed, will cause the controller to:

determine if a condition has been satisfied; and

responsive to determining the condition has been satisfied, cause execution of descriptors in the layer descriptor list to branch from a first descriptor to a second descriptor.

5. The neural network processor of claim 1 , wherein the layer descriptor list further comprises a synchronization descriptor which, when executed by the controller, will cause the controller to synchronize the one or more neurons.

6. The neural network processor of claim 1 , wherein the layer descriptor list further comprises a configuration descriptor which, when executed by the controller, modifies a configuration state of the neural network processor.

7. The neural network processor of claim 1 , wherein the layer descriptor list further comprises a host communication descriptor which, when executed by the controller, will cause the controller to transmit data to the host computing device.

8. A computer-implemented method, comprising:

storing a layer descriptor list in a memory of a neural network module, the layer descriptor list defining configurations for layers of a deep neural network, the layer descriptor list comprising

at least one memory-to-memory (M2M) descriptor, and

at least one operation descriptor;

retrieving the layer descriptor list and configuring at least one neuron of the neural network module based on the layer descriptor list;

executing, by the at least one neuron, the at least one M2M descriptor of the retrieved layer descriptor list to perform a M2M operation for obtaining data to be operated on by the one or more neurons from a memory of a host computing device; and

executing, by the at least one neuron, the at least one operation descriptor by way of the neural network module to cause the one or more neurons to perform an operation on the data.

9. The computer-implemented method of claim 8 , wherein the at least one operation descriptor comprises a field specifying the operation to be performed on the data by the one or more neurons, and wherein the operation comprises

an additive combining operation,

a scalar multiply and add operation,

a convolution operation,

a deconvolution operation,

a max pooling operation, or

a fully connected layer operation.

10. The computer-implemented method of claim 8 , wherein the at least one operation descriptor comprises a field specifying a type of activation function to be used by the one or more neurons.

11. The computer-implemented method of claim 8 , wherein the at least one operation descriptor comprises a field specifying a mathematical precision to be utilized by the operation.

12. The computer-implemented method of claim 8 , wherein the at least one operation descriptor comprises microcode for configuring the neural network module for performing the operation.

13. The computer-implemented method of claim 8 , wherein the layer descriptor list further comprises a host communication descriptor which, when executed by the controller, will cause the controller to interrupt or signal the host computing device and transmit data to the host computing device.

14. The computer-implemented method of claim 8 , wherein the layer descriptor list further comprises a branch descriptor which, when executed, will cause the neural network module to:

determine if a condition has been satisfied; and

responsive to determining the condition has been satisfied, cause execution of descriptors in the layer descriptor list to branch from a first descriptor to a second descriptor.

15. A neural network processor, comprising:

one or more neurons;

a first memory device storing a layer descriptor list comprising an ordered list of descriptors defining configurations for layers of a neural network; and

a controller configured to

execute a first descriptor in the layer descriptor list to obtain data to be operated on by the one or more neurons, and

execute a second descriptor in the layer descriptor list to cause the one or more neurons to perform an operation on the data.

16. The neural network processor of claim 15 , wherein the first descriptor comprises a field specifying the operation, and wherein the operation comprises

an additive combining operation,

a scalar multiply and add operation,

a convolution operation,

a deconvolution operation,

a max pooling operation, or

a fully connected layer operation.

17. The neural network processor of claim 15 , wherein the first descriptor comprises a field specifying a type of activation function to be used by the one or more neurons.

18. The neural network processor of claim 15 , wherein the layer descriptor list further comprises a descriptor which, when executed by the controller, will cause the controller to:

determine if a condition has been satisfied; and

branch the execution of the descriptors in the layer descriptor list responsive to determining the condition has been satisfied.

19. The neural network processor of claim 15 , wherein the layer descriptor list further comprises a descriptor which, when executed by the controller, modifies a configuration state of the neural network processor.

20. The neural network processor of claim 15 , wherein the layer descriptor list further comprises a descriptor which, when executed by the controller, will cause the

controller to synchronize the one or more neurons.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2018
From: AMBARDEKAR, AMOL ASHOK; CEDOLA, KENT D.; WALL, LARRY MARVIN; BOBROV, BORIS; PETRE, GEORGE; MCBRIDE, CHAD BALLING
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 045511/0970 →
Continuity (2)
Provisional Application 62486432 · Apr 17, 2017
Related Publication 20180300614A1 · Oct 18, 2018