IP Library Granted Patent US 11,476,869
Granted Patent B2
US 11,476,869 · App. 15/953,330 · Granted Oct 18, 2022

Dynamically partitioning workload in a deep neural network module to reduce power consumption

Inventors: Amol Ashok Ambardekar (Redmond, WA); Boris Bobrov (Kirkland, WA); Chad Balling McBride (North Bend, WA); George Petre (Redmond, WA); Kent D. Cedola (Bellevue, WA); Larry Marvin Wall (Seattle, WA)
Assignee: Microsoft Technology Licensing, LLC
H03M7/3059G06F1/324G06F1/3275G06F3/0604G06F3/067G06F3/0631G06F9/30087G06F9/3836G06F9/3887G06F9/46G06F9/4881G06F12/0207G06F12/0238G06F12/08G06F12/0862G06F12/10G06F13/1673G06F13/1689G06F13/28G06F15/8007G06F17/15G06N3/04G06N3/049G06N3/0454G06N3/06G06N3/063G06N3/0635G06N3/08G06N3/10H03M7/6005H03M7/6011H03M7/70H04L45/04H04L67/02H04L67/1001G06F2209/484G06F2209/485G06F2212/657H03M7/46H04L45/50Y02D10/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,476,869
App. No.
15/953,330
Granted
Oct 18, 2022
Kind
B2
Abstract

A deep neural network (DNN) module is disclosed that can dynamically partition neuron workload to reduce power consumption. The DNN module includes neurons and a group partitioner and scheduler unit. The group partitioner and scheduler unit divides a workload for the neurons into partitions in order to maximize the number of neurons that can simultaneously process the workload. The group partitioner and scheduler unit then assigns a group of neurons to each of the partitions. The groups of neurons in the DNN module process the workload in their assigned partition to generate a partial output value. The neurons in each group can then sum their partial output values to generate a final output value for the workload. The neurons can be powered down once the groups of neurons have completed processing their assigned workload to reduce power consumption.

Claims (43)

1. A neural network processor, comprising:

a plurality of neurons; and

a group partitioner and scheduler configured to:

divide a workload for the neural network processor into a plurality of partitions based on a quantity of the plurality of neurons, and

assign a group of the neurons to each of the plurality of partitions to maximize a total number of the plurality of neurons that simultaneously process the workload while reducing power consumption; and

wherein the neurons within each group of neurons are configured to:

process the workload in an assigned partition to generate a partial output value by

performing a convolution operation on a partition containing a portion of an input volume and a portion of a weight volume where the partition comprises an input frame defined by a set of kernels, a number of channels per kernel, a height, and a width, and

performing the convolution operation on overlapping intervals defined by strides in two dimensions; and

sum partial output values generated by the neurons in each group of neurons to generate an output value for the workload.

2. The neural network processor of claim 1 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the depth dimension.

3. The neural network processor of claim 1 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the height dimension.

4. The neural network processor of claim 1 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the width dimension.

5. The neural network processor of claim 1 , wherein the workload is divided into a plurality of partitions such that the number of neurons that can simultaneously process the workload is maximized.

6. The neural network processor of claim 1 , wherein the plurality of neurons are powered down following generation of the output values for the workload.

7. A neural network processor, comprising:

a buffer storing an input volume and a weight volume;

a plurality of neurons; and

a group partitioner and scheduler configured to

partition the input volume and the weight volume into a plurality of partitions based on a quantity of the plurality of neurons, and

assign a group of the neurons to each of the plurality of partitions to maximize a total number of the plurality of neurons that simultaneously process a workload while reducing power consumption; and

wherein the neurons within each group of neurons are configured to:

process the workload in an assigned partition to generate a partial output value by

performing a convolution operation on a partition containing a portion of an input volume and a portion of a weight volume where the partition comprises an input frame defined by a set of kernels, a number of channels per kernel, a height, and a width, and

performing the convolution operation on overlapping intervals defined by strides in two dimensions; and

sum partial output values generated by the neurons in each group of neurons to generate an output value for the workload.

8. The neural network processor of claim 7 , wherein the workload is divided into a plurality of partitions such that the number of neurons that can simultaneously process the workload is maximized.

9. The neural network processor of claim 7 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the depth dimension.

10. The neural network processor of claim 7 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the height dimension.

11. The neural network processor of claim 7 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the width dimension.

12. The neural network processor of claim 7 , wherein the plurality of neurons are powered down following generation of the output values for the workload.

13. A computer-implemented method, comprising:

dividing a workload for a neural network processor into a plurality of partitions based on a quantity of neurons of the neural network processor;

assigning a group of neurons of the neural network processor to each of the plurality of partitions to maximize a total number of the group of neurons that simultaneously process the workload while reducing power consumption;

processing, by way of the group of neurons, the workload in an assigned partition to generate a partial output value by:

performing a convolution operation on a partition containing a portion of an input volume and a portion of a weight volume where the partition comprises an input frame defined by a set of kernels, a number of channels per kernel, a height, and a width, and

performing the convolution operation on overlapping intervals defined by strides in two dimensions; and

summing partial output values generated by the neurons in each group of neurons to generate an output value for the workload.

14. The computer-implemented method of claim 13 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the depth dimension.

15. The computer-implemented method of claim 13 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the height dimension.

16. The computer-implemented method of claim 13 , wherein the workload comprises an input volume and a weight volume having height, width, and depth dimensions, and wherein the workload is partitioned along the width dimension.

17. The computer-implemented method of claim 13 , wherein the workload is divided into a plurality of partitions such that the number of neurons that can simultaneously process the workload is maximized.

18. The computer-implemented method of claim 13 , further comprising powering down the group of neurons following generation of the output values for the workload.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2018
From: AMBARDEKAR, AMOL ASHOK; BOBROV, BORIS; MCBRIDE, CHAD BALLING; PETRE, GEORGE; CEDOLA, KENT D.; WALL, LARRY MARVIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 045539/0922 →
Continuity (2)
Provisional Application 62486432 · Apr 17, 2017
Related Publication 20180300616A1 · Oct 18, 2018
Cited By (1)
US 12,307,294