Optimized placement for efficiency for accelerated deep learning
Techniques in optimized placement for efficiency for accelerated deep learning provide improvements in one or more of accuracy, performance, and energy efficiency. An array of processing elements comprising a portion of a neural network accelerator performs flow-based computations on wavelets of data. Each processing element comprises a compute element to execute programmed instructions using the data and a router to route the wavelets. The routing is in accordance with virtual channel specifiers of the wavelets and controlled by routing configuration information of the router. A software stack determines optimized placement based on a description of a neural network. The determined placement is used to configure the routers including usage of the respective colors. The determined placement is used to configure the compute elements including the respective programmed instructions each is configured to execute.
1 . A method comprising:
extracting a model from a neural network description;
computing delays based on convergent nodes of the extracted model;
determining routing to implement data communication based on arcs of the extracted model;
determining accelerator configuration information usable to configure a deep learning accelerator to provide a trained model, wherein the accelerator configuration information indicates delay buffer placement based on the delays and on the routing, and wherein the deep learning accelerator comprises a fabric and a plurality of processing elements enabled to communicate packets with each other via the fabric in accordance with a plurality of communication pathways identifiable by respective virtual channel identifiers; and
configuring the deep learning accelerator based on the accelerator configuration information.
2 . The method of claim 1 , wherein the determining the routing ignores interactions between routes.
3 . The method of claim 2 , further comprising scanning results based on the determining the routing to produce hotspot information to repeat the determining the routing in accordance therewith.
4 . The method of claim 1 , wherein the determining the routing ignores coloring and bandwidth interactions with other routes.
5 . A non-transitory computer-readable medium comprising one or more instructions encoded thereon that, when executed by one or more processors, cause the one or more processors to perform actions comprising:
extracting a model from a neural network description;
computing delays based on convergent nodes of the extracted model;
determining routing to implement data communication based on arcs of the extracted model;
determining accelerator configuration information usable to configure a deep learning accelerator to provide a trained model, wherein the accelerator configuration information indicates delay buffer placement based on the delays and on the routing, and wherein the deep learning accelerator comprises a fabric and a plurality of processing elements enabled to communicate packets with each other via the fabric in accordance with a plurality of communication pathways identifiable by respective virtual channel identifiers; and
configuring the deep learning accelerator based on the accelerator configuration information.
6 . The non-transitory computer-readable medium of claim 5 , wherein the determining the routing ignores interactions between routes.
7 . The non-transitory computer-readable medium of claim 6 , further comprising scanning results based on the determining the routing to produce hotspot information to repeat the determining the routing in accordance therewith.
8 . The non-transitory computer-readable medium of claim 5 , wherein the determining the routing ignores coloring and bandwidth interactions with other routes.
9 . A deep learning accelerator comprising:
a fabric; and
circuitry configured as a plurality of processing elements that is enabled to communicate packets with each other via the fabric in accordance with a plurality of communication pathways identifiable by respective virtual channel identifiers, wherein the circuitry is configured to:
extract a model from a neural network description;
compute delays based on convergent nodes of the extracted model;
determine routing to implement data communication based on arcs of the extracted model;
determine accelerator configuration information usable to configure the deep learning accelerator to provide a trained model, wherein the accelerator configuration information indicates delay buffer placement based on the delays and on the routing; and
configure the deep learning accelerator based on the accelerator configuration information.
10 . The system of claim 9 , wherein the circuitry, when determining the routing, is configured to ignore interactions between routes.
11 . The system of claim 10 , wherein the circuitry is further configured to scan results based on the determining the routing to produce hotspot information to repeat the determining the routing in accordance therewith.
12 . The system of claim 9 , wherein the circuitry, when determining the routing, is configured to ignore coloring and bandwidth interactions with other routes.