IP Library › Granted Patent US 11,816,568
Granted Patent B2
US 11,816,568 · App. 17/016,908 · Granted Nov 14, 2023

Optimizing execution of a neural network based on operational performance parameters

Inventors: Sek Meng Chai (Princeton, NJ); Jagadeesh Kandasamy (Cupertino, CA)
Assignee: Latent AI, Inc.
G06N3/08G06F16/9024G06F17/18G06N3/04G06N3/10G06N5/04H04L41/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,816,568
App. No.
17/016,908
Granted
Nov 14, 2023
Kind
B2
Abstract

The disclosed embodiments relate to a system that optimizes execution of a DNN based on operational performance parameters. During operation, the system collects the operational performance parameters from the DNN during operation of the DNN, wherein the operational performance parameters include parameters associated with operating conditions for the DNN, parameters associated with resource utilization during operation of the DNN, and parameters associated with accuracy of results produced by the DNN. Next, the system uses the operational performance parameters to update the DNN model to improve performance and efficiency during execution of the DNN.

Claims (40)

1. A method for optimizing execution of a deep neural network (DNN) based on operational performance parameters, comprising:

collecting the operational performance parameters from the DNN during operation of the DNN, wherein:

the operational performance parameters include parameters associated with operating conditions for the DNN, parameters associated with resource utilization during operation of the DNN, and parameters associated with accuracy of results produced by the DNN; and

the operational performance parameters further include profiling data, which identifies pathways within the DNN that are activated while the DNN performs inference-processing operations; and

using the operational performance parameters to update the DNN model to improve performance and efficiency during execution of the DNN.

2. The method of claim 1 , wherein the method further comprises deploying and executing the updated DNN model at a location in a hierarchy of computing nodes, wherein the location is determined based on a global system-level optimization.

3. The method of claim 2 , wherein the operational performance parameters include information that is used to optimize overall network bandwidth within the hierarchy of computing nodes in which the DNN operates.

4. The method of claim 1 , wherein the profiling data is used to synthesize additional training data, which is used to train the updated DNN model to improve robustness.

5. The method of claim 1 , wherein while executing the DNN, a runtime engine for the DNN selectively activates pathways in the DNN to facilitate computationally efficient inference-processing operations.

6. The method of claim 1 , wherein the operational performance parameters are analyzed to determine coefficients for regularizer terms in a loss function that is used to train the updated DNN model, wherein the regularizer terms include a quantization term, which represents differences between pre-quantization and post-quantization weight values in the DNN, and a magnitude term, which represents magnitudes of the weight values.

7. The method of claim 1 , wherein a runtime engine for the DNN uses a policy generated using the operational performance parameters to achieve two or more of the following objectives:

maximizing classification accuracy of the DNN;

minimizing computational operations performed while executing the DNN;

minimizing power consumption of a device, which is executing the DNN; and

minimizing latency involved in executing the DNN to produce an output.

8. The method of claim 1 , wherein the updated DNN model is comprised of a plurality of DNN models trained simultaneously based on the operational performance parameters.

9. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for optimizing execution of a deep neural network (DNN) based on operational performance parameters, the method comprising:

collecting the operational performance parameters from the DNN during operation of the DNN, wherein:

the operational performance parameters include parameters associated with operating conditions for the DNN, parameters associated with resource utilization during operation of the DNN, and parameters associated with accuracy of results produced by the DNN; and

the operational performance parameters further include profiling data, which identifies pathways within the DNN that are activated while the DNN performs inference-processing operations; and

using the operational performance parameters to update the DNN model to improve performance and efficiency during execution of the DNN.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the method further comprises deploying and executing the updated DNN model at a location in a hierarchy of computing nodes, wherein the location is determined based on a global system-level optimization.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the operational performance parameters include information that is used to optimize overall network bandwidth within the hierarchy of computing nodes in which the DNN operates.

12. The non-transitory computer-readable storage medium of claim 9 , wherein the profiling data is used to synthesize additional training data, which is used to train the updated DNN model to improve robustness.

13. The non-transitory computer-readable storage medium of claim 9 , wherein while executing the DNN, a runtime engine for the DNN selectively activates pathways in the DNN to facilitate computationally efficient inference-processing operations.

14. The non-transitory computer-readable storage medium of claim 9 , wherein the operational performance parameters are analyzed to determine coefficients for regularizer terms in a loss function that is used to train the updated DNN model, wherein the regularizer terms include a quantization term, which represents differences between pre-quantization and post-quantization weight values in the DNN, and a magnitude term, which represents magnitudes of the weight values.

15. The non-transitory computer-readable storage medium of claim 9 , wherein a runtime engine for the DNN uses a policy generated using the operational performance parameters to achieve two or more of the following objectives:

maximizing classification accuracy of the DNN;

minimizing computational operations performed while executing the DNN;

minimizing power consumption of a device, which is executing the DNN; and

minimizing latency involved in executing the DNN to produce an output.

16. The non-transitory computer-readable storage medium of claim 9 , wherein the updated DNN model is comprised of a plurality of DNN models trained simultaneously based on the operational performance parameters.

17. A system that optimizes execution of a deep neural network (DNN) based on operational performance parameters, comprising:

at least one processor and at least one associated memory; and

a processing mechanism that executes on the at least one processor, wherein during operation, the processing mechanism:

collects the operational performance parameters from the DNN during operation of the DNN, wherein:

the operational performance parameters include parameters associated with operating conditions for the DNN, parameters associated with resource utilization during operation of the DNN, and parameters associated with accuracy of results produced by the DNN; and

the operational performance parameters further include profiling data, which identifies pathways within the DNN that are activated while the DNN performs inference-processing operations; and

uses the operational performance parameters to update the DNN model to improve performance and efficiency during execution of the DNN.

18. The system of claim 17 , wherein the processing mechanism also deploys and executes the updated DNN model at a location in a hierarchy of computing nodes, wherein the location is determined based on a global system-level optimization.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2020
From: CHAI, SEK MENG; KANDASAMY, JAGADEESH
To: LATENT AI, INC.
Reel/Frame 053860/0239 →
Continuity (3)
Provisional Application 63018236 · Apr 30, 2020
Provisional Application 62900311 · Sep 13, 2019
Related Publication 20210081789A1 · Mar 18, 2021