IP Library Granted Patent US 12,079,723
Granted Patent B2
US 12,079,723 · App. 18/183,515 · Granted Sep 3, 2024

Optimizing neural network structures for embedded systems

Inventors: Harsimran Singh Sidhu (Fremont, CA); Paras Jagdish Jain (Cupertino, CA); Daniel Paden Tomasello (Los Altos Hills, CA); Forrest Nelson Iandola (San Jose, CA)
Assignee: Tesla, Inc.
G06N3/08G05B13/027G05D1/0088G05D1/0214G05D1/0221G06F9/45533G06N3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,079,723
App. No.
18/183,515
Granted
Sep 3, 2024
Kind
B2
Abstract

A model training and implementation pipeline trains models for individual embedded systems. The pipeline iterates through multiple models and estimates the performance of the models. During a model generation stage, the pipeline translates the description of the model together with the model parameters into an intermediate representation in a language that is compatible with a virtual machine. The intermediate representation is agnostic or independent to the configuration of the target platform. During a model performance estimation stage, the pipeline evaluates the performance of the models without training the models. Based on the analysis of the performance of the untrained models, a subset of models is selected. The selected models are then trained and the performance of the trained models are analyzed. Based on the analysis of the performance of the trained models, a single model is selected for deployment to the target platform.

Claims (22)

1. A method for generating a machine-learned model comprising:

generating an untrained model;

generating an intermediate representation of the untrained model, the intermediate representation in an intermediate language compatible with a virtual machine;

evaluating the performance of the untrained model, wherein evaluating the performance includes at least one of determining a latency in applying the untrained model in a target system, determining a frequency at which the untrained model can be applied in the target system, determining an amount of resources used by the untrained model, and determining an amount of power consumed by the target system using the untrained model;

iteratively generating and evaluating new untrained models, a new untrained model generated based on a performance of a previous model;

selecting a subset of models based on a performance of the generated new untrained models;

training the selected subset of models, thereby generating trained models;

evaluating an accuracy for each of the trained models; and

selecting a trained model based on the performance evaluation of the trained models for deployment to the target system.

2. The method of claim 1 , wherein determining an amount of resources used by the untrained model comprises:

determining a number of floating point operations used by the untrained model when implemented with default kernels.

3. The method of claim 1 , wherein determining an amount of resources used by the untrained model comprises:

determining a number of floating point operations used by the untrained model when implemented with optimized kernels.

4. The method of claim 1 , wherein determining an amount of resources used by the untrained model comprises:

determining a total amount of memory used by the untrained model; and

determining a total memory bandwidth used by the untrained model.

5. The method of claim 1 , wherein determining an amount of resources used by the untrained model comprises:

determining an amount of memory used by the untrained model after parameters and variables used by the untrained model have been scheduled; and

determining a memory bandwidth used by the untrained model after the parameters and variables used by the untrained model have been scheduled and after operations of the untrained model have been scheduled.

6. The method of claim 1 , further comprising generating an intermediate representation of the trained models.

7. The method of claim 1 , wherein selecting the subset of models comprises electing a first subset of models that perform with at least a specified performance.

8. The method of claim 4 , further comprising reducing a number of untrained models based on heuristics to identify the models to be trained.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: SIDHU, HARSIMRAN SINGH; JAIN, PARAS JAGDISH; TOMASELLO, DANIEL PADEN; IANDOLA, FORREST NELSON
To: DEEPSCALE, INC.
Reel/Frame 067554/0965 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: DEEPSCALE, INC.
To: TESLA, INC.
Reel/Frame 067555/0148 →
Continuity (3)
Division 16522411 · Jul 25, 2019
Provisional Application 62703837 · Jul 26, 2018
Related Publication 20230289599A1 · Sep 14, 2023