IP Library Granted Patent US 11,615,304
Granted Patent B1
US 11,615,304 · App. 16/809,644 · Granted Mar 28, 2023

Quantization aware training by constraining input

Inventor: Malhar Palkar (Cupertino, CA)
Assignee: Ambarella International LP
G06N3/08G06N3/04G06V10/82G06V20/52G06T2207/20084G06T2207/30232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,615,304
App. No.
16/809,644
Granted
Mar 28, 2023
Kind
B1
Abstract

A method of generating a quantized neural network comprises (i) receiving a neural network model, (ii) modifying the neural network model by quantizing input of at least convolution layers of the neural network model, and (iii) training the modified neural network model using a dataset that is representative of one or more desired inferences.

Claims (25)

1. A method of generating a quantized neural network comprising:

receiving a neural network model comprising a plurality of layers having a first input and output container format;

modifying the neural network model by quantizing inputs of at least convolution layers of the neural network model from said first input and output container format to a second container format based on a range of input data defined by a minimum input value and a maximum input value; and

training the modified neural network model using a dataset that is representative of one or more desired inferences.

2. The method according to claim 1 , wherein said neural network model comprises a directed acyclic graph.

3. The method according to claim 1 , wherein modifying the neural network model comprises inserting pseudo-quantization operations to quantize the inputs of the convolution layers of the neural network model.

4. The method according to claim 3 , wherein said pseudo-quantization operations comprise a gradient that allows training of said minimum input value and said maximum input value defining said range of input data for quantizing said inputs of said convolution layers of the neural network model.

5. The method according to claim 4 , wherein said minimum input value and said maximum input value for quantizing said inputs of said convolution layers of the neural network model are determined from said dataset used to train said neural network model.

6. The method according to claim 4 , wherein said pseudo-quantization operations take into account a data container format of an edge device on which said neural network model will be executed.

7. The method according to claim 3 , wherein said training comprises a quantization aware training process.

8. The method according to claim 1 , wherein said convolution layers are implemented with Conv2D class operators.

9. The method according to claim 1 , further comprising programming at least one edge device with a quantized neural network model and weights determined during the training of the modified neural network model.

10. The method according to claim 9 , wherein programming the at least one edge device comprises burning the quantized neural network model and the weights into a die of the at least one edge device.

11. The method according to claim 1 , wherein the quantized neural network generates one or more inferences about an input by performing one or more computer vision operations.

12. The method according to claim 11 , wherein said input is generated by a sensor.

13. The method according to claim 12 , wherein said sensor comprises a video camera.

14. The method according to claim 12 , wherein said sensor and said quantized neural network are configured as part of an edge device.

15. The method according to claim 12 , wherein said sensor and said quantized neural network are configured as part of at least one of a battery-powered device or a battery-powered security camera.

16. An apparatus comprising:

a sensor to generate a data input; and

a processor to generate one or more outputs in response to said data input based upon one or more inferences made by executing a pre-trained neural network model comprising a plurality of layers having a first input and output container format, wherein said pre-trained neural network model was trained by (i) modifying a neural network model by quantizing inputs of at least convolution layers of the neural network model from said first input and output container format to a second container format based on a range of input data defined by a minimum input value and a maximum input value and (ii) training the modified neural network model using a dataset that is representative of one or more desired inferences.

17. The apparatus according to claim 16 , wherein said sensor comprises a video camera and the pre-trained neural network model generates the one or more inferences about said data input by performing one or more computer vision operations.

18. The apparatus according to claim 16 , wherein said sensor and said processor are configured as part of an edge device.

19. The apparatus according to claim 16 , wherein said sensor and said processor are configured as part of a battery-powered device.

20. The apparatus according to claim 16 , wherein said sensor and said processor are configured as part of a battery-powered security camera.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2020
From: PALKAR, MALHAR
To: AMBARELLA INTERNATIONAL LP
Reel/Frame 052225/0186 →
Cited By (2)
US 12,632,711 US 12,737,609