IP Library › Granted Patent US 12,632,712
Granted Patent B2
US 12,632,712 · App. 17/961,453 · Granted May 19, 2026

Method and electronic device for quantizing DNN model

Inventors: Tejpratap Venkata Subbu Lakshmi Gollanapalli (Karnataka, IN); Arun Abraham (Karnataka, IN); Raja Kumar (Karnataka, IN); Pradeep Nelahonne Shivamurthappa (Karnataka, IN); Vikram Nelvoy Rajendiran (Karnataka, IN); Prasen Kumar Sharma (Karnataka, IN)
Assignee: Samsung Electronics Co., Ltd.
G06N3/0495G06N3/045G06N3/08G06N3/084G06V10/82G06N3/063G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,712
App. No.
17/961,453
Granted
May 19, 2026
Kind
B2
Abstract

Various embodiments of the disclosure disclose a method for quantizing a Deep Neural Network (DNN) model in an electronic device. The method includes: estimating, by the electronic device, an activation range of each layer of the DNN model using self-generated data (e.g. retro image, audio, video, etc.) and/or a sensitive index of each layer of the DNN model; quantizing, by the electronic device, the DNN model based on the activation range and/or the sensitive index; and allocating, by the electronic device, a dynamic bit precision for each channel of each layer of the DNN model to quantize the DNN model.

Claims (44)

1 . A method for quantizing a Deep Neural Network (DNN) model in an electronic device comprising a quantization engine, wherein the quantization engine comprises processing circuitry and/or executable program instructions, operably connected to the memory and the processor, the method comprising:

estimating, by the quantization engine, at least one of an activation range of each layer of the DNN model using self-generated data and a sensitive index of each layer of the DNN model; and

quantizing, by the quantization engine, the DNN model based on the at least one of the activation range and the sensitive index

wherein estimating, by the quantization engine, the activation range of each layer of the DNN model using the self-generated data comprises:

determining, by the quantization engine, a plurality of random images, wherein each random image of the plurality of random images comprises uniform distribution data across the images;

passing, by the quantization engine, each random image into the DNN model;

determining, by the quantization engine, weight distributions of the DNN model for each random image after each layer of the DNN model;

determining, by the quantization engine, layer statistics of the DNN model for each random image after each layer of the DNN model, wherein the layer statistics of the DNN model comprises at least one of a mean and a variance;

determining, by the quantization engine, a difference between pre-stored layer statistics of the DNN model and the determined layer statistics of the DNN model;

determining, by the quantization engine, whether the difference is less than a threshold; and

performing, by the quantization engine, one of:

generating the data using at least one of the layer statistics of the DNN model and the weight distributions of the DNN model in response to determining that the difference is less than the threshold, or

executing back propagation in the DNN model in response to determining that the difference is greater than or equal to the threshold.

2 . The method as claimed in claim 1 , wherein the self-generated data is generated based on at least one of layer statistics of the DNN model and weight distributions of the DNN model.

3 . The method as claimed in claim 1 , wherein the DNN model quantizes weights and activation at lower bit precision without access to at least one of training dataset and validation dataset to obtain a compression of the DNN model and a fast inference of the DNN model.

4 . The method as claimed in claim 1 , wherein the self-generated data is a plurality of retro data images, wherein the plurality of retro data images are equivalent or represent all features of the DNN model.

5 . The method as claimed in claim 1 , wherein a Z-score provides a difference between two distributions.

6 . The method as claimed in claim 1 , wherein the self-generated data is a plurality of retro data images, wherein the plurality of retro data images are equivalent or represent all features of the DNN model.

7 . The method as claimed in claim 1 , wherein estimating, by the quantization engine, the sensitive index of each layer of the DNN model comprises:

determining, by the quantization engine, an accuracy using per-channel quantization and per-tensor quantization schemes at each layer of the DNN model; and

determining, by the quantization engine, a sensitive index of each layer corresponding to per-channel quantization and per-tensor quantization.

8 . The method as claimed in claim 1 , wherein quantizing, by the quantization engine, the DNN model based the sensitive index comprises:

determining, by the quantization engine, an optimal quantization for each layer of the DNN model based on the sensitive index to quantize the DNN model, wherein the sensitive index comprises a minimum sensitivity; and

applying, by the quantization engine, the optimal quantization for each layer of the DNN model.

9 . The method as claimed in claim 1 , wherein the sensitive index is determined a using a Kullback-Leibler Divergence.

10 . The method as claimed in claim 8 , wherein the optimal quantization comprises at least one of per-channel quantization, and per-tensor quantization for the DNN model.

11 . The method as claimed in claim 1 , wherein the sensitive index is used to reduce a search space of each layer of the DNN model from exponential to linear, wherein n is a number of layers in the DNN model.

12 . The method as claimed in claim 1 , wherein the method comprises:

allocating, by the quantization engine, a dynamic bit precision for each channel of each layer of the DNN model, wherein the allocated dynamic bit precision minimizes and/or reduces overall quantization noise of each layer.

13 . An electronic device configured to quantize a Deep Neural Network (DNN) model, the electronic device comprising:

a memory;

a processor comprising processing circuitry; and

a quantization engine comprising processing circuitry and/or executable program instructions, operably connected to the memory and the processor, configured to:

estimate at least one of an activation range of each layer of the DNN model using self-generated data and a sensitive index of each layer of the DNN model;

determine a plurality of random images, wherein each random image of the plurality of random images comprises uniform distribution data across the images;

pass each random image into the DNN model;

determine weight distributions of the DNN model for each random image after each layer of the DNN model;

determine layer statistics of the DNN model for each random image after each layer of the DNN model, wherein the layer statistics of the DNN model comprises at least one of a mean and a variance;

determine a difference between pre-stored layer statistics of the DNN model and the determined layer statistics of the DNN model;

determine whether the difference is less than a threshold;

perform, one of:

generate the data using at least one of the layer statistics of the DNN model and the weight distributions of the DNN model in response to determining that the difference is less than the threshold; or

execute back propagation in the DNN model in response to determining that the difference is greater than or equal to the threshold; and

quantize the DNN model based on the at least one of the activation range and the sensitive index.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2022
From: GOLLANAPALLI, TEJPRATAP VENKATA SUBBU LAKSHMI; ABRAHAM, ARUN; KUMAR, RAJA; NELAHONNE SHIVAMURTHAPPA, PRADEEP; RAJENDIRAN, VIKRAM NELVOY; SHARMA, PRASEN KUMAR
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061341/0007 →
Priority Claims (3)
IN 202041044037 · Apr 9, 2021 · national
IN 202141039024 · Aug 27, 2021 · national
IN 202041044037 · Jan 25, 2022 · national
Continuity (2)
Continuation PCTKR2022005122 · Apr 8, 2022
Related Publication 20230068381A1 · Mar 2, 2023
References Cited (45)
US 11494657B2 · Sather · 2022 [cited by examiner]
US 11568251B1 · Palkar · 2023 [cited by examiner]
US 20160328646A1 · Lin · 2016 [cited by examiner]
US 20180046896A1 · Yu · 2018 [cited by examiner]
US 20190122116A1 · Choi · 2019 [cited by examiner]
US 20190251444A1 · Alakuijala et al. · 2019 [cited by applicant]
US 20190354842A1 · Louizos · 2019 [cited by examiner]
US 20190354865A1 · Reisser · 2019 [cited by examiner]
US 20200026986A1 · Ha · 2020 [cited by examiner]
US 20200097823A1 · Chen · 2020 [cited by examiner]
US 20200218962A1 · Lee et al. · 2020 [cited by applicant]
US 20200234112A1 · Wang et al. · 2020 [cited by applicant]
US 20200242473A1 · De Vangel · 2020 [cited by examiner]
US 20210089922A1 · Lu · 2021 [cited by examiner]
US 20210133278A1 · Fang · 2021 [cited by examiner]
US 20210201117A1 · Ha · 2021 [cited by examiner]
US 20210218414A1 · Malhotra · 2021 [cited by examiner]
US 20210224658A1 · Mathew · 2021 [cited by examiner]
US 20220044109A1 · Donnelly · 2022 [cited by examiner]
US 20220044114A1 · Sriram · 2022 [cited by examiner]
US 20220083855A1 · Choi · 2022 [cited by examiner]
US 20220101118A1 · Yan · 2022 [cited by examiner]
US 20220101133A1 · Ardywibowo · 2022 [cited by examiner]
US 20220164666A1 · Liu · 2022 [cited by examiner]
US 20220343162A1 · Lee · 2022 [cited by examiner]
CN 108537322A · 2018 [cited by examiner]
CN 111027684 · 2020 [cited by applicant]
CN 111783961 · 2020 [cited by applicant]
CN 111931906 · 2020 [cited by applicant]
CN 112016674 · 2020 [cited by applicant]
KR 1020200086581 · 2020 [cited by applicant]
WO WO2020160787A1 · 2020 [cited by examiner]
Cohen et al., “Lightweight Compression Of Neural Network Feature Tensors for Collaborative Intelligence,” 2020 IEEE International Conference on Multimedia and Expo (ICME), London, UK, 2020, pp. 1-6, doi:10.1109/ICME4628… [cited by examiner]
Ding et al., “Quantized deep neural networks for energy efficient hardware-based inference,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jeju, Korea (South), 2018, pp. 1-8, doi:10.1109/ASPDA… [cited by examiner]
Hubara et al., “Improving post training neural quantization: Layer-wise calibration and integer programming.” arXiv preprint arXiv: 2006.10518, 2020, doi:10.48550/arXiv.2006.10518. (Year: 2020). [cited by examiner]
Chen et al., “Deep neural network quantization via layer-wise optimization using limited training data.” Proceedings of the AAAI Conference on Artificial Intelligence. vol. 33. No. 01., 2019, doi:10.1609/aaai.v33i01.330… [cited by examiner]
Qiu et al., “Deep Quantization: Encoding Convolutional Activations with Deep Generative Model,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 4085-4094, doi:10.1109… [cited by examiner]
Seo et al., “Efficient Weights Quantization of Convolutional Neural Networks Using Kernel Density Estimation based Non-uniform Quantizer” Applied Sciences 9, No. 12: 2559, 2019, doi:10.3390/app9122559. (Year: 2019). [cited by examiner]
Amir Gholami, et al., “A Survey of Quantization Methods for Efficient Neural Network Inference”, arXiv:2103.13630v1, Mar. 2021, 29 pages. [cited by applicant]
Tej Pratap Gvsl, et al., “Hybrid and Non-Uniform Quantization Methods Using Retro Synthesis Data for Efficient Inference”, arXiv:2012.13716v1, Dec. 2020, 14 pages. [cited by applicant]
Raghuraman Krishnamoorthi, “Quantizing deep convolutional networks for efficient inference: A whitepaper”, arXiv preprint arXiv:1806.08342, Jun. 2018, 36 pages. [cited by applicant]
Dongqing Zhang, et al., “LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks”, In the European Conference on Computer Vision (ECCV), Sep. 2018, 18 pages. [cited by applicant]
International Search Report for PCT/KR2022/005122 dated Jul. 22, 2022, 4 pages. [cited by applicant]
Written Opinion of the ISA for PCT/KR2022/005122 dated Jul. 22, 2022, 4 pages. [cited by applicant]
Examination Report dated Oct. 28, 2022 issued by the India Patent Office for Indian Patent Application No. 202041044037. [cited by applicant]