IP Library › Granted Patent US 12,591,776
Granted Patent B2
US 12,591,776 · App. 17/520,326 · Granted Mar 31, 2026

Electronic apparatus and method for controlling thereof

Inventors: Dongsoo Lee (Suwon-si, KR); Sejung Kwon (Suwon-si, KR); Byeoungwook Kim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/082G06F18/40G06V10/751
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,776
App. No.
17/520,326
Filed
Nov 5, 2021
Granted
Mar 31, 2026
Kind
B2
Art Unit
2128
USPC
706/15
Abstract

An electronic apparatus, including a memory configured to store weight data used for computation of a neural network model; and a processor configured to: identify, from among weight values included in the weight data, at least one weight value having a size less than or equal to a threshold value, quantize remaining weight values other than the identified at least one weight value to obtain first quantized data including quantized values corresponding to the remaining weight values, identify, from among the quantized values, a quantized value closest to a predetermined value, obtain second quantized data including a quantized value corresponding to the at least one weight value based on the quantized value closest to the predetermined value, and store the first quantized data and the second quantized data in the memory.

Claims (68)

1 . An electronic apparatus comprising:

a memory configured to store weight data used for computation of a neural network model; and

a processor configured to:

identify, from among weight values included in the weight data, at least one weight value having a size less than or equal to a threshold value,

quantize remaining weight values other than the identified at least one weight value to obtain first quantized data including quantized values corresponding to the remaining weight values,

identify, from among the quantized values, a quantized value closest to a predetermined value,

obtain second quantized data including a quantized value corresponding to the at least one weight value based on the quantized value closest to the predetermined value,

store the first quantized data and the second quantized data in the memory,

quantize the neural network model by modifying the at least one weight value based on the second quantized data and modifying the remaining weight values based on the first quantized data to obtain a quantized neural network model, and

obtain output data by providing input data to the quantized neural network model.

2 . The electronic apparatus of claim 1 , wherein the processor is further configured to identify the quantized value closest to the predetermined value as the quantized value corresponding to the at least one weight value.

3 . The electronic apparatus of claim 1 , wherein the processor is further configured to:

based on the quantization of the remaining weight values, obtain the first quantized data including a plurality of scaling factors and bit values corresponding to the remaining weight values,

obtain a plurality of computational values based on computation using the plurality of scaling factors,

identify, from among the plurality of computational values, a computational value closest to the predetermined value, and

obtain the second quantized data based on the identified computational value.

4 . The electronic apparatus of claim 3 , wherein the processor is further configured to identify the computational value closest to the predetermined value as the quantized value corresponding to the at least one weight value.

5 . The electronic apparatus of claim 3 , wherein the processor is further configured to:

obtain the plurality of computational values by adding values obtained by multiplying each of the plurality of scaling factors by +1 or −1, and

identify the computational value closest to the predetermined value among the plurality of computational values.

6 . The electronic apparatus of claim 3 , wherein the processor is further configured to:

identify a first plurality of computations which output the computational value closest to the predetermined value among a second plurality of computations using the plurality of scaling factors,

identify, from among the first plurality of computations, a computation which outputs a computational value having a same code as a code of the at least one weight value, and

obtain the second quantized data based on the identified computation.

7 . The electronic apparatus of claim 3 , wherein the processor is further configured to:

identify codes of the plurality of scaling factors included in a computation, among a plurality of computations using the plurality of scaling factors, to output the computational value closest to the predetermined value, and

obtain the second quantized data based on the codes.

8 . The electronic apparatus of claim 1 , wherein the processor is further configured to:

receive a user input to set a pruning rate, and

identify the at least one weight value based on sizes of the weight values.

9 . The electronic apparatus of claim 1 , wherein the predetermined value is zero.

10 . A method for controlling an electronic apparatus, the method comprising:

identifying, from among weight values included in weight data, at least one weight value having a size less than or equal to a threshold value;

quantizing remaining weight values other than the identified at least one weight value to obtain first quantized data including quantized values corresponding to the remaining weight values;

identifying, from among the quantized values, a quantized value closest to a predetermined value;

obtaining second quantized data including a quantized value corresponding to the at least one weight value based on the quantized value closest to the predetermined value; and

storing the first quantized data and the second quantized data,

quantizing the neural network model by modifying the at least one weight value based on the second quantized data and modifying the remaining weight values based on the first quantized data to obtain a quantized neural network model, and

obtaining output data by providing input data to the quantized neural network model.

11 . The method of claim 10 , wherein the obtaining the second quantized data comprises identifying the quantized value closest to the predetermined value as the quantized value corresponding to the at least one weight value.

12 . The method of claim 10 , wherein the obtaining the first quantized data comprises, based on the quantization of the remaining weight values, obtaining the first quantized data including a plurality of scaling factors and bit values corresponding to the remaining weight values, and

wherein the obtaining the second quantized data comprises:

obtaining a plurality of computational values based on computation using the plurality of scaling factors;

identifying, from among the plurality of computational values, a computational value closest to the predetermined value; and

obtaining the second quantized data based on the identified computational value.

13 . The method of claim 12 , wherein the obtaining the second quantized data comprises identifying the computational value closest to the predetermined value as the quantized value corresponding to the at least one weight value.

14 . The method of claim 12 , wherein the identifying the computational value comprises:

obtaining the plurality of computational values by adding values obtained by multiplying each of the plurality of scaling factors by +1 or −1; and

identifying the computational value closest to the predetermined value among the plurality of computational values.

15 . The method of claim 12 , wherein the obtaining the second quantized data comprises:

identifying a first plurality of computations which output the computational value closest to the predetermined value among a second plurality of computations using the plurality of scaling factors;

identifying, from among the first plurality of computations, a computation which outputs a computational value having a same code as a code of the at least one weight value; and

obtaining the second quantized data based on the identified computation.

16 . A method for controlling an electronic apparatus, the method comprising:

obtaining a weight values included in weight data of a neural network model;

selecting, from among the weight values, a first weight value to be pruned from the weight values;

quantizing remaining weight values other than the first weight value to obtain quantized weight values corresponding to the remaining weight values;

identifying, from among the quantized weight values, a smallest quantized weight value;

based on the smallest quantized weight value, determining a first quantized weight value corresponding to the first weight value;

storing the quantized weight values and the first quantized weight value;

quantize the neural network model by modifying the at least one weight value based on the second quantized data and modifying the remaining weight values based on the first quantized data to obtain a quantized neural network model; and

performing a computation using the quantized neural network model.

17 . The method of claim 16 , wherein the smallest quantized weight value is determined as the first quantized weight value.

18 . The method of claim 16 , wherein the first weight value is selected based on a result of a comparison between the first weight value and a threshold value.

19 . The method of claim 18 , wherein the threshold value is set based on a user input.

20 . The method of claim 19 , wherein the threshold value is determined based on a pruning rate set based on the user input.

21 . The electronic device of claim 1 , further comprising normalizing the plurality of weight values based on fisher values corresponding to the plurality of weight values,

wherein the quantized values are obtained based on the normalized plurality of weight values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2021
From: LEE, DONGSOO; KWON, SEJUNG; KIM, BYEOUNGWOOK
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058039/0448 →
Priority Claims (1)
KR 10-2020-0058669 · May 15, 2020 · national
Continuity (2)
Continuation In Part PCTKR2021001766 · Feb 10, 2021
Related Publication 20220058487A1 · Feb 24, 2022
References Cited (117)
US 5809490A · Guiver et al. · 1998 [cited by applicant]
US 8086052B2 · Toth et al. · 2011 [cited by applicant]
US 9400955B2 · Garimella · 2016 [cited by applicant]
US 9916531B1 · Zivkovic et al. · 2018 [cited by applicant]
US 10115393B1 · Kumar · 2018 [cited by examiner]
US 10410098B2 · Nealis et al. · 2019 [cited by applicant]
US 10417525B2 · Ji et al. · 2019 [cited by applicant]
US 11074072B2 · Nealis et al. · 2021 [cited by applicant]
US 11321609B2 · Choi et al. · 2022 [cited by applicant]
US 11593586B2 · Ji et al. · 2023 [cited by applicant]
US 11693658B2 · Nealis et al. · 2023 [cited by applicant]
US 11875268B2 · Ji et al. · 2024 [cited by applicant]
US 20060251330A1 · Toth et al. · 2006 [cited by applicant]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160328645A1 · Lin et al. · 2016 [cited by applicant]
US 20160328647A1 · Lin et al. · 2016 [cited by applicant]
US 20170286830A1 · El-Yaniv et al. · 2017 [cited by applicant]
US 20170308789A1 · Langford et al. · 2017 [cited by applicant]
US 20170337471A1 · Kadav et al. · 2017 [cited by applicant]
US 20180046903A1 · Yao et al. · 2018 [cited by applicant]
US 20180082181A1 · Brothers et al. · 2018 [cited by applicant]
US 20180107925A1 · Choi et al. · 2018 [cited by applicant]
US 20180239992A1 · Chalfin et al. · 2018 [cited by applicant]
US 20180300603A1 · Ambardekar et al. · 2018 [cited by applicant]
US 20180300614A1 · Ambardekar et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180341857A1 · Lee et al. · 2018 [cited by applicant]
US 20190042935A1 · Deisher · 2019 [cited by applicant]
US 20190042948A1 · Lee et al. · 2019 [cited by applicant]
US 20190057308A1 · Cho et al. · 2019 [cited by applicant]
US 20190087744A1 · Schiemenz · 2019 [cited by applicant]
US 20190332903A1 · Nealis et al. · 2019 [cited by applicant]
US 20190347554A1 · Choi et al. · 2019 [cited by applicant]
US 20190392253A1 · Ji et al. · 2019 [cited by applicant]
US 20200012926A1 · Murata · 2020 [cited by applicant]
US 20200082269A1 · Gao et al. · 2020 [cited by applicant]
US 20200097818A1 · Li et al. · 2020 [cited by applicant]
US 20200143249A1 · Georgiadis · 2020 [cited by applicant]
US 20200184318A1 · Minezawa et al. · 2020 [cited by applicant]
US 20200293893A1 · Georgiadis · 2020 [cited by examiner]
US 20200302292A1 · Tseng et al. · 2020 [cited by applicant]
US 20210182077A1 · Chen · 2021 [cited by examiner]
US 20210232894A1 · Yamada et al. · 2021 [cited by applicant]
US 20210373886A1 · Nealis et al. · 2021 [cited by applicant]
US 20230140474A1 · Ji et al. · 2023 [cited by applicant]
US 20230359461A1 · Nealis et al. · 2023 [cited by applicant]
CN 1857001A · 2006 [cited by applicant]
CN 103744620A · 2014 [cited by applicant]
CN 105447498A · 2016 [cited by applicant]
CN 106203624A · 2016 [cited by applicant]
CN 107967515A · 2018 [cited by applicant]
CN 108734285A · 2018 [cited by applicant]
CN 107103113B · 2019 [cited by applicant]
JP 2019032729A · 2019 [cited by applicant]
JP 2020009048A · 2020 [cited by applicant]
KR 100248072B1 · 2000 [cited by applicant]
KR 1020060027795A · 2006 [cited by applicant]
KR 100662516B1 · 2006 [cited by applicant]
KR 100790900B1 · 2008 [cited by applicant]
KR 1020160034814A · 2016 [cited by applicant]
KR 1020180013674A · 2018 [cited by applicant]
KR 1020180043154A · 2018 [cited by applicant]
KR 1020180120967A1 · 2018 [cited by applicant]
KR 1020190054449A · 2019 [cited by applicant]
KR 1020190093932A · 2019 [cited by applicant]
KR 1020190130443A · 2019 [cited by applicant]
KR 1020190130455A · 2019 [cited by applicant]
KR 1020200049422A · 2020 [cited by applicant]
KR 1020200052200A · 2020 [cited by applicant]
WO 2016173488A1 · 2016 [cited by applicant]
WO 2017049496A1 · 2017 [cited by applicant]
WO 2019008752A1 · 2019 [cited by applicant]
WO 2019088657A1 · 2019 [cited by applicant]
WO 2020075433A1 · 2020 [cited by applicant]
Zhang, et al., “LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks”, ECCV2018, 2018 (Year: 2018). [cited by examiner]
Hubara, et al., “Quantized neural networks: training neural networks with low precision weights and activations”, Journal of Machine Learning Research 18 (2018) 1-30, 2018 (Year: 2018). [cited by examiner]
Nagel, et al., “A white paper on neural network quantization”, arXiv:2016.08295v1 [cs. LG] Jun. 15, 2021 (Year: 2021). [cited by examiner]
Communication issued Feb. 26, 2024 by the European Patent Office for EP Patent Application No. 20819392.0. [cited by applicant]
Office Action issued on Mar. 30, 2024 by the Chinese Patent Office in corresponding CN Patent Application No. 202010493206.8. [cited by applicant]
Notice Of Allowance issued Apr. 26, 2024 issued by the Chinese Patent Office in corresponding CN Patent Application No. 201980025238.0. [cited by applicant]
Office Action issued on May 9, 2024 by the U.S. Patent Office in corresponding U.S. Appl. No. 17/044,104. [cited by applicant]
Hsiao, Yu-Shun, Yun-Chen Lo, and Ren-Shuo Liu. “FlexNet: Neural networks with inherent inference-time bitwidth flexibility”, In Proceedings of International Symposium on Microarchitecture, 2018, 2 pages. [cited by applicant]
Wang, Kuan, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. “Haq: Hardware-aware automated quantization with mixed precision”, In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8612… [cited by applicant]
Guo, Yiwen, Anbang Yao, Hao Zhao, and Yurong Chen. “Network sketching: Exploiting binary structure in deep cnns”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5955-5963, Jun. 7, … [cited by applicant]
Office Action issued on May 14, 2024 by the Korean Patent Office in corresponding KR Patent Application No. 10-2019-0066396. [cited by applicant]
Communication dated May 27, 2022, issued by the European Patent Office in counterpart European Application No. 20819392.0. [cited by applicant]
Yang et al., “Deploy Large-Scale Deep Neural Networks in Resource Constrained IoT Devices with Local Quantization Region,” arxiv.org, Cornell University Library, May 24, 2018, XP081234517, Total 8 pages. [cited by applicant]
Communication dated Jun. 22, 2022, issued by the United States Patent and Trademark Office in U.S. Appl. No. 16/876,688. [cited by applicant]
Office Action issued Jan. 23, 2023 by the United States Patent and Trademark Office in counterpart U.S. Appl. No. 16/876,688. [cited by applicant]
International Search Report (PCT/ISA/210) dated May 28, 2021 issued by the International Searching Authority in International Application No. PCT/KR2021/001766. [cited by applicant]
Written Opinion (PCT/ISA/237) dated May 28, 2021 issued by the International Searching Authority in International Application No. PCT/KR2021/001766. [cited by applicant]
International Search Report (PCT/ISA/210) and Written Opinion (PCT/ISA/237) dated Apr. 18, 2019 issued by the International Searching Authority in International Application No. PCT/KR2019/000142. [cited by applicant]
International Search Report (PCT/ISA/210) and Written Opinion (PCT/ISA/237) dated Aug. 14, 2020 issued by the International Searching Authority in International Application No. PCT/KR2020/006411. [cited by applicant]
Communication dated Jun. 24, 2021 issued by the European Patent Office in application No. 19800295.8. [cited by applicant]
Zhou, S., et al., “IQNN: Training Quantized Neural Networks with Iterative Optimizations”, International Conference on Image Analysis and Processing, 17th International Conference, Sep. 9-13, 2013, pp. 688-695, XP047451… [cited by applicant]
Ji, Z., et al., “Reducing Weight Precision of Convolutional Neural Networks towards Large-scale On-chip Image Recognition”, Proceedings of SPIE, IEEE, vol. 9496, May 20, 2015, 9 pages, XP060055095. [cited by applicant]
Pilipović, R., et al., “Compression of convolutional neural networks”, 17th International Symposium INFOTEH-JAHORINA, Mar. 21-23, 2018, 6 pages, XP033335103. [cited by applicant]
Li, H., et al., “Towards a Deeper Understanding of Training Quantized Neural Networks”, Aug. 1, 2017, 9 pages, XP055590849. [cited by applicant]
Park, E., et al., “Weighted-Entropy-based Quantization for Deep Neural Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition, pp. 7197-7205. [cited by applicant]
Dong, Y., et al., “Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization”, arXiv:1708.01001v1 [cs.CV], Aug. 3, 2017, pp. 1-12. [cited by applicant]
Xu, C., et al., “Alternating Multi-Bit Quantization for Recurrent Neural Networks”, arXiv:1802.00150v1 [cs.LG], Feb. 1, 2018, pp. 1-13. [cited by applicant]
Zhu, C., et al., “Trained Ternary Quantization”, arXiv:1612.01064v3 [cs.LG], ICLR, Feb. 23, 2017, pp. 1-10. [cited by applicant]
Courbariaux, M., et al., “BinaryConnect: Training Deep Neural Networks with binary weights during propagations”, https://github.com/MatthieuCourbariaux/BinaryConnect, pp. 1-9. [cited by applicant]
Office Action issued Oct. 19, 2024 by the Chinese Patent Office for Chinese Patent Application No. 202010493206.8. [cited by applicant]
Office Action issued Dec. 6, 2024 by the Korean Patent Office for Korean Patent Application No. 10-2020-0058669. [cited by applicant]
Office Action issued Jan. 10, 2025 for U.S. Appl. No. 17/044,104. [cited by applicant]
Office Action, issuance date: Oct. 19, 2022, issued for U.S. Appl. No. 16/876,688. [cited by applicant]
Communication dated Nov. 24, 2023, issued by the Korean Intellectual Property Office in counterpart Korean Application No. 10-2018-0053185. [cited by applicant]
Communication dated Dec. 26, 2023, issued by the National Intellectual Property Administration, PRC in counterpart Chinese Application No. 201980025238.0. [cited by applicant]
Communication dated Feb. 12, 2024, issued by the European Patent Office in counterpart European Application No. 19800295.8. [cited by applicant]
Office Action dated Feb. 8, 2024, issued by the United States Patent and Trademark Office in counterpart U.S. Appl. No. 17/044,104. [cited by applicant]
Communication dated Jan. 16, 2025 issued by the China National Intellectual Property Administration in Chinese Patent Application No. 202010493206.8. [cited by applicant]
Communication dated Jan. 31, 2025 issued by the European Patent Office in European Patent Application No. 19800295.8. [cited by applicant]
Communication issued on Jun. 25, 2025 from the European Patent Office for European Patent Application No. 19800295.8. [cited by applicant]
Communication issued on Apr. 10, 2025 from the China National Intellectual Property Administration for Chinese Patent Application No. 202010493206.8. [cited by applicant]
Communication issued on Jun. 16, 2025 from the European Patent Office for European Patent Application No. 19800295.8. [cited by applicant]
Communication issued Oct. 6, 2025 by the European Patent Office in European Patent Application No. 20819392.0. [cited by applicant]