IP Library › Granted Patent US 12,340,247
Granted Patent B2
US 12,340,247 · App. 17/745,320 · Granted Jun 24, 2025

Electronic device for processing neural network model and method of operating the same

Inventors: Junhyuk Lee (Suwon-si, KR); Hyunbin Park (Suwon-si, KR); Seungjin Yang (Suwon-si, KR); Jin Choi (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06F9/5016G06N3/08H04L5/0064
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,247
App. No.
17/745,320
Granted
Jun 24, 2025
Kind
B2
Abstract

An electronic device is provided. The electronic device includes a memory, and a processor including a resource management unit and a neural processing unit. The processor may be configured to obtain an execution request for a specific function operating based on a specific neural network model, identify an available bandwidth of the memory through the resource management unit, and quantize the specific neural network model based on the available bandwidth of the memory through the neural processing unit.

Claims (57)

1. An electronic device comprising:

memory storing one or more computer programs; and

one or more processors, including a resource management unit and a neural processing unit, communicatively coupled to the memory,

wherein the one or more computer programs include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to:

obtain an execution request for a specific function operating based on a specific neural network model,

identify an available bandwidth of the memory through the resource management unit,

quantize the specific neural network model based on the available bandwidth of the memory through the neural processing unit,

identify a section including the available bandwidth of the memory from a predefined table, and

select a bit depth corresponding to the identified section as an activation bit depth.

2. The electronic device of claim 1 , wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to identify the available bandwidth of the memory based on a predetermined period.

3. The electronic device of claim 1 , wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to identify the available bandwidth of the memory based on obtaining the execution request for the specific function.

4. The electronic device of claim 1 , wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to identify the available bandwidth of the memory based on obtaining a processing request for a specific layer of the specific neural network model.

5. The electronic device of claim 1 ,

wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to:

identify a first section including the available bandwidth of the memory from the predefined table based on the available bandwidth of the memory exceeding a first threshold, and

select a first bit depth corresponding to the first section as the activation bit depth, and

wherein the first bit depth is a 16-bit depth.

6. The electronic device of claim 5 ,

wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to:

identify a second section including the available bandwidth of the memory from the predefined table based on the available bandwidth of the memory being less than or equal to the first threshold and exceeding a second threshold, and

select a second bit depth corresponding to the second section as the activation bit depth, and

wherein the second bit depth is an 8-bit depth.

7. The electronic device of claim 6 ,

wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to:

identify a third section including the available bandwidth of the memory from the predefined table based on the available bandwidth of the memory being less than or equal to the second threshold and exceeding a third threshold, and

select a third bit depth corresponding to the third section as the activation bit depth, and

wherein the third bit depth is a 6-bit depth.

8. The electronic device of claim 1 , wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to:

quantize the specific neural network model by performing dynamic quantization on a specific activation bit depth of the specific neural network model based on the activation bit depth while maintaining a specific weight bit depth of the specific neural network model through the neural processing unit, and

perform a multiply and accumulate (MAC) operation on the quantized specific neural network model using the specific activation bit depth.

9. The electronic device of claim 1 , wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to:

quantize the specific neural network model by performing dynamic quantization on a specific activation bit depth of a specific layer of the specific neural network model based on the activation bit depth while maintaining a specific weight bit depth of the specific layer of the specific neural network model through the neural processing unit, and

perform a multiply and accumulate (MAC) operation on the specific layer of the quantized specific neural network model using the specific activation bit depth.

10. A method of operating an electronic device, the method comprising:

obtaining an execution request for a specific function operating based on a specific neural network model;

identifying an available bandwidth of a memory of the electronic device through a resource management unit included in a processor of the electronic device;

quantizing the specific neural network model based on the available bandwidth of the memory through a neural processing unit included in the processor;

identifying a section including the available bandwidth of the memory from a predefined table; and

selecting a bit depth corresponding to the identified section as an activation bit depth.

11. The method of claim 10 , wherein the identifying of the available bandwidth of the memory includes identifying the available bandwidth of the memory based on a predetermined period.

12. The method of claim 10 , wherein the identifying of the available bandwidth of the memory includes identifying the available bandwidth of the memory based on obtaining the execution request for the specific function.

13. The method of claim 10 , wherein the identifying of the available bandwidth of the memory includes identifying the available bandwidth of the memory based on obtaining a processing request for a specific layer of the specific neural network model.

14. The method of claim 10 ,

wherein the selecting of the activation bit depth includes:

identifying a first section including the available bandwidth of the memory from the predefined table based on the available bandwidth of the memory exceeding a first threshold, and

selecting a first bit depth corresponding to the first section as the activation bit depth, and

wherein the first bit depth is a 16-bit depth.

15. The method of claim 14 ,

wherein the selecting of the activation bit depth includes:

identifying a second section including the available bandwidth of the memory from the predefined table based on the available bandwidth of the memory being less than or equal to the first threshold and exceeding a second threshold, and

selecting a second bit depth corresponding to the second section as the activation bit depth, and

wherein the second bit depth is an 8-bit depth.

16. The method of claim 15 ,

wherein the selecting of the activation bit depth includes:

identifying a third section including the available bandwidth of the memory from the predefined table based on the available bandwidth of the memory being less than or equal to the second threshold and exceeding a third threshold, and

selecting a third bit depth corresponding to the third section as the activation bit depth, and

wherein the third bit depth is a 6-bit depth.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: LEE, JUNHYUK; PARK, HYUNBIN; YANG, SEUNGJIN; CHOI, JIN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 059920/0308 →
Priority Claims (2)
KR 10-2021-0119368 · Sep 7, 2021 · national
KR 10-2021-0142991 · Oct 25, 2021 · national
Continuity (2)
Continuation PCTKR2022005700 · Apr 21, 2022
Related Publication 20230072337A1 · Mar 9, 2023
References Cited (28)
US 6167088A · Sethuraman · 2000 [cited by applicant]
US 11561833B1 · Heaton · 2023 [cited by examiner]
US 11580719B2 · Desappan · 2023 [cited by examiner]
US 20060190650A1 · Burchard · 2006 [cited by examiner]
US 20160328647A1 · Lin et al. · 2016 [cited by applicant]
US 20190102673A1 · Leibovich et al. · 2019 [cited by applicant]
US 20190147324A1 · Martin · 2019 [cited by applicant]
US 20190147364A1 · Alt et al. · 2019 [cited by applicant]
US 20190392300A1 · Weber et al. · 2019 [cited by applicant]
US 20200107025A1 · Karczewicz · 2020 [cited by examiner]
US 20200125926A1 · Choudhury et al. · 2020 [cited by applicant]
US 20200184320A1 · Croxford · 2020 [cited by examiner]
US 20210160499A1 · Wang et al. · 2021 [cited by applicant]
US 20220092346A1 · Jones · 2022 [cited by examiner]
US 20220391776A1 · Mody · 2022 [cited by examiner]
US 20230168921A1 · Kim · 2023 [cited by examiner]
CN 104965763A · 2015 [cited by applicant]
CN 110070181A · 2019 [cited by applicant]
CN 111340226A · 2020 [cited by applicant]
EP 3767548A1 · 2021 [cited by applicant]
KR 1020210009584A · 2021 [cited by applicant]
Wang et al., “HAQ: Hardware-Aware Automated Quantization with Mixed Precision”, arXiv: 1811. 08886v3, Apr. 2019 [Jul. 7, 2022]. <https://arxiv.org/pdf/1811.08886v3.pdf> pp. 2-4; and figure 2. [cited by applicant]
International Search Report dated Jul. 22, 2022, issued in International Application No. PCT/KR2022/005700. [cited by applicant]
Written Opinion dated Jul. 22, 2022, issued in International Application No. PCT/KR2022/005700. [cited by applicant]
Manuele Rusci et al., Memory-Driven Mixed Low Precision Quantization For Enabling Deep Network Inference On Microcontrollers, May 30, 2019, XP081366031. [cited by applicant]
Marios Fournarakis et al., In-Hindsight Quantization Range Estimation for Quantized Training, May 10, 2021, XP081960724. [cited by applicant]
Tian Huang et al., Adaptive Precision Training for Resource Constrained Devices, Dec. 23, 2020, XP081845418. [cited by applicant]
European Search Report dated Sep. 10, 2024, issued in European Application No. 22867486.7. [cited by applicant]