IP Library Granted Patent US 12,197,926
Granted Patent B2
US 12,197,926 · App. 17/097,501 · Granted Jan 14, 2025

Dynamic loading neural network inference at DRAM/on-bus SRAM/serial flash for power optimization

Inventors: Chih-Hsiang Hsiao (Hsinchu, TW); Chia-Feng Hsu (Hsinchu, TW)
Assignee: MEDIATEK INC.
G06F9/44521G06F1/3275G06F13/28G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,926
App. No.
17/097,501
Granted
Jan 14, 2025
Kind
B2
Abstract

Aspects of the disclosure provide a method and an apparatus for executing a program, e.g., a neural network (NN) inference. For example, the apparatus can include an executor and a dynamic loading agent. The executor can be coupled to a second memory, and be configured to execute a portion of the NN inference loaded on the second memory from a first memory that stores the NN inference, and to generate a signal based on a progress of the execution of the NN inference. The dynamic loading agent can be coupled to the executor, the first memory and the second memory, and be configured to load a next portion of the NN inference stored in the first memory to the second memory and to manage power supplied to the first memory based on the signal from the executor and an inference executing scheme stored in the second memory.

Claims (28)

1. An apparatus for executing a program, comprising:

an executor that is configured to execute a portion of the program loaded on a second memory from a first memory that stores the program and to generate a signal based on a progress of the execution of the program; and

a dynamic loading agent coupled to the executor, the first memory, and the second memory, the dynamic loading agent being configured to load a next portion of the program stored in the first memory to the second memory and to manage power supplied to the first memory based on the signal from the executor and an executing scheme stored in the second memory.

2. The apparatus of claim 1 , wherein the dynamic loading agent manages the power supplied to the first memory by powering on/off the first memory, configuring an operation mode of the first memory, or scaling a voltage and/or a frequency applied to the first memory.

3. The apparatus of claim 1 , wherein the executing scheme includes a script, a rule, or a model.

4. The apparatus of claim 3 , wherein the rule generates a script for a certain layer of the program based on input tensors from a previous layer of the program.

5. The apparatus of claim 1 , wherein the signal includes an operation, a node identification (ID), a tensor, a kernel, and/or time of the program.

6. The apparatus of claim 1 , wherein the second memory is a tightly-coupled memory (TCM) or a static random access memory (SRAM).

7. The apparatus of claim 1 , wherein the first memory is a dynamic random access memory (DRAM), an on-bus SRAM, or a serial flash memory.

8. The apparatus of claim 1 , wherein:

the executor is a central processing unit (CPU) and the dynamic loading agent is a microcontroller unit (MCU) or a CPU exception, or the executor is a deep learning accelerator (DLA) and the dynamic loading agent is an MCU.

9. The apparatus of claim 1 , wherein the dynamic loading agent and the executor are included on a single chip.

10. The apparatus of claim 1 , further comprising a direct memory access (DMA) controller coupled to the dynamic loading agent, the first memory, and the second memory, wherein the dynamic loading agent is further configured to instruct the DMA controller to load the next portion of the program stored in the first memory to the second memory.

11. The apparatus of claim 1 , wherein the executing scheme is man-made or created by an offline optimizer or an online/runtime optimizer.

12. A method for executing a program, comprising:

loading a portion of the program from a first memory that stores the program to a second memory;

executing the portion of the program loaded on the second memory, and generating a signal based on a progress of the execution of the program; and

loading a next portion of the program stored in the first memory to the second memory and managing power supplied to the first memory based on the signal and an executing scheme stored in the second memory.

13. The method of claim 12 , wherein managing the power supplied to the first memory includes powering on/off the first memory, configuring an operation mode of the first memory, or scaling a voltage and/or a frequency applied to the first memory.

14. The method of claim 12 , wherein the executing scheme includes a script, a rule, or a model.

15. The method of claim 14 , wherein the rule generates a script for a certain layer of the program based on input tensors from a previous layer of the program.

16. The method of claim 12 , wherein the signal includes an operation, a node ID, a tensor, a kernel and/or time of the program.

17. The method of claim 12 , wherein the second memory is a TCM or an SRAM.

18. The method of claim 12 , wherein the first memory is a DRAM, an on-bus SRAM, or a serial flash memory.

19. The method of claim 12 , wherein the executing scheme is man-made or created by an offline optimizer or an online/runtime optimizer.

20. The method of claim 12 , wherein executing the portion of the program and loading a next portion of the program stored in the first memory to the second memory and managing power supplied to the first memory are performed asynchronously.

21. The method of claim 12 , wherein the second memory is more power efficient than the first memory.

22. The apparatus of claim 1 , wherein the second memory is more power efficient than the first memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2020
From: HSIAO, CHIH-HSIANG; HSU, CHIA-FENG
To: MEDIATEK INC.
Reel/Frame 054413/0981 →
Continuity (2)
Provisional Application 63047939 · Jul 3, 2020
Related Publication 20220004399A1 · Jan 6, 2022
References Cited (25)
US 11615322B1 · Thomas · 2023 [cited by examiner]
US 20180300601A1 · Cedola · 2018 [cited by examiner]
US 20190147332A1 · Lagudu · 2019 [cited by examiner]
US 20190370632A1 · Hashemi et al. · 2019 [cited by applicant]
US 20200257627A1 · Chamarty · 2020 [cited by examiner]
US 20210019652A1 · Gadelrab · 2021 [cited by examiner]
US 20220004399A1 · Hsiao et al. · 2022 [cited by applicant]
CN 108475235A · 2018 [cited by applicant]
CN 109389214A · 2019 [cited by applicant]
CN 109919312A · 2019 [cited by applicant]
TW 201225087A · 2012 [cited by applicant]
TW 201835918A · 2018 [cited by applicant]
TW 202004497A · 2020 [cited by applicant]
TW 110998486A · 2020 [cited by applicant]
TW 202203036A · 2022 [cited by applicant]
Combined Taiwanese Office Action and Search Report issued Jul. 7, 2022, in corresponding Taiwanese Patent Application No. 110118831 (with English Translation of Category of Cited Documents) , 8 pages. [cited by applicant]
Extended European Search Report issued Nov. 12, 2021 in European Patent Application No. 21183203.5, 14 pages. [cited by applicant]
Thierry Moreau, et al., “VTA: An Open Hardware-Software Stack for Deep Learning” Cornell University Library, XP81115846, Jul. 11, 2018, pp. 1-19. [cited by applicant]
Lita Yang, et al., “SRAM Voltage Scaling for Energy-Efficient Convolutional Neural Networks” 18 [cited by applicant]
Udit Gupta, et al., “MASR: A Modular Accelerator for Sparse RNNs” 28 [cited by applicant]
Combined Taiwanese Office Action and Search Report issued Mar. 8, 2022 in Taiwanese Patent Application No. 110118831 (with English translation of categories of cited documents), 9 pages. [cited by applicant]
European Office Action issued on Mar. 30, 2023 in European Patent Application No. 21183203.5, 8 pages. [cited by applicant]
Koppula Skanda et al: “EDEN Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM”, Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, Micro … [cited by applicant]
Thierry Moreau et al., “VTA: An Open Hardware-Software Stack for Deep Learning”, ARXIV, Jul. 11, 2018, pp. 1-19. ,Jul. 11, 2018. [cited by applicant]
Yang, Lita et al., “SRAM Voltage Scaling for Energy-Efficient Convolutional Neural Networks”, Proceedings of the Eighteenth International Symposium On Quality Electronic Design (ISQED), 2017 IEEE, pp. 7-12. ,2017. [cited by applicant]