IP Library › Granted Patent US 12,651,172
Granted Patent B2
US 12,651,172 · App. 18/323,908 · Granted Jun 9, 2026

Static scheduling and dynamic scheduling for compiler-hinted and self-scheduling multi-engine artificial intelligence (AI) processing unit system

Inventors: Chieh-Fang Teng (Hsinchu, TW); En-Jui Chang (Hsinchu, TW); Chih Chung Cheng (Hsinchu, TW)
Assignee: MEDIATEK INC.
G06N3/10G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,172
App. No.
18/323,908
Granted
Jun 9, 2026
Kind
B2
Abstract

Aspects of the present disclosure provide an apparatus. For example, the apparatus can include a compiler configured to compile a neural network (NN) model to generate a plurality of operations/threads and determine whether each of the operations/threads is compute bound or memory bound, and a memory coupled to the compiler and configured to store the operations/threads. The apparatus can also include a thread scheduler coupled to the memory and configured to schedule the operations/threads of the NN model. The apparatus can also include a multi-engine processing unit that includes a plurality of compute units (CUs), and an executor coupled between the thread scheduler and the multi-engine processing unit. The executor can be configured to allocate the operations/threads of the NN model and activate a number of the CUs of the multi-engine processing unit for each of the operations/threads based on whether the operation/thread is compute bound or memory bound.

Claims (31)

1 . An apparatus for optimizing resource allocation in a multi-engine artificial intelligence (AI) processing unit, comprising:

a compiler configured to compile a neural network (NN) model to generate a plurality of operations assigned to threads for execution, and determine whether each of the operations and threads is compute bound or memory bound;

a memory coupled to the compiler, the memory configured to store the operations and threads of the NN model;

a thread scheduler coupled to the memory, the thread scheduler configured to schedule the operations and threads of the NN model;

a multi-engine AI processing unit that includes a plurality of physical compute units (CUs); and

an executor coupled between the thread scheduler and the multi-engine AI processing unit, the executor configured to allocate the operations and threads of the NN model and activate a number of the physical CUs of the multi-engine AI processing unit for each of the operations and threads, based on runtime performance of the multi-engine AI processing unit, and based on whether the operation and/or thread is compute bound or memory bound.

2 . The apparatus of claim 1 , further comprising:

a performance monitor coupled to the executor, the performance monitor configured to monitor runtime performance of the multi-engine processing unit and/or runtime performance of a network coupled to the multi-engine processing unit,

wherein the executor is further configured to change the number of the CUs that are activated based on the runtime performance of the multi-engine processing unit and/or the runtime performance of the network.

3 . The apparatus of claim 2 , wherein the executor is configured to change the number of the physical CUs that are activated by activating one of the physical CUs that are not activated or deactivating one of the physical CUs that are activated at a time.

4 . The apparatus of claim 2 , wherein the runtime performance of the multi-engine processing unit includes throughput of the physical CUs.

5 . The apparatus of claim 2 , wherein the runtime performance of the network includes input/output bandwidth between the network and the multi-engine processing unit.

6 . The apparatus of claim 1 , wherein the multi-engine processing unit further includes a buffer that is coupled to and shared by at least two of the physical CUs.

7 . The apparatus of claim 1 , wherein the thread scheduler is configured to schedule the operation and threads by maximizing uRate of the multi-engine processing unit while meeting thread execution constraints.

8 . The apparatus of claim 1 , wherein the compiler determines that one of the operations and threads is the compute bound if a compute cycle of the operation or thread is greater than a memory read/write (R/W) cycle of the operation or thread, or is the memory bound if the compute cycle is not greater than the memory R/W cycle.

9 . The apparatus of claim 1 , wherein the memory includes a queue.

10 . An apparatus for optimizing resource allocation in a multi-engine artificial intelligence (AI) processing unit, comprising:

a compiler configured to compile a neural network (NN) model to generate a plurality of operations assigned to threads for execution;

a memory coupled to the compiler, the memory configured to store the operations and threads of the NN model;

a thread scheduler coupled to the memory, the thread scheduler configured to schedule the operations and threads of the NN model;

a multi-engine AI processing unit that includes a plurality of physical compute units (CUs);

a performance monitor configured to monitor runtime performance of the multi-engine AI processing unit and/or runtime performance of a network coupled to the multi-engine AI processing unit; and

an executor coupled between the thread scheduler and the multi-engine AI processing unit, and the performance monitor, the executor configured to allocate the operations and threads of the NN model, activate a number of the physical CUs of the multi-engine AI processing unit for each of the operations and threads, and change the number of physical CUs that are activated based on the runtime performance of the multi-engine AI processing unit and/or the runtime performance of the network, and based on whether the operation and/or thread is determined to be compute bound or memory bound.

11 . The apparatus of claim 10 , wherein the executor is configured to change the number of the physical CUs that are activated by activating one of the physical CUs that are not activated or deactivating one of the physical CUs that are activated at a time.

12 . The apparatus of claim 10 , wherein the runtime performance of the multi-engine processing unit includes throughput of the physical CUs.

13 . The apparatus of claim 10 , wherein the runtime performance of the network includes input/output bandwidth between the network and the multi-engine processing unit.

14 . The apparatus of claim 10 , wherein the compiler is further configured to determine whether each of the operations and threads is compute bound or memory bound, and the executor is configured to activate the number of the physical CUs of the multi-engine processing unit for each of the operations and threads based on whether the operation or thread is compute bound or memory bound.

15 . The apparatus of claim 10 , wherein the multi-engine processing unit further includes a buffer that is coupled to and shared by at least two of the physical CUs.

16 . The apparatus of claim 10 , wherein the thread scheduler is configured to schedule the operation/threads by maximizing uRate of the multi-engine processing unit while meeting thread execution constraints.

17 . The apparatus of claim 10 , wherein the compiler determines that one of the operations and threads is the compute bound if a compute cycle of the operation or thread is greater than a memory read/write (R/W) cycle of the operation or thread, or is the memory bound if the compute cycle is not greater than the memory R/W cycle.

18 . The apparatus of claim 10 , wherein the memory includes a queue.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2023
From: TENG, CHIEH-FANG; CHANG, EN-JUI; CHENG, CHIH CHUNG
To: MEDIATEK INC.
Reel/Frame 063789/0944 →
Continuity (2)
Provisional Application 63385215 · Nov 29, 2022
Related Publication 20240177019A1 · May 30, 2024
References Cited (30)
US 11593628B2 · Febbo · 2023 [cited by examiner]
US 11599777B2 · Ben-Avi · 2023 [cited by examiner]
US 11829440B2 · Narayanamoorthy · 2023 [cited by examiner]
US 12079734B1 · Zheng · 2024 [cited by examiner]
US 12400106B1 · Diamant · 2025 [cited by examiner]
US 20140089699A1 · O'Connor · 2014 [cited by examiner]
US 20160379109A1 · Chung · 2016 [cited by examiner]
US 20170205863A1 · Lee · 2017 [cited by applicant]
US 20180046900A1 · Dally · 2018 [cited by examiner]
US 20190065281A1 · Bernat · 2019 [cited by examiner]
US 20210042610A1 · Chang · 2021 [cited by examiner]
US 20210191770A1 · Rao · 2021 [cited by examiner]
US 20210271960A1 · Raha · 2021 [cited by examiner]
US 20210279557A1 · Febbo · 2021 [cited by applicant]
US 20220035679A1 · Sunwoo · 2022 [cited by examiner]
US 20220036163A1 · Kuo · 2022 [cited by applicant]
US 20220206850A1 · Paul · 2022 [cited by examiner]
US 20220244984A1 · Lee · 2022 [cited by examiner]
US 20230409387A1 · Gupta · 2023 [cited by examiner]
CA 3069779C · 2021 [cited by examiner]
CN 104375899B · 2016 [cited by examiner]
CN 118114729A · 2024 [cited by examiner]
EP 3396533B1 · 2022 [cited by examiner]
EP 4208786B9 · 2025 [cited by examiner]
TW 202121169A · 2021 [cited by applicant]
WO WO2021030376A1 · 2021 [cited by examiner]
Chinese language office action dated Dec. 9, 2024, issued in application No. TW 112146320. [cited by applicant]
Chien-Hung Lin et al. ; A 3.4-to-13.3TOPS/W 3.6TOPS Dual-Core Deep-Learning Accelerator for Versatile AI Applications in 7nm 5G Smartphone SoC ; IEEE International Solid-State Circuits Conference ; 2020 ; 3 Pages. [cited by applicant]
Samuel Williams et al. ; Roofline: An Insightful Visual Performance Model for Multicore Architectures ; Communication of the ACM, vol. 52, No. 4 ; Apr. 2009 ; 12 Pages. [cited by applicant]
Extended European Search Report dated Mar. 19, 2024, issued in application No. EP 23211692.1. [cited by applicant]