IP Library › Granted Patent US 12,579,419
Granted Patent B1
US 12,579,419 · App. 19/336,529 · Granted Mar 17, 2026

System and SoC comprising heterogeneous processors capable of processing artificial intelligence and method thereof

Inventor: Jeong Gyu Jang (Suwon-si, KR)
Assignee: DEEPX CO., LTD.
G06N3/063G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,419
App. No.
19/336,529
Granted
Mar 17, 2026
Kind
B1
Abstract

According to an example of the present disclosure, a system is provided. The system may include: a plurality of heterogeneous processors configured to perform inference using a neural network model comprising a plurality of layers; a memory for storing the neural network model, wherein each layer of the neural network model comprises a plurality of parameters having different bit-widths; and a controller for assigning an operation for an arbitrary layer of the neural network model to an arbitrary processor among the plurality of heterogeneous processors. An operation for a first portion of the arbitrary layer of the neural network model may be assigned to a first processor among the plurality of heterogeneous processors. An operation for a second portion of the arbitrary layer of the neural network model may be assigned to a second processor among the plurality of heterogeneous processors. A first bit-width of a first parameter for the first portion and a second bit-width of a second parameter for the second portion may be different from each other.

Claims (57)

1 . A system, comprising:

a plurality of heterogeneous processors configured to perform inference using a neural network model comprising a plurality of layers;

a memory configured to store the neural network model, wherein each layer of the neural network model comprises a plurality of parameters having different bit-widths; and

a controller configured to assign an operation for an arbitrary layer of the neural network model to an arbitrary processor among the plurality of heterogeneous processors,

wherein an operation for a first portion of the arbitrary layer of the neural network model is assigned to a first processor among the plurality of heterogeneous processors,

wherein an operation for a second portion of the arbitrary layer of the neural network model is assigned to a second processor among the plurality of heterogeneous processors,

wherein a first bit-width of a first parameter for the first portion and a second bit-width of a second parameter for the second portion are different,

wherein at least a portion of the plurality of parameters is pruned according to a pruning ratio, and

wherein, when the operation for the arbitrary layer is assigned to the arbitrary processor, the pruning ratio is increased when a computational load of the arbitrary processor is higher than a reference value and is decreased when the computational load is lower than the reference value.

2 . The system of claim 1 , wherein the plurality of parameters comprises at least one of a feature map, a weight, a kernel, and an activation map.

3 . The system of claim 1 , wherein the plurality of heterogeneous processors comprises at least one of a neural processing unit (NPU), a central processing unit (CPU), a graphic processing unit (GPU), and a Micro Controller Unit (MCU).

4 . The system of claim 1 , wherein

the first bit-width is an integer of 4 bits to 8 bits, and

the second bit-width is a floating point of 16 bits to 32 bits.

5 . The system of claim 1 , wherein

the first processor is an NPU, and

the second processor is one of a CPU, a GPU, and an MCU.

6 . The system of claim 1 , wherein

an operation for a first layer of the neural network model is assigned to the first processor, and

an operation for a second layer of the neural network model is assigned to the second processor.

7 . The system of claim 6 , wherein a bit-width of a parameter for the first layer and a bit-width of a parameter for the second layer are different.

8 . The system of claim 1 , wherein at least a portion of the plurality of parameters is quantized.

9 . The system of claim 8 , wherein the quantization comprises 8-bit quantization, 4-bit quantization, and sub-4-bit quantization.

10 . The system of claim 1 , wherein the pruning comprises;

applying a method of replacing 80% of weights with 0 (Prune 80%) to a first subset of the plurality of parameters,

applying a method of replacing 70% of weights with 0 (Prune 70%) to a second subset of the plurality of parameters, and

applying a method of replacing 60% of weights with 0 (Prune 60%) to a third a subset of the plurality of parameters.

11 . The system of claim 1 , wherein, for each layer of the neural network model, the controller is further configured to

calculate an objective function for each of the plurality of heterogeneous processors, the objective function combining a performance loss of the neural network due to lightening and at least one of processing time, computational load, chip temperature, and power consumption,

select, as a main target processor for the arbitrary layer, a processor of the plurality of heterogeneous processors, the selected processor being a processor that minimizes the objective function, and

select, as at least one sub target processor for the arbitrary layer, at least one other processor of the plurality of heterogeneous processors, the selected at least one other processor being a processor that has a larger value of the objective function than the main target processor.

12 . A System on Chip (SoC), comprising:

a plurality of heterogeneous processors configured to perform inference using a neural network model comprising a plurality of layers;

a memory configured to store the neural network model, wherein each layer of the neural network model comprises a plurality of parameters having different bit-widths; and

a controller configured to assign an operation for an arbitrary layer of the neural network model to an arbitrary processor among the plurality of heterogeneous processors,

wherein an operation for a first portion of the arbitrary layer of the neural network model is assigned to a first processor among the plurality of heterogeneous processors,

wherein an operation for a second portion of the arbitrary layer of the neural network model is assigned to a second processor among the plurality of heterogeneous processors,

wherein a first bit-width of a first parameter for the first portion and a second bit-width of a second parameter for the second portion are different,

wherein at least a portion of the plurality of parameters is pruned according to a pruning ratio, and

wherein, when the operation for the arbitrary layer is assigned to the arbitrary processor, the pruning ratio is increased when a computational load of the arbitrary processor is higher than a reference value and is decreased when the computational load is lower than the reference value.

13 . The SoC of claim 12 , wherein the plurality of parameters comprises at least one of a feature map, a weight, a kernel, and an activation map.

14 . The SoC of claim 12 , wherein the plurality of heterogeneous processors comprises at least one of a neural processing unit (NPU), a central processing unit (CPU), a graphic processing unit (GPU), and a Micro Controller Unit (MCU).

15 . The SoC of claim 12 , wherein

the first processor is an NPU, and

the second processor is one of a CPU, a GPU, and an MCU.

16 . The SoC of claim 12 , wherein

an operation for a first layer of the neural network model is assigned to the first processor, and

an operation for a second layer of the neural network model is assigned to the second processor.

17 . An electronic device, comprising:

a plurality of heterogeneous processors for performing inference using a neural network model comprising a plurality of layers;

a controller for assigning an operation for a current layer of the neural network model to one of the plurality of heterogeneous processors,

wherein the layers include a first layer comprising parameters having a first bit-width and a second layer comprising parameters having a second bit-width different from the first bit-width,

wherein the controller selects at least one first processor from among the plurality of heterogeneous processors to perform an operation for the first layer of the neural network model based on the parameters of the first layer, the selected first processor including a first main target processor and at least one first sub target processor, the operation for the first layer assigned to the at least one first sub target processor if, due to being occupied by another work, the first main target processor cannot perform the operation for the first layer or if an abnormal condition is present in the electronic device, and

wherein the controller selects at least one second processor from among the plurality of heterogeneous processors to perform an operation for the second layer of the neural network model based on the parameters of the second layer, the selected second processor including a second main target processor and at least one second sub target processor, the operation for the second layer assigned to the at least one second sub target processor if, due to being occupied by another work, the second main target processor cannot perform the operation for the second layer or if the abnormal condition is present in the electronic device.

18 . The electronic device of claim 17 , further comprising a memory for storing the neural network model.

19 . The electronic device of claim 17 , wherein the abnormal condition includes situations where an internal temperature of the electronic device rises above a reference value.

20 . The electronic device of claim 17 , wherein the abnormal condition includes situations where a remaining battery level of the electronic device falls below a reference value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2025
From: JANG, JEONG GYU
To: DEEPX CO., LTD.
Reel/Frame 072335/0679 →
Priority Claims (1)
KR 10-2025-0037583 · Mar 24, 2025 · national
References Cited (8)
US 20230297519A1 · Kim · 2023 [cited by examiner]
CN 110852439A · 2020 [cited by examiner]
KR 1020200111948A · 2020 [cited by applicant]
KR 1020210084123A · 2021 [cited by applicant]
KR 1020220049759A · 2022 [cited by applicant]
Risso, Matteo, Alessio Burrello, Giuseppe Maria Sarda, Luca Benini, Enrico Macii, Massimo Poncino, Marian Verhelst, and Daniele Jahier Pagliari. “Precision-aware latency and energy balancing on multi-accelerator platfor… [cited by examiner]
Dobiáš, Petr, Emmanuel Casseau, and Oliver Sinnen. “Comparison of Enhancing Methods for Primary/Backup Approach Meant for Fault Tolerant Scheduling.” PhD diss., Univ Rennes, Inria, CNRS, IRISA, France, 2021. (Year: 2021… [cited by examiner]
Chen, Le, Dahu Feng, Erhu Feng, Rong Zhao, Yingrui Wang, Yubin Xia, Haibo Chen, and Pinjie Xu. “HeteroLLM: Accelerating Large Language Model Inference on Mobile SoCs platform with Heterogeneous AI Accelerators.” arXiv p… [cited by examiner]
Cited By (1)
US 12,682,236