IP Library › Granted Patent US 12,505,335
Granted Patent B2
US 12,505,335 · App. 19/260,485 · Granted Dec 23, 2025

Processing heterogeneous generative artificial intelligence models

Inventor: Lok Won Kim (Yongin-si, KR)
Assignee: DEEPX CO., LTD.
G06N3/0475
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,335
App. No.
19/260,485
Granted
Dec 23, 2025
Kind
B2
Abstract

A device includes a memory of a capacity to store a generative neural network model with first parameters. The device also includes a neural processing unit that generates a response corresponding to an input query utilizing the generative neural network model stored in the memory. The neural processing unit may store execution code of the generative neural network model compiled to process speculative decoding.

Claims (38)

1 . A device comprising:

a first memory of a first capacity and configured to store first parameters of a first generative neural network model;

a second memory of a second capacity and configured to store second parameters of a second generative neural network model, a number of the second parameters is larger than a number of the first parameters; and

a neural processing unit configured to:

read at least a subset of the first parameters to execute the first generative neural network model,

generate a response corresponding to an input query by processing the input query using the first generative neural network model,

read at least a subset of the second parameters to execute the second generative neural network model,

perform an accept or rejection operation on the response using the second generative neural network model alternately with processing of the input query using the first generative neural network model to perform speculative decoding,

generate attention parameters from at least a portion of the input query processed by the first generative neural network model, and

perform the accept and rejection operation by the second generative neural network model using the attention parameters.

2 . The device of claim 1 , wherein the fust-neural processing unit is further configured to store at least part of first execution code for executing the first generative neural network model.

3 . The device of claim 2 , wherein the neural processing unit is configured to store at least part of second execution code for executing the second generative neural network model.

4 . The device of claim 1 , wherein the neural processing unit further comprises an internal memory configured to communicate with the first memory and a controller circuit configured to control operations of the fust-neural processing unit.

5 . The device of claim 4 , wherein first execution code for executing the first generative neural network model is stored in the internal memory or in the controller circuit.

6 . The device of claim 1 , further comprising a second memory of a second capacity configured to store second parameters of the second generative neural network model.

7 . The device of claim 1 , wherein the response comprises at least a token generated by executing the first generative neural network model.

8 . The device of claim 1 , wherein the neural processing unit is configured to operate in a low-power mode before receiving the input query.

9 . The device of claim 1 , wherein the neural processing unit is configured to store execution states of the first generative neural network model and the second generative neural network model.

10 . The device of claim 9 , wherein the states comprise at least one of program counters, pointers to attention parameters, and statuses of operations of the first generative neural network model and the second generative neural network model.

11 . The device of claim 9 , wherein the neural processing unit is configured to prefetch the at least subset of the second parameters during generating of the response using the first generative neural network model.

12 . The device of claim 1 , further comprising:

a first communication bus between the first memory and the neural processing unit, the first communication bus configured to transmit the at least subset of the first parameters to the neural processing unit; and

a second communication bus between the second memory and the neural processing unit, the second communication bus separate from the first communication bus and configured to transmit the at least subset of the second parameters.

13 . The device of claim 12 , wherein the first communication bus has a wider data path compared to the second communication bus.

14 . The device of claim 12 , wherein the second communication bus has more channels or a higher clock frequency compared to the first communication bus.

15 . The device of claim 1 , wherein the generating of the attention parameters is skipped when using the second generative neural network model.

16 . The device of claim 1 , wherein the generated attention parameters are stored in the first memory and transferred to the second memory for reading to perform the accept and rejection operation by the second generative neural network model via direct memory-to-memory transfer.

17 . The device of claim 1 , wherein the neural processing unit is further configured to:

establish a direct mapping of the generated attention parameters for access during performance of the accept and rejection operation using the generated attention parameters.

18 . A device comprising:

a memory configured to store a first generative neural network model and a second generative neural network model having a larger number of parameters than the first generative neural network model; and

a neural processing unit configured to:

perform speculative decoding by executing the first generative neural network model and the second generative neural network model in a time-divisional manner, generate a response corresponding to an input query by reading the first generative neural network model from the memory and executing the first generative neural network model on the input query,

generate attention parameters from at least a portion of the input query Processed by the first generative neural network model, and

perform an accept and rejection operation by the second generative neural network model using the attention parameters, the neural processing unit comprising:

a processing core circuit configured to receive input integer parameters,

a vector core circuit and a scalar core circuit configured to receive input floating-point parameters, and

a number format conversion circuit configured to convert between the integer parameters and the floating-point parameters to process the generative neural network model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2025
From: KIM, LOK WON
To: DEEPX CO., LTD.
Reel/Frame 071611/0937 →
Priority Claims (1)
KR 10-2024-0017704 · Feb 5, 2024 · national
Continuity (2)
Continuation PCTKR2025001719 · Feb 5, 2025
Related Publication 20250335752A1 · Oct 30, 2025
References Cited (22)
US 20230079074A1 · Cella et al. · 2023 [cited by applicant]
US 20240320433A1 · Lott · 2024 [cited by examiner]
US 20240354346A1 · Lott · 2024 [cited by examiner]
US 20240362468A1 · Lee · 2024 [cited by examiner]
US 20250021761A1 · Santhanam · 2025 [cited by examiner]
US 20250124255A1 · Bergner · 2025 [cited by examiner]
US 20250200358A1 · Lee · 2025 [cited by examiner]
US 20250209271A1 · Imanigooghari · 2025 [cited by examiner]
US 20250231989A1 · Lott · 2025 [cited by examiner]
US 20250245430A1 · Jeon · 2025 [cited by examiner]
US 20250245530A1 · Goel · 2025 [cited by examiner]
KR 1020220036980A · 2022 [cited by applicant]
KR 1020220067871A · 2022 [cited by applicant]
KR 1020220160814A · 2022 [cited by applicant]
KR 102481428B1 · 2022 [cited by applicant]
Kim, Lok-Won. “DeepX: Deep learning accelerator for restricted boltzmann machine artificial neural networks.” IEEE transactions on neural networks and learning systems 29.5 (2017): 1441-1453. (Year: 2018). [cited by examiner]
Liu, Xiaoxuan, et al. “Online speculative decoding.” arXiv preprint arXiv:2310.07177 v2 (2023). (Year: 2023). [cited by examiner]
Park, Brian Changeun. “Adaptive Speculative Decoding for Large Language Models.” (2024): i-73 (Year: 2024). [cited by examiner]
Hooper, Coleman, et al. “SPEED: Speculative Pipelined Execution for Efficient Decoding.” arXiv e-prints (Jan. 2024): arXiv-2310 v2. (Year: 2024). [cited by examiner]
Le, Hung. “Memory and attention in deep learning.” arXiv preprint arXiv:2107.01390 (2021). (Year: 2021). [cited by examiner]
HernÃndez, AdriÃn, and Josà © M. Amigà [cited by examiner]
International Search Report of PCT/KR2025/001719 mailed on May 2, 2025. [cited by applicant]