IP Library › Granted Patent US 12,639,298
Granted Patent B2
US 12,639,298 · App. 18/769,510 · Granted May 26, 2026

Data processing acceleration apparatus

Inventors: Juhyun Kim (Yongin-si, KR); Jin Yeong Kim (Yongin-si, KR)
Assignee: XCENA Inc.
G06F16/2453G06F13/382G06F13/4221G06F16/2237G06F2213/0026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,298
App. No.
18/769,510
Granted
May 26, 2026
Kind
B2
Abstract

A data processing acceleration apparatus is disclosed. The data processing acceleration apparatus comprises a hardware accelerator, a memory device, and a memory controller configured to control the hardware accelerator and the memory device. And the memory controller includes a plurality of computing units, a first interface configured to communicate with a host processor based on a first protocol, a second interface configured to communicate with the memory device based on a second protocol, and a third interface configured to communicate with the hardware accelerator based on a third protocol.

Claims (42)

1 . A data processing acceleration apparatus, comprising:

a hardware accelerator, which is at least one of a Neural Processing Unit (NPU) or a Graphic Processing Unit (GPU) configured to execute a Large Language Model (LLM);

a memory device configured to store a vector database, which includes parameters of the LLM, latent variables of the LLM, and embedding vector data associated with information not used for training the LLM, wherein the memory device comprises a first memory device and a second memory device; and

a memory controller configured to control the hardware accelerator and the second memory device, wherein the memory controller comprises:

a plurality of computing units comprising a plurality of independent cores, each of which is configured to perform an independent pointer search, and accessory devices, each of which includes a hardware device configured to accelerate calculation of vector similarity;

a first interface, which is a Compute eXpress Link (CXL) interface configured to communicate with a host processor based on a first protocol, which is a CXL protocol;

a second interface, which is a Dual Data Rate (DDR) interface configured to communicate with the second memory device based on a second protocol, which is a DDR protocol; and

a third interface, which is a Peripheral Component Interconnect Express (PCIe) interface or a Universal Chiplet Interconnect Express (UCIe) interface configured to communicate with the hardware accelerator based on a third protocol, which is a PCIe protocol or a UCIe protocol,

wherein the first memory device is a first dual in-line memory module (DIMM)-based memory which is configured to directly communicate with the host processor based on the DDR protocol and is configured to indirectly communicate with the hardware accelerator via the memory controller based on the CXL protocol,

wherein the second memory device is a second DIMM-based memory which is configured to indirectly communicate with the host processor via the memory controller and is configured to directly communicate with the memory controller that is configured to directly communicate with the hardware accelerator based on the PCIe protocol or the UCIe protocol,

wherein the second memory device has a higher latency than the first memory device, and

wherein, while a result according to an LLM algorithm is being generated by the hardware accelerator, another request for a vector search is directly transmitted to the hardware accelerator without passing through the host processor;

wherein the memory controller is configured to:

receive a query embedding from the host processor through the first interface;

perform a vector search process for the embedding vector data included in the vector database based on the query embedding; and

transmit some of the parameters of the LLM, some of the latent variables of the LLM, a result of the vector search process, and the query embedding to the hardware accelerator through the third interface, and

the hardware accelerator is configured to generate an output of the LLM based on some of the parameters of the LLM, some of the latent variables of the LLM, the result of the vector search process, and the query embedding.

2 . The data processing acceleration apparatus according to claim 1 , wherein the memory controller is configured to:

receive a query embedding from the host processor through the first interface;

perform a vector search process on a vector database stored in the memory device based on the query embedding; and

transmit a result of the vector search process to the host processor through the first interface, or transmit the result of the vector search process and the query embedding to the hardware accelerator through the third interface.

3 . The data processing acceleration apparatus according to claim 2 , wherein the hardware accelerator is configured to generate an output of a Large Language Model (LLM) based on the result of the vector search process and the query embedding, and the memory controller is configured to:

receive the output of the LLM from the hardware accelerator through the third interface;

and transmit the output of the LLM to the host processor through the first interface.

4 . The data processing acceleration apparatus according to claim 1 , wherein an operation of updating vector data and deleting some vector data for the vector database is performed by the host processor.

5 . The data processing acceleration apparatus according to claim 1 , further comprising a storage device, wherein

the memory controller further includes a fourth interface configured to communicate with the storage device based on a fourth protocol,

the storage device stores a vector database, and

the memory device operates as a cache for the storage device.

6 . The data processing acceleration apparatus according to claim 5 , further comprising:

an additional hardware accelerator;

an additional memory device; and

an additional memory controller configured to control the additional hardware accelerator and the additional memory device,

wherein the additional memory controller includes: a plurality of additional computing units;

a fifth interface configured to communicate with the host processor based on the first protocol;

a sixth interface configured to communicate with the additional memory device based on the second protocol;

a seventh interface configured to communicate with the additional hardware accelerator based on the third protocol; and

an eighth interface configured to communicate with the storage device based on the fourth protocol, and the additional memory device operates as a cache for the storage device.

7 . The data processing acceleration apparatus according to claim 6 , wherein, if at least one of the memory controller or the additional memory controller performs a vector search process on a vector database stored in the storage device, vector data stored in the vector database operates in a read-only mode.

8 . The data processing acceleration apparatus according to claim 6 , wherein, if the host processor updates a vector database stored in the storage device, the memory controller operates as a master device, and the additional memory controller operates as a subordinate device,

the host processor updates the vector database stored in the storage device through the memory controller, and

reading and writing operations with respect to the vector database stored in the storage device of the additional memory controller are prohibited.

Assignments (2)
CHANGE OF NAME Recorded Jan 13, 2025
From: METISX CO., LTD.
To: XCENA INC.
Reel/Frame 069877/0477 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: KIM, JUHYUN; KIM, JIN YEONG
To: METISX CO., LTD.
Reel/Frame 067957/0550 →
Priority Claims (2)
KR 10-2023-0090249 · Jul 12, 2023 · national
KR 10-2023-0133627 · Oct 6, 2023 · national
Continuity (1)
Related Publication 20250021551A1 · Jan 16, 2025
References Cited (9)
US 20220137865A1 · Lee et al. · 2022 [cited by applicant]
US 20240221738A1 · Garg · 2024 [cited by examiner]
US 20240330341A1 · Xie · 2024 [cited by examiner]
US 20240406166A1 · Bell · 2024 [cited by examiner]
US 20240419706A1 · Gutierrez · 2024 [cited by examiner]
US 20250156356A1 · Paul · 2025 [cited by examiner]
KR 1020220056986A · 2022 [cited by applicant]
Choi et al.; “Advancing Processor-in-Memory (PIM) Technology: A Response to the Computational Demands of Large Language Models (LLM)”; SoC Platform Research Center, Semiconductor and Display R&D Division Korea Electroni… [cited by applicant]
“Design and Practice of Al-oriented General-purpose Vector Database Systems”; Vector Database for AI; <https://medium.com/vector-database;> Nov. 8, 2021; pp. 1-13. [cited by applicant]