IP Library › Granted Patent US 12,591,431
Granted Patent B2
US 12,591,431 · App. 18/539,012 · Granted Mar 31, 2026

Artificial intelligence processing apparatus, and data prefetching device and method for artificial intelligence processor

Inventor: Hyun-Mi Kim (Daejeon, KR)
Assignee: Electronics and Telecommunications Research Institute
G06F9/30036G06F12/0862
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,431
App. No.
18/539,012
Granted
Mar 31, 2026
Kind
B2
Abstract

Disclosed herein are a prefetching device and method for an artificial intelligence processor. The prefetching method includes prefetching data, stored in external off-chip memory, into internal on-chip memory in the artificial intelligence processor, and storing information including an address value and a total amount of matrix operation data in at least one control and status register, as a kernel program is executed, extracting a matrix operation instruction among instructions provided from an instruction cache of the off-chip memory, determining whether prefetching is enabled based on a result of extracting the matrix operation instruction, as prefetching is enabled, determining a number of blocks to be prefetched based on the information stored in the at least one control and status register, and determining a bus burst value corresponding to the determined number of blocks and transmitting the bus burst value as a data request signal through a bus interface.

Claims (57)

1 . An artificial intelligence processing apparatus, comprising:

an off-chip memory;

a main central processing unit configured to execute an artificial intelligence program; and

one or more artificial intelligence processors,

wherein each of the one or more artificial intelligence processors comprises:

a processing core;

an on-chip memory; and

a prefetcher configured to load data stored in the off-chip memory into the on-chip memory;

wherein the processing core is implemented as a pair of a floating-point operation-based tensor processing unit configured to perform a matrix operation and a central processing unit configured to perform a normal operation.

2 . The artificial intelligence processing apparatus of claim 1 , wherein the on-chip memory comprises:

an instruction cache configured to store an instruction of the artificial intelligence program; and

a data cache configured to store operation data of the artificial intelligence program.

3 . The artificial intelligence processing apparatus of claim 2 , wherein the prefetcher comprises at least one control and status register configured to store information including an address value and a total amount of matrix operation data.

4 . The artificial intelligence processing apparatus of claim 3 , wherein the main central processing unit executes:

a compiler configured to create machine code by optimizing the artificial intelligence program; and

runtime software configured to execute the created machine code,

wherein information extracted by the compiler and the runtime software is recorded in the at least one control and status register.

5 . The artificial intelligence processing apparatus of claim 4 , wherein the main central processing unit records the information extracted by the compiler and the runtime software in the control and status register through an Advanced Peripheral Bus (APB) interface.

6 . The artificial intelligence processing apparatus of claim 4 , wherein the compiler performs:

extracting a matrix operation represented by a nested loop in the artificial intelligence program;

performing tiling to allocate the extracted matrix operation data to each of multiple processing cores;

generating a matrix operation instruction based on a result of tiling; and

creating machine code dedicated for the multiple processing cores from the matrix operation instruction.

7 . The artificial intelligence processing apparatus of claim 4 , wherein the runtime software performs:

separating a kernel program by decoding the artificial intelligence program;

allocating a dynamic memory;

allocating an address of the matrix operation data; and

setting the address of the matrix operation data in the control and status register.

8 . The artificial intelligence processing apparatus of claim 3 , wherein the prefetcher performs:

as a kernel program is executed, extracting a matrix operation instruction among instructions provided from the instruction cache;

determining whether prefetching is enabled based on a result of extracting the matrix operation instruction;

as prefetching is enabled, determining a number of blocks to be prefetched; and

determining a bus burst value corresponding to the determined number of blocks and transmitting the bus burst value as a data request signal through a bus interface.

9 . The artificial intelligence processing apparatus of claim 8 , wherein the prefetcher further performs:

receiving a Program Counter (PC) value of the central processing unit and adjusting the program counter value so that a difference between an address value of the artificial intelligence program, read from the instruction cache, and the program counter value is not increased to a certain distance or more.

10 . The artificial intelligence processing apparatus of claim 8 , wherein the prefetcher is configured to, when determining whether prefetching is enabled, determine that prefetching is enabled only when a first matrix operation instruction is extracted.

11 . The artificial intelligence processing apparatus of claim 8 , wherein the prefetcher is configured to, when determining the number of blocks, determine the number of blocks based on the address value and the total amount of data of the matrix operation data, stored in the control and status register.

12 . The artificial intelligence processing apparatus of claim 8 , wherein the prefetcher is configured to, when determining the bus burst value, determine the bus burst value based on a size of one block of the data cache and a data bandwidth of a bus interface.

13 . A prefetching device for an artificial intelligence processor, comprising:

at least one control and status register configured to prefetch data, stored in an external off-chip memory, into an internal on-chip memory in an artificial intelligence processor, and to store information including an address value and a total amount of matrix operation data;

a matrix operation discrimination unit configured to, as a kernel program is executed, extract a matrix operation instruction among instructions provided from an instruction cache of the off-chip memory;

a prefetching/non-prefetching determination unit configured to determine whether prefetching is enabled based on a result of extracting the matrix operation instruction;

a prefetch block number determination unit configured to, as prefetching is enabled, determine a number of blocks to be prefetched based on the information stored in the at least one control and status register; and

a request signal generation unit configured to determine a bus burst value corresponding to the determined number of blocks and transmit the bus burst value as a data request signal through a bus interface.

14 . The prefetching device of claim 13 , wherein the matrix operation discrimination unit is configured to receive a Program Counter (PC) value of a central processing unit of the artificial intelligence processor and adjust the program counter value so that a difference between an address value of the artificial intelligence program, read from the instruction cache, and the program counter value is not increased to a certain distance or more.

15 . The prefetching device of claim 13 , wherein the prefetching/non-prefetching determination unit determines that prefetching is enabled only when a first matrix operation instruction is extracted.

16 . The prefetching device of claim 13 , wherein the prefetch block number determination unit determines the number of blocks to be prefetched based on the address value and the total amount of matrix operation data stored in the control and status register.

17 . The prefetching device of claim 13 , wherein the request signal generation unit determines the bus burst value based on a size of one block of a data cache in the off-chip memory and a data bandwidth of a bus interface.

18 . A prefetching method for an artificial intelligence processor, comprising:

prefetching data, stored in an external off-chip memory, into an internal on-chip memory in the artificial intelligence processor, and storing information including an address value and a total amount of matrix operation data in at least one control and status register;

as a kernel program is executed, extracting a matrix operation instruction among instructions provided from an instruction cache of the off-chip memory;

determining whether prefetching is enabled based on a result of extracting the matrix operation instruction;

as prefetching is enabled, determining a number of blocks to be prefetched based on the information stored in the at least one control and status register; and

determining a bus burst value corresponding to the determined number of blocks and transmitting the bus burst value as a data request signal through a bus interface.

19 . The prefetching method of claim 18 , wherein:

determining the number of blocks to be prefetched comprises determining the number of blocks to be prefetched based on the address value and the total amount of the matrix operation data stored in the control and status register, and

transmitting as the data request signal comprises determining the bus burst value based on a size of one block of a data cache of the off-chip memory and a data bandwidth of the bus interface.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2023
From: KIM, HYUN-MI
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 065863/0242 →
Priority Claims (2)
KR 10-2022-0176165 · Dec 15, 2022 · national
KR 10-2023-0086849 · Jul 5, 2023 · national
Continuity (1)
Related Publication 20240201992A1 · Jun 20, 2024
References Cited (29)
US 5724599A · Balmer · 1998 [cited by examiner]
US 11720364B2 · Alam et al. · 2023 [cited by applicant]
US 11809981B1 · Jain · 2023 [cited by examiner]
US 20020034971A1 · Chang · 2002 [cited by examiner]
US 20110199391A1 · Olsson · 2011 [cited by examiner]
US 20130246708A1 · Ono · 2013 [cited by examiner]
US 20160055089A1 · Kim et al. · 2016 [cited by applicant]
US 20190361811A1 · Saeki et al. · 2019 [cited by applicant]
US 20210192328A1 · Ross · 2021 [cited by applicant]
US 20210397557A1 · Lee et al. · 2021 [cited by applicant]
US 20220067524A1 · Mathaikutty · 2022 [cited by examiner]
US 20220100813A1 · Lagudu · 2022 [cited by examiner]
US 20220171708A1 · Lee et al. · 2022 [cited by applicant]
US 20220261622A1 · Norrie et al. · 2022 [cited by applicant]
US 20220284658A1 · Müller · 2022 [cited by examiner]
US 20220365695A1 · Zhang · 2022 [cited by examiner]
US 20220413866A1 · Nathella · 2022 [cited by examiner]
US 20230009375A1 · Lu · 2023 [cited by examiner]
US 20230281270A1 · Ito · 2023 [cited by examiner]
US 20240069914A1 · George · 2024 [cited by examiner]
US 20240143457A1 · Sity et al. · 2024 [cited by applicant]
KR 1020140132424A · 2014 [cited by applicant]
KR 1020200047551A · 2020 [cited by applicant]
KR 1020210122043A · 2021 [cited by applicant]
KR 1020210123435A · 2021 [cited by applicant]
KR 1020210157624A · 2021 [cited by applicant]
KR 1020220074702A · 2022 [cited by applicant]
KR 1020220076325A · 2022 [cited by applicant]
Dong, “Simple but Effective Heterogeneous Main Memory with On-Chip Memory Controller Support”, Nov. 2010, IEEE (Year: 2010). [cited by examiner]