IP Library Granted Patent US 12,197,886
Granted Patent B2
US 12,197,886 · App. 18/191,737 · Granted Jan 14, 2025

Neural processing device and method for converting data thereof

Inventor: Jinwook Oh (Seongnam-si, KR)
Assignee: Rebellions Inc.
G06F5/00G06F3/0604G06F3/0656G06F5/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,886
App. No.
18/191,737
Granted
Jan 14, 2025
Kind
B2
Abstract

A neural processing device and a method for converting data thereof are provided. The neural processing device comprises a first compute unit configured to receive first input data in first precision and generate first output data in the first precision by performing calculations, a second compute unit configured to receive second input data in second precision which is different from the first precision and generate second output data in the second precision by performing calculation, and a first converting buffer configured to receive and store the first output data, generate the second input data by converting the first output data into the second precision, and transmit the second input data to the second compute unit.

Claims (46)

1. A neural processing device comprising:

a first processing element (PE) array configured to receive first input data in first precision and generate first output data in the first precision by performing two-dimensional matrix multiplication in the first precision;

a second PE array configured to receive second input data in second precision which is different from the first precision and generate second output data in the second precision by performing one-dimensional calculations in the second precision; and

a first converting buffer connected to the first PE array and the second PE array and configured to receive and store the first output data, generate the second input data by converting the first output data into the second precision, and transmit the second input data to the second PE array,

wherein the first converting buffer comprises:

an input converting register configured to receive the first output data in the first precision and convert the first output data into the second input data in the second precision,

a storage configured to receive the second input data from the input converting register and store the second input data, and

an output register configured to receive the second input data from the storage and transmit the second input data to the second PE array.

2. The neural processing device of claim 1 , further comprising:

a second converting buffer configured to convert the first input data in the second precision into the first precision and provide the converted first input data to the first PE array.

3. The neural processing device of claim 2 , further comprising:

a memory unit configured to store the second output data in the second precision and provide the first input data in the second precision to the second converting buffer.

4. The neural processing device of claim 1 , further comprising:

a third converting buffer configured to receive the second output data in the second precision from the second PE array and convert the second output data into the first precision.

5. The neural processing device of claim 1 , wherein the first converting buffer has a first in first out (FIFO) structure.

6. The neural processing device of claim 1 , wherein the first PE array has a coarse-grained reconfigurable array (CGRA) structure.

7. The neural processing device of claim 1 , wherein a number of the first output data is different from a number of the second input data.

8. A method for converting data of a neural processing device, comprising:

receiving first input data in first precision by a first processing element (PE) array;

generating first output data in the first precision by the first PE array;

generating, by a first converting buffer connected to the PE array and a second PE array, second input data by receiving the first output data and converting the first output data into second precision different from the first precision,

wherein generating the second input data comprises:

receiving the first output data by an input register,

storing the first output data in a storage,

transmitting the first output data to an output converting register,

generating the second input data by converting the first output data into the second precision, and

outputting the second input data;

receiving the second input data in the second precision by the second PE array; and

generating second output data in the second precision by the second PE array.

9. The method for converting data of the neural processing device of claim 8 , wherein

the first PE array receives i pieces of first input data in the first precision and generates j pieces of the first output data in the first precision.

10. The method for converting data of the neural processing device of claim 9 , wherein

the first converting buffer has a FIFO structure, and

the first converting buffer receives the j pieces of first output data and converts the j pieces of the first output data into k pieces of the second input data in the second precision.

11. The method for converting data of the neural processing device of claim 10 , wherein

the second PE array performs one-dimensional calculations, receives the k pieces of the second output data, and generates the second output data.

12. The method for converting data of the neural processing device of claim 11 , wherein generating the second output data comprises:

outputting the second output data in the second precision by the second PE array;

transmitting, by a first FIFO buffer, h pieces of the second output data to an L 0 memory; and

storing the h pieces of the second output data in the L 0 memory.

13. The method for converting data of the neural processing device of claim 10 , wherein converting into the k pieces of the second input data in the second precision comprises:

converting the j pieces of the first output data in the first precision into j pieces of the second input data in the second precision; and

generating k pieces of the second input data in the second precision by merging the j pieces of the second input data in the second precision.

14. The method for converting data of the neural processing device of claim 8 , wherein receiving the first input data comprises:

transmitting h pieces of the first input data in the first precision by an L 0 memory; and

receiving the first input data and outputting i pieces of the first input data in the first precision by a first FIFO buffer.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded May 22, 2025
From: REBELLIONS INC.; SAPEON KOREA INC.
To: REBELLIONS INC.
Reel/Frame 071357/0522 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: OH, JINWOOK
To: REBELLIONS INC.
Reel/Frame 063137/0285 →
Priority Claims (1)
KR 10-2022-0041152 · Apr 1, 2022 · national
Continuity (1)
Related Publication 20230315336A1 · Oct 5, 2023
References Cited (16)
US 11494163B2 · Mellempudi · 2022 [cited by examiner]
US 20190347553A1 · Lo · 2019 [cited by examiner]
US 20190392287A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20210064976A1 · Sun · 2021 [cited by examiner]
JP WO2020008643A1 · 2020 [cited by applicant]
KR 1020200122256A · 2020 [cited by applicant]
KR 102258566B1 · 2021 [cited by applicant]
Venkataramani, Swagath, et al. “RaPiD: AI accelerator for ultra-low precision training and inference.” 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2021. (Year: 2021). [cited by examiner]
Oh, Jinwook, et al. “A 3.0 TFLOPS 0.62 V scalable processor core for high compute utilization AI training and inference.” 2020 IEEE Symposium on VLSI Circuits. IEEE, 2020. (Year: 2020). [cited by examiner]
Venkataramani, Swagath, et al. “Efficient AI system design with cross-layer approximate computing.” Proceedings of the IEEE 108.12 (2020): 2232-2250. (Year: 2020). [cited by examiner]
Shukla, Sunil, et al. “A scalable multi-TeraOPS core for AI training and inference.” IEEE Solid-State Circuits Letters 1.12 (2018): 217-220. (Year: 2018). [cited by examiner]
Judd, Patrick, et al. “Proteus: Exploiting numerical precision variability in deep neural networks.” Proceedings of the 2016 International Conference on Supercomputing. 2016. (Year: 2016). [cited by examiner]
Moons, Bert, and Marian Verhelst. “An energy-efficient precision-scalable ConvNet processor in 40-nm CMOS.” IEEE Journal of solid-state Circuits 52.4 (2016): 903-914. (Year: 2016). [cited by examiner]
Lee, Sae Kyu, et al. “A 7-nm four-core mixed-precision AI chip with 26.2-TFLOPS hybrid-FP8 training, 104.9-TOPS INT4 inference, and workload-aware throttling.” IEEE Journal of Solid-State Circuits 57.1 (2021): 182-197. … [cited by examiner]
Office Action for KR 10-2022-0041152 by Korean Intellectual Property Office dated Mar. 21, 2024. [cited by applicant]
Office Action for KR 10-2022-0041152 by Korean Intellectual Property Office dated Nov. 26, 2024. [cited by applicant]