IP Library › Granted Patent US 12,182,532
Granted Patent B1
US 12,182,532 · App. 18/751,662 · Granted Dec 31, 2024

Mixed-precision multiply-and-accumulation tree structure to maximize memory bandwidth usage for computational acceleration of generative large language model

Inventor: Jung-Hoon Kim (Hwaseong-si, KR)
Assignee: HyperAccel Co., Ltd.
G06F7/483
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,182,532
App. No.
18/751,662
Granted
Dec 31, 2024
Kind
B1
Abstract

Provided is a mixed-precision multiply-and-accumulation (MAC) tree structure to maximize memory bandwidth usage for computational acceleration of a generative large language model. A MAC tree-based operator may include a plurality of floating-point (FP) multipliers connected in parallel and configured to process a multiplication operation on data delivered from an external memory; a plurality of first converters configured to convert output of each of the plurality of FP multipliers from floating point to fixed point; a fixed-point (FXP) adder tree connected to the plurality of first converters and configured to process summation of multiplication results of the plurality of FP multipliers; an FXP accumulator configured to accumulate output of the FXP adder tree; and a second converter configured to convert output of the FXP accumulator from the fixed point to the floating point.

Claims (32)

1. A multiply-and-accumulation (MAC) tree-based operator comprising:

a plurality of floating-point (FP) multipliers connected in parallel and configured to process a multiplication operation on data delivered from an external memory;

a plurality of first converters configured to convert output of each of the plurality of FP multipliers from floating point to fixed point;

a fixed-point (FXP) adder tree connected to the plurality of first converters and configured to process summation of multiplication results of the plurality of FP multipliers;

an FXP accumulator configured to accumulate output of the FXP adder tree; and

a second converter configured to convert output of the FXP accumulator from the fixed point to the floating point,

wherein the MAC tree-based operator corresponds to one of a plurality of MAC tree-based operators included in a hardware accelerator for acceleration of an artificial intelligence (AI) model, and

at least one of the number of the plurality of MAC tree-based operators included in the hardware accelerator and the number of the plurality of FP multipliers included in the MAC tree-based operator is determined based on a memory bandwidth provided for the hardware accelerator.

2. The MAC tree-based operator of claim 1 , wherein the plurality of MAC tree-based operators is configured to perform a matrix multiplication operation for at least one partition among a plurality of partitions that implements the AI model.

3. The MAC tree-based operator of claim 2 , wherein the external memory includes a high bandwidth memory in which the at least one partition is stored and a local memory unit included in the hardware accelerator.

4. The MAC tree-based operator of claim 1 , wherein each of the plurality of FP multipliers comprises:

a mixed-precision FXP exponent adder for addition of the exponent; and

a mixed-precision FXP mantissa multiplier for multiplication of the mantissa.

5. The MAC tree-based operator of claim 1 , wherein each of the plurality of FP multipliers is configured to process multiplication between a first operand and a second operand with the same bit precision in response to a high-precision mode being selected and to compute a first result value.

6. A multiply-and-accumulation (MAC) tree-based operator comprising:

a plurality of floating-point (FP) multipliers connected in parallel and configured to process a multiplication operation on data delivered from an external memory;

a plurality of first converters configured to convert output of each of the plurality of FP multipliers from floating point to fixed point;

a fixed-point (FXP) adder tree connected to the plurality of first converters and configured to process summation of multiplication results of the plurality of FP multipliers;

an FXP accumulator configured to accumulate output of the FXP adder tree; and

a second converter configured to convert output of the FXP accumulator from the fixed point to the floating point,

wherein each of the plurality of FP multipliers is configured to simultaneously process first multiplication between a first operand with first bit precision and a (2-1)-th operand with second bit precision and second multiplication between the first operand and a (2-2)-th operand with the second bit precision in response to a high-performance mode being selected and to simultaneously compute a first result value of the first multiplication and a second result value of the second multiplication.

7. The MAC tree-based operator of claim 6 , wherein:

the first bit precision includes 16-bit precision, and

the second bit precision includes 8-bit precision.

8. An operating method of a multiply-and-accumulation (MAC) tree-based operator, wherein the MAC tree-based operator comprises a plurality of floating-point (FP) multipliers connected in parallel, a plurality of first converters connected to the plurality of FP multipliers, a fixed-point (FXP) adder tree, an FXP accumulator, and a second converter, and the method comprises:

processing, using the plurality of FP multipliers, a multiplication operation on data delivered from an external memory;

converting, using the plurality of first converters, a result of multiplication operation of each of the plurality of FP multipliers from floating point to fixed point;

processing, using the FXP adder tree, summation of the converted result of the plurality of FP multipliers;

accumulating, using the FXP accumulator, output of the FXP adder tree; and

converting, using the second converter, output of the FXP accumulator from the fixed point to the floating point, and

the MAC tree-based operator corresponds to one of a plurality of MAC tree-based operators included in a hardware accelerator for acceleration of an artificial intelligence (AI) model, and

at least one of the number of the plurality of MAC tree-based operators included in the hardware accelerator and the number of the plurality of FP multipliers included in the MAC tree-based operator is determined based on a memory bandwidth provided for the hardware accelerator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2024
From: KIM, JUNG-HOON
To: HYPERACCEL CO., LTD.
Reel/Frame 067813/0116 →
Priority Claims (1)
KR 10-2023-0082645 · Jun 27, 2023 · national
Cited By (1)
US 12,730,636