IP Library › Granted Patent US 12,591,410
Granted Patent B2
US 12,591,410 · App. 17/690,014 · Granted Mar 31, 2026

Data processing method for processing unit, electronic device and computer readable storage medium

Inventors: YuFei Zhang (Shanghai, CN); Dacheng Liang (Shanghai, CN)
Assignee: Shanghai Biren Technology Co., Ltd.
G06F7/483G06F7/5443G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,410
App. No.
17/690,014
Granted
Mar 31, 2026
Kind
B2
Abstract

Embodiments of the present disclosure relate to a method for data processing and to the field of computers. The method comprises: in the m th clock cycle, determining a first exponent value; in the m+1 th clock cycle, inputting two n-dimensional vectors into the processing unit to determine n second exponent values; determining a maximum value among the first exponent value and the determined n second exponent values; determining whether a target second exponent value exists among the n second exponent values, an absolute value of a difference between the target second exponent value and the maximum value being greater than or equal to a first threshold; and in response to determining that the target second exponent value exists in n second exponent values, not performing a multiply operation of two floating point numbers corresponding to the target second exponent value in the processing unit, during the m+1 th clock cycle.

Claims (29)

1 . A data processing method for systolic array, wherein the systolic array is an implementation of a general matrix multiply arithmetic logic unit and comprises a plurality of data processing units, each of the data processing units comprises a first source register, a second source register, a destination register and a calculation unit including a multiplier and an adder, a processor is used to control each of the data processing units for performing following steps, comprising:

in m th clock cycle, activating a timer to send a first clock control signal to the first source register and the second source register, so that in the m th clock cycle, the first source register inputs a first n-dimensional vector into the calculation unit which is a hardware logic component, and the second source register inputs a second n-dimensional vector into the calculation unit, performing a first dot product operation on the first n-dimensional vector and the second n-dimensional vector by the calculation unit, and determining a first exponent value of a floating point number of a result of the first dot product operation performed, and storing the first exponent value, where m is a positive integer greater than or equal to 1;

in m+1 th clock cycle, activating the timer to send a second clock control signal to the first source register and the second source register, so that in the m+1 th clock cycle, the first source register inputs a third n-dimensional vector into the calculation unit, and the second source register inputs fourth n-dimensional vector into the calculation unit, determining n second exponent values in the calculation unit, the second exponent values being obtained by adding the exponent of a floating point number in one of the third n-dimensional vector and the fourth n-dimensional vector with the exponent of a corresponding floating point number in the other of the two n-dimensional vectors, n being a positive integer greater than or equal to 1;

determining a maximum value among the first exponent value and the determined n second exponent values;

using a number of reserved bits of a mantissa of the floating point number stored in the destination register as a first threshold;

determining whether a target second exponent value exists among the n second exponent values, an absolute value of a difference between the target second exponent value and the maximum value being greater than or equal to the first threshold; and

in response to determining that the target second exponent value exists in n second exponent values, sending a signal to the multiplier of the calculation unit to cause the multiplier to bypass a multiply operation of two floating point numbers corresponding to the target second exponent value, during the m+1 th clock cycle.

2 . The data processing method according to claim 1 , wherein the processor is used to performing following steps, comprising: in response to determining that the target second exponent value does not exist among the n second exponent values, the multiply operation is performed for all data in the two n-dimensional vectors during the m+1 th clock cycle.

3 . The data processing method according to claim 1 , wherein the data processing unit comprises the destination register; and

an absolute value of a difference between the first threshold and the number of bits of the mantissa of the floating point number stored in the destination register is smaller than a second threshold.

4 . The data processing method according to claim 1 , wherein the first exponent value is obtained in the following manner:

in the m th clock cycle, performing the first dot product operation to obtain a floating point number of a first dot product operation result;

performing normalization processing for the floating point number of the first dot product operation result to obtain the first exponent value.

5 . The data processing method according to claim 4 , wherein performing the first dot product operation to obtain the floating point number of the first dot product operation result comprises:

multiplying i data from the first n-dimensional vector with corresponding i data from the second n-dimensional vector, respectively, to obtain i floating point numbers, where i=0, 1, 2 . . . n;

performing exponent matching processing on at least two floating point numbers among the i floating point numbers and the floating point numbers obtained by the dot product operation in m−1 th clock cycle;

summing the exponent matching processed floating point numbers, to obtain a floating point number of the first dot product operation result in the m th clock cycle.

6 . The data processing method according to claim 2 , wherein the exponent value of the floating point number of the result of the operation performed in the m+1 th clock cycle is determined as the first exponent value for comparison with the n second exponent values determined in m+2 th clock cycle.

7 . An electronic device, comprising:

a systolic array, comprising a plurality of data processing units, wherein each of the data processing units comprises a first source register, a second source register, destination register and a calculation unit including a multiplier and an adder, wherein the systolic array is an implementation of a general matrix multiply arithmetic logic unit;

a processor; and

a memory coupled to the processor, the memory having instructions stored therein, wherein the processor is configured to execute the instructions to control each of the data processing units to:

in m th clock cycle, activate a timer to send a first clock control signal to the first source register and the second source register, so that in the m th clock cycle, the first source register inputs a first n-dimensional vector into the calculation unit which is a hardware logic component, and the second source register inputs a second n-dimensional vector into the calculation unit, perform a first dot product operation on the first n-dimensional vector and the second n-dimensional vector by the calculation unit, and determine a first exponent value of a floating point number of a result of the first dot product operation performed, and store the first exponent value, where m is a positive integer greater than or equal to 1;

in m+1 th clock cycle, activate the timer to send a second clock control signal to the first source register and the second source register, so that in the m+1 th clock cycle, the first source register inputs a third n-dimensional vector into the calculation unit, and the second source register inputs a fourth n-dimensional vector into the calculation unit, determine n second exponent values in the calculation unit, and store the n second exponent values, the second exponent values being obtained by adding the exponent of a floating point number in one of the third n-dimensional vector and the fourth n-dimensional vector with the exponent of a corresponding floating point number in the other of the two n-dimensional vectors, n being a positive integer greater than or equal to 1;

determine a maximum value among the first exponent value and the determined n second exponent values;

use a number of reserved bits of a mantissa of the floating point number stored in a destination register as a first threshold;

determine whether a target second exponent value exists among the n second exponent values, an absolute value of a difference between the target second exponent value and the maximum value being greater than or equal to the first threshold; and

in response to determining that the target second exponent value exists in n second exponent values, send a signal to the multiplier of the calculation unit to cause the multiplier to bypass a multiply operation of two floating point numbers corresponding to the target second exponent value, during the m+1 th clock cycle.

8 . A non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions, when executed by a machine, implementing the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2022
From: ZHANG, YUFEI; LIANG, DACHENG
To: SHANGHAI BIREN TECHNOLOGY CO.,LTD
Reel/Frame 059202/0599 →
Priority Claims (1)
CN 202110258250.5 · Mar 9, 2021 · national
Continuity (1)
Related Publication 20220291901A1 · Sep 15, 2022
References Cited (9)
US 20210349694A1 · Ulrich · 2021 [cited by examiner]
US 20220269753A1 · Finch · 2022 [cited by examiner]
US 20230297337A1 · Mohamed Awad · 2023 [cited by examiner]
US 20250348278A1 · Ravgad · 2025 [cited by examiner]
WO WO2020191417A2 · 2020 [cited by examiner]
Gilani, Syed Zohaib et al. “Energy-efficient floating-point arithmetic for digital signal processors.” 2011 Conference Record of the Forty Fifth Asilomar Conference on Signals, Systems and Computers (ASILOMAR) (2011): 1… [cited by examiner]
A. Hagiescu, M. Langhammer, B. Pasca, P. Colangelo, J. Thong and N. Ilkhani, “BFLOAT MLP Training Accelerator for FPGAs,” 2019 International Conference on ReConFigurable Computing and FPGAs (ReConFig), Cancun, Mexico, 2… [cited by examiner]
Deierling, Kevin. “What Is a DPU? . . . And what's the difference between a DPU, a CPU and a GPU?”. May 20, 2020. Nvidia. https://blogs.nvidia.com/blog/whats-a-dpu-data-processing-unit/ (Year: 2020). [cited by examiner]
Patterson, David A., Hennessy, John L. 2012. Computer Architecture, Fifth Edition: A Quantitative Approach (5th. ed.). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA. (Year: 2012). [cited by examiner]