IP Library Patent Application 19055819
Patent Application
App. No. 19/055,819

LUT-FREE HARDWARE BASED SOFTMAX ACCELERATOR

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/055,819
Abstract

SoftMax operation is one part of a deep neural network (DNN). Because computing SoftMax is complex and time-consuming, the SoftMax operation can limit the overall execution latency of the DNN. To address this issue, an in-line data path is added to pass output data from a matrix-to-matrix multiplication core to a hardware SoftMax accelerator. During a denominator phase of the SoftMax operation, the SoftMax accelerator can operate in-line to produce a denominator value using output values generated by the matrix-to-matrix multiplication core and received over the in-line data path. During a numerator phase of the SoftMax operation, the SoftMax accelerator can calculate SoftMax outputs using output values generated by the matrix-to-matrix multiplication core and retrieved from a memory. In other words, the SoftMax accelerator can produce partial results while the matrix-to-matrix multiplication is in-flight to cut down overall latency and reduce memory transactions.

Claims (56)

1 . An apparatus, comprising:

a memory;

a matrix-to-matrix multiplication core; and

a SoftMax accelerator having:

an in-line data path to receive output data from the matrix-to-matrix multiplication core;

a compute block having one or more registers, one or more arithmetic components, and a finite state machine; and

a memory controller to provide the output data received via the in-line data path to the compute block.

2 . The apparatus of claim 1 , wherein the memory controller is further to write one or more SoftMax outputs generated by the compute block to the memory.

3 . The apparatus of claim 1 , wherein the memory controller is further to receive an instruction to perform at least a part of SoftMax in-line with the matrix-to-matrix multiplication core.

4 . The apparatus of claim 1 , wherein the memory controller is further to receive an instruction to perform SoftMax, the instruction having a variable to indicate standalone execution or at least partial in-line execution.

5 . The apparatus of claim 1 , wherein the memory controller is further to receive an instruction to perform SoftMax, the instruction having one or more of:

an output base address of the memory at which one or more SoftMax outputs generated by the compute block is to be written;

an input base address of the memory from which the output data generated by the matrix-to-matrix multiplication core is to be read; and

a value specifying an amount of the output data from the in-line data path that the SoftMax accelerator is to process.

6 . The apparatus of claim 1 , wherein the finite state machine has an idle phase, a denominator phase, and a numerator phase.

7 . The apparatus of claim 1 , wherein:

the one or more arithmetic components comprises an adder; and

the finite state machine disables a processing path having the adder in response to determining the adder has performed a number of additions.

8 . The apparatus of claim 1 , wherein:

the one or more registers comprise a register to store an accumulated value.

9 . The apparatus of claim 1 , wherein:

the one or more registers comprise a register to store a reciprocal of an accumulated value.

10 . The apparatus of claim 1 , wherein:

the one or more arithmetic components comprise one or more of: an exponential, an adder, a reciprocal, and a multiplier.

11 . The apparatus of claim 1 , wherein:

the one or more registers comprise a register to store an accumulated value, and a further register to store a scale factor for the accumulated value.

12 . The apparatus of claim 1 , wherein:

the one or more arithmetic components comprise a component which operates with a floating-point format where an exponent part is represented by 9 bits or 10 bits.

13 . The apparatus of claim 1 , wherein:

the one or more arithmetic components comprise an exponential and an adder;

the exponential outputs an exponent value and an unnormalized mantissa value; and

the adder receives the exponent value, the unnormalized mantissa value, and an accumulated value.

14 . The apparatus of claim 13 , wherein:

the exponent value is represented by 9 bits or 10 bits; and

the accumulated value has a floating-point format where an exponent part is represented by 9 bits or 10 bits.

15 . The apparatus of claim 13 , wherein the unnormalized mantissa value is represented by 26 bits.

16 . A method, comprising:

receiving an output value from a matrix-to-matrix multiplication core over an in-line data path;

computing an exponential based on the output value;

adding, the exponential to an accumulated value stored in a register;

computing a reciprocal value of the accumulated value and storing the reciprocal value in the register;

reading a stored output value generated by the matrix-to-matrix multiplication core from a memory;

computing a further exponential based on the stored output value;

multiplying the further exponential by the reciprocal value to produce a SoftMax output; and

writing the SoftMax output to the memory.

17 . The method of claim 16 , further comprising:

receiving an instruction to perform at least a part of SoftMax in-line with the matrix-to-matrix multiplication core.

18 . The method of claim 16 , further comprising:

receiving an instruction to perform SoftMax, the instruction comprising one or more of:

an output base address for writing SoftMax outputs to a memory;

an input base address from which output data generated by the matrix-to-matrix multiplication core is to be read from the memory; and

a value specifying a number of output values from the in-line data path to process.

19 . One or more non-transitory computer-readable media storing instructions, that when executed on a processor, cause the processor to:

transmit an instruction to a matrix-to-matrix multiplication core to perform multiplication of an input tensor and a further input tensor and store an output tensor in a memory; and

transmit a further instruction to a SoftMax accelerator invoking the SoftMax accelerator to perform at least a part of SoftMax in-line with the matrix-to-matrix multiplication core.

20 . The one or more non-transitory computer-readable media of claim 19 , wherein the further instruction comprises a value specifying a number of output values received from an in-line data path from the matrix-to-matrix multiplication core to be processed by the SoftMax accelerator.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2025
From: INTEL CORPORATION
To: ALTERA CORPORATION
Reel/Frame 072423/0899 →
SECURITY INTEREST Recorded Sep 12, 2025
From: ALTERA CORPORATION
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 073431/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2025
From: INTEL CORPORATION
To: ALTERA CORPORATION
Reel/Frame 072704/0307 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2025
From: PILLAI, KAMLESH; PASCA, BOGDAN; LANGHAMMER, MARTIN
To: INTEL CORPORATION
Reel/Frame 070242/0593 →