IP Library Granted Patent US 10,929,760
Granted Patent B1
US 10,929,760 · App. 16/420,028 · Granted Feb 23, 2021

Architecture for table-based mathematical operations for inference acceleration in machine learning

Inventors: Avinash Sodani (San Jose, CA); Ulf Hanebutte (Gig Harbor, WA); Chia-Hsin Chen (Santa Clara, CA)
Assignee: Marvell Asia Pte, Ltd.
G06N5/04G06F1/0307G06F7/483G06F15/7821G06F17/17G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,929,760
App. No.
16/420,028
Granted
Feb 23, 2021
Kind
B1
Abstract

A processing unit to support inference acceleration for machine learning (ML) comprises an inline post processing unit configured to accept and maintain one or more lookup tables for performing each of one or more non-linear mathematical operations. The inline post processing unit is further configured to accept data from a set of registers maintaining output from a processing block instead of streaming the data from an on-chip memory (OCM), perform the one or more non-linear mathematical operations on elements of the data from the processing block via their corresponding lookup tables, and stream post processing result of the one or more non-linear mathematical operations back to the OCM after the one or more non-linear mathematical operations are complete.

Claims (48)

1. A processing unit to support inference acceleration for machine learning (ML), comprising:

an inline post processing unit configured to

accept and maintain one or more lookup tables for performing each of one or more non-linear mathematical operations;

accept data from a set of registers maintaining a final output from a processing block instead of streaming the data from an on-chip memory (OCM);

divide the one or more non-linear mathematical operations into multiple sections, where each section is represented by a curve that is extrapolated based on a specific lookup table;

determine a value of the one or more non-linear mathematical operations by referencing the specific lookup table for each section of the multiple sections associated with an input value to perform the one or more non-linear mathematical operations on elements of the data from the processing block;

stream post processing result of the one or more non-linear mathematical operations back to the OCM after the one or more non-linear mathematical operations are complete.

2. The processing unit of claim 1 further comprising one or more of:

an inline rectified linear unit (ReLU) configured to perform a rectified linear operation on the final output from the processing block;

an inline quantization unit configured to perform a quantization operation the final output from the processing block.

3. The processing unit of claim 1 , wherein:

the inline post processing unit is configured to utilize multiple lookup tables to approximate and implement one of the one or more non-linear mathematical operations via piece-wise linear approximation.

4. The processing unit of claim 1 , wherein:

the one or more of the non-linear mathematical operations is a logarithmic operation for floating-point input values.

5. The processing unit of claim 4 , wherein:

the inline post processing unit is configured to conduct an input range check on a floating-point input value to the logarithmic operation and return an error indication if the input value is non-positive.

6. The processing unit of claim 4 , wherein:

the inline post processing unit is configured to implement the logarithmic operation for a floating-point input value by adopting a floating number expression of the input value as exponent and mantissa portions and using the exponent and mantissa values of the input value for the computation.

7. The processing unit of claim 6 , wherein:

the inline post processing unit is configured to implement the logarithmic operation for the floating-point input value by utilizing multiple lookup tables and a Taylor series expansion of different portions of the floating number expression of the input value, respectively.

8. The processing unit of claim 7 , wherein:

the inline post processing unit is configured to implement the logarithmic operation on the mantissa portion of the input value using a single lookup table via one same index value.

9. The processing unit of claim 6 , wherein:

the inline post processing unit is configured to implement the logarithmic operation for the floating-point input value by utilizing multiple lookup tables for different portions of the floating number expression of the input value, respectively, without a Taylor series expansion.

10. The processing unit of claim 9 , wherein:

the inline post processing unit is configured to implement the logarithmic operation for the floating-point input value by replacing the Taylor series expansion with a table lookup operation.

11. A method to support inference acceleration for machine learning (ML), comprising:

accepting and maintaining one or more lookup tables for performing each of one or more non-linear mathematical operations;

accepting data from a set of registers maintaining a final output from a processing block instead of streaming the data from an on-chip memory (OCM);

performing the one or more non-linear mathematical operations on elements of the data from the processing block via a corresponding lookup associated therewith from the one or more lookup tables, wherein the performing includes:

dividing the one or more non-linear mathematical operations into multiple sections, where each section is represented by a curve that is extrapolated based on a specific lookup table;

determining a value of the one or more non-linear mathematical operations by referencing the specific lookup table corresponding to a section from the multiple sections associated with an input value;

streaming post processing result of the one or more non-linear mathematical operations back to the OCM after the one or more non-linear mathematical operations are complete.

12. The method of claim 11 , further comprising:

utilizing multiple lookup tables to approximate and implement one of the one or more non-linear mathematical operations via piece-wise linear approximation.

13. The method of claim 11 , wherein:

one of the one or more non-linear mathematical operations is a logarithmic operation for floating-point input values.

14. The method of claim 13 , further comprising:

conducting an input range check on a floating-point input value to the logarithmic operation and return an error indication if the input value is non-positive.

15. The method of claim 13 , further comprising:

implementing the logarithmic operation for a floating-point input value by adopting a floating number expression of the input value as exponent and mantissa portions and using the exponent and mantissa values of the input value for the computation.

16. The method of claim 15 , further comprising:

implementing the logarithmic operation for the floating-point input value by utilizing multiple lookup tables and a Taylor series expansion of different portions of the floating number expression of the input value, respectively.

17. The method of claim 16 , further comprising:

implementing the logarithmic operation on the mantissa portion of the input value using a single lookup table via one same index value.

18. The method of claim 15 , further comprising:

implementing the logarithmic operation for the floating-point input value by utilizing multiple lookup tables for different portions of the floating number expression of the input value, respectively, without a Taylor series expansion.

19. The method of claim 18 , further comprising: implementing the logarithmic operation for the floating-point input value by replacing the Taylor series expansion with a table lookup operation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: SODANI, AVINASH; HANEBUTTE, ULF; CHEN, CHIA-HSIN
To: MARVELL SEMICONDUCTOR, INC.
Reel/Frame 054540/0385 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: MARVELL SEMICONDUCTOR, INC.
To: MARVELL INTERNATIONAL LTD.
Reel/Frame 054540/0396 →
Continuity (2)
Continuation In Part 16226559 · Dec 19, 2018
Provisional Application 62675076 · May 22, 2018