IP Library › Granted Patent US 11,551,148
Granted Patent B2
US 11,551,148 · App. 16/862,549 · Granted Jan 10, 2023

System and method for INT9 quantization

Inventors: Avinash Sodani (San Jose, CA); Ulf Hanebutte (Gig Harbor, WA); Chia-Hsin Chen (Santa Clara, CA)
Assignee: Marvell Asia Pte Ltd
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,148
App. No.
16/862,549
Granted
Jan 10, 2023
Kind
B2
Abstract

A method of converting a data stored in a memory from a first format to a second format is disclosed. The method includes extending a number of bits in the data stored in a double data rate (DDR) memory by one bit to form an extended data. The method further includes determining whether the data stored in the DDR is signed or unsigned data. Moreover, responsive to determining that the data is signed, a sign value is added to the most significant bit of the extended data and the data is copied to lower order bits of the extended data. Responsive to determining that the data is unsigned, the data is copied to lower order bits of the extended data and the most significant bit is set to an unsigned value, e.g., zero. The extended data is stored in an on-chip memory (OCM) of a processing tile of a machine learning computer array.

Claims (35)

1. A method of converting a data stored in a memory from a first format to a second format for machine learning (ML) operations, the method comprising:

extending a number of bits in the data stored in a double data rate (DDR) memory by one bit to form an extended data;

determining whether the data stored in the DDR memory is signed or unsigned data;

responsive to determining that the data is signed, adding a sign value to the most significant bit of the extended data and copying the data to lower order bits of the extended data;

responsive to determining that the data is unsigned, copying the data to lower order bits of the extended data and setting the most significant bit to an unsigned value; and

storing the extended data in an on-chip memory (OCM) of a processing tile of a machine learning computer array.

2. The method of claim 1 , wherein the data is an unsigned integer.

3. The method of claim 1 , wherein the data is a signed integer.

4. The method of claim 1 , wherein the data is 8 bits and wherein the extended data is 9 bits.

5. The method of claim 1 , wherein the extended data is an int9 data.

6. The method of claim 1 further comprising:

tracking whether the data stored in the DDR memory is signed or unsigned; and

scheduling appropriate instructions for the extended data based on whether the data is signed or unsigned.

7. The method of claim 6 further comprising performing an arithmetic logic unit (ALU) operation on the extended data as an operand.

8. The method of claim 7 further comprising storing a result of the operation in the OCM of the processing tile of the machine learning computer array.

9. The method of claim 8 further comprising storing the result stored in the OCM into the DDR memory.

10. The method of claim 9 further comprising: prior to storing the result in the DDR memory, adjusting a value of the result to a maximum value of a range for the data if the value of the result exceeds the maximum value and adjusting the value of the result to a minimum value of the range for the data if the value of the result is lower than a minimum value of the range for the data.

11. The method of claim 9 further comprising dropping the most significant bit of the result before storing the result in the DDR memory from the OCM.

12. The method of claim 1 , wherein the data stored in the DDR memory is an integer representation of the data for a floating point data.

13. The method of claim 12 , wherein the floating point data is scaled and quantized to form the data in the first format.

14. The method of claim 13 , wherein a first scaling value is used for converting the floating point data to an int8 format and wherein a second scaling value is used for converting the floating point data to a uint8 format.

15. A system comprising:

a double data rate (DDR) memory configured to store integer data in a first format; and

a machine learning processing unit comprising a plurality of processing tiles, wherein each processing tile comprises:

an on-chip memory (OCM) configured to accept and maintain an extended data that is converted from the integer data in the first format from the DDR memory, for various ML operations, wherein the extended data includes one additional bit in comparison to the integer data in the first format, and wherein the most significant bit of the extended data is signed if the integer data in the first format is signed and wherein the most significant bit of the extended data is set to an unsigned value if the integer data in the first format is unsigned and wherein the least significant bits of the extended data is the same as the integer data in the first format.

16. The system of claim 15 , wherein the integer data in the first format is either int8 or uint8.

17. The system of claim 15 , wherein the extended data is int9.

18. The system of claim 15 , wherein whether the integer data in the first format stored in the DDR memory is signed or unsigned is tracked and wherein appropriate instructions are scheduled depending whether the integer data in the first format is signed or unsigned.

19. The system of claim 15 , wherein the extended data is an operand for an operation.

20. The system of claim 19 , wherein a result of the operation is stored in the OCM.

21. The system of claim 20 , wherein the result of the operation that is stored in the OCM is further stored in the DDR memory.

22. The system of claim 21 , wherein a value of the result is adjusted to a maximum value of a range for the integer data in the first format if the value of the result exceeds the maximum value and adjusting the value of the result to a minimum value of the range for the integer data in the first format if the value of the result is lower than a minimum value of the range for the data, prior to storing the result in the DDR memory.

23. The system of claim 21 , wherein the most significant bit of the result is dropped before storing the result in the DDR memory.

24. The system of claim 15 , wherein the integer data in the first format is integer representation of a floating point data, and wherein the floating point data is scaled and quantized to form the integer data in the first format.

25. The system of claim 24 , wherein a first scaling value is used for converting the floating point data to an int8 format and wherein a second scaling value is used for converting the floating point data to a uint8 format.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2022
From: SODANI, AVINASH; HANEBUTTE, ULF; CHEN, CHIA-HSIN
To: MARVELL SEMICONDUCTOR, INC.
Reel/Frame 060216/0725 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2022
From: MARVELL SEMICONDUCTOR, INC.
To: MARVELL ASIA PTE LTD
Reel/Frame 060216/0735 →
Continuity (1)
Related Publication 20210342734A1 · Nov 4, 2021