IP Library Granted Patent US 10,410,098
Granted Patent B2
US 10,410,098 · App. 15/494,710 · Granted Sep 10, 2019

Compute optimizations for neural networks

Inventors: Kevin Nealis (San Jose, CA); Anbang Yao (Beijing, CN); Xiaoming Chen (Shanghai, CN); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Sara S. Baghsorkhi (San Jose, CA); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Cordova, CA)
Assignee: Intel Corporation
G06K9/66G06F9/3001G06F9/3851G06F9/3887G06F9/3893G06F9/46G06K9/00973G06N3/0445G06N3/0454G06N3/063G06N3/084G06T1/20G06F2207/4824
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,410,098
App. No.
15/494,710
Filed
Apr 24, 2017
Granted
Sep 10, 2019
Kind
B2
Art Unit
2183
USPC
712/221
Abstract

One embodiment provides for a compute apparatus to perform machine learning operations, the apparatus comprising a decode unit to decode a single instruction into a decoded instruction that specifies multiple operands including an input value and a quantized weight value associated with a neural network and an arithmetic logic unit including a barrel shifter, an adder, and an accumulator register, wherein to execute the decoded instruction, the barrel shifter is to shift the input value by the quantized weight value to generate a shifted input value and the adder is to add the shifted input value to a value stored in the accumulator register and update the value stored in the accumulator register.

Claims (24)

1. A compute apparatus to perform machine learning operations, the apparatus comprising:

a decode unit to decode a single instruction into a decoded instruction that specifies multiple operands including an input value and a quantized weight value associated with a neural network, wherein the quantized weight value is an exponent value for a neural network weight and the neural network weight is constrained to a power of a base value; and

an arithmetic logic unit including a barrel shifter, an adder, and an accumulator register, wherein to execute the decoded instruction, the barrel shifter is to shift the input value by the quantized weight value to generate a shifted input value and the adder is to add the shifted input value to a value stored in the accumulator register and update the value stored in the accumulator register.

2. The compute apparatus as in claim 1 , additionally including an output register to store an output value of the single instruction.

3. The compute apparatus as in claim 1 , wherein the neural network weight is constrained to a power of two value.

4. The compute apparatus as in claim 3 , wherein an exponent associated with the quantized weight value is input to the barrel shifter.

5. The compute apparatus as in claim 4 , wherein the input value is a multiple bit input value.

6. The compute apparatus as in claim 1 , wherein the compute apparatus includes multiple arithmetic logic units configured as a single instruction multiple data compute unit.

7. The compute apparatus as in claim 6 , wherein the single instruction multiple data compute unit is to perform operations for multiple threads of a single instruction multiple thread compute architecture.

8. The compute apparatus as in claim 1 , wherein the compute apparatus is a system on a chip integrated circuit including a media processor and a vision processor.

9. The compute apparatus as in claim 8 , wherein the media processor is to decode multiple simultaneous video streams and output the multiple decoded video streams to an on-chip memory.

10. The compute apparatus as in claim 9 , wherein the vision processor is to parse a decoded video stream to perform processing operations on frames of the decoded video stream via a trained image recognition model associated with the neural network.

11. A method of performing machine learning operations, the method comprising:

decoding a single instruction specifying multiple operands, the operands specifying data including an input value and a quantized weight value of a neural network wherein the quantized weight value is an exponent value for a neural network weight and the neural network weight is constrained to a power of a base value;

issuing the single instruction for execution within a compute unit of a general-purpose graphics processing unit; and

responsive to the execution of the single instruction, generating a result based on shifting the input value by the quantized weight value of the neural network via barrel shifter logic and adding the shifted value to a value stored in an accumulation register.

12. The method as in claim 11 , wherein the neural network weight is constrained to a power of two value.

13. A data processing system comprising:

a general-purpose graphics processing unit comprising a decode unit to decode a single instruction into a decoded instruction that specifies multiple operands including an input value and a quantized weight value associated with a neural network, wherein the quantized weight value is an exponent value for a neural network weight, the neural network weight constrained to a power of a base value, and an arithmetic logic unit including a barrel shifter, an adder, and an accumulator register, wherein to execute the decoded instruction, the barrel shifter is to shift the input value by the quantized weight value to generate a shifted input value and the adder is to add the shifted input value to a value stored in the accumulator register and update the value stored in the accumulator register; and

a memory coupled with the general-purpose graphics processing unit.

14. The data processing system as in claim 13 , the general-purpose graphics processing unit including an output register to store an output value of the single instruction.

15. The data processing system as in claim 13 , wherein the neural network weight is constrained to a power of two value.

16. The data processing system as in claim 15 , wherein an exponent associated with the quantized weight value is input to the barrel shifter.

17. The data processing system as in claim 16 , wherein the input value is a multiple bit input value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2017
From: BAGHSORKHI, SARA S.; YAO, ANBANG; CHEN, XIAOMING; OULD-AHMED-VALL, ELMOUSTAPHA; NEALIS, KEVIN; NURVITADHI, ERIKO; VEMBU, BALAJI; GALOPPO VON BORRIES, NICOLAS C.; BARIK, RAJKISHORE; LIN, TSUNG-HAN; SINHA, KAMAL
To: INTEL CORPORATION
Reel/Frame 043611/0180 →
Continuity (1)
Related Publication 20180307950A1 · Oct 25, 2018
Cited By (2)
US 12,327,179 US 12,591,776