IP Library › Granted Patent US 11,360,767
Granted Patent B2
US 11,360,767 · App. 17/305,355 · Granted Jun 14, 2022

Instructions and logic to perform floating point and integer operations for machine learning

Inventors: Himanshu Kaul (Portland, OR); Mark A. Anders (Hillsboro, OR); Sanu K. Mathew (Hillsboro, OR); Anbang Yao (Beijing, CN); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA); Xiaoming Chen (Shanghai, CN); Tatiana Shpeisman (Menlo Park, CA); Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Kamal Sinha (Rancho Cordova, CA); Balaji Vembu (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Vasanth Ranganathan (El Dorado Hills, CA); Sanjeev Jahagirdar (Folsom, CA)
Assignee: Intel Corporation
G06F9/3001G06F7/483G06F7/5443G06F9/30014G06F9/30036G06F9/3851G06N3/0445G06N3/0454G06N3/063G06N3/08G09G5/393G06F9/3013G06F9/30025G06F17/16G06F2207/3824G06N20/00G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,360,767
App. No.
17/305,355
Granted
Jun 14, 2022
Kind
B2
Abstract

A processing apparatus is provided comprising a multiprocessor having a multithreaded architecture. The multiprocessor can execute at least one single instruction to perform parallel mixed precision matrix operations. In one embodiment the apparatus includes a memory interface and an array of multiprocessors coupled to the memory interface. At least one multiprocessor in the array of multiprocessors is configured to execute a fused multiply-add instruction in parallel across multiple threads.

Claims (43)

1. A graphics processing unit (GPU) comprising:

a plurality of memory controllers;

cache memory coupled with the plurality of memory controllers;

a graphics multiprocessor coupled with the cache memory and the plurality of memory controllers, the graphics multiprocessor having a single instruction, multiple thread (SIMT) architecture, wherein the graphics multiprocessor includes:

a register file; and

circuitry coupled with the register file, the circuitry including a first core to perform a mixed precision matrix operation and a second core to perform, in response to a single instruction, multiple compute operations, wherein the multiple compute operations include a first operation to perform a fused multiply-add and a second operation to apply a rectified linear unit function to a result of the first operation.

2. The GPU as in claim 1 , wherein the first operation and the second operation are single instruction multiple data (SIMD) operations.

3. The GPU as in claim 1 , wherein the multiple compute operations are performed on input in a 16-bit floating-point format having a 1-bit sign and an 8-bit exponent.

4. The GPU as in claim 3 , wherein the second core includes a dynamic precision processing resource that is configurable to automatically convert input in a 32-bit floating point format to the 16-bit floating-point format in conjunction with execution of the single instruction.

5. The GPU as in claim 4 , wherein the dynamic precision processing resource includes configuration circuitry to dynamically configure a precision of functional circuitry within the dynamic precision processing resource.

6. The GPU as in claim 5 , wherein the configuration circuitry is to configure a first stage of the dynamic precision processing resource to operate at a first precision and a second stage of the dynamic precision processing resource to operate at a second precision.

7. The GPU as in claim 1 , wherein the rectified linear unit function is an activation function associated with a first neuron of a neural network.

8. A system comprising:

a memory device;

a graphics processing unit (GPU) comprising a plurality of memory controllers coupled with the memory device, cache memory coupled with the plurality of memory controllers, and a graphics multiprocessor coupled with the cache memory and the plurality of memory controllers, the graphics multiprocessor having a single instruction, multiple thread (SIMT) architecture, wherein the graphics multiprocessor includes:

a register file; and

circuitry coupled with the register file, the circuitry including a first core to perform a mixed precision matrix operation and a second core to perform, in response to a single instruction, multiple compute operations, wherein the multiple compute operations include a first operation to perform a fused multiply-add and a second operation to apply a rectified linear unit function to a result of the first operation.

9. The system as in claim 8 , wherein the first operation and the second operation are single instruction multiple data (SIMD) operations.

10. The system as in claim 8 , wherein the multiple compute operations are performed on input in a 16-bit floating-point format having a 1-bit sign and an 8-bit exponent.

11. The system as in claim 10 , wherein the second core includes a dynamic precision processing resource that is configurable to automatically convert input in a 32-bit floating point format to the 16-bit floating-point format in conjunction with execution of the single instruction.

12. The system as in claim 11 , wherein the dynamic precision processing resource includes configuration circuitry to dynamically configure a precision of functional circuitry within the dynamic precision processing resource.

13. The system as in claim 12 , wherein the configuration circuitry is to configure a first stage of the dynamic precision processing resource to operate at a first precision and a second stage of the dynamic precision processing resource to operate at a second precision.

14. The system as in claim 8 , wherein the rectified linear unit function is an activation function associated with a first neuron of a neural network.

15. A method comprising:

fetching a single instruction from a cache memory of a graphics processing unit (GPU); loading operands associated with the single instruction into a register file of the GPU;

decoding the single instruction into a decoded instruction;

executing the decoded instruction via circuitry coupled with the register file, the circuitry including a first core and a second core, the first core to perform a mixed precision matrix operation and the second core to execute the decoded instruction to perform multiple compute operations; and

performing the multiple compute operations via the second core, wherein the multiple compute operations include a first operation to perform a fused multiply-add and a second operation to apply a rectified linear unit function to a result of the first operation.

16. The method as in claim 15 , wherein the first operation and the second operation are single instruction multiple data (SIMD) operations.

17. The method as in claim 15 , further comprising performing the multiple compute operations on input in a 16-bit floating-point format having a 1-bit sign and an 8-bit exponent.

18. The method as in claim 17 , further comprising automatically converting input in a 32-bit floating point format to the 16-bit floating-point format in conjunction with execution of the decoded instruction.

19. The method as in claim 18 , further comprising dynamically configuring a precision of functional circuitry within the second core in conjunction with performing the fused multiply-add.

20. The method as in claim 19 , wherein dynamically configuring the precision of the functional circuitry includes configuring a first stage of the functional circuitry to operate at a first precision and a second stage of the functional circuitry to operate at a second precision.

21. An apparatus comprising:

means for fetching a single instruction from a cache memory of a processing unit;

means for loading operands associated with the single instruction;

means for decoding the single instruction into a decoded instruction;

means for executing the decoded instruction by performing multiple compute operations; and

means for performing the multiple compute operations via a first operation to perform a fused multiply-add and a second operation to apply a rectified linear unit function to a result of the first operation.

22. The apparatus as in claim 21 , wherein the first operation and the second operation are single instruction multiple data (SIMD) operations.

23. The apparatus as in claim 21 , further comprising means for performing the multiple compute operations on input in a 16-bit floating-point format having a 1-bit sign and an 8-bit exponent.

24. The apparatus as in claim 21 , wherein the apparatus is a general purpose graphics processing unit (GPGPU).

25. The apparatus as in claim 21 , wherein the apparatus is a compute accelerator to accelerate operations associated with a neural network.

Continuity (7)
Continuation 17169232 · Feb 5, 2021
Continuation 17115989 · Dec 9, 2020
Continuation 16432402 · Jun 5, 2019
Continuation 15819152 · Nov 21, 2017
Continuation 15787129 · Oct 18, 2017
Provisional Application 62491699 · Apr 28, 2017
Related Publication 20220019431A1 · Jan 20, 2022
Cited By (20)
US 12,198,222 US 12,204,487 US 12,210,477 US 12,217,053 US 12,242,414 US 12,293,431 US 12,321,310 US 12,361,600 US 12,386,779 US 12,411,695 US 12,493,922 US 12,554,674 US 12,561,276 US 12,561,277 US 12,572,997 US 12,670,121 US 12,688,146 US 12,730,759 US 12,737,317 US 12,737,318