IP Library Granted Patent US 11,010,659
Granted Patent B2
US 11,010,659 · App. 15/495,020 · Granted May 18, 2021

Dynamic precision for neural network compute operations

Inventors: Kamal Sinha (Rancho Cordova, CA); Balaji Vembu (Folsom, CA); Eriko Nurvitadhi (Hillsboro, OR); Nicolas C. Galoppo Von Borries (Portland, OR); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA); Xiaoming Chen (Shanghai, CN); Anbang Yao (Beijing, CN); Tatiana Shpeisman (Menlo Park, CA); Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Farshad Akhbari (Chandler, AZ); Narayan Srinivasa (Portland, OR); Feng Chen (Shanghai, CN); Dukhwan Kim (San Jose, CA); Nadathur Rajagopalan Satish (Santa Clara, CA); John C. Weast (Portland, OR); Mike B. MacPherson (Portland, OR); Linda L. Hurd (Cool, CA); Vasanth Ranganathan (El Dorado Hills, CA); Sanjeev S. Jahagirdar (Folsom, CA)
Assignee: INTEL CORPORATION
G06N3/063G06F1/3287G06F1/3293G06F9/30014G06F9/30036G06F15/76G06F15/78G06N3/04G06N3/0445G06N3/0454G06N3/08G06N3/084G06T1/20G06T15/005G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,010,659
App. No.
15/495,020
Granted
May 18, 2021
Kind
B2
Abstract

In an example, an apparatus comprises a compute engine comprising a high precision component and a low precision component; and logic, at least partially including hardware logic, to receive instructions in the compute engine; select at least one of the high precision component or the low precision component to execute the instructions; and apply a gate to at least one of the high precision component or the low precision component to execute the instructions. Other embodiments are also disclosed and claimed.

Claims (29)

1. A graphics multiprocessor comprising:

an instruction cache to receive a stream of instructions;

an instruction unit to dispatch instructions in the stream of instructions for execution;

a compute engine comprising a fused component block comprising a high precision single instruction, multiple data (SIMD) processor which operates at a high speed and a low precision floating point single instruction, multiple data (SIMD) processor which operates at a low speed;

a shared cache memory communicatively coupled to the high-precision component and low-precision component; and

a processor to:

receive instructions from the instruction unit for execution in the compute engine;

select, based at least in part on a compute requirement of the instructions, at least one of the high precision SIMD processor or the low precision SIMD processor to execute the instructions; and

apply a gate to at least one of the high precision SIMD processor or the low precision SIMD processor to execute the instructions.

2. The graphics multiprocessor of claim 1 , wherein:

the gate comprises a clock gate.

3. The graphics multiprocessor of claim 1 , wherein:

the gate comprises a power gate.

4. A computer-based method, comprising:

receiving instructions from an instruction unit for execution in a compute engine, the compute engine comprising a fused component block comprising a high precision single instruction, multiple data (SIMD) processor which operates at a high speed and a low precision floating point single instruction, multiple data (SIMD) processor which operates at a low speed;

selecting, based at least in part on a compute requirement of the instructions, at least one of the high precision SIMD processor or the low precision SIMD processor to execute the instructions; and

applying a gate to at least one of the high precision SIMD processor or the low precision SIMD processor to execute the instructions.

5. The method of claim 4 , wherein:

the gate comprises a clock gate.

6. The method of claim 4 , wherein:

the gate comprises a power gate.

7. A non-transitory computer-readable medium comprising instructions which, when executed by a processor, cause the processor to:

receive instructions from an instruction unit for execution in a compute engine, the compute engine comprising a fused component block comprising a high precision single instruction, multiple data (SIMD) processor which operates at a high speed and a low precision floating point single instruction, multiple data (SIMD) processor which operates at a low speed;

select, based at least in part on a compute requirement of the instructions, at least one of the high precision SIMD processor or the low precision SIMD processor to execute the instructions; and

apply a gate to at least one of the high precision SIMD processor or the low precision SIMD processor to execute the instructions.

8. The computer-readable medium of claim 7 , wherein:

the gate comprises a clock gate.

9. The computer-readable medium of claim 7 , wherein:

the gate comprises a power gate.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2017
From: SINHA, KAMAL; VEMBU, BALAJI; NURVITADHI, ERIKO; GALOPPO VON BORRIES, NICOLAS C.; BARIK, RAJKISHORE; LIN, TSUNG-HAN; RAY, JOYDEEP; TANG, PING T.; STRICKLAND, MICHAEL S.; CHEN, XIAOMING; YAO, ANBANG; SHPEISMAN, TATIANA; APPU, ABHISHEK R.; KOKER, ALTUG; AKHBARI, FARSHAD; SRINIVASA, NARAYAN; CHEN, FENG; KIM, DUKHWAN; SATISH, NADATHUR RAJAGOPALAN; WEAST, JOHN C.; MACPHERSON, MIKE B.; HURD, LINDA L.; RANGANATHAN, VASANTH; JAHAGIRDAR, SANJEEV
To: INTEL CORPORATION
Reel/Frame 042821/0640 →
Continuity (1)
Related Publication 20180307971A1 · Oct 25, 2018