IP Library Granted Patent US 10,521,349
Granted Patent B2
US 10,521,349 · App. 16/277,267 · Granted Dec 31, 2019

Extend GPU/CPU coherency to multi-GPU cores

Inventors: Chandrasekaran Sakthivel (Sunnyvale, CA); Prasoonkumar Surti (Folsom, CA); John C. Weast (Portland, OR); Sara S. Baghsorkhi (San Jose, CA); Justin E. Gottschlich (Santa Clara, CA); Abhishek R. Appu (El Dorado Hills, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Joydeep Ray (Folsom, CA); Narayan Srinivasa (Portland, OR); Feng Chen (Shanghai, CN); Ben J. Ashbaugh (Folsom, CA); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Rancho Cordova, CA); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsom, CA); Altug Koker (El Dorado Hills, CA)
Assignee: INTEL CORPORATION
G06F12/0837G06F12/0815G06N3/0445G06N3/0454G06N3/063G06N3/08G06N3/084G06N3/088G06N20/00G06T1/20G06F2212/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,521,349
App. No.
16/277,267
Granted
Dec 31, 2019
Kind
B2
Abstract

In an example, an apparatus comprises a plurality of processing unit cores, a plurality of cache memory modules associated with the plurality of processing unit cores, and a machine learning model communicatively coupled to the plurality of processing unit cores, wherein the plurality of cache memory modules share cache coherency data with the machine learning model. Other embodiments are also disclosed and claimed.

Claims (34)

1. An apparatus, comprising a processor to:

compute a first three-dimensional convolution across a set of three frames comprising a first frame, a second frame, and a third frame in an image sequence;

store a convolutional data set for the second frame and the third frame in a memory;

receive a fourth frame in the image sequence;

use the convolutional data set for the second frame and the third frame to compute a second three-dimensional convolution across a set of three frames comprising the second frame, the third frame, and the fourth frame; and

delete the convolutional data set from the memory in response to a determination that the convolutional data set has no dependencies in subsequent convolutional operations.

2. The apparatus of claim 1 , the processor comprising a plurality of execution units (EUs) including a first type of execution unit and a second type of execution unit.

3. The apparatus of claim 2 , wherein:

the first type of execution unit and the second type of execution unit reside on a common die.

4. The apparatus of claim 1 , wherein the memory comprises a cache memory communicatively coupled to the processor.

5. The apparatus of claim 4 , the processor to issue an instruction to mark a specific cache line in the cache memory for eviction in response to a determination that data in the cache line is no longer required for a subsequent convolutional operation.

6. A computer-based method comprising:

computing, in a processor, a first three-dimensional convolution across a set of three frames comprising a first frame, a second frame, and a third frame in an image sequence;

storing a convolutional data set for the second frame and the third frame in memory communicatively coupled to the processor;

receiving, in the processor, a fourth frame in the image sequence;

using, in the processor, the convolutional data set for the second frame and the third frame to compute a second three-dimensional convolution across a set of three frames comprising the second frame, the third frame, and the fourth frame; and

deleting the convolutional data set from the memory in response to a determination that the convolutional data set has no dependencies in subsequent convolutional operations.

7. The method of claim 6 , the processor comprising a plurality of execution units (EUs) including a first type of execution unit and a second type of execution unit.

8. The method of claim 7 , wherein the first type of execution unit and the second type of execution unit reside on a common die.

9. The method of claim 6 , wherein the memory comprises a cache memory communicatively coupled to the processor.

10. The method of claim 9 , further comprising issuing an instruction to mark a specific cache line in the cache memory for eviction in response to a determination that data in the cache line is no longer required for a subsequent convolutional operation.

11. An electronic device, comprising:

a processor to:

compute a first three-dimensional convolution across a set of three frames comprising a first frame, a second frame, and a third frame in an image sequence;

store a convolutional data set for the second frame and the third frame in memory;

receive a fourth frame in the image sequence;

use the convolutional data set for the second frame and the third frame to compute a second three-dimensional convolution across a set of three frames comprising the second frame, the third frame, and the fourth frame; and

delete the convolutional data set from the memory in response to a determination that the convolutional data set has no dependencies in subsequent convolutional operations; and

a computer readable memory communicatively coupled to the processor.

12. The electronic device of claim 11 , the processor comprising a plurality of execution units (EUs) including a first type of execution unit and a second type of execution unit.

13. The electronic device of claim 12 , wherein:

the first type of execution unit and the second type of execution unit reside on a common die.

14. The electronic device of claim 11 , wherein the memory comprises a cache memory communicatively coupled to the processor.

15. The electronic device of claim 14 , the processor to issue an instruction to mark a specific cache line in the cache memory for eviction in response to a determination that data in the cache line is no longer required for a subsequent convolutional operation.

Continuity (2)
Continuation 15489149 · Apr 17, 2017
Related Publication 20190243764A1 · Aug 8, 2019