IP Library Granted Patent US 10,261,903
Granted Patent B2
US 10,261,903 · App. 15/489,149 · Granted Apr 16, 2019

Extend GPU/CPU coherency to multi-GPU cores

Inventors: Chandrasekaran Sakthivel (Sunnyvale, CA); Prasoonkumar Surti (Folsom, CA); John C. Weast (Portland, OR); Sara S. Baghsorkhi (San Jose, CA); Justin E. Gottschlich (Santa Clara, CA); Abhishek R. Appu (El Dorado Hills, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Joydeep Ray (Folsom, CA); Narayan Srinivasa (Portland, OR); Feng Chen (Shanghai, CN); Ben J. Ashbaugh (Folsom, CA); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Rancho Cordova, CA); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsom, CA); Altug Koker (El Dorado Hills, CA)
Assignee: INTEL CORPORATION
G06F12/0837G06N3/08G06N20/00G06T1/20G06F2212/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,261,903
App. No.
15/489,149
Granted
Apr 16, 2019
Kind
B2
Abstract

In an example, an apparatus comprises a plurality of processing unit cores, a plurality of cache memory modules associated with the plurality of processing unit cores, and a machine learning model communicatively coupled to the plurality of processing unit cores, wherein the plurality of cache memory modules share cache coherency data with the machine learning model. Other embodiments are also disclosed and claimed.

Claims (37)

1. An apparatus, comprising a processor to:

compute a first three-dimensional convolution across a set of three frames comprising a first frame, a second frame, and a third frame in an image sequence;

store a convolutional data set for the second frame and the third frame in a memory;

receive a fourth frame in the image sequence; and

use the convolutional data set for the second frame and the third frame to compute a second three-dimensional convolution across a set of three frames comprising the second frame, the third frame, and the fourth frame.

2. The apparatus of claim 1 , the processor comprising a plurality of execution units (EUs)

including a first type of execution unit and a second type of execution unit.

3. The apparatus of claim 2 , wherein:

the first type of execution unit and the second type of execution unit reside on a common die.

4. A computer-based method comprising:

computing, in a processor, a first three-dimensional convolution across a set of three frames comprising a first frame, a second frame, and a third frame in an image sequence;

storing a convolutional data set for the second frame and the third frame in memory communicatively coupled to the processor;

receiving, in the processor, a fourth frame in the image sequence; and

using, in the processor, the convolutional data set for the second frame and the third frame to compute a second three-dimensional convolution across a set of three frames comprising the second frame, the third frame, and the fourth frame.

5. The method of claim 4 , the processor comprising a plurality of execution units (EUs) including a first type of execution unit and a second type of execution unit.

6. The method of claim 5 , wherein the first type of execution unit and the second type of execution unit reside on a common die.

7. The method of claim 5 , further comprising deleting the convolutional data set from the memory in response to a determination that the convolutional data set has no dependencies in subsequent convolutional operations.

8. The method of claim 7 , wherein the memory comprises a cache memory communicatively coupled to the processor.

9. The method of claim 8 , further comprising issuing an instruction to mark a specific cache line in the cache memory for eviction in response to a determination that data in the cache line is no longer required for a subsequent convolutional operation.

10. An electronic device, comprising:

a processor to:

compute a first three-dimensional convolution across a set of three frames comprising a first frame, a second frame, and a third frame in an image sequence;

store a convolutional data set for the second frame and the third frame in memory;

receive a fourth frame in the image sequence; and

use the convolutional data set for the second frame and the third frame to compute a second three-dimensional convolution across a set of three frames comprising the second frame, the third frame, and the fourth frame; and

a computer readable memory communicatively coupled to the processor.

11. The electronic device of claim 10 the processor comprising a plurality of execution units (EUs) including a first type of execution unit and a second type of execution unit.

12. The electronic device of claim 11 , wherein:

the first type of execution unit and the second type of execution unit reside on a common die.

13. The electronic device of claim 11 , the processor to:

delete the convolutional data set from the memory in response to a determination that the convolutional data set has no dependencies in subsequent convolutional operations.

14. The electronic device of claim 13 , wherein the memory comprises a cache memory communicatively coupled to the processor.

15. The electronic device of claim 14 , the processor to issue an instruction to mark a specific cache line in the cache memory for eviction in response to a determination that data in the cache line is no longer required for a subsequent convolutional operation.

16. The apparatus of claim 2 , the processor to:

delete the convolutional data set from the memory in response to a determination that the convolutional data set has no dependencies in subsequent convolutional operations.

17. The apparatus of claim 16 , wherein the memory comprises a cache memory communicatively coupled to the processor.

18. The apparatus of claim 17 , the processor to issue an instruction to mark a specific cache line in the cache memory for eviction in response to a determination that data in the cache line is no longer required for a subsequent convolutional operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2017
From: SAKTHIVEL, CHANDRASEKARAN; SURTI, PRASOONKUMAR; WEAST, JOHN C.; BAGHSORKHI, SARA S.; GOTTSCHLICH, JUSTIN E.; APPU, ABHISHEK R.; GALOPPO VON BORRIES, NICOLAS C.; RAY, JOYDEEP; SRINIVASA, NARAYAN; CHEN, FENG; ASHBAUGH, BEN J.; BARIK, RAJKISHORE; LIN, TSUNG-HAN; SINHA, KAMAL; NURVITADHI, ERIKO; VEMBU, BALAJI; KOKER, ALTUG
To: INTEL CORPORATION
Reel/Frame 042702/0186 →
Continuity (1)
Related Publication 20180300246A1 · Oct 18, 2018
Cited By (3)
US 12,541,454 US 12,619,376 US 12,645,490