IP Library › Granted Patent US 11,036,545
Granted Patent B2
US 11,036,545 · App. 16/355,565 · Granted Jun 15, 2021

Graphics systems and methods for accelerating synchronization using fine grain dependency check and scheduling optimizations based on available shared memory space

Inventors: Subramaniam Maiyuran (Gold River, CA); Varghese George (Folsom, CA); Altug Koker (El Dorado Hills, CA); Aravindh Anantaraman (Folsom, CA); SungYe Kim (Folsom, CA); Valentin Andrei (San Jose, CA); Joydeep Ray (Folsom, CA)
Assignee: Intel Corporation
G06F9/4881G06F9/30109G06F9/3838G06F9/3877G06F9/5016G06F12/0837
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,036,545
App. No.
16/355,565
Granted
Jun 15, 2021
Kind
B2
Abstract

Accelerated synchronization operations using fine grain dependency check are disclosed. A graphics multiprocessor includes a plurality of execution units and synchronization circuitry that is configured to determine availability of at least one execution unit. The synchronization circuitry to perform a fine grain dependency check of availability of dependent data or operands in shared local memory or cache when at least one execution unit is available.

Claims (15)

1. A graphics multiprocessor, comprising:

a plurality of processing resources; and

synchronization circuitry that is configured to determine availability of at least one processing resource and when at least one processing resource is available, to perform a fine grain dependency check of availability of a minimum data size of dependent data or operands in shared local memory or cache to start a computation before a complete data set for the computation is available.

2. The graphics multiprocessor of claim 1 , further comprising:

a plurality of registers, wherein the synchronization circuitry is further configured to cause the loading of the dependent data or operands into registers if the dependent data or operands are available in shared local memory or cache.

3. The graphics multiprocessor of claim 2 , wherein the synchronization circuitry is further configured to cause the loading of the dependent data or operands from the registers into an available processing resource to start and process the computation.

4. The graphics multiprocessor of claim 3 , wherein the dependent data or operands are loaded into tile registers.

5. A method for accelerated synchronization operations using fine grain dependency check of a graphics processing unit (GPU), comprising:

determining, with a synchronization circuitry of the GPU, availability of at least one processing resource; and

given availability of at least one processing resource, performing with the synchronization circuitry a fine grain dependency check of availability of a minimum data size of dependent data or operands in shared local memory or cache to start a computation before a complete data set for the computation is available.

6. The method of claim 5 , further comprising:

if the dependent data or operands are available in shared local memory or cache, loading with the synchronization circuitry, the dependent data or operands into a register file.

7. The method of claim 6 , further comprising:

loading the dependent data or operands from the register file into an available processing resource or processing engine to start and process the computation.

8. The method of claim 7 , wherein the dependent data or operands are loaded into a tile register.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2019
From: MAIYURAN, SUBRAMANIAM; GEORGE, VARGHESE; KOKER, ALTUG; ANANTARAMAN, ARAVINDH; KIM, SUNGYE; ANDREI, VALENTIN; RAY, JOYDEEP
To: INTEL CORPORATION
Reel/Frame 049579/0216 →
Continuity (1)
Related Publication 20200293369A1 · Sep 17, 2020