IP Library › Granted Patent US 12,450,680
Granted Patent B2
US 12,450,680 · App. 17/484,619 · Granted Oct 21, 2025

Out-of-order execution of graphics processing unit texture sampler operations

Inventors: Carlos Nava Rodriguez (Folsom, CA); Benjamin Pletcher (Mather, CA); Yoav Harel (Carmichael, CA); Bret Martin (Folsom, CA); Sudarshanram Shetty (Portland, OR)
Assignee: Intel Corporation
G06T1/20G06T1/60G06T15/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,680
App. No.
17/484,619
Granted
Oct 21, 2025
Kind
B2
Abstract

Embodiments described herein are generally directed facilitating out-of-order execution of GPU texture sampler operations. An embodiment of a method includes a texture sampler of a GPU maintaining (i) a latency queue operable to store information regarding a set of transactions associated with each of multiple texture sampler operations and (ii) multiple virtual channel (VC) queues each operable to store information regarding transactions for a respective single texture sampler operation at a time. Out-of-order processing of the texture sampler operations is facilitated by making use of the latency queue and the VC queues. For example, during a transaction processing interval, the availability of data in a cache for the transactions associated with each of the VC queues may be determined. A VC queue may be selected based on the determined availability of data. A transaction associated with a head of the selected VC queue may then be processed.

Claims (42)

1. A graphics processing unit (GPU) comprising:

a level 1 (L1) cache; and

a texture sampler, coupled to the L1 cache, including (i) a latency queue operable to store information regarding a set of transactions associated with each of a plurality of texture sampler operations and (ii) a plurality of virtual channel (VC) queues each operable to store information regarding transactions for a respective single texture sampler operation at a time, wherein the texture sampler is operable to, during a transaction processing interval, facilitate out-of-order processing of the plurality of texture sampler operations by:

determining availability of data in the L1 cache associated with the texture sampler for the transactions associated with each of the plurality of VC queues;

selecting a VC queue of the plurality of VC queues based on the determined availability of data including prioritizing a particular VC queue of the plurality of VC queues containing information regarding transactions representing all transactions for the respective single texture sampler operation over another VC queue of the plurality of VC queues containing information regarding transactions representing less than all transactions for the respective single texture sampler operation; and

processing a transaction associated with a head of the selected VC queue.

2. The GPU of claim 1 , wherein the texture sampler is further operable to:

dequeue information regarding a transaction associated with a head of the latency queue that is part of the set of transactions associated with a particular texture sampler operation of the plurality of texture sampler operations; and

enqueue the information regarding the transaction at a tail of a VC queue of the plurality of VC queues currently storing information regarding transactions associated with the particular texture sampler operation.

3. The GPU of claim 1 , wherein the texture sampler is further operable to continue to process subsequent transactions associated with the head of the selected VC queue until all of the transactions for the respective single texture sampler operation have been completed.

4. The GPU of claim 1 , wherein said selecting a VC queue of the plurality of VC queues based on the determined availability of data includes determining data is available in the L1 cache for at least a threshold number of the transactions associated with the selected VC queue.

5. The GPU of claim 4 , wherein the threshold number is 8.

6. The GPU of claim 1 , wherein the plurality of VC queues comprises 8 VC queues per 32 threads.

7. The GPU of claim 1 , wherein each of the plurality of VC queues includes 16 entries.

8. A method comprising:

maintaining within a texture sampler of a graphics processing unit (i) a latency queue operable to store information regarding a set of transactions associated with each of a plurality of texture sampler operations and (ii) a plurality of virtual channel (VC) queues each operable to store information regarding transactions for a respective single texture sampler operation at a time;

during a transaction processing interval, facilitating out-of-order processing of the plurality of texture sampler operations by:

determining availability of data in a cache associated with the texture sampler for the transactions associated with each of the plurality of VC queues;

selecting a VC queue of the plurality of VC queues based on the determined availability of data including prioritizing a particular VC queue of the plurality of VC queues containing information regarding transactions representing all transactions for the respective single texture sampler operation over another VC queue of the plurality of VC queues containing information regarding transactions representing less than all transactions for the respective single texture sampler operation; and

processing a transaction associated with a head of the selected VC queue.

9. The method of claim 8 , further comprising during the transaction processing interval:

dequeuing information regarding a transaction associated with a head of the latency queue that is part of the set of transactions associated with a particular texture sampler operation of the plurality of texture sampler operations; and

enqueuing the information regarding the transaction at a tail of a VC queue of the plurality of VC queues currently storing information regarding transactions associated with the particular texture sampler operation.

10. The method of claim 8 , further comprising continuing to process subsequent transactions associated with the head of the selected VC queue until all of the transactions for the respective single texture sampler operation have been completed.

11. The method of claim 8 , wherein said selecting a VC queue of the plurality of VC queues based on the determined availability of data includes determining data is available in the cache for at least a threshold number of the transactions associated with the selected VC queue.

12. The method of claim 11 , wherein the threshold number is 8.

13. The method of claim 8 , wherein the plurality of VC queues comprises 8 VC queues per 32 threads.

14. The method of claim 8 , wherein each of the plurality of VC queues includes 16 entries.

15. A texture sampler of a graphics processing unit, the texture sampler comprising:

a latency queue operable to store information regarding a set of transactions associated with each of a plurality of texture sampler operations; and

a plurality of virtual channel (VC) queues each operable to store information regarding transactions for a respective single texture sampler operation at a time, wherein the texture sampler is operable to, during a transaction processing interval, facilitate out-of-order processing of the plurality of texture sampler operations by:

determining availability of data in a level 1 (L1) cache associated with the texture sampler for the transactions associated with each of the plurality of VC queues;

selecting a VC queue of the plurality of VC queues based on the determined availability of data including prioritizing a particular VC queue of the plurality of VC queues containing information regarding transactions representing all transactions for the respective single texture sampler operation over another VC queue of the plurality of VC queues containing information regarding transactions representing less than all transactions for the respective single texture sampler operation; and

processing a transaction associated with a head of the selected VC queue.

16. The texture sampler of claim 15 , wherein the texture sampler is further operable to:

dequeue information regarding a transaction associated with a head of the latency queue that is part of the set of transactions associated with a particular texture sampler operation of the plurality of texture sampler operations; and

enqueue the information regarding the transaction at a tail of a VC queue of the plurality of VC queues currently storing information regarding transactions associated with the particular texture sampler operation.

17. The texture sampler of claim 15 , wherein the texture sampler is further operable to continue to process subsequent transactions associated with the head of the selected VC queue until all of the transactions for the respective single texture sampler operation have been completed.

18. The texture sampler of claim 15 , wherein said selecting a VC queue of the plurality of VC queues based on the determined availability of data includes determining data is available in the L1 cache for at least a threshold number of the transactions associated with the selected VC queue.

19. The texture sampler of claim 18 , wherein the threshold number is 8.

20. The texture sampler of claim 15 , wherein the plurality of VC queues comprises 8 VC queues per 32 threads.

21. The texture sampler of claim 15 , wherein each of the plurality of VC queues includes 16 entries.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: NAVA RODRIGUEZ, CARLOS; PLETCHER, BENJAMIN; HAREL, YOAV; MARTIN, BRET; SHETTY, SUDARSHANRAM
To: INTEL CORPORATION
Reel/Frame 058151/0846 →
Continuity (1)
Related Publication 20230111571A1 · Apr 13, 2023
References Cited (9)
US 11443479B1 · Yeung · 2022 [cited by examiner]
US 20010049771A1 · Tischler · 2001 [cited by examiner]
US 20040036692A1 · Alcorn · 2004 [cited by examiner]
US 20150379670A1 · Koker · 2015 [cited by examiner]
US 20160055833A1 · Havlir · 2016 [cited by examiner]
US 20180165790A1 · Schneider · 2018 [cited by examiner]
US 20200294180A1 · Koker · 2020 [cited by examiner]
US 20210097750A1 · Woop · 2021 [cited by examiner]
Notification of Publication for CN 202211022141.4, Mar. 28, 2023, 74 pages. [cited by applicant]