IP Library › Granted Patent US 12,743,741
Granted Patent B2
US 12,743,741 · App. 18/519,844 · Granted Sep 22, 2026

Systems and methods for exploiting queues and transitional storage for improved low-latency high-bandwidth on-die data retrieval

Inventors: Aravindh Anantaraman (Folsom, CA); Altug Koker (El Dorado Hills, CA); Varghese George (Folsom, CA); Subramaniam Maiyuran (Gold River, CA); SungYe Kim (Folsom, CA); Valentin Andrei (San Jose, CA)
Assignee: INTEL CORPORATION
G06T1/20G06F9/3802G06F9/3877G06T1/60G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,741
App. No.
18/519,844
Granted
Sep 22, 2026
Kind
B2
Abstract

Apparatuses including general-purpose graphics processing units and graphics multiprocessors that exploit queues or transitional buffers for improved low-latency high-bandwidth on-die data retrieval are disclosed. In one embodiment, a graphics multiprocessor includes at least one compute engine to provide a request, a queue or transitional buffer, and logic coupled to the queue or transitional buffer. The logic is configured to cause a request to be transferred to a queue or transitional buffer for temporary storage without processing the request and to determine whether the queue or transitional buffer has a predetermined amount of storage capacity.

Claims (15)

1 . A graphics processor comprising:

a processing resource to provide a request;

a queue or transitional buffer to function as cache memory and to provide on-die or on-chip data retrieval for the graphics processor; and

logic on-die or on chip with the queue or transitional buffer, the logic is configured to cause the request of the processing resource to be transferred to the queue or transitional buffer for temporary storage instead of transferring the request to off-chip memory when the queue or transitional buffer has a predetermined amount of storage capacity, wherein the request is transferred to the queue or transitional buffer without being processed instead of being transferred to the off- chip memory when the queue or the transitional buffer has the predetermined amount of storage capacity that is greater than a threshold level, wherein the logic is further configured to transfer the request to one or more of a different queue, a different transitional buffer, or a next level of memory when the queue or transitional buffer lacks the predetermined amount of storage capacity, wherein the logic is further configured to determine whether the queue or transitional buffer has the predetermined amount of storage capacity including greater than a threshold level of storage capacity.

2 . The graphics processor of claim 1 , wherein the logic comprises a cache controller or memory interface controller, wherein the queue or transitional buffer is located in any location on-die or on chip with the logic and the processing resource, wherein the queue or transitional buffer provides functionality of cache memory.

3 . A method comprising:

causing, by logic associated with processing circuitry of a computing device, a request associated with a processing resource to be transferred to a queue or transitional buffer for temporary storage instead of transferring the request to off-chip memory when the queue or transitional buffer has a predetermined amount of storage capacity, wherein the request is transferred to the queue or transitional buffer without being processed instead of being transferred to the off- chip memory when the queue or the transitional buffer has the predetermined amount of storage capacity that is greater than a threshold level;

transferring the request to one or more of a different queue, a different transitional buffer, or a next level of memory when the queue or transitional buffer lacks the predetermined amount of storage capacity; and

determining whether the queue or transitional buffer has the predetermined amount of storage capacity including greater than a threshold level of storage capacity.

4 . The method of claim 3 , wherein the logic comprises a cache controller or memory interface controller, wherein the queue or transitional buffer is located in any location on-die or on chip with the logic and the processing resource, wherein the queue or transitional buffer provides functionality of cache memory, wherein the processing circuitry comprises graphics processing circuitry.

5 . At least one non-transitory computer-readable medium having stored thereon instructions which, when executed, cause a computing device to perform operations comprising:

causing, by logic associated with processing circuitry of the computing device, a request associated with a processing resource to be transferred to a queue or transitional buffer for temporary storage instead of transferring the request to off-chip memory when the queue or transitional buffer has a predetermined amount of storage capacity, wherein the request is transferred to the queue or transitional buffer without being processed instead of being transferred to the off- chip memory when the queue or the transitional buffer has the predetermined amount of storage capacity that is greater than a threshold level;

transferring the request to one or more of a different queue, a different transitional buffer, or a next level of memory when the queue or transitional buffer lacks the predetermined amount of storage capacity; and

determining whether the queue or transitional buffer has the predetermined amount of storage capacity including greater than a threshold level of storage capacity.

6 . The non-transitory computer-readable medium of claim 5 , wherein the logic comprises a cache controller or memory interface controller, wherein the queue or transitional buffer is located in any location on-die or on chip with the logic and the processing resource, wherein the queue or transitional buffer provides functionality of cache memory, wherein the processing circuitry comprises graphics processing circuitry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2026
From: INTEL CORPORATION
To: INTEL PRODUCTS IP LLC
Reel/Frame 075991/0754 →
Continuity (3)
Continuation 17544826 · Dec 7, 2021
Continuation 16355250 · Mar 15, 2019
Related Publication 20240169466A1 · May 23, 2024
References Cited (29)
US 5353410A · Macon et al. · 1994 [cited by applicant]
US 6900812B1 · Morein · 2005 [cited by applicant]
US 7185033B2 · Jain · 2007 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10482043B2 · Nygren · 2019 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20140189256A1 · Kranich et al. · 2014 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180285116A1 · Vembu et al. · 2018 [cited by applicant]
US 20200294178A1 · Anantaraman et al. · 2020 [cited by applicant]
CN 113424156A · 2021 [cited by applicant]
EP 3382539A1 · 2018 [cited by applicant]
EP 3396529A1 · 2018 [cited by applicant]
WO 2016167915A1 · 2016 [cited by applicant]
WO 2020190455 · 2020 [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US/2020/019529, 11 pages, May 13, 2020. [cited by applicant]
International Preliminary Report on Patentability for PCT Application No. PCT/US/2020/019529, 9 pages, Sep. 30, 2021. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/544,826, mailed Sep. 7, 2023, 7 pages. [cited by applicant]
Anonymous: “Method for a lazy write-update policy for multiprocessor systems” (Oct. 23, 2002), XP013004939, 9 pages. [cited by applicant]
Notification of Publication for EP 20714753.9, Dec. 22, 2021, 1 page. [cited by applicant]
Communication pursuant to Article 94(3) EPC for EP20714753.9 mailed Mar. 13, 2024, 11 pages. [cited by applicant]