IP Library › Granted Patent US 12,625,814
Granted Patent B2
US 12,625,814 · App. 17/484,634 · Granted May 12, 2026

Graphics processor memory access architecture with address sorting

Inventors: Joydeep Ray (Folsom, CA); Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Aditya Navale (Folsom, CA); Varghese George (Folsom, CA); Vasanth Ranganathan (El Dorado Hills, CA); Fangwen Fu (Folsom, CA); Ben J. Ashbaugh (Folsom, CA); Vidhya Krishnan (Folsom, CA); Sabareesh Ganapathy (Bangalore, IN); Prathamesh Raghunath Shinde (Folsom, CA)
Assignee: Intel Corporation
G06F12/0842G06F1/14G06F7/08G06F2212/1016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,814
App. No.
17/484,634
Granted
May 12, 2026
Kind
B2
Abstract

One embodiment provides a graphics processor including a processing resource including a register file, memory, a cache, and load/store/cache circuitry to process load, store, and prefetch messages from the processing resource. The circuitry will sort received memory access messages into address sorted lists of reads and writes. The circuitry schedules a first set of address sorted requests from a first request buffer for a first period of time, then schedules a second set of address sorted requests from a second request buffer for a second period of time.

Claims (62)

1 . A graphics processor comprising:

a processing resource including a register file;

a memory device;

a cache coupled with the processing resource and the memory device; and

circuitry to process memory access messages received from the processing resource, wherein to process the memory access messages, the circuitry is configured to:

schedule a batch of multiple address-sorted read requests from a read request buffer during a first period of clock cycles without regard to receipt order of the memory access messages associated with the multiple address-sorted read requests; and

schedule a batch of multiple address-sorted write requests from a write request buffer during a second period of clock cycles without regard to receipt order of the memory access messages associated with the multiple address-sorted write requests, the read request buffer distinct from the write request buffer, the circuitry to alternately schedule batches of multiple requests from the read request buffer and the write request buffer to reduce a memory switching penalty occurrence.

2 . The graphics processor as in claim 1 , wherein the circuitry is configured to:

store received memory access messages to a memory access request buffer; and

sort the memory access messages in the memory access request buffer to generate an address-sorted list of memory access messages.

3 . The graphics processor as in claim 2 , wherein the memory access messages indicate to transfer data between the register file and the memory device or between the memory device and the cache.

4 . The graphics processor as in claim 3 , wherein memory access messages include a load message to transfer data from the memory device and a store message to transfer data from the register file.

5 . The graphics processor as in claim 4 , wherein the circuitry is configured to:

generate a set of address-sorted read requests in response to the load message;

generate a set of address-sorted write requests in response to the store message;

write the set of address-sorted read requests to the read request buffer; and

write the set of address-sorted write requests to the write request buffer.

6 . The graphics processor as in claim 5 , wherein the load message includes a response length to indicate an amount of data to write to the register file in response to the load message.

7 . The graphics processor as in claim 6 , wherein the circuitry is configured to:

determine that the load message has a response length of zero; and

transfer data between the memory device and the cache in response to the load message without transferring data to the register file.

8 . The graphics processor as in claim 7 , wherein the circuitry is configured to:

determine whether the load message and the store message each include a reference to a same memory address; and

schedule the load message and the store message in order of receipt in response to the determination.

9 . The graphics processor as in claim 8 , wherein to schedule the load message and the store message in order of receipt, the circuitry is to:

determine that the load message has an earlier time of receipt relative to the store message;

write the set of address-sorted read requests to the read request buffer; and

write the set of address-sorted write requests to the write request buffer after the set of address-sorted read requests are scheduled from the read request buffer.

10 . The graphics processor as in claim 1 , wherein the circuitry is configured to submit one or more access requests to the cache for each of the multiple address-sorted read requests and the multiple address-sorted write requests.

11 . The graphics processor as in claim 10 , further comprising cache control circuitry associated with the cache, wherein the cache control circuitry is to:

read a tag and a cache control setting associated with an access request submitted to the cache; and

service the access request from the cache or the memory device based on the tag and the cache control setting.

12 . The graphics processor as in claim 1 , further comprising a surface state cache to store surface state parameters for a memory surface stored on the memory device.

13 . The graphics processor as in claim 12 , wherein the circuitry is configured to read the surface state parameters from the surface state cache and submit one or more access requests to the cache for each of the multiple address-sorted read requests and the multiple address-sorted write requests that are to a location within bounds of the memory surface, the circuitry to determine the bounds of the memory surface based on the surface state parameters.

14 . The graphics processor as in claim 13 , wherein surface is a two-dimensional surface and the circuitry is to submit multiple access requests to the cache in response to a single memory access message to access a two-dimensional block of data within the two-dimensional surface.

15 . The graphics processor as in claim 14 , wherein the surface state parameters include a tiling format for the surface and the circuitry is to determine the multiple access requests to submit to the cache based on the tiling format for the surface.

16 . The graphics processor as in claim 15 , wherein the circuitry is configured to:

determine the tiling format for the surface via the surface state cache;

determine, based on the tiling format, a mapping between a cacheline of the cache and a row of the two-dimensional block of data; and

submit one or more access requests to the cache for each cacheline associated with the row of the two-dimensional block of data, the one or more access requests determined based on the mapping between the cacheline of the cache and the row of the two-dimensional block of data.

17 . A method comprising:

receiving a memory access request at circuitry of a graphics processor, wherein the circuitry is configured to perform memory load and store operations and the memory access request is received from a processing resource of the graphics processor;

storing the memory access request to a memory access request buffer within the circuitry;

retrieving a set of multiple memory access requests from the memory access request buffer;

sorting the set of multiple memory access requests into an address-sorted list of memory access requests without regard to receipt order of the set of multiple memory access requests;

storing read requests in the address-sorted list of memory access requests to a read request buffer within the circuitry;

storing write requests in the address-sorted list of memory access requests to a write request buffer within the circuitry, the write request buffer distinct from the read request buffer; and

alternately scheduling batches of multiple requests from the read request buffer and the write request buffer to reduce a memory switching penalty occurrence.

18 . The method as in claim 17 , further comprising:

determining whether a read request and a write request reference a same memory address; and

scheduling the read request and the write request in order of receipt in response to the determination, wherein the read request is a request to prefetch data to a cache according to a cache control setting associated with the read request.

19 . A data processing system comprising:

a memory device; and

one or more processors coupled with the memory device, the one or more processors to execute instructions stored on the memory device, the instructions to cause the one or more processors to:

store received memory access messages to a memory access request buffer;

sort memory access messages in the memory access request buffer to generate an address-sorted list of memory access messages;

write one or more memory read requests to a read request buffer in response to a first memory access message in the address-sorted list of memory access messages;

write one or more memory write requests to a write request buffer in response to a second memory access message in the address-sorted list of memory access messages, the write request buffer distinct from the read request buffer;

via circuitry configured to process memory access messages triggered in response to instructions executed by the one or more processors:

schedule multiple address-sorted read requests from a read request buffer during a first period of clock cycles without regard to receipt order of the memory access messages associated with the multiple address-sorted read requests; and

schedule multiple address-sorted write requests from the write request buffer during a second period of clock cycles without regard to receipt order of the memory access messages associated with the multiple address-sorted write requests, the circuitry to alternately schedule batches of multiple requests from the read request buffer and the write request buffer to reduce a memory switching penalty occurrence.

20 . The data processing system as in claim 19 , wherein the memory access messages include a load message to transfer data from the memory device to a cache or a register file and a store message to transfer data from a register file to the memory device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2022
From: RAY, JOYDEEP; APPU, ABHISHEK R.; KOKER, ALTUG; NAVALE, ADITYA; GEORGE, VARGHESE; RANGANATHAN, VASANTH; FU, FANGWEN; ASHBAUGH, BEN J.; KRISHNAN, VIDHYA; GANAPATHY, SABAREESH; SHINDE, PRATHAMESH RAGHUNATH
To: INTEL CORPORATION
Reel/Frame 059949/0668 →
Continuity (1)
Related Publication 20230104845A1 · Apr 6, 2023
References Cited (14)
US 6999091B2 · Saxena et al. · 2006 [cited by applicant]
US 8578097B2 · Kim et al. · 2013 [cited by applicant]
US 10733688B2 · Ray et al. · 2020 [cited by applicant]
US 20150046662A1 · Heinrich · 2015 [cited by examiner]
US 20150277802A1 · Oikarinen · 2015 [cited by examiner]
US 20190065209A1 · Mishra · 2019 [cited by examiner]
US 20190205284A1 · Fleming · 2019 [cited by examiner]
US 20200065028A1 · Keil · 2020 [cited by examiner]
US 20210097641A1 · Iyer · 2021 [cited by examiner]
JP 2002208026A · 2002 [cited by examiner]
Notification of Publication for CN Application No. 202211019025.7, 4 pages, Apr. 13, 2023. [cited by applicant]
Lueh et al., “C-for-Metal: High Performance SIMD Programming on Intel GPUs”, Jan. 26, 2021, 13 pages. [cited by applicant]
Blackmore, M., “A Quantitative Analysis of Memory Controller Page Policies”, Thesis from Portland State University, 2013, 58 pages. [cited by applicant]
Chang et al., “Improving DRAM Performance by Parallelizing Refreshes with Accesses”, Carnegie Mellon University, Intel Labs, Feb. 15, 2014, 12 pages. [cited by applicant]