IP Library › Granted Patent US 11,481,219
Granted Patent B2
US 11,481,219 · App. 16/868,793 · Granted Oct 25, 2022

Store prefetches for dependent loads in a processor

Inventors: Mohit Karve (Austin, TX); Edmund Joseph Gieske (Cedar Park, TX); George W. Rohrbaugh, III (Charlotte, VT)
Assignee: International Business Machines Corporation
G06F9/3802G06F9/30043G06F12/0875G06F2212/452
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,219
App. No.
16/868,793
Granted
Oct 25, 2022
Kind
B2
Abstract

An information handling system, method, and processor that detects a store instruction for data in a processor where the store instruction is a reliable indicator of a future load for the data; in response to detecting the store instruction, sends a prefetch request to memory for an entire cache line containing the data referenced in the store instruction, and preferably only the single cache line containing the data; and receives, in response to the prefetch request, the entire cache line containing the data referenced in the store instruction.

Claims (61)

1. A method of processing data in an information handling system, comprising:

detecting a store instruction for data in a processor where the store instruction has been designated as an indicator of a future load for the data;

in response to detecting the store instruction, sending a prefetch request to memory for an entire cache line containing the data referenced in the store instruction; and

receiving, in response to the prefetch request, the entire cache line containing the data referenced in the store instruction.

2. The method according to claim 1 , wherein sending the prefetch request to memory is for only a single, entire cache line.

3. The method according to claim 1 , wherein the data referenced by the store instruction is not for an entire cache line.

4. The method according to claim 1 , further comprising:

performing a store read operation for the entire cache line containing the data referenced in the store instruction; and

using portions of the entire cache line retrieved from memory.

5. The method according to claim 4 , wherein using portions of the entire cache line retrieved from memory comprises:

receiving the entire cache line from memory;

overwriting a first portion of the entire cache line received from memory that contains the data referenced in the store instruction; and

keeping a second portion of the entire cache line containing the data not referenced in the store instruction.

6. The method according to claim 1 , further comprising:

processing a subsequent load instruction after the store instruction, where the load instruction is for the same data referenced in the store instruction; and

using the data referenced in the store instruction for the subsequent load instruction.

7. The method according to claim 1 , wherein the store instruction references a designated stack access register.

8. The method according to claim 1 , wherein the prefetch request is to transmit the entire cache line to cache.

9. The method according to claim 1 , further comprising:

setting a flag in response to detecting the store instruction; and

sending the prefetch request in response to reading the flag.

10. The method according to claim 4 , wherein the store read uses the cache line fetched to cache in response to the prefetch request.

11. An information handling system, comprising:

a memory subsystem;

a processor; and

one or more data caches having circuitry and logic to hold data for use by the processor,

the processor comprising:

an instruction fetch unit having circuitry and logic to fetch instructions for the processor, including store and load instructions;

a memory controller having circuitry and logic to manage the store and load instructions; and

a load store unit having circuitry and logic to execute store and load instructions, the load store unit having a prefetcher,

wherein the processor is configured to:

detect a store instruction for data where the store instruction has been preset as an indicator of a future load for the data;

in response to detecting the store instruction, send a prefetch request to the memory subsystem for an entire cache line containing the data referenced in the store instruction; and

receive in the one or more data caches, in response to the prefetch request, the entire cache line containing the data referenced in the store instruction.

12. The system according to claim 11 , wherein the prefetcher is configured to send the prefetch request to the memory subsystem for only a single, entire cache line.

13. The system according to claim 11 , wherein the data referenced by the store instruction is not for an entire cache line.

14. The system according to claim 1 , wherein the processor is further configured to:

perform a store read operation for the entire cache line containing the data referenced in the store instruction;

use only a portion of the entire cache line retrieved from memory.

15. The system according to claim 14 , wherein the processor is further configured to:

receive the entire cache line from the memory subsystem;

overwrite a first portion of the entire cache line received from the memory subsystem that contains the data referenced in the store instruction; and

keep a second portion of the entire cache line containing the data not referenced in the store instruction.

16. The system according to claim 11 , wherein the processor is further configured to:

process a subsequent load instruction after the store instruction, where the load instruction is for the same data referenced in the store instruction; and

use the data referenced in the store instruction for the subsequent load instruction.

17. The system according to claim 11 , wherein processor is configured to process the store instruction where the store instruction references a designated stack access register.

18. The system according to claim 11 , wherein the processor is further configured to process the prefetch request where the prefetch request contains instructions to transmit the entire cache line to cache.

19. The system according to claim 11 , further comprising:

a decode unit that is configured to detect the store instruction, and in response to detecting the store instruction setting a flag; and

the prefetcher is configured to send the prefetch request in response to reading the flag.

20. A processor comprising:

an instruction fetch unit having circuitry and logic to fetch instructions for the processor, including store and load instructions;

a load store unit having circuitry and logic to execute store and load instructions, the load store unit having a prefetcher;

a memory controller having circuitry and logic to manage the store and load instructions;

one or more data caches having circuitry and logic to hold data for use by the processor; and

a computer readable storage medium having program instructions, the program instructions executable by the processor,

wherein the program instructions when executed by the processor cause the processor to:

detect a store instruction for a stack access to a designated stack access register;

in response to detecting the store instruction for a stack access, send a prefetch request to a memory subsystem for an entire cache line containing the data referenced in the store instruction; and

receive in the one or more data caches, in response to the prefetch request, the entire cache line containing the data referenced in the store instruction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2020
From: KARVE, MOHIT; GIESKE, EDMUND JOSEPH; ROHRBAUGH, GEORGE W., III
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052600/0667 →
Continuity (1)
Related Publication 20210349722A1 · Nov 11, 2021