IP Library Granted Patent US 12,360,771
Granted Patent B2
US 12,360,771 · App. 17/241,726 · Granted Jul 15, 2025

Rescheduling a load instruction based on past replays

Inventor: Jonathan C. Masters (Boston, MA)
Assignee: Red Hat, Inc.
G06F9/3836G06F9/30043G06F9/34G06F9/3802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,360,771
App. No.
17/241,726
Granted
Jul 15, 2025
Kind
B2
Abstract

Rescheduling a load instruction based on past replays is disclosed. A load replay predictor of a processor device determines, at a first time, that a load instruction is scheduled to be executed by a load store unit to load data from a memory location. The load replay predictor accesses load replay data associated with a previous replay of the load instruction and, based on the load replay data, causes the load instruction to be rescheduled.

Claims (47)

1. A method comprising:

receiving, by a load store unit of a processor device, a load instruction to load data from a memory location, the load instruction identified in a cache line of a cache memory, the cache line comprising one or more processor instructions and cache line metadata associated with the one or more processor instructions, the one or more processor instructions comprising the load instruction;

determining at a first time, by a load replay predictor of the processor device, that the load instruction is scheduled to be executed by the load store unit;

accessing, by the load replay predictor from the cache line metadata, load replay data associated with a previous replay of the load instruction;

based on the load replay data, causing, by the load replay predictor, the load instruction to be rescheduled for subsequent execution;

sending, by the load replay predictor to the load store unit, a first message indicating that the load instruction is not to be executed;

in response to causing the load instruction to be rescheduled for subsequent execution, causing, by the processor device at a second time subsequent to the first time, execution of the load instruction;

receiving, by the load replay predictor from a component of the processor device, a second message indicating that the load instruction was executed, wherein the second message comprises a memory address from which the load instruction is to load the data; and

based on the second message, updating, by the load replay predictor, the load replay data to indicate that the load instruction was executed.

2. The method of claim 1 wherein the load instruction is a programmatic instruction from an executable program.

3. The method of claim 1 wherein the load instruction is an internal load instruction generated by the processor device.

4. The method of claim 1 further comprising:

determining, at a time prior to the first time, that the load instruction is scheduled to be executed to load the data from the memory location;

executing the load instruction to load the data from the memory location;

subsequently determining that the load instruction must be replayed; and

generating the load replay data based on determining that the load instruction must be replayed.

5. The method of claim 4 further comprising:

storing, by the load replay predictor, the load replay data in the cache line metadata that corresponds to the load instruction.

6. A processor device comprising:

a load store unit to:

receive a load instruction to load data from a memory location, the load instruction identified in a cache line of a cache memory, the cache line comprising one or more processor instructions and cache line metadata associated with the one or more processor instructions, the one or more processor instructions comprising the load instruction; and

a load replay predictor to:

determine, at a first time, that the load instruction is scheduled to be executed by the load store unit;

access, from the cache line metadata, load replay data associated with a previous replay of the load instruction;

based on the load replay data, cause the load instruction to be rescheduled for subsequent execution;

send, to the load store unit, a first message indicating that the load instruction is not to be executed;

in response to causing the load instruction to be rescheduled for subsequent execution, cause, at a second time subsequent to the first time, execution of the load instruction;

receive, from a component of the processor device, a second message indicating that the load instruction was executed, wherein the second message comprises a memory address from which the load instruction is to load the data; and

based on the second message, update the load replay data to indicate that the load instruction was executed.

7. The processor device of claim 6 wherein the load instruction is a programmatic instruction from an executable program.

8. The processor device of claim 6 wherein the load replay predictor is further to:

determine, at a time prior to the first time, that the load instruction is scheduled to be executed to load the data from the memory location;

execute the load instruction to load the data from the memory location;

subsequently determine that the load instruction must be replayed; and

generate the load replay data based on determining that the load instruction must be replayed.

9. The processor device of claim 8 wherein the load replay predictor is further to:

store the load replay data in the cache line metadata that corresponds to the load instruction.

10. A non-transitory computer-readable storage medium including executable instructions to cause a processor device to:

receive, by a load store unit of the processor device, a load instruction to load data from a memory location, the load instruction identified in a cache line of a cache memory, the cache line comprising one or more processor instructions and cache line metadata associated with the one or more processor instructions, the one or more processor instructions comprising the load instruction;

determining at a first time, by a load replay predictor of the processor device, that the load instruction is scheduled to be executed by the load store unit;

accessing, by the load replay predictor from the cache line metadata, load replay data associated with a previous replay of the load instruction;

based on the load replay data, causing, by the load replay predictor, the load instruction to be rescheduled for subsequent execution;

sending, by the load replay predictor to the load store unit, a first message indicating that the load instruction is not to be executed;

in response to causing the load instruction to be rescheduled for subsequent execution, causing, by the processor device at a second time subsequent to the first time, execution of the load instruction;

receiving, by the load replay predictor from a component of the processor device, a second message indicating that the load instruction was executed, wherein the second message comprises a memory address from which the load instruction is to load the data; and

based on the second message, updating, by the load replay predictor, the load replay data to indicate that the load instruction was executed.

11. The non-transitory computer-readable storage medium of claim 10 , wherein the load instruction is a programmatic instruction from an executable program.

Assignments (2)
CHANGE OF NAME Recorded Mar 3, 2026
From: RED HAT, INC.
To: RED HAT, LLC
Reel/Frame 074913/0759 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2021
From: MASTERS, JONATHAN C.
To: RED HAT, INC.
Reel/Frame 056056/0443 →
Continuity (1)
Related Publication 20220342672A1 · Oct 27, 2022
References Cited (22)
US 5136697A · Johnson · 1992 [cited by examiner]
US 6438673B1 · Jourdan · 2002 [cited by examiner]
US 7707391B2 · Musoll · 2010 [cited by examiner]
US 7861066B2 · Dhodapkar · 2010 [cited by examiner]
US 7930485B2 · Fertig · 2011 [cited by examiner]
US 8966232B2 · Tran · 2015 [cited by applicant]
US 9367455B2 · Eckert et al. · 2016 [cited by applicant]
US 10127046B2 · Col et al. · 2018 [cited by applicant]
US 10936319B2 · Srinivasan · 2021 [cited by examiner]
US 20020152368A1 · Nakamura · 2002 [cited by examiner]
US 20030126406A1 · Hammarlund · 2003 [cited by examiner]
US 20030177338A1 · Luick · 2003 [cited by examiner]
US 20030208665A1 · Peir · 2003 [cited by examiner]
US 20050268046A1 · Heil · 2005 [cited by examiner]
US 20120117362A1 · Bhargava · 2012 [cited by examiner]
US 20190266096A1 · Lee et al. · 2019 [cited by applicant]
Memik, Gokhan & Reinman, Glenn & Mangione-Smith, William. Precise instruction scheduling. Journal of Instruction—Level Parallelism. 7. 1-29. (Year: 2005). [cited by examiner]
“Cache Architecture and Design”<https://www.cs.swarthmore.edu/˜kwebb/cs31/f18/memhierarchy/caching.html> (Year: 2021). [cited by examiner]
R. E. Kessler, E. J. McLellan and D. A. Webb, “The Alpha 21264 microprocessor architecture,” Proceedings International Conference on Computer Design. VLSI in Computers and Processors, pp. 90-95 (Year: 1998). [cited by examiner]
Alves, Ricardo, et al., “Minimizing Replay under Way-Prediction,” http://uu.diva-portal.org/smash/get/diva2:1316465/FULLTEXT01.pdf, 2019, 10 pages. [cited by applicant]
Kim, Ilhyun, et al., “Understanding Scheduling Replay Schemes,” 10th International Symposium on High Performance Computer Architecture (HPCA'04), 2004, 12 pages. [cited by applicant]
Yoaz, Adi, et al., “Speculation Techniques for Improving Load Related Instruction Scheduling,” ACM Computer Architecture News. 27. 42-53. 10.1109/ISCA.1999.765938, Feb. 1999, 13 pages. [cited by applicant]