IP Library › Granted Patent US 12,547,407
Granted Patent B2
US 12,547,407 · App. 18/399,959 · Granted Feb 10, 2026

Return address stack with branch mispredict recovery

Inventors: James Youngsae Cho (Los Gatos, CA); Rabin Sugumar (Sunnyvale, CA)
Assignee: Akeana, Inc.
G06F9/3806G06F9/3861G06F9/3844
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,407
App. No.
18/399,959
Filed
Dec 29, 2023
Granted
Feb 10, 2026
Kind
B2
Art Unit
2183
USPC
712/239
Abstract

Techniques for providing a return address stack with branch mispredict recovery are disclosed. A processor core is accessed. The processor core includes a return address stack (RAS), a local cache hierarchy, and branch prediction logic. RAS state information, including a write pointer, a read pointer, and a RAS count, is sent to a branch execution unit. One or more call instructions are detected in an instruction stream. The detecting generates a predicted return address for each of the one or more call instructions which are pushed on the RAS. The pushing is directed by the write pointer. One or more return instructions are recognized in the instruction stream. The write pointer and the read pointer for the RAS are updated, based on information from the branch execution unit. The predicted return address for each of the one or more return instructions is popped from the RAS.

Claims (49)

1 . A processor-implemented method for predicting addresses comprising:

accessing a processor core, wherein the processor core includes a return address stack (RAS), a local cache hierarchy, and branch prediction logic, wherein the processor core is coupled to a memory system;

sending, to a branch execution unit, a write pointer, a read pointer, and a RAS count;

detecting one or more call instructions in an instruction stream, wherein the detecting generates a predicted return address for each of the one or more call instructions;

pushing, on the RAS, the predicted return address for each of the one or more call instructions, wherein the pushing is directed by the write pointer;

recognizing one or more return instructions in the instruction stream;

updating the write pointer and the read pointer for the RAS, wherein the updating is based on a misprediction signal from the branch execution unit and wherein the updating includes manipulating the RAS count;

popping, from the RAS, the predicted return address for each of the one or more return instructions wherein the popping is directed by the read pointer; and

rolling back, by the branch execution unit, the write pointer, the read pointer, and the RAS count to a previous value.

2 . The method of claim 1 wherein the sending is accomplished every cycle.

3 . The method of claim 2 wherein a pipeline flush is executed or a branch instruction is mispredicted, and wherein the branch instruction is not a call or a return instruction.

4 . The method of claim 3 further comprising returning, to the RAS, the write pointer, the read pointer, and the RAS count which were backed up.

5 . The method of claim 2 further comprising adjusting, by the branch execution unit, the write pointer, the read pointer, and the RAS count, wherein a call instruction from the one or more call instructions is mispredicted.

6 . The method of claim 5 further comprising returning, to the RAS, the write pointer, the read pointer, and the RAS count which were adjusted.

7 . The method of claim 2 further comprising adjusting the write pointer, the read pointer, and the RAS count, wherein a return instruction from the one or more return instructions is mispredicted.

8 . The method of claim 7 further comprising returning, to the RAS, the write pointer, the read pointer, and the RAS count which were adjusted.

9 . The method of claim 1 wherein the processor core executes one or more instructions out of order.

10 . The method of claim 9 wherein the pushing further comprises incrementing the RAS count as part of the manipulating.

11 . The method of claim 10 wherein the popping further comprises decrementing the RAS count as part of the manipulating.

12 . The method of claim 11 further comprising ignoring the one or more return instructions, wherein the RAS count is not zero.

13 . The method of claim 1 wherein the pushing further comprises updating, in the RAS, a next pointer field indexed by the write pointer with contents of the read pointer.

14 . The method of claim 13 further comprising updating the read pointer with contents of the write pointer.

15 . The method of claim 14 further comprising incrementing the write pointer.

16 . The method of claim 1 wherein the popping further comprises updating the read pointer with contents of a next pointer field indexed by the read pointer.

17 . The method of claim 1 wherein the predicted return address is the address of one of the one or more call instructions+4 bytes.

18 . The method of claim 1 further comprising initiating the RAS, wherein the initiating sets the write pointer and the read pointer to an initial value.

19 . The method of claim 1 wherein the RAS comprises eight entries.

20 . The method of claim 1 wherein the RAS comprises a Last-In-First-Out (LIFO) memory element.

21 . The method of claim 1 wherein the pushing includes storing data in the RAS at a location specified by the write pointer.

22 . The method of claim 1 wherein the popping includes reading data which was last stored in the RAS, wherein location of the data which was last stored is specified by the read pointer.

23 . A computer program product embodied in a non-transitory computer readable medium for predicting addresses, the computer program product comprising code which causes one or more processors to generate semiconductor logic for:

accessing a processor core, wherein the processor core includes a return address stack (RAS), a local cache hierarchy, and branch prediction logic, wherein the processor core is coupled to a memory system;

sending, to a branch execution unit, a write pointer, a read pointer, and a RAS count;

detecting one or more call instructions in an instruction stream, wherein the detecting generates a predicted return address for each of the one or more call instructions;

pushing, on the RAS, the predicted return address for each of the one or more call instructions, wherein the pushing is directed by the write pointer;

recognizing one or more return instructions in the instruction stream;

updating the write pointer and the read pointer for the RAS, wherein the updating is based on a misprediction signal from the branch execution unit and wherein the updating includes manipulating the RAS count;

popping, from the RAS, the predicted return address for each of the one or more return instructions wherein the popping is directed by the read pointer; and

rolling back, by the branch execution unit, the write pointer, the read pointer, and the RAS count to a previous value.

24 . An apparatus for predicting addresses comprising:

a processor core coupled to a memory wherein the processor core and the memory are used to perform operations comprising:

accessing the processor core, wherein the processor core includes a return address stack (RAS), a local cache hierarchy, and branch prediction logic, wherein the processor core is coupled to a memory system;

sending, to a branch execution unit, a write pointer, a read pointer, and a RAS count;

detecting one or more call instructions in an instruction stream, wherein the detecting generates a predicted return address for each of the one or more call instructions;

pushing, on the RAS, the predicted return address for each of the one or more call instructions, wherein the pushing is directed by the write pointer;

recognizing one or more return instructions in the instruction stream;

updating the write pointer and the read pointer for the RAS, wherein the updating is based on a misprediction signal from the branch execution unit and wherein the updating includes manipulating the RAS count;

popping, from the RAS, the predicted return address for each of the one or more return instructions wherein the popping is directed by the read pointer; and

rolling back, by the branch execution unit, the write pointer, the read pointer, and the RAS count to a previous value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2025
From: CHO, JAMES YOUNGSAE; SUGUMAR, RABIN
To: AKEANA, INC.
Reel/Frame 070940/0846 →
Continuity (18)
Provisional Application 63605620 · Dec 4, 2023
Provisional Application 63602514 · Nov 24, 2023
Provisional Application 63547574 · Nov 7, 2023
Provisional Application 63547404 · Nov 6, 2023
Provisional Application 63546769 · Nov 1, 2023
Provisional Application 63545961 · Oct 27, 2023
Provisional Application 63542797 · Oct 6, 2023
Provisional Application 63526009 · Jul 11, 2023
Provisional Application 63521365 · Jun 16, 2023
Provisional Application 63471283 · Jun 6, 2023
Provisional Application 63467335 · May 18, 2023
Provisional Application 63463371 · May 2, 2023
Provisional Application 63462542 · Apr 28, 2023
Provisional Application 63444619 · Feb 10, 2023
Provisional Application 63439761 · Jan 18, 2023
Provisional Application 63436133 · Dec 30, 2022
Provisional Application 63436144 · Dec 30, 2022
Related Publication 20240220267A1 · Jul 4, 2024
References Cited (26)
US 6253315B1 · Yeh · 2001 [cited by examiner]
US 6560696B1 · Hummel · 2003 [cited by examiner]
US 6898699B2 · Jourdan · 2005 [cited by examiner]
US 6910124B1 · Sinharoy · 2005 [cited by examiner]
US 6934809B2 · Tremblay et al. · 2005 [cited by applicant]
US 6973563B1 · Sander · 2005 [cited by examiner]
US 7506105B2 · Al-Sukhni et al. · 2009 [cited by applicant]
US 10013356B2 · Chou · 2018 [cited by applicant]
US 10671394B2 · Britto et al. · 2020 [cited by applicant]
US 10929948B2 · Benthin et al. · 2021 [cited by applicant]
US 11288405B2 · Belgarric et al. · 2022 [cited by applicant]
US 11403099B2 · Cerny et al. · 2022 [cited by applicant]
US 11403225B2 · Zheng et al. · 2022 [cited by applicant]
US 11429529B2 · Hornung et al. · 2022 [cited by applicant]
US 11442863B2 · Shulyak et al. · 2022 [cited by applicant]
US 11474130B2 · Lentz et al. · 2022 [cited by applicant]
US 11486911B2 · Tuncer et al. · 2022 [cited by applicant]
US 20110320790A1 · Dieffenderfer · 2011 [cited by examiner]
US 20140123286A1 · Fischer · 2014 [cited by examiner]
US 20140281394A1 · Smith · 2014 [cited by examiner]
US 20140317390A1 · Demongeot · 2014 [cited by examiner]
US 20140331028A1 · Demongeot · 2014 [cited by examiner]
US 20220004639A1 · Yardi et al. · 2022 [cited by applicant]
US 20220029780A1 · Dafali · 2022 [cited by applicant]
US 20220197657A1 · Soundararajan et al. · 2022 [cited by applicant]
WO 2022117687A1 · 2022 [cited by applicant]