IP Library › Granted Patent US 12,493,466
Granted Patent B1
US 12,493,466 · App. 18/645,272 · Granted Dec 9, 2025

Microprocessor that builds inconsistent loop that iteration count unrolled loop multi-fetch block macro-op cache entries

Inventors: John G. Favor (San Francisco, CA); Michael N. Michael (Folsom, CA)
Assignee: Ventana Micro Systems Inc.
G06F9/30065G06F9/321G06F9/3846
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,466
App. No.
18/645,272
Granted
Dec 9, 2025
Kind
B1
Abstract

A microprocessor includes a prediction unit (PRU) that predicts a sequence of fetch blocks (FBlks) in a program instruction stream, a macro-op (MOP) cache (MOC) that comprises MOC entries (MEs), a fusion engine. An ME holds MOPs into which architectural instructions of one or more FBlks are decoded. The PRU detects a loop body ME within the program instruction stream, accumulates loop iteration count information about a series of instances of a loop on the loop body ME in the program instruction stream, updates a consistency counter of the loop body ME while accumulating the loop iteration count information, and in response to detecting that the consistency counter has reached a threshold, instructs the fusion engine to use F copies of the MOPs of the loop body ME to build in the MOC an unrolled loop multi-FBlk ME; F is a loop unroll factor that is at least two.

Claims (146)

1 . A microprocessor, comprising:

a prediction unit (PRU) that predicts a sequence of fetch blocks (FBlks) in a program instruction stream, wherein a FBlk comprises a sequential run of architectural instructions;

a macro-op (MOP) cache (MOC) that comprises MOC entries (MEs), wherein a ME holds MOPs into which architectural instructions of one or more FBlks are decoded; and

a fusion engine;

wherein the PRU is configured to:

detect a loop body ME within the program instruction stream;

accumulate loop iteration count information about a series of instances of a loop on the loop body ME in the program instruction stream;

update a consistency counter of the loop body ME while accumulating the loop iteration count information;

in response to detecting that the consistency counter has reached a threshold, instruct the fusion engine to use F copies of the MOPs of the loop body ME to build in the MOC an unrolled loop multi-FBlk ME (MF-ME) (ULP-MF-ME), wherein F is a loop unroll factor that is at least two.

2 . The microprocessor of claim 1 ,

wherein the loop iteration count information comprises a minimum loop iteration count observed among the series of instances.

3 . The microprocessor of claim 2 ,

wherein the PRU is configured to select F based on the minimum loop iteration count.

4 . The microprocessor of claim 1 ,

wherein the loop body ME comprises a minimum loop iteration count that is initialized to a maximum value; and

wherein for each current instance of the series of instances of the loop on the loop body ME, wherein to update the consistency counter, the PRU:

when the loop iteration count of the current instance of the loop on the loop body ME is less than the minimum loop iteration count of the loop body ME:

resets the consistency counter; and

updates the minimum loop iteration count to the loop iteration count of the current instance; and

otherwise:

increments the consistency counter.

5 . The microprocessor of claim 1 ,

wherein the loop iteration count information comprises a histogram of frequencies over a range of loop iteration counts observed in the series of instances.

6 . The microprocessor of claim 5 ,

wherein for each instance of the series of instances of the loop on the loop body ME, to update the consistency counter the PRU increments the consistency counter.

7 . The microprocessor of claim 5 ,

wherein the PRU is configured to select F based on the histogram.

8 . The microprocessor of claim 5 ,

wherein the PRU is configured to select a loop iteration count from the range of loop iteration counts of the histogram.

9 . The microprocessor of claim 8 ,

wherein to select a loop iteration count from the range of loop iteration counts of the histogram, the PRU:

selects as the loop iteration count a smallest of the loop iteration counts of the range having a frequency greater than a predetermined value.

10 . The microprocessor of claim 8 ,

wherein to selecting a loop iteration count from the range of loop iteration counts of the histogram, the PRU:

selects as the loop iteration count a loop iteration count from the range of loop iteration counts of the histogram based on variability of the frequencies of the histogram.

11 . The microprocessor of claim 5 ,

wherein the loop body ME comprises an unrolled loop iteration count;

wherein the PRU is configured to:

populate the unrolled loop iteration count based on F and the histogram when the ULP-MF-ME is built; and

in response to a hit in the MOC on the ULP-MF-ME, fetch from the MOC a number of copies of the MOPs of the ULP-MF-ME equal to the unrolled loop iteration count.

12 . The microprocessor of claim 11 ,

wherein the PRU is configured to, after fetching from the MOC a number of copies of the MOPs of the ULP-MF-ME equal to the unrolled loop iteration count:

in response to a multiple-hit in the MOC on the ULP-MF-ME and on the loop body ME immediately subsequent to the hit in the MOC on the ULP-MF-ME, fetch the MOPs of the loop body ME from the MOC until falling out of the loop.

13 . The microprocessor of claim 1 ,

wherein the PRU is configured to:

select a loop iteration count using the loop iteration count information; and

populate the unrolled loop iteration count with a floor function of a quotient of the selected loop iteration count and F.

14 . The microprocessor of claim 1 ,

wherein the PRU is configured to:

select a loop iteration count using the loop iteration count information; and

populate the unrolled loop iteration count with a ceiling function of a quotient of the selected loop iteration count and F.

15 . The microprocessor of claim 1 ,

wherein the PRU is configured to:

select one loop iteration count from the range of loop iteration counts of the histogram; and

populate the unrolled loop iteration count with a ceiling function of a quotient of the selected one loop iteration count and F.

16 . The microprocessor of claim 1 ,

wherein the PRU is configured to, prior to said detecting the loop body ME, instruct the fusion engine to build in the MOC a sequential multi-FBlk ME (MF-ME) (SEQ-MF-ME) using the MOPs of a sequence of MEs in the program instruction stream;

wherein the loop body ME is the SEQ-MF-ME.

17 . The microprocessor of claim 16 ,

wherein the PRU is configured to retain the loop body ME in the MOC co-resident with the ULP-MF-ME when allocating the ULP-MF-ME in the MOC.

18 . The microprocessor of claim 17 ,

wherein the PRU is configured to set an indicator in the ULP-MF-ME to distinguish the ULP-MF-ME from the loop body ME when allocating the ULP-MF-ME in the MOC.

19 . The microprocessor of claim 1 ,

wherein the threshold is software configurable.

20 . The microprocessor of claim 1 ,

wherein the threshold is dynamically varied based on recent characteristics of the program instruction stream.

21 . The microprocessor of claim 1 ,

wherein the F copies of the MOPs of the loop body ME have a total number of MOPs;

wherein the ULP-MF-ME has a second number of MOPs; and

wherein to use the F copies of the MOPs of the loop body ME to build in the MOC the ULP-MF-ME, the fusion engine:

fuses the F copies of the MOPs of the loop body ME into the MOPs of the ULP-MF-ME such that the second number is fewer than the total number.

22 . A method, comprising:

in a microprocessor comprising:

a prediction unit (PRU) that predicts a sequence of fetch blocks (FBlks) in a program instruction stream, wherein a FBlk comprises a sequential run of architectural instructions; and

a macro-op (MOP) cache (MOC) that comprises MOC entries (MEs), wherein a ME holds MOPs into which architectural instructions of one or more FBlks are decoded;

detecting a loop body ME within the program instruction stream;

accumulating loop iteration count information about a series of instances of a loop on the loop body ME in the program instruction stream;

updating a consistency counter of the loop body ME while said accumulating the loop iteration count information;

in response to detecting that the consistency counter has reached a threshold, using F copies of the MOPs of the loop body ME to build in the MOC an unrolled loop multi-FBlk ME (MF-ME) (ULP-MF-ME), wherein F is a loop unroll factor that is at least two.

23 . The method of claim 22 ,

wherein the loop iteration count information comprises a minimum loop iteration count observed among the series of instances.

24 . The method of claim 23 , further comprising:

selecting F based on the minimum loop iteration count.

25 . The method of claim 22 ,

wherein the loop body ME comprises a minimum loop iteration count that is initialized to a maximum value; and

wherein for each current instance of the series of instances of the loop on the loop body ME, said updating the consistency counter comprises:

when the loop iteration count of the current instance of the loop on the loop body ME is less than the minimum loop iteration count of the loop body ME:

resetting the consistency counter; and

updating the minimum loop iteration count to the loop iteration count of the current instance; and

otherwise:

incrementing the consistency counter.

26 . The method of claim 22 ,

wherein the loop iteration count information comprises a histogram of frequencies over a range of loop iteration counts observed in the series of instances.

27 . The method of claim 26 ,

wherein for each instance of the series of instances of the loop on the loop body ME, said updating the consistency counter comprises incrementing the consistency counter.

28 . The method of claim 26 , further comprising:

selecting F based on the histogram.

29 . The method of claim 26 , further comprising:

selecting a loop iteration count from the range of loop iteration counts of the histogram.

30 . The method of claim 29 ,

wherein said selecting a loop iteration count from the range of loop iteration counts of the histogram comprises:

selecting as the loop iteration count a smallest of the loop iteration counts of the range having a frequency greater than a predetermined value.

31 . The method of claim 29 ,

wherein said selecting a loop iteration count from the range of loop iteration counts of the histogram comprises:

selecting as the loop iteration count a loop iteration count from the range of loop iteration counts of the histogram based on variability of the frequencies of the histogram.

32 . The method of claim 26 , further comprising:

wherein the loop body ME comprises an unrolled loop iteration count;

populating the unrolled loop iteration count based on F and the histogram when the ULP-MF-ME is built; and

in response to a hit in the MOC on the ULP-MF-ME, fetching from the MOC a number of copies of the MOPs of the ULP-MF-ME equal to the unrolled loop iteration count.

33 . The method of claim 32 , further comprising:

after said fetching from the MOC a number of copies of the MOPs of the ULP-MF-ME equal to the unrolled loop iteration count:

in response to a multiple-hit in the MOC on the ULP-MF-ME and on the loop body ME immediately subsequent to the hit in the MOC on the ULP-MF-ME, fetching the MOPs of the loop body ME from the MOC until falling out of the loop.

34 . The method of claim 22 , further comprising:

selecting a loop iteration count using the loop iteration count information; and

populating the unrolled loop iteration count with a floor function of a quotient of the selected loop iteration count and F.

35 . The method of claim 22 , further comprising:

selecting a loop iteration count using the loop iteration count information; and

populating the unrolled loop iteration count with a ceiling function of a quotient of the selected loop iteration count and F.

36 . The method of claim 22 , further comprising:

selecting one loop iteration count from the range of loop iteration counts of the histogram; and

populating the unrolled loop iteration count with a ceiling function of a quotient of the selected one loop iteration count and F.

37 . The method of claim 22 , further comprising:

prior to said detecting the loop body ME, building in the MOC a sequential multi-FBlk ME (MF-ME) (SEQ-MF-ME) using the MOPs of a sequence of MEs in the program instruction stream;

wherein the loop body ME is the SEQ-MF-ME.

38 . The method of claim 37 , further comprising:

retaining the loop body ME in the MOC co-resident with the ULP-MF-ME when allocating the ULP-MF-ME in the MOC.

39 . The method of claim 38 , further comprising:

setting an indicator in the ULP-MF-ME to distinguish the ULP-MF-ME from the loop body ME when allocating the ULP-MF-ME in the MOC.

40 . The method of claim 22 ,

wherein the threshold is software configurable.

41 . The method of claim 22 , further comprising:

dynamically varying the threshold based on recent characteristics of the program instruction stream.

42 . The method of claim 22 ,

wherein the F copies of the MOPs of the loop body ME have a total number of MOPs;

wherein the ULP-MF-ME has a second number of MOPs; and

wherein said using the F copies of the MOPs of the loop body ME to build in the MOC the ULP-MF-ME comprises:

fusing the F copies of the MOPs of the loop body ME into the MOPs of the ULP-MF-ME such that the second number is fewer than the total number.

43 . A non-transitory computer-readable medium having instructions stored thereon that are capable of causing or configuring a microprocessor comprising:

a prediction unit (PRU) that predicts a sequence of fetch blocks (FBlks) in a program instruction stream, wherein a FBlk comprises a sequential run of architectural instructions;

a macro-op (MOP) cache (MOC) that comprises MOC entries (MEs), wherein a ME holds MOPs into which architectural instructions of one or more FBlks are decoded; and

a fusion engine;

wherein the PRU is configured to:

detect a loop body ME within the program instruction stream;

accumulate loop iteration count information about a series of instances of a loop on the loop body ME in the program instruction stream;

update a consistency counter of the loop body ME while accumulating the loop iteration count information;

in response to detecting that the consistency counter has reached a threshold, instruct the fusion engine to use F copies of the MOPs of the loop body ME to build in the MOC an unrolled loop multi-FBlk ME (MF-ME) (ULP-MF-ME), wherein F is a loop unroll factor that is at least two.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2024
From: FAVOR, JOHN G.; MICHAEL, MICHAEL N.
To: VENTANA MICRO SYSTEMS INC.
Reel/Frame 067941/0951 →
Continuity (4)
Continuation In Part 18380152 · Oct 13, 2023
Continuation In Part 18380150 · Oct 13, 2023
Continuation In Part 18240249 · Aug 30, 2023
Continuation In Part 18240249 · Aug 30, 2023
References Cited (76)
US 4161026A · Wilhite · 1979 [cited by applicant]
US 5724565A · Dubey · 1998 [cited by examiner]
US 6473832B1 · Ramagopal et al. · 2002 [cited by applicant]
US 6539541B1 · Geva · 2003 [cited by applicant]
US 6799263B1 · Morris et al. · 2004 [cited by applicant]
US 7496726B1 · Nussbaum · 2009 [cited by applicant]
US 7590825B2 · Krimer et al. · 2009 [cited by applicant]
US 7681019B1 · Favor · 2010 [cited by applicant]
US 7797517B1 · Favor · 2010 [cited by applicant]
US 7814298B1 · Thaik et al. · 2010 [cited by applicant]
US 7870369B1 · Nelson et al. · 2011 [cited by applicant]
US 7941607B1 · Thaik et al. · 2011 [cited by applicant]
US 7949854B1 · Thaik et al. · 2011 [cited by applicant]
US 7953933B1 · Thaik et al. · 2011 [cited by applicant]
US 7953961B1 · Thaik et al. · 2011 [cited by applicant]
US 7987342B1 · Thaik et al. · 2011 [cited by applicant]
US 8024522B1 · Favor et al. · 2011 [cited by applicant]
US 8032710B1 · Ashcraft et al. · 2011 [cited by applicant]
US 8037285B1 · Thaik et al. · 2011 [cited by applicant]
US 8103831B2 · Rappoport et al. · 2012 [cited by applicant]
US 8370609B1 · Favor et al. · 2013 [cited by applicant]
US 8499293B1 · Ashcraft et al. · 2013 [cited by applicant]
US 8930679B2 · Day et al. · 2015 [cited by applicant]
US 9524164B2 · Olson et al. · 2016 [cited by applicant]
US 10579535B2 · Rappoport et al. · 2020 [cited by applicant]
US 10896044B2 · Evers et al. · 2021 [cited by applicant]
US 20020087794A1 · Jouppi et al. · 2002 [cited by applicant]
US 20030140245A1 · Dahan · 2003 [cited by applicant]
US 20050257037A1 · Elwood et al. · 2005 [cited by applicant]
US 20080005547A1 · Papakipos et al. · 2008 [cited by applicant]
US 20080126771A1 · Chen et al. · 2008 [cited by applicant]
US 20120311308A1 · Xekalakis et al. · 2012 [cited by applicant]
US 20130212352A1 · Anderson · 2013 [cited by examiner]
US 20140006698A1 · Chappell et al. · 2014 [cited by applicant]
US 20140007061A1 · Perkins et al. · 2014 [cited by applicant]
US 20140143494A1 · Whalley · 2014 [cited by examiner]
US 20140229719A1 · Smeets et al. · 2014 [cited by applicant]
US 20150100762A1 · Jacobs · 2015 [cited by applicant]
US 20150149747A1 · Lee · 2015 [cited by examiner]
US 20170139706A1 · Chou et al. · 2017 [cited by applicant]
US 20180095752A1 · Kudaravalli et al. · 2018 [cited by applicant]
US 20190130102A1 · Johnson et al. · 2019 [cited by applicant]
US 20190163902A1 · Reid et al. · 2019 [cited by applicant]
US 20190188142A1 · Rappoport et al. · 2019 [cited by applicant]
US 20190196833A1 · Bouzguarrou et al. · 2019 [cited by applicant]
US 20190303161A1 · Nassi et al. · 2019 [cited by applicant]
US 20200082280A1 · Orion et al. · 2020 [cited by applicant]
US 20200110610A1 · Lapeyre et al. · 2020 [cited by applicant]
US 20200125498A1 · Betts et al. · 2020 [cited by applicant]
US 20200150967A1 · Ishii et al. · 2020 [cited by applicant]
US 20200372129A1 · Gupta · 2020 [cited by applicant]
US 20210026770A1 · Ishii et al. · 2021 [cited by applicant]
US 20210064533A1 · Schinzler et al. · 2021 [cited by applicant]
US 20210124586A1 · Bouzguarrou et al. · 2021 [cited by applicant]
US 20210397452A1 · Ireland et al. · 2021 [cited by applicant]
US 20220067155A1 · Favor et al. · 2022 [cited by applicant]
US 20220107807A1 · Schinzler et al. · 2022 [cited by applicant]
US 20220342671A1 · Schinzler et al. · 2022 [cited by applicant]
US 20230305962A1 · Dutta · 2023 [cited by applicant]
US 20230367600A1 · Dutta · 2023 [cited by applicant]
US 20240045610A1 · Favor et al. · 2024 [cited by applicant]
US 20250077438A1 · Favor et al. · 2025 [cited by applicant]
Burtscher, Martin et al. “Load Value Prediction Using Prediction Outcome Histories.” 1999 Technical Report CU-CS-873-98. Department of Computer Science, University of Colorado. pp. 1-9. [cited by applicant]
Appendix to the specification, 25 Pages, Mail room date Jul. 23, 2007, Doc code Appendix, referred to as “Appendix A” at col. 4, lines 46-47 of U.S. Pat. No. 7,987,342 to Thaik et al. issued Jul. 26, 2011; downloaded Ju… [cited by applicant]
Appendix to the specification, 28 Pages, Mail room date Jul. 23, 2007, Doc code Appendix, referred to as “Appendix B” at col. 4, lines 48-49 of U.S. Pat. No. 7,987,342 to Thaik et al. issued Jul. 26, 2011; downloaded Ju… [cited by applicant]
White Paper. “Security Analysis of AMD Predictive Store Forwarding.” Advanced Micro Devices, Inc. (AMD). Aug. 2023. pp. 1-7. [cited by applicant]
Liu, Chang et al. “Uncovering and Exploiting AMD Speculative Memory Access Predictors for Fun and Profit.” 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Mar. 2-6, 2024. pp. 1-15. [cited by applicant]
Rotenberg, Eric et al. “Trace Cache: a Low Latency Approach to High Bandwidth Instruction Fetching.” [cited by applicant]
Ren, Xida et al. “I see Dead μops: Leaking Secrets via Intel/AMD Micro-Op Caches.” [cited by applicant]
Kotra, Jagadish B. et al. “Improving the Utilization of Micro-operation Caches in x86 Processors.” [cited by applicant]
Slechta, Brian et al. “Dynamic Optimization of Micro-Operations.” HPCA '03: Proceedings of the 9th International Symposium on High-Performance Computer Architecture. Feb. 2003. Pages 1-12. [cited by applicant]
Petric, Vlad et al. “Reno: A Rename-Based Instruction Optimizer.” ACM SIGARCH Computer Architecture News, vol. 33, Issue 2. May 2005. pp. 98-109. [cited by applicant]
Patel, Sanjay J et al. “rePLay: A Hardware Framework for Dynamic Optimization.” IEEE Transactions on Computers, vol. 50, No. 6. Jun. 2001. pp. 590-608. [cited by applicant]
Moody, Logan et al. “Speculative Code Compaction: Eliminating Dead Code via Speculative Microcode Transformations.” 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) Chicago, IL. 2022. pp. 162-180. [cited by applicant]
Behar, Michael et al. “Trace Cache Sampling Filter.” ACM Transactions on Computer Systems. Feb. 2007. pp. 1-10. [cited by applicant]
Friendly, Daniel Holmes et al. “Putting the fill unit to work: dynamic optimizations for trace cache microprocessors.” MICRO 31: Proceedings of the 31st Annual ACM/IEEE International Symposium on Microarchitecture. Nov.… [cited by applicant]
Cited By (1)
US 12,693,864