IP Library Granted Patent US 12,282,430
Granted Patent B1
US 12,282,430 · App. 18/380,152 · Granted Apr 22, 2025

Macro-op cache data entry pointers distributed as initial pointers held in tag array and next pointers held in data array for efficient and performant variable length macro-op cache entries

Inventors: John G. Favor (San Francisco, CA); Michael N. Michael (Folsom, CA)
Assignee: Ventana Micro Systems Inc.
G06F12/0864G06F2212/452
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,430
App. No.
18/380,152
Granted
Apr 22, 2025
Kind
B1
Abstract

A microprocessor includes a macro-op (MOP) cache (MOC) comprising a set-associative MOC tag RAM (MTR) and a MOC data RAM (MDR) managed as a pool of MDR entries. A MOC entry (ME) comprises one MTR entry and one or more MDR entries that hold the MOPs of the ME. The MDR entries of the ME have a program order. Each MDR entry holds MOPs and a next MDR entry pointer. Each MTR entry holds initial MDR entry pointers and specifies the number of the MDR entries of the ME. During ME allocation, the MOC populates the MDR entry pointers to point to the MDR entries based on the program order. In response to an access that hits upon an MTR entry, the MOC fetches the MDR entries according to the program order initially using the initial pointers and subsequently using the next pointers.

Claims (123)

1. A microprocessor, comprising:

a macro-op (MOP) cache (MOC) comprising:

a MOC tag RAM (MTR) arranged as a set-associative cache of MTR entries; and

a MOC data RAM (MDR) managed as a pool of MDR entries;

wherein a MOC entry (ME) comprises:

one MTR entry; and

one or more MDR entries that hold the MOPs of the ME, wherein the MDR entries of the ME have a program order;

wherein each MDR entry is configured to hold:

one or more MOPs of the ME; and

a next MDR entry pointer;

wherein each MTR entry is configured to hold:

a length that specifies the number of the MDR entries of the ME; and

one or more initial MDR entry pointers;

wherein the MOC is configured to:

during allocation of a ME into the MOC, populate the one or more initial MDR entry pointers and the next MDR entry pointers to point to the one or more MDR entries of the ME based on the program order; and

in response to an access of the MOC that hits upon the MTR entry of an ME, fetch the MDR entries of the ME according to the program order:

initially using the one or more initial MDR entry pointers of the hit upon MTR entry; and

subsequently using the next MDR entry pointers of fetched MDR entries of the ME, until all the MDR entries of the ME have been fetched from the MDR.

2. The microprocessor of claim 1 ,

wherein the one or more initial MDR entry pointers is a predetermined number; and

wherein the number of MDR entries of a ME is constrained by a size of the pool of MDR entries but is not constrained by the predetermined number of the one or more initial MDR entry pointers because of the next MDR entry pointers.

3. The microprocessor of claim 1 ,

wherein each MDR entry is configured to hold up to a predetermined number Q of MOPs,

wherein Q is greater than zero.

4. The microprocessor of claim 3 ,

wherein Q is three.

5. The microprocessor of claim 1 ,

wherein the MDR comprises a pipeline having a fetch latency of S clock cycles, wherein S is at least two;

wherein the one or more initial MDR entry pointers comprise at least S initial MDR entry pointers; and

wherein the S initial MDR entry pointers are used to fetch first through Sth MDR entries of the ME during first through Sth adjacent clock cycles such that pipeline bubbles are avoided during transition from fetching the MDR entries of the ME initially using the one or more initial MDR entry pointers provided by the hit upon MTR entry to fetching the remaining MDR entries of the ME using the next MDR entry pointers provided by the MDR.

6. The microprocessor of claim 1 ,

wherein the one or more initial MDR entry pointers comprise B initial MDR entry pointers for concurrently fetching B MDR entries from the MDR, wherein B is greater than one.

7. The microprocessor of claim 6 ,

wherein the MDR comprises B read ports for concurrently fetching the B MDR entries.

8. The microprocessor of claim 6 ,

wherein the number of MDR entries of a ME need not be allocated in quanta of B MDR entries but is instead allocatable in quanta of less than B MDR entries.

9. The microprocessor of claim 1 ,

wherein the MDR comprises a pipeline having a fetch latency of S clock cycles, wherein S is at least two;

wherein the one or more initial MDR entry pointers comprise S*B initial MDR entry pointers; and

wherein the S*B initial MDR entry pointers are used to fetch first through S*Bth MDR entries of the ME during first through Sth adjacent clock cycles such that pipeline bubbles are avoided during transition from fetching the MDR entries of the ME initially using the one or more initial MDR entry pointers provided by the hit upon MTR entry to fetching the remaining MDR entries of the ME using the next MDR entry pointers provided by the MDR.

10. The microprocessor of claim 1 ,

wherein each MDR entry is configured to hold up to a predetermined number Q of MOPs,

wherein Q is greater than one.

11. The microprocessor of claim 10 ,

wherein Q is three.

12. The microprocessor of claim 1 ,

wherein the MOC allocates the MDR entries of the pool such that the initial MDR entry pointers and the next MDR entry pointers may point to any MDR entry of the pool.

13. The microprocessor of claim 1 , further comprising:

a fusion engine configured to:

fuse MOPs decoded from architectural instructions of a plurality of fetch blocks into the MOPs of a ME; and

allocate the ME into the MOC;

wherein a fetch block comprises a sequential run of architectural instructions in a program instruction stream.

14. A method, comprising:

in a microprocessor comprising:

a macro-op (MOP) cache (MOC) comprising:

a MOC tag RAM (MTR) arranged as a set-associative cache of MTR entries; and

a MOC data RAM (MDR) managed as a pool of MDR entries;

wherein a MOC entry (ME) comprises:

one MTR entry; and

one or more MDR entries that hold the MOPs of the ME, wherein the MDR entries of the ME have a program order;

wherein each MDR entry is configured to hold:

one or more MOPs of the ME; and

a next MDR entry pointer;

wherein each MTR entry is configured to hold:

a length that specifies the number of the MDR entries of the ME; and

one or more initial MDR entry pointers;

during allocation of a ME into the MOC, populating by the MOC the one or more initial MDR entry pointers and the next MDR entry pointers to point to the one or more MDR entries of the ME based on the program order; and

in response to an access of the MOC that hits upon the MTR entry of an ME, fetching by the MOC the MDR entries of the ME according to the program order:

initially using the one or more initial MDR entry pointers of the hit upon MTR entry; and

subsequently using the next MDR entry pointers of fetched MDR entries of the ME, until all the MDR entries of the ME have been fetched from the MDR.

15. The method of claim 14 ,

wherein the one or more initial MDR entry pointers is a predetermined number; and

wherein the number of MDR entries of a ME is constrained by a size of the pool of MDR entries but is not constrained by the predetermined number of the one or more initial MDR entry pointers because of the next MDR entry pointers.

16. The method of claim 14 ,

wherein each MDR entry is configured to hold up to a predetermined number Q of MOPs,

wherein Q is greater than zero.

17. The microprocessor of claim 16 ,

wherein Q is three.

18. The method of claim 14 ,

wherein the MDR comprises a pipeline having a fetch latency of S clock cycles, wherein S is at least two; and

wherein the one or more initial MDR entry pointers comprise at least S initial MDR entry pointers;

the method further comprising:

using the S initial MDR entry pointers to fetch first through Sth MDR entries of the ME during first through Sth adjacent clock cycles such that pipeline bubbles are avoided during transition from fetching the MDR entries of the ME initially using the one or more initial MDR entry pointers provided by the hit upon MTR entry to fetching the remaining MDR entries of the ME using the next MDR entry pointers provided by the MDR.

19. The method of claim 14 ,

wherein the one or more initial MDR entry pointers comprise B initial MDR entry pointers for concurrently fetching B MDR entries from the MDR, wherein B is greater than one.

20. The method of claim 19 ,

wherein the MDR comprises B read ports for concurrently fetching the B MDR entries.

21. The method of claim 19 ,

wherein the number of MDR entries of a ME need not be allocated in quanta of B MDR entries but is instead allocatable in quanta of less than B MDR entries.

22. The method of claim 14 ,

wherein the MDR comprises a pipeline having a fetch latency of S clock cycles, wherein S is at least two; and

wherein the one or more initial MDR entry pointers comprise S*B initial MDR entry pointers; and

the method further comprising:

using the S*B initial MDR entry pointers to fetch first through S*Bth MDR entries of the ME during first through Sth adjacent clock cycles such that pipeline bubbles are avoided during transition from fetching the MDR entries of the ME initially using the one or more initial MDR entry pointers provided by the hit upon MTR entry to fetching the remaining MDR entries of the ME using the next MDR entry pointers provided by the MDR.

23. The method of claim 14 ,

wherein each MDR entry is configured to hold up to a predetermined number Q of MOPs,

wherein Q is greater than one.

24. The method of claim 10 ,

wherein Q is three.

25. The method of claim 14 ,

wherein the MOC allocates the MDR entries of the pool such that the initial MDR entry pointers and the next MDR entry pointers may point to any MDR entry of the pool.

26. The method of claim 14 , further comprising:

fusing, by a fusion engine, MOPs decoded from architectural instructions of a plurality of fetch blocks into the MOPs of a ME; and

allocating, by the fusion engine, the ME into the MOC;

wherein a fetch block comprises a sequential run of architectural instructions in a program instruction stream.

27. A non-transitory computer-readable medium having instructions stored thereon that are capable of causing or configuring a microprocessor comprising:

a macro-op (MOP) cache (MOC) comprising:

a MOC tag RAM (MTR) arranged as a set-associative cache of MTR entries; and

a MOC data RAM (MDR) managed as a pool of MDR entries;

wherein a MOC entry (ME) comprises:

one MTR entry; and

one or more MDR entries that hold the MOPs of the ME, wherein the MDR entries of the ME have a program order;

wherein each MDR entry is configured to hold:

one or more MOPs of the ME; and

a next MDR entry pointer;

wherein each MTR entry is configured to hold:

a length that specifies the number of the MDR entries of the ME; and

one or more initial MDR entry pointers;

wherein the MOC is configured to:

during allocation of a ME into the MOC, populate the one or more initial MDR entry pointers and the next MDR entry pointers to point to the one or more MDR entries of the ME based on the program order; and

in response to an access of the MOC that hits upon the MTR entry of an ME, fetch the MDR entries of the ME according to the program order:

initially using the one or more initial MDR entry pointers of the hit upon MTR entry; and

subsequently using the next MDR entry pointers of fetched MDR entries of the ME, until all the MDR entries of the ME have been fetched from the MDR.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2025
From: FAVOR, JOHN G.; MICHAEL, MICHAEL N.
To: VENTANA MICRO SYSTEMS INC.
Reel/Frame 072308/0956 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: FAVOR, JOHN G.; MICHAEL, MICHAEL N.
To: VENTANA MICRO SYSTEMS INC.
Reel/Frame 065545/0331 →
Continuity (1)
Continuation In Part 18240249 · Aug 30, 2023
References Cited (43)
US 7590825B2 · Krimer et al. · 2009 [cited by applicant]
US 7681019B1 · Favor · 2010 [cited by applicant]
US 7797517B1 · Favor · 2010 [cited by applicant]
US 7814298B1 · Thaik et al. · 2010 [cited by applicant]
US 7870369B1 · Nelson et al. · 2011 [cited by applicant]
US 7941607B1 · Thaik et al. · 2011 [cited by applicant]
US 7949854B1 · Thaik et al. · 2011 [cited by applicant]
US 7953933B1 · Thaik et al. · 2011 [cited by applicant]
US 7953961B1 · Thaik et al. · 2011 [cited by applicant]
US 7987342B1 · Thaik et al. · 2011 [cited by applicant]
US 8032710B1 · Ashcraft et al. · 2011 [cited by applicant]
US 8037285B1 · Thaik et al. · 2011 [cited by applicant]
US 8103831B2 · Rappoport et al. · 2012 [cited by applicant]
US 8370609B1 · Favor et al. · 2013 [cited by applicant]
US 8499293B1 · Ashcraft et al. · 2013 [cited by applicant]
US 8930679B2 · Day et al. · 2015 [cited by applicant]
US 9524164B2 · Olson et al. · 2016 [cited by applicant]
US 10579535B2 · Rappoport et al. · 2020 [cited by applicant]
US 20120311308A1 · Xekalakis et al. · 2012 [cited by applicant]
US 20150100762A1 · Jacobs · 2015 [cited by applicant]
US 20170139706A1 · Chou et al. · 2017 [cited by applicant]
US 20190188142A1 · Rappoport · 2019 [cited by examiner]
US 20190303161A1 · Nassi et al. · 2019 [cited by applicant]
US 20200110610A1 · Lapeyre · 2020 [cited by examiner]
US 20200125498A1 · Betts et al. · 2020 [cited by applicant]
US 20210026770A1 · Ishii · 2021 [cited by examiner]
US 20220107807A1 · Schinzler · 2022 [cited by examiner]
US 20230305962A1 · Dutta · 2023 [cited by examiner]
US 20250077438A1 · Favor et al. · 2025 [cited by applicant]
Slechta, Brian et al. “Dynamic Optimization of Micro-Operations.” HPCA '03: Proceedings of the 9th International Symposium on High-Performance Computer Architecture. Feb. 2003. pp. 1-12. [cited by applicant]
Petric, Vlad et al. “RENO: A Rename-Based Instruction Optimizer.” ACM SIGARCH Computer Architecture News, vol. 33, Issue 2. May 2005. pp. 98-109. [cited by applicant]
Patel, Sanjay J. et al. “rePLay: A Hardware Framework for Dynamic Optimization.” IEEE Transactions on Computers, vol. 50, No. 6. Jun. 2001. pp. 590-608. [cited by applicant]
Moody, Logan et al. “Speculative Code Compaction: Eliminating Dead Code via Speculative Microcode Transformations.” 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) Chicago, IL. 2022. pp. 162-180. [cited by applicant]
Behar, Michael et al. “Trace Cache Sampling Filter.” ACM Transactions on Computer Systems. Feb. 2007. pp. 1-10. [cited by applicant]
Friendly, Daniel Holmes et al. “Putting the fill unit to work: dynamic optimizations for trace cache microprocessors.” MICRO 31: Proceedings of the 31st Annual ACM/IEEE International Symposium on Microarchitecture. Nov.… [cited by applicant]
Rotenberg, Eric et al. “Trace Cache: a Low Latency Approach to High Bandwidth Instruction Fetching.” [cited by applicant]
Ren, Xida et al. “I see Dead μops: Leaking Secrets via Intel/AMD Micro-Op Caches.” [cited by applicant]
Kotra, Jagadish B. et al. “Improving the Utilization of Micro-operation Caches in x86 Processors.” [cited by applicant]
Burtscher, Martin et al. “Load Value Prediction Using Prediction Outcome Histories.” 1999 Technical Report CU-CS-873-98. Department of Computer Science, University of Colorado. pp. 1-9. [cited by applicant]
Appendix to the specification, 25 Pages, Mail room date Jul. 23, 2007, Doc code Appendix, referred to as “Appendix A” at col. 4, lines 46-47 of U.S. Pat. No. 7,987,342 to Thaik et al. issued Jul. 26, 2011; downloaded Ju… [cited by applicant]
Appendix to the specification, 28 Pages, Mail room date Jul. 23, 2007, Doc code Appendix, referred to as “Appendix B” at col. 4, lines 48-49 of U.S. Pat. No. 7,987,342 to Thaik et al. issued Jul. 26, 2011; downloaded Ju… [cited by applicant]
White Paper. “Security Analysis of AMD Predictive Store Forwarding.” Advanced Micro Devices, Inc. (AMD). Aug. 2023. pp. 1-7. [cited by applicant]
Liu, Chang et al. “Uncovering and Exploiting AMD Speculative Memory Access Predictors for Fun and Profit.” 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Mar. 2-6, 2024. pp. 1-15. [cited by applicant]
Cited By (1)
US 12,693,864