IP Library › Granted Patent US 12,373,207
Granted Patent B2
US 12,373,207 · App. 17/325,067 · Granted Jul 29, 2025

Implementing a micro-operation cache with compaction

Inventors: Jagadish B. Kotra (Austin, TX); John Kalamatianos (Boxborough, MA)
Assignee: Advanced Micro Devices, Inc.
G06F9/223G06F9/3016G06F12/0875G06F2212/452
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,207
App. No.
17/325,067
Granted
Jul 29, 2025
Kind
B2
Abstract

Systems, apparatuses, and methods for compacting multiple groups of micro-operations into individual cache lines of a micro-operation cache are disclosed. A processor includes at least a decode unit and a micro-operation cache. When a new group of micro-operations is decoded and ready to be written to the micro-operation cache, the micro-operation cache determines which set is targeted by the new group of micro-operations. If there is a way in this set that can store the new group without evicting any existing group already stored in the way, then the new group is stored into the way with the existing group(s) of micro-operations. Metadata is then updated to indicate that the new group of micro-operations has been written to the way. Additionally, the micro-operation cache manages eviction and replacement policy at the granularity of micro-operation groups rather than at the granularity of cache lines.

Claims (41)

1. A processor, comprising:

a decode unit comprising circuitry configured to decode instructions from an instruction stream into a first group of micro-operations; and

a micro-operation cache comprising circuitry configured to:

identify a victim compacted cache line for storing the first group of micro-operations; and

evict at least one micro-operation, but fewer than all micro-operations, from the victim compacted cache line.

2. The processor as recited in claim 1 , wherein to identify the victim compacted cache line for storing the first group of micro-operations, the micro-operation cache is further configured to determine an existing compacted cache line comprising one or more micro-operations is not available that has sufficient space to store the first group of micro-operations.

3. The processor as recited in claim 1 , wherein to evict the at least one micro-operation from the victim compacted cache line, the micro-operation cache is configured to evict fewer than all micro-operations from the victim compacted cache line to form sufficient space to store the first group of micro-operations.

4. The processor as recited in claim 1 , wherein to identify the victim compacted cache line for storing the first group of micro-operations, the micro-operation cache is configured to determine a targeted set of the micro-operation cache includes a victim way that has a second group of micro-operations that when evicted provides sufficient space to store the first group of micro-operations.

5. The processor as recited in claim 1 , wherein the micro-operation cache is further configured to determine that evicting the at least one micro-operation will provide sufficient space to store the first group of micro-operations.

6. The processor as recited in claim 1 , wherein the micro-operation cache is further configured to:

merge the first group of instructions in the victim compacted cache line with at least one additional micro-operation in the victim compacted cache line.

7. The processor as recited in claim 4 , wherein the micro-operation cache is further configured to update metadata of the victim way to account for eviction of the second group of micro-operations and fill of the first group of micro-operations.

8. A method, comprising:

decoding, by decode unit circuitry, instructions from an instruction stream into a first group of micro-operations;

identifying, by a micro-operation cache, a victim compacted cache line for storing the first group of micro-operations; and

evicting, by the micro-operation cache, at least one micro-operation, but fewer than all micro-operations, from the victim compacted cache line.

9. The method as recited in claim 8 , wherein identifying the victim compacted cache line for storing the first group of micro-operations comprises determining an existing compacted cache line comprising one or more micro-operations is not available that has sufficient space to store the first group of micro-operations.

10. The method as recited in claim 8 , wherein evicting the at least one micro-operation from the victim compacted cache line comprises evicting the at least one micro-operation, but fewer than all micro-operations, from the victim compacted cache line to form sufficient space to store the first group of micro-operations.

11. The method as recited in claim 8 , wherein identifying the victim compacted cache line for storing the first group of micro-operations comprises:

determining, by the micro-operation cache, a targeted set of the micro-operation cache includes a victim way that has a second group of micro-operations that when evicted provides sufficient space to store the first group of micro-operations, wherein the victim way is the victim compacted cache line.

12. The method as recited in claim 8 , further comprising merging the first group of micro-operations with at least one other micro-operation in a given compacted cache line of the micro-operation cache.

13. The method as recited in claim 11 , further comprising:

evicting, by the micro-operation cache, the second group of instructions from the victim way; and

merging, by the micro-operation cache, the first group of instructions in the victim way to store the first group of instructions and at least one additional micro-operation in the victim way.

14. The method as recited in claim 12 , further comprising updating, by the micro-operation cache, metadata of the given compacted cache line to identify both the first group of micro-operations and the at least one other micro-operation.

15. A system, comprising:

a memory; and

a processor coupled to the memory;

wherein the processor comprises circuitry configured to:

receive an instruction stream from the memory;

decode instructions from the instruction stream into a first group of micro-operations;

identify a victim compacted cache line for storing the first group of micro-operations; and

evict at least one micro-operation, but fewer than all micro-operations, from the victim compacted cache line.

16. The system as recited in claim 15 , wherein to identify the victim compacted cache line for storing the first group of micro-operations, the processor is further configured to determine an existing compacted cache line comprising one or more micro-operations is not available that has sufficient space to store the first group of micro-operations.

17. The system as recited in claim 15 , wherein to evict the at least one micro-operation from the victim compacted cache line, the processor is further configured to evict the at least one micro-operation, but fewer than all micro-operations, from the victim compacted cache line to form sufficient space to store the first group of micro-operations.

18. The system as recited in claim 15 , wherein to identify the victim compacted cache line for storing the first group of micro-operations, the processor is further configured to:

determine a targeted set that includes a victim way that has a second group of micro-operations that when evicted provides sufficient space to store the first group of micro-operations, wherein the victim way is the victim compacted cache line.

19. The system as recited in claim 18 , wherein the processor is further configured to determine the second group of instructions is a minimum sized group of instructions in the targeted set that when evicted provides sufficient space to store the first group of micro-operations.

20. The system as recited in claim 18 , wherein the processor is further configured to:

evict the second group of instructions from the victim way; and

merge the first group of instructions in the victim way to store the first group of instructions and at least one additional micro-operation in the victim way.

Continuity (2)
Continuation 16297358 · Mar 8, 2019
Related Publication 20210279054A1 · Sep 9, 2021
References Cited (39)
US 4021779A · Gardner · 1977 [cited by applicant]
US 4455604A · Ahlstrom et al. · 1984 [cited by applicant]
US 4498132A · Ahlstrom et al. · 1985 [cited by applicant]
US 4642757A · Sakamoto · 1987 [cited by applicant]
US 4901235A · Vora et al. · 1990 [cited by applicant]
US 5036453A · Renner et al. · 1991 [cited by applicant]
US 5574927A · Scantlin · 1996 [cited by applicant]
US 5649112A · Yeager et al. · 1997 [cited by applicant]
US 5671356A · Wang · 1997 [cited by applicant]
US 5796972A · Johnson et al. · 1998 [cited by applicant]
US 5845102A · Miller et al. · 1998 [cited by applicant]
US 6141740A · Mahalingaiah et al. · 2000 [cited by applicant]
US 7095342B1 · Hum et al. · 2006 [cited by applicant]
US 7165169B2 · Henry et al. · 2007 [cited by applicant]
US 7734873B2 · Lauterbach et al. · 2010 [cited by applicant]
US 7743232B2 · Shen et al. · 2010 [cited by applicant]
US 8195886B2 · Ozer · 2012 [cited by examiner]
US 8370609B1 · Favor et al. · 2013 [cited by applicant]
US 8782374B2 · Rappoport et al. · 2014 [cited by applicant]
US 9244686B2 · Henry et al. · 2016 [cited by applicant]
US 9262327B2 · Steely, Jr. · 2016 [cited by examiner]
US 10719447B2 · Akenine-Moller · 2020 [cited by examiner]
US 11016763B2 · Kotra et al. · 2021 [cited by applicant]
US 11681620B2 · Scrbak · 2023 [cited by examiner]
US 20060010310A1 · Bean et al. · 2006 [cited by applicant]
US 20070083735A1 · Glew · 2007 [cited by applicant]
US 20100235576A1 · Guthrie · 2010 [cited by examiner]
US 20120260066A1 · Henry et al. · 2012 [cited by applicant]
US 20130024647A1 · Gove · 2013 [cited by examiner]
US 20160283391A1 · Nilsson · 2016 [cited by examiner]
US 20170278215A1 · Appu · 2017 [cited by examiner]
US 20170286110A1 · Agron · 2017 [cited by examiner]
US 20180095753A1 · Bai et al. · 2018 [cited by applicant]
US 20190042453A1 · Basak · 2019 [cited by examiner]
US 20200019406A1 · Kalamatianos et al. · 2020 [cited by applicant]
EP 0098494A2 · 1984 [cited by applicant]
EP 0178671A2 · 1986 [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2008/08802, dated Oct. 6, 2008, 14 pages. [cited by applicant]
Rice et al. “A Formal Model for SIMD Computation”, Proceedings., 2nd Symposium on the Frontiers of Massively Parallel Computation, Oct. 1988, pp. 601-607. [cited by applicant]