IP Library › Granted Patent US 12,625,704
Granted Patent B2
US 12,625,704 · App. 18/648,835 · Granted May 12, 2026

Low power late-selected caches using a set-prediction history

Inventors: David A. Hrusecky (Cedar Park, TX); Wolfgang Penth (Holzgerlingen, DE)
Assignee: International Business Machines Corporation
G06F9/30047G06F9/3806G06F11/1407
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,704
App. No.
18/648,835
Granted
May 12, 2026
Kind
B2
Abstract

A method, computer program product, and computer system for reading data stored in a set associative cache. A cache read instruction that did not read the cache after being previously launched is relaunched after an effective address (EA) of the instruction was ascertained. A hash of the ascertained EA (EAHash) and a class congruence class (CCC) is determined from the ascertained EA. A search is performed for a match of the EAHash and CCC of the ascertained EA to the EAHash and CCC, respectively, of an instruction whose EAHash, CCC, and set are stored in an instruction history stream. If the match is found, only read enables associated with the stored set of the match, which is a read enable of only one class of one address group in the cache, are activated. If the match is not found, all read enables of the one address group are activated.

Claims (55)

1 . A computer-implemented method for reading data stored in a set associative cache, said method comprising:

relaunching, by one or more processors of a computer system, a cache read instruction that did not read the set associative cache after being previously launched, wherein an effective address (EA) of the instruction was ascertained prior to said relaunching and after the instruction was previously launched, wherein the set associative cache comprises G address groups encompassing all of the cache's stored data, wherein each address group comprises S sets and C classes, wherein S mod C=0 and each class comprises S/C sets, wherein each class has a read enable, and wherein G is at least 1, S is at least 2, and C is at least 1;

determining, by the one or more processors, a hash of the ascertained EA (EAHash) and a cache congruence class (CCC) from the ascertained EA; and

searching, by the one or more processors, for a match of the EAHash and a CCC of the ascertained EA to an EAHash and a CCC, respectively, of an instruction whose EAHash, CCC, and set are stored in an instruction history stream,

wherein if the match is found from said searching, then the stored set of the match is referred to as an inferred set which is indicative of a read enable of only one class of one address group, and the one or more processors activate only read enables associated with the inferred set which indicates the read enable of the only one class of the one address group in the cache; and

wherein if the match is not found from said searching, then activating, by the one or more processors, all read enables of the one address group in the cache.

2 . The method of claim 1 , wherein C is at least 2, wherein the match is found from the search, and wherein said activating only read enables associated with the inferred set comprises:

inferring a class associated with the inferred set;

address decoding the relaunched instruction to determine the G address groups;

logically combining the inferred class and the classes of the determined G address groups to identify the one class of the one address group in the cache; and

activating the read enable of only the one class of the one address group in the cache.

3 . The method of claim 2 , wherein said logically combining is implemented via use of multiple AND gates comprising one AND gate for each class of each address group.

4 . The method of claim 1 , wherein each class of each address group includes at least one SRAM having global bit lines, wherein each class of each address group includes at least 2 subclasses, and wherein the method further comprises prior to execution of a read of the cache at the inferred set:

identifying, by the one or more processors, a subclass of the one class of the one address group such that the subclass includes the inferred set; and

precharging, by the one or more processors, only global bit lines associated with the identified subclass.

5 . The method of claim 4 , wherein C=2, wherein the C classes consist of a class of even numbered sets and a class of odd numbered sets, wherein the inferred set consists of a least significant bit and remaining upper bits, wherein said inferring the class of the inferred set utilizes the least significant bit, and wherein said identifying the subclass of the one class of the one address group utilizes the remaining upper bits.

6 . The method of claim 1 , wherein the instruction history stream is a dynamically changing data buffer of constant depth K that stores an array of data for each processed instruction of K previously processed instructions, wherein the arrays of data are sequentially ordered in the buffer according to a latest time of entry into the buffer of the processed instructions such that each new processed instruction entering the buffer results in the instruction having the earliest time of entry into the buffer being dropped out of the buffer, wherein each stored array includes an EAHash, CCC, and predicted set (SETP) determined set of a respective processed instruction that entered the buffer, and wherein K is at least 2.

7 . The method of claim 6 , wherein K is in a range of 3 to 5.

8 . The method of claim 6 , wherein each stored array further includes a valid bit (V) selected from the group consisting of 1 or 0 denoting that the processed instruction is valid or invalid, respectively, and wherein the valid bit is set to 1 for each processed instruction entering the buffer.

9 . The method of claim 8 , wherein in response to a determination that a CCC and a SETP determined set of a cache write instruction respectively matches a CCC and a SETP determined set in one stored array in the buffer and that the valid bit of the one stored array is 1, setting the valid bit to 0 for the one stored array.

10 . The method of claim 6 , wherein an invalid SETP determined set in an array of one instruction in the instruction history stream is indicative of a SETPmiss due to the one instruction having attempted to read non-existent data from a cache line in the cache at the invalid SETP determined set.

11 . The method of claim 6 , wherein said searching results in a multihit SETP determined set match due to a hit on two different set values.

12 . The method of claim 1 , wherein said relaunching comprises relaunching the cache read instruction from a load launch queue that includes the ascertained EA.

13 . The method of claim 1 , wherein the match is not found from the search.

14 . The method of claim 1 , wherein the match is found from the search, and wherein the method further comprises:

obtaining, by the one or more processors, an actual set of the relaunched instruction from a predicted set (SETP) array; and

determining, by the one or more processors, that the inferred set is not equal to the actual set so that data read from the cache is incorrect and cannot be used and in response, performing, by the one or more processors, an auto-correct process that mitigates incorrect data having been read from the cache.

15 . A computer program product, comprising one or more computer readable hardware storage devices having computer readable program code stored therein, said program code containing instructions executable by one or more processors of a computer system to implement a computer-implemented method for reading data stored in a set associative cache, said method comprising:

relaunching, by the one or more processors, a cache read instruction that did not read the set associative cache after being previously launched, wherein an effective address (EA) of the instruction was ascertained prior to said relaunching and after the instruction was previously launched, wherein the set associative cache comprises G address groups encompassing all of the cache's stored data, wherein each address group comprises S sets and C classes, wherein S mod C=0 and each class comprises S/C sets, wherein each class has a read enable, and wherein G is at least 1, S is at least 2, and C is at least 1;

determining, by the one or more processors, a hash of the ascertained EA (EAHash) and a cache congruence class (CCC) from the ascertained EA;

searching, by the one or more processors, for a match of the EAHash and a CCC of the ascertained EA to an EAHash and a CCC, respectively, of an instruction whose EAHash, CCC, and set are stored in an instruction history stream,

wherein if the match is found from said searching, then the stored set of the match is referred to as an inferred set which is indicative of a read enable of only one class of one address group, and the one or more processors activate only read enables associated with the inferred set which indicates the read enable of the only one class of the one address group in the cache; and

wherein if the match is not found from said searching, then activating, by the one or more processors, all read enables of the one address group in the cache.

16 . The computer program product of claim 15 , wherein C is at least 2, and wherein the match is found from the search, and wherein said activating only read enables associated with the inferred set comprises:

inferring a class associated with the inferred set;

address decoding the relaunched instruction to determine the G address groups;

logically combining the inferred class and the classes of the determined G address groups to identify the one class of the one address group in the cache; and

activating the read enable of only the one class of the one address group in the cache.

17 . The computer program product of claim 15 , wherein each class of each address group includes at least one SRAM having global bit lines, wherein each class of each address group includes at least 2 subclasses, and wherein the method further comprises prior to execution of a read of the cache at the inferred set:

identifying, by the one or more processors, a subclass of the one class of the one address group such that the subclass includes the inferred set; and

precharging, by the one or more processors, only global bit lines associated with the identified subclass.

18 . A computer system, comprising one or more processors, one or more memories, and one or more computer readable hardware storage devices, said one or more hardware storage devices containing program code executable by the one or more processors via the one or more memories to implement a computer-implemented method for reading data stored in a set associative cache, said method comprising:

relaunching, by the one or more processors, a cache read instruction that did not read the set associative cache after being previously launched, wherein an effective address (EA) of the instruction was ascertained prior to said relaunching and after the instruction was previously launched, wherein the set associative cache comprises G address groups encompassing all of the cache's stored data, wherein each address group comprises S sets and C classes, wherein S mod C=0 and each class comprises S/C sets, wherein each class has a read enable, and wherein G is at least 1, S is at least 2, and C is at least 1;

determining, by the one or more processors, a hash of the ascertained EA (EAHash) and a cache congruence class (CCC) from the ascertained EA;

searching, by the one or more processors, for a match of the EAHash and a CCC of the ascertained EA to an EAHash and a CCC, respectively, of an instruction whose EAHash, CCC, and set are stored in an instruction history stream,

wherein if the match is found from said searching, then the stored set of the match is referred to as an inferred set which is indicative of a read enable of only one class of one address group, and the one or more processors activate only read enables associated with the inferred set which indicates the read enable of the only one class of the one address group in the cache; and

wherein if the match is not found from said searching, then activating, by the one or more processors, all read enables of the one address group in the cache.

19 . The computer system of claim 18 , wherein C is at least 2, and wherein the match is found from the search, and wherein said activating only read enables associated with the inferred set comprises:

inferring a class associated with the inferred set;

address decoding the relaunched instruction to determine the G address groups;

logically combining the inferred class and the classes of the determined G address groups to identify the one class of the one address group in the cache; and

activating the read enable of only the one class of the one address group in the cache.

20 . The computer system of claim 18 , wherein each class of each address group includes at least one SRAM having global bit lines, wherein each class of each address group includes at least 2 subclasses, and wherein the method further comprises prior to execution of a read of the cache at the inferred set:

identifying, by the one or more processors, a subclass of the one class of the one address group such that the subclass includes the inferred set; and

precharging, by the one or more processors, only global bit lines associated with the identified subclass.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2024
From: HRUSECKY, DAVID A.; PENTH, WOLFGANG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 067253/0213 →
Continuity (1)
Related Publication 20250335198A1 · Oct 30, 2025
References Cited (20)
US 5418922A · Liu · 1995 [cited by examiner]
US 6356990B1 · Aoki · 2002 [cited by applicant]
US 6418525B1 · Charney · 2002 [cited by examiner]
US 7475192B2 · Correale, Jr. · 2009 [cited by applicant]
US 10042770B2 · Chadha · 2018 [cited by applicant]
US 10157137B1 · Jain · 2018 [cited by examiner]
US 11281586B2 · Liu · 2022 [cited by applicant]
US 20090094435A1 · Lu · 2009 [cited by examiner]
US 20130262777A1 · Ghai · 2013 [cited by examiner]
US 20140181407A1 · Crum · 2014 [cited by examiner]
US 20150234664A1 · Kim · 2015 [cited by examiner]
US 20170286119A1 · Al Sheikh · 2017 [cited by examiner]
US 20210240631A1 · Joo · 2021 [cited by examiner]
US 20230063976A1 · Fernsler · 2023 [cited by examiner]
IP.com No. IPCOM000196384D, Management of Dynamically Resizable Data Processing Units based on Application State Predictive Estimation, IP.com Electronic Publication Date: Jun. 2, 2010, 6 pages. [cited by applicant]
IP.com No. IPCOM000216961D, System and Method for Recovering Global Branch Prediction Information Using Address Offset Information, IP.com Electronic Publication Date: Apr. 25, 2012, 5 pages. [cited by applicant]
IP.com No. IPCOM000263479D, Value Prediction Implementation, IP.com Electronic Publication Date: Sep. 3, 2020, 5 pages. [cited by applicant]
Jalili, M. et al., Reducing Load Latency with Cache Level Prediction, arXiv:2103.14808v1 [cs.AR] Mar. 27, 2021, 12 pages. [cited by applicant]
Nicolaescu, D. et al., Reducing Data Cache Energy Consumption via Cached Load/Store Queue, ISLPED'03, Aug. 25-27, 2003, Seoul, Korea; Copyright 2003 ACM 1-58113-682-X/03/0008, 6 pages. [cited by applicant]
Wang, L. et al., Way Prediction Set-Associative Data Cache for Low Power Digital Signal Processors, he Key Laboratory of Information Technology for Autonomous Underwater Vehicles, Chinese Academy of Sciences Institute o… [cited by applicant]