IP Library › Granted Patent US 12,499,048
Granted Patent B2
US 12,499,048 · App. 18/527,293 · Granted Dec 16, 2025

Prefetcher engine configuration selection with multi-armed bandit

Inventors: Shie Mannor (Haifa, IL); Ariel Szapiro (Kfar Netter, IL); Gil Levy (Hod Hasharon, IL); Arye Albahari (Kiryat Motzkin, IL); Gaby Diengott (Kadima-Tzoran, IL); Elad Alon (Tel Aviv, IL); Sagi Lahav (Kiryat Bialik, IL); Amir Rosen (Haifa, IL)
Assignee: Mellanox Technologies, Ltd
G06F12/0862G06F12/0837
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,048
App. No.
18/527,293
Granted
Dec 16, 2025
Kind
B2
Abstract

In one embodiment, a system includes prefetcher engines to predict next memory access addresses of a memory from which to load data to a cache during execution of a software application, and load the data from the predicted next memory access addresses to the cache during execution of the software application, and a processor to control the prefetcher engines according to configurations of the prefetcher engines selected by a machine learning agent in exploration phases and in exploitation phases during execution of the software application, and execute the machine learning agent to select from a pruned set of configurations to control the prefetcher engines in the exploration phases, perform measurements on the system during execution of the machine learning agent, and execute the machine learning agent to select from the configurations to maximize potential rewards from controlling the prefetcher engines in the exploitation phases based on the performed measurements.

Claims (38)

1 . A system, comprising:

a plurality of prefetcher engines to:

predict next memory access addresses of a memory from which to load data to a cache during execution of a software application; and

load the data from the predicted next memory access addresses to the cache during execution of the software application; and

a processor to:

reduce an original set of configurations of the prefetcher engines to yield a reduced set of configurations of the prefetcher engines using a first configuration reduction method including defining groups of the prefetcher engines such that for each one of the defined groups a respective setting is to be applied equally to the prefetcher engines of the one group;

reduce the reduced set of configurations of the prefetcher engines to yield a pruned set of configurations using a second configuration reduction method including defining an order of at least some of the configurations;

control the prefetcher engines according to configurations of the prefetcher engines selected by a machine learning agent in exploration phases and in exploitation phases during execution of the software application; and

execute the machine learning agent to select from the pruned set of configurations to control the prefetcher engines in the exploration phases;

perform measurements on the system during execution of the machine learning agent; and

execute the machine learning agent to select from the pruned set of configurations to maximize potential rewards from controlling the prefetcher engines in the exploitation phases based on the performed measurements.

2 . The system according to claim 1 , wherein the processor is to execute the machine learning agent to select from the pruned set of configurations to maximize potential rewards from controlling the prefetcher engines in the exploitation phases and minimize potential losses of reward from controlling the prefetcher engines in the exploration phases, based on the performed measurements.

3 . The system according to claim 1 , wherein the machine learning agent is a multi-armed bandit machine learning agent.

4 . The system according to claim 1 , wherein each of the prefetcher engines is to selectively provide different levels of aggressiveness, the original set of configurations of the prefetcher engines providing different aggressiveness configurations of the prefetcher engines.

5 . The system according to claim 1 , wherein the performed measurements include executed instructions per cycle.

6 . The system according to claim 1 , wherein the processor is to compute a potential reward of selecting a given configuration of the prefetcher engines based on an average of previous performance scores from previously controlling the prefetcher engines according to the given configuration.

7 . The system according to claim 6 , wherein the processor is to compute the potential reward of selecting the given configuration of the prefetcher engines based on an aged average of the previous performance scores.

8 . The system according to claim 1 , wherein the processor is to compute a potential reward of selecting a given configuration of the prefetcher engines based on maximizing executed instructions per cycle.

9 . The system according to claim 8 , wherein the processor is to compute the potential reward of selecting the given configuration of the prefetcher engines based on minimizing memory transactions per cycle.

10 . The system according to claim 1 , wherein the processor is to compute a potential reward of selecting a given configuration of the prefetcher engines based on previous performance scores from previously controlling the prefetcher engines according to the given configuration, the previous performance scores being based on any two or more of the following: executed instructions per cycle; memory transactions per cycle; power cost per memory transaction; average core frequency; average core power; power budget; and measured temperature.

11 . A method, comprising:

predicting next memory access addresses of a memory from which to load data to a cache during execution of a software application;

loading the data from the predicted next memory access addresses to the cache during execution of the software application;

reducing an original set of configurations of prefetcher engines to yield a reduced set of configurations of the prefetcher engines using a first configuration reduction method including defining groups of the prefetcher engines such that for each one of the defined groups a respective setting is to be applied equally to the prefetcher engines of the one group;

reducing the reduced set of configurations of the prefetcher engines to yield a pruned set of configurations using a second configuration reduction method including defining an order of at least some of the configurations;

controlling the prefetcher engines according to configurations of the prefetcher engines selected by a machine learning agent in exploration phases and in exploitation phases during execution of the software application;

executing the machine learning agent to select from the pruned set of configurations to control the prefetcher engines in the exploration phases;

performing measurements during execution of the machine learning agent; and

executing the machine learning agent to select from the pruned set of configurations to maximize potential rewards from controlling the prefetcher engines in the exploitation phases based on the performed measurements.

12 . The method according to claim 11 , wherein the executing the machine learning agent processor includes executing the machine learning agent to select from the pruned set of configurations to maximize potential rewards from controlling the prefetcher engines in the exploitation phases and minimize potential losses of reward from controlling the prefetcher engines in the exploration phases, based on the performed measurements.

13 . The method according to claim 11 , wherein the machine learning agent is a multi-armed bandit machine learning agent.

14 . The method according to claim 11 , wherein each of the prefetcher engines is to selectively provide different levels of aggressiveness, the original set of configurations of the prefetcher engines providing different aggressiveness configurations of the prefetcher engines.

15 . The method according to claim 11 , wherein the performed measurements include executed instructions per cycle.

16 . The method according to claim 11 , further comprising computing a potential reward of selecting a given configuration of the prefetcher engines based on an average of previous performance scores from previously controlling the prefetcher engines according to the given configuration.

17 . The method according to claim 16 , wherein the computing includes computing the potential reward of selecting the given configuration of the prefetcher engines based on an aged average of the previous performance scores.

18 . The method according to claim 11 , further comprising computing a potential reward of selecting a given configuration of the prefetcher engines based on maximizing executed instructions per cycle.

19 . The method according to claim 18 , wherein the computing includes computing the potential reward of selecting the given configuration of the prefetcher engines based on minimizing memory transactions per cycle.

20 . The method according to claim 11 , further comprising computing a potential reward of selecting a given configuration of the prefetcher engines based on previous performance scores from previously controlling the prefetcher engines according to the given configuration, the previous performance scores being based on any two or more of the following: executed instructions per cycle; memory transactions per cycle; power cost per memory transaction; average core frequency; average core power; power budget; and measured temperature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2023
From: MANNOR, SHIE; SZAPIRO, ARIEL; LEVY, GIL; ALBAHARI, ARYE; DIENGOTT, GABY; ALON, ELAD; LAHAV, SAGI; ROSEN, AMIR
To: MELLANOX TECHNOLOGIES, LTD.
Reel/Frame 065759/0531 →
Continuity (1)
Related Publication 20250181509A1 · Jun 5, 2025
References Cited (28)
US 11804050B1 · Milletari et al. · 2023 [cited by applicant]
US 12222875B1 · Huberty et al. · 2025 [cited by applicant]
US 20140304477A1 · Hughes et al. · 2014 [cited by applicant]
US 20160283970A1 · Ghavamzadeh et al. · 2016 [cited by applicant]
US 20200117608A1 · Thompto et al. · 2020 [cited by applicant]
US 20210089472A1 · Ishii et al. · 2021 [cited by applicant]
US 20210374523A1 · Gottin et al. · 2021 [cited by applicant]
US 20220197856A1 · Khasawneh · 2022 [cited by examiner]
US 20220374367A1 · Fang et al. · 2022 [cited by applicant]
US 20230236977A1 · Dev et al. · 2023 [cited by applicant]
US 20250139439A1 · Abts · 2025 [cited by applicant]
US 20250238376A1 · Castorina · 2025 [cited by applicant]
CN 113986774A · 2022 [cited by applicant]
EP 3486785A1 · 2019 [cited by applicant]
WO 20171766442A1 · 2017 [cited by applicant]
WO 2017189033A1 · 2017 [cited by applicant]
WO 2023088535A1 · 2023 [cited by applicant]
Rahman et al., “Maximizing Hardware Prefetch Effectiveness with Machine Learning,” Proceedings of the ACM/IEEE Conference on High Performance Computing and Communications, pp. 1-7, year 2015. [cited by applicant]
Liao et al., “Machine Learning-Based Prefetch Optimization for Data Center Applications,” Conference Paper, SC 09, pp. 1-11, Nov. 2009. [cited by applicant]
Eris et al., “Puppeteer: A Random Forest Based Manager for Hardware Prefetchers Across the Memory Hierarchy,” ACM Transactions on Architecture and Code Optimization, vol. 20, No. 1, Article 19, pp. 1-25, Dec. 2022. [cited by applicant]
Wikipedia, “Greedy Algorithm,” pp. 1-6, Aug. 14, 2023. [cited by applicant]
Gerogiannis et al., “Micro-Armed Bandit: Lightweight & Reusable Reinforcement Learning for Microarchitecture Decision-Making,” Conference Paper, Micro '23, pp. 1-16, Nov. 2023. [cited by applicant]
Szapiro et al., U.S. Appl. No. 18/527,294, filed Dec. 3, 2023. [cited by applicant]
Rosen et al., U.S. Appl. No. 18/527,296, filed Dec. 3, 2023. [cited by applicant]
Rosen et al., U.S. Appl. No. 18/527,295, filed Dec. 3, 2023. [cited by applicant]
Rosen et al., U.S. Appl. No. 18/527,297, filed Dec. 3, 2023. [cited by applicant]
U.S. Non Final Office Action U.S. Appl. No. 18/623,099, dated Apr. 4, 2025. [cited by applicant]
U.S. Non Final Office Action U.S. Appl. No. 18/623,103, dated May 13, 2025. [cited by applicant]