IP Library › Granted Patent US 12,737,232
Granted Patent B2
US 12,737,232 · App. 18/527,297 · Granted Sep 15, 2026

Multi-armed bandit improvement

Inventors: Amir Rosen (Haifa, IL); Shie Mannor (Haifa, IL); Gil Levy (Hod Hasharon, IL); Arye Albahari (Kiryat Motzkin, IL); Ariel Szapiro (Kfar Netter, IL)
Assignee: Mellanox Technologies, Ltd
G06F9/505G06F9/30047G06F12/0862
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,232
App. No.
18/527,297
Granted
Sep 15, 2026
Kind
B2
Abstract

In one embodiment, a system includes a processor to control a resource according to policies selected by a multi-armed bandit machine learning agent in exploration phases and in exploitation phases, and execute the multi-armed bandit machine learning agent to select from the policies to control the resource in the exploration phases according to probabilities to explore corresponding one of the policies, wherein the probabilities include different probabilities, perform measurements on the system during execution of the multi-armed bandit machine learning agent, and execute the multi-armed bandit machine learning agent to select from the policies to maximize potential rewards from controlling the resource in exploitation phases based on the performed measurements, and a memory to store data used by the processor.

Claims (49)

1 . A system, comprising:

a processor to:

control a resource according to policies selected by a multi-armed bandit machine learning agent in exploration phases and in exploitation phases; and

execute the multi-armed bandit machine learning agent to select from the policies to control the resource in the exploration phases according to probabilities to explore corresponding one of the policies, wherein the probabilities include different probabilities;

perform measurements on the system during execution of the multi-armed bandit machine learning agent; and

execute the multi-armed bandit machine learning agent to select from the policies to maximize potential rewards from controlling the resource in exploitation phases based on the performed measurements; and

a memory to store data used by the processor.

2 . The system according to claim 1 , wherein:

the resource includes prefetcher engines to predict next memory access addresses of the memory, each of the prefetcher engines being configurable to provide a level of aggressiveness;

the policies correspond to different configurations of the prefetcher engines; and

the measurements include executed instructions per cycle.

3 . The system according to claim 1 , further comprising processing circuitry to compute the probabilities to explore with the corresponding policies based on prior knowledge gained from controlling the resource according to the policies while executing benchmark applications.

4 . The system according to claim 3 , wherein the processing circuitry is comprised in the processor.

5 . The system according to claim 3 , wherein the processing circuitry is to compute the probability to explore with a given policy of the policies based on prior knowledge of an extent of success and failure of the given policy while executing the benchmark applications.

6 . The system according to claim 5 , wherein the processing circuitry is to compute the probability to explore with the given policy such that the probability to explore with the given policy optimizes a loss function representing a system parameter to be optimized.

7 . The system according to claim 6 , wherein the loss function includes a parameter which compares: (a) a first value of a quality metric when the given policy is applied during execution of a given benchmark application; with (b) a second value of the quality metric of the best policy, which is one of the policies providing a highest value of the quality metric among the policies applied during execution of the given benchmark application.

8 . The system according to claim 7 , wherein the loss function includes any one or more of the following parameters:

an empiric parameter which measures the average or distribution of a number of times ones of the policies have to be explored in order to identify their relative quality compared with other ones of the policies for the given benchmark application;

a weight of the given benchmark application with respect to other ones of the benchmark applications; and

a number of checkpoints in the given benchmark application.

9 . The system according to claim 1 , wherein the processing circuitry is to compute the different probabilities to explore with the corresponding policies such that the probabilities to explore optimize a loss function representing a system parameter to be optimized.

10 . The system according to claim 9 , wherein the loss function includes:

a first term that represents the loss of benefit due to the time taken during exploration for the multi-armed bandit machine learning agent to find ones of the policies that are better than other ones of the policies; and

a second term that represents the loss of benefit due to exploration with sub-optimal policies.

11 . The system according to claim 10 , wherein the processing circuitry is to compute the different probabilities to explore with the corresponding policies based on an expression formed by comparing a derivative of the loss function to zero.

12 . A method, comprising:

controlling a resource according to policies selected by a multi-armed bandit machine learning agent in exploration phases and in exploitation phases; and

executing the multi-armed bandit machine learning agent to select from the policies to control the resource in the exploration phases according to probabilities to explore corresponding one of the policies, wherein the probabilities include different probabilities;

performing measurements on the system during execution of the multi-armed bandit machine learning agent; and

executing the multi-armed bandit machine learning agent to select from the policies to maximize potential rewards from controlling the resource in exploitation phases based on the performed measurements; and

storing data used by the processor.

13 . The method according to claim 12 , wherein:

the resource includes prefetcher engines;

each of the prefetcher engines is configurable to provide a level of aggressiveness;

the policies correspond to different configurations of the prefetcher engines; and

the measurements include executed instructions per cycle.

14 . The method according to claim 12 , further comprising computing the probabilities to explore with the corresponding policies based on prior knowledge gained from controlling the resource according to the policies while executing benchmark applications.

15 . The method according to claim 14 , wherein the computing includes computing the probability to explore with a given policy of the policies based on prior knowledge of an extent of success and failure of the given policy while executing the benchmark applications.

16 . The method according to claim 15 , wherein the computing includes computing the probability to explore with the given policy such that the probability to explore with the given policy optimizes a loss function representing a system parameter to be optimized.

17 . The method according to claim 16 , wherein the loss function includes a parameter which compares: (a) a first value of a quality metric when the given policy is applied during execution of a given benchmark application; with (b) a second value of the quality metric of the best policy, which is one of the policies providing a highest value of the quality metric among the policies applied during execution of the given benchmark application.

18 . The method according to claim 17 , wherein the loss function includes any one or more of the following parameters:

an empiric parameter which measures the average or distribution of a number of times ones of the policies have to be explored in order to identify their relative quality compared with other ones of the policies for the given benchmark application;

a weight of the given benchmark application with respect to other ones of the benchmark applications; and

a number of checkpoints in the given benchmark application.

19 . The method according to claim 12 , further comprising computing the different probabilities to explore with the corresponding policies such that the probabilities to explore optimize a loss function representing a system parameter to be optimized.

20 . The method according to claim 19 , wherein the loss function includes:

a first term that represents the loss of benefit due to the time taken during exploration for the multi-armed bandit machine learning agent to find ones of the policies that are better than other ones of the policies; and

a second term that represents the loss of benefit due to exploration with sub-optimal policies.

21 . The method according to claim 20 , wherein the computing includes computing the different probabilities to explore with the corresponding policies based on an expression formed by comparing a derivative of the loss function to zero.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2023
From: ROSEN, AMIR; MANNOR, SHIE; LEVY, GIL; ALBAHARI, ARYE; SZAPIRO, ARIEL
To: MELLANOX TECHNOLOGIES, LTD.
Reel/Frame 065759/0699 →
Continuity (1)
Related Publication 20250181411A1 · Jun 5, 2025
References Cited (24)
US 11804050B1 · Milletari et al. · 2023 [cited by applicant]
US 12222875B1 · Huberty et al. · 2025 [cited by applicant]
US 20140304477A1 · Hughes et al. · 2014 [cited by applicant]
US 20160283970A1 · Ghavamzadeh et al. · 2016 [cited by applicant]
US 20210089472A1 · Ishii et al. · 2021 [cited by applicant]
US 20210374523A1 · Gottin et al. · 2021 [cited by applicant]
US 20220197856A1 · Khasawneh et al. · 2022 [cited by applicant]
US 20220374367A1 · Fang et al. · 2022 [cited by applicant]
US 20230236977A1 · Dev et al. · 2023 [cited by applicant]
US 20250139439A1 · Abts · 2025 [cited by applicant]
WO 2017189033A1 · 2017 [cited by applicant]
WO 2023088535A1 · 2023 [cited by applicant]
Rahman et al., “Maximizing Hardware Prefetch Effectiveness with Machine Learning,” Proceedings of the ACM/IEEE Conference on High Performance Computing and Communications, pp. 1-7, year 2015. [cited by applicant]
Liao et al., “Machine Learning-Based Prefetch Optimization for Data Center Applications,” Conference Paper, SC 09, pp. 1-11, Nov. 2009. [cited by applicant]
Eris et al., “Puppeteer: A Random Forest Based Manager for Hardware Prefetchers Across the Memory Hierarchy,” ACM Transactions on Architecture and Code Optimization, vol. 20, No. 1, Article 19, pp. 1-25, Dec. 2022. [cited by applicant]
Wikipedia, “Greedy Algorithm,” pp. 1-6, Aug. 14, 2023. [cited by applicant]
Gerogiannis et al., “Micro-Armed Bandit: Lightweight & Reusable Reinforcement Learning for Microarchitecture Decision-Making,” Conference Paper, MICRO '23, pp. 1-16, Nov. 2023. [cited by applicant]
Mannor et al., U.S. Appl. No. 18/527,293, filed Dec. 3, 2023. [cited by applicant]
Szapiro et al., U.S. Appl. No. 18/527,294, filed Dec. 3, 2023. [cited by applicant]
Rosen et al., U.S. Appl. No. 18/527,296, filed Dec. 3, 2023. [cited by applicant]
Rosen et al., U.S. Appl. No. 18/527,295, filed Dec. 3, 2023. [cited by applicant]
US Non Final Office Action U.S. Appl. No. 18/527,293, dated Apr. 10, 2025. [cited by applicant]
US Non Final Office Action U.S. Appl. No. 18/623,103, dated May 13, 2025. [cited by applicant]
US Non Final Office Action U.S. Appl. No. 18/623,099, dated Apr. 4, 2025. [cited by applicant]