IP Library Granted Patent US 12,608,233
Granted Patent B2
US 12,608,233 · App. 18/016,982 · Granted Apr 21, 2026

System and method for recommending computing resources

Inventors: Gabriel Martin (Quintana de la Sirena, ES); Max Alt (San Francisco, CA)
Assignee: Advanced Micro Devices, Inc.
G06F9/5038G06F9/5072G06F2209/5019
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,233
App. No.
18/016,982
Granted
Apr 21, 2026
Kind
B2
Abstract

A system and method for recommending computing resources for processing jobs in a distributed computing environment with multiple heterogeneous computing resources are disclosed. Training applications or jobs are executed and measured on different computing resources and on different configurations of the computing resources to establish a database of performance metrics. A matrix of application features and computing resource features is created and populated with performance data. Machine learning may be used to create and update multiple recommendation engines based on the matrix that are cross-validated and merged to form final performance estimators. The performance estimators are applied to new applications and determine which existing applications are most similar and which resources to recommend.

Claims (48)

1 . A method for recommending computing resources in a distributed computing system, the method comprising:

(a) gathering performance data by running a plurality of training computing jobs on a plurality of different computing devices in the distributed computing system;

(b) storing the performance data in a database;

(c) performing feature extraction on the gathered performance data to identify job features of the training jobs and system features of the computing devices;

(d) creating a matrix of the job features and the system features provided with the extracted performance data;

(e) generate performance estimates using a plurality of recommendation algorithms;

(f) cross validating the performance estimates to identify best estimates; and

(g) recommend computing resources for a next computing job using the best estimates.

2 . The method of claim 1 , further comprising:

executing a next computing job on the recommended computing resources; and

repeating steps (a)-(f) using new performance data generated from the execution of the next computing job.

3 . The method of claim 1 , further comprising aggregating the performance data in the database by job, app, queue, and app-queue.

4 . The method of claim 1 , wherein the performance data comprises instruction per second metrics.

5 . The method of claim 1 , wherein one of the features is cloud versus bare metal.

6 . The method of claim 1 , further comprising using similar queues to provide data for queues where insufficient performance data exists in the database to perform steps (c)-(f).

7 . The method of claim 1 , wherein at least one of the plurality of recommendation algorithms is K-nearest neighbors (KNN).

8 . The method of claim 1 , wherein the plurality of recommendation algorithms comprises content-based algorithms and collaborative filtering using different feature sets and different performance metrics.

9 . The method of claim 1 , wherein the performance data comprises job parameters comprising measures of CPU boundness, GPU boundness, MPI boundness, memory boundness, and network boundness.

10 . The method of claim 9 , wherein the performance data further comprises latency and bandwidth.

11 . The method of claim 1 , wherein the job features comprise libraries used, datasets used, application domain, and application metadata.

12 . The method of claim 1 , wherein the system features comprise a number of cores, a core type, an available memory amount, and a maximum core clock speed.

13 . The method of claim 1 , wherein the system features comprise storage type and network type.

14 . The method of claim 1 , wherein the system features comprise IO throughput.

15 . The method of claim 1 , wherein the system features comprise memory bandwidth and latency.

16 . The method of claim 1 , wherein the system features comprise benchmark scores.

17 . A system having at least one processor, the system configured to:

gather performance data by running a plurality of training computing jobs on a plurality of different computing devices in a distributed computing system;

store the performance data in a database;

perform feature extraction on the gathered performance data to identify job features of the training jobs and system features of the computing devices;

create a matrix of the job features and the system features provided with the extracted performance data;

generate performance estimates based on a plurality of recommendation algorithms;

cross-validate the performance estimates to identify best estimates; and

recommend computing resources for a next computing job using the best estimates.

18 . A system having at least one processor, the system configured to:

analyze gathered performance data to create aggregated performance values, wherein the gathered performance data is obtained by running a plurality of test jobs on a plurality of different subsets of computing resources;

perform feature extraction on the aggregated performance values to identify job features of the test jobs and system features of the computing resources;

populate a matrix with the job features and the system features based on the aggregated performance values;

create a recommender using the matrix;

characterize a new job to determine which of the test jobs is most similar based on resource utilization; and

select via the recommender an optimal subset of the computing resources to execute the new job based on the aggregated performance values for the most similar test job.

19 . The system of claim 18 , further configured to:

execute the new job;

gather new performance data for the new job;

create new aggregated performance values based on the gathered new performance data;

update the matrix with the new aggregated performance values; and

update the recommender using the updated matrix.

20 . The system of claim 18 , further configured to select via the recommender a subset of system resources that comprises more of a specific type of computing resource in response to the new job showing intensive utilization of the specific type of computing resource.

21 . The system of claim 18 , wherein if the new job's performance on a first computing system characterizes similarly to one of the plurality of test job's performance on the first computing system, then the recommender uses the test job's performance on a second computing system to estimate a performance of the new job on the second computing system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2023
From: CORE SCIENTIFIC, INC.; CORE SCIENTIFIC OPERATING COMPANY
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 063170/0088 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: HERNANDEZ, GABRIEL MARTIN; ALT, MAX
To: CORE SCIENTIFIC, INC
Reel/Frame 063132/0265 →
Continuity (2)
Provisional Application 63054458 · Jul 21, 2020
Related Publication 20230281051A1 · Sep 7, 2023
References Cited (12)
US 10841236B1 · Jin · 2020 [cited by examiner]
US 20180113742A1 · Chung · 2018 [cited by examiner]
US 20180365576A1 · Guttmann · 2018 [cited by applicant]
US 20190123973A1 · Jeuk · 2019 [cited by examiner]
US 20190171483A1 · Santhar · 2019 [cited by applicant]
US 20190208009A1 · Prabhakaran · 2019 [cited by applicant]
US 20210124614A1 · Gupta · 2021 [cited by examiner]
CN 104077212A · 2014 [cited by examiner]
CN 105975345B · 2019 [cited by examiner]
Taghavi et al., Compute Job Memory Recommnder System Using Machine Learning, ACM, Aug. 2016, 8 pages. [cited by examiner]
Prekopesak et al., Cross-validation: the illusion of reliable performance estimation, Budapest University of Technology and Economics, May 2012, 6 pages. (Year: 2012). [cited by examiner]
International Search Report and Written Opinion dated Oct. 22, 2021 for International Patent Application No. PCT/US2021/042439. [cited by applicant]