IP Library › Granted Patent US 12,737,278
Granted Patent B2
US 12,737,278 · App. 17/850,847 · Granted Sep 15, 2026

Detecting and optimizing program workload inefficiencies at runtime

Inventors: Szymon Migacz (Santa Clara, CA); Pawel Morkisz (San Jose, CA); Alex Fit-Florea (Los Altos Hills, CA); Maciej Bala (Warsaw, PL); Jakub Zakrzewski (Warsaw, PL); Trivikram Krishnamurthy (Los Altos, CA); Nitin Nitin (Emeryville, CA); Sangkug Lym (San Jose, CA); Shang Wang (Oakville, CA); Chenhan Yu (Mountain House, CA); Alexandre Milesi (Strasbourg, FR)
Assignee: NVIDIA Corporation
G06F11/3612
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,278
App. No.
17/850,847
Granted
Sep 15, 2026
Kind
B2
Abstract

Methods and systems for comparing information obtained during execution of a workload to a set of inefficiency patterns, and determining the workload includes a potential inefficiency when the information matches at least one of the set of inefficiency patterns.

Claims (83)

1 . A method comprising:

executing a workload on a computing system;

monitoring execution of the workload to obtain information by intercepting one or more calls to one or more operations of a plurality of operations caused to be executed by the workload;

detecting during execution of the workload that the one or more operations, which are called by the one or more calls, include functions includes a potential inefficiency if the information matches at least one of a predetermined set of inefficiency patterns stored in memory or an obsolete function; and

if the potential inefficiency is detected, performing, by the computing system, at least one action to address the potential inefficiency.

2 . The method of claim 1 , further comprising:

obtaining a metric of utilization of a hardware platform executing the workload, the information comprising the metric.

3 . The method of claim 1 ,

the information comprising an identifier of at least one of the one or more operations.

4 . The method of claim 1 , wherein the information comprises at least one parameter included in the one or more calls.

5 . The method of claim 1 , further comprising:

detecting that the one or more operations have been performed as part of the workload as the workload executes, the information comprising at least one value generated by the one or more operations.

6 . The method of claim 1 , further comprising:

patching instructions to the one or more operations to be performed as part of the workload as the workload executes; and

performing the instructions after the one or more operations are called during the execution of the workload, the instructions causing the information to be compared to the predetermined set of inefficiency patterns.

7 . The method of claim 1 , further comprising:

determining the at least one action to address the potential inefficiency.

8 . The method of claim 7 , further comprising:

using machine learning to determine the at least one action.

9 . The method of claim 7 , further comprising:

providing the at least one action to a user.

10 . The method of claim 7 , further comprising:

automatically implementing the at least one action.

11 . The method of claim 10 , wherein the at least one action is automatically implemented before the execution completes.

12 . The method of claim 7 , further comprising:

determining an amount of performance improvement achievable by implementing the at least one action.

13 . The method of claim 7 , further comprising:

producing a first modified workload by implementing the at least one action;

identifying at least one improvement to the first modified workload by executing the first modified workload; and

producing a second modified workload by implementing the at least one improvement.

14 . The method of claim 1 , wherein the predetermined set of inefficiency patterns comprises a plurality of inefficiency patterns, a plurality of potential inefficiencies comprise the potential inefficiency, the plurality of inefficiency patterns is associated with a plurality of priority values, and the method further comprises:

determining the workload includes the plurality of potential inefficiencies when the information matches one or more of the predetermined set of inefficiency patterns;

determining a plurality of actions that address the plurality of potential inefficiencies;

sorting the plurality of actions by the plurality of priority values associated with the plurality of inefficiency patterns that matched the plurality of potential inefficiencies addressed by the plurality of actions; and

providing the sorted plurality of actions to a user.

15 . The method of claim 1 , wherein the workload executes within a hardware and software domain, and

the predetermined set of inefficiency patterns comprises at least one predefined inefficiency pattern that is specific to the hardware and software domain.

16 . The method of claim 1 , wherein the one or more operations comprise at least one of a function or a procedure.

17 . A system comprising:

at least one circuit to:

automatically select a workload;

execute the workload;

monitor the workload to obtain information by at least intercepting one or more calls to one or more operations of a plurality of operations caused to be executed by the workload;

detect during execution of the workload that the one or more operations, which are called by the one or more calls, include one or more potential inefficiencies if at least one inefficiency pattern from a set of predetermined inefficiency patterns matches at least a portion of the information, wherein one or more inefficiency patterns in the set of predetermined inefficiency patterns identify at least one obsolete function; and

if the one or more potential inefficiencies are detected, perform at least one action to address the one or more potential inefficiencies.

18 . The system of claim 17 , wherein the at least one circuit is to determine the at least one action that addresses the one or more potential inefficiencies.

19 . The system of claim 18 , wherein the at least one circuit is to at least one of provide the at least one action to a user or automatically implement the at least one action.

20 . The system of claim 18 , wherein the at least one circuit is to determine an amount of performance improvement achievable by implementing the at least one action.

21 . The system of claim 17 , wherein the at least one circuit is to automatically select the workload from a plurality of workloads that were performed previously.

22 . The system of claim 21 , wherein the at least one circuit is to automatically select the workload based on information included in the plurality of workloads.

23 . The system of claim 17 , wherein the at least one circuit is to repeat automatically selecting the workload, comparing the information, and identifying the one or more potential inefficiencies after a predetermined amount of time.

24 . The system of claim 17 , wherein the at least one circuit is to implement a first execution environment that is separate from a second execution environment in which the workload was performed previously, and

the workload is executed in the first execution environment.

25 . The system of claim 17 , wherein the at least one circuit is to enter an issue corresponding to at least one of the one or more potential inefficiencies into an issue tracking system.

26 . The system of claim 17 , wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a first system for performing simulation operations;

a second system for performing digital twin operations;

a third system for performing light transport simulation;

a fourth system for performing collaborative content creation for 3D assets;

a fifth system for performing deep learning operations;

a sixth system implemented using an edge device;

a seventh system implemented using a robot;

an eighth system for performing conversational Artificial Intelligence operations;

a nineth system for generating synthetic data;

a tenth system incorporating one or more virtual machines (VMs);

an eleventh system implemented at least partially in a data center;

a twelfth system implemented at least partially using cloud computing resources;

a thirteenth system for implementing a web-hosted service for detecting program workload inefficiencies; or

an application as an application programming interface (“API”).

27 . One or more processors, comprising:

circuitry to:

execute a workload;

monitor the workload to obtain information by at least intercepting one or more calls to one or more operations of a plurality of operations caused to be executed by the workload; and,

perform at least one action to address one or more potential inefficiencies detected for the one or more operations, which are called by the one or more calls, during execution of the workload, by comparing the information to at least one of a predetermined set of inefficiency patterns or an obsolete function.

28 . The one or more processors of claim 27 , wherein the circuitry is further to import performance analyzer instructions into workload instructions for the workload such that executing the workload instructions executes the performance analyzer instructions, and

wherein the one or more potential inefficiencies are identified via execution of the performance analyzer instructions.

29 . The one or more processors of claim 28 , wherein the circuitry is further to:

patch a portion of the performance analyzer instructions to the one or more operations to be performed by the workload instructions as the workload instructions execute; and

perform the portion of the performance analyzer instructions after the one or more operations is called by the workload instructions, the portion causing the circuitry to identify at least one of one or more potential inefficiencies.

30 . The one or more processors of claim 27 , wherein the circuitry is further to perform at least one of a static analysis of workload instructions for the workload or a telemetry measurement analysis of at least one hardware component of a computing system on which the workload instructions execute.

31 . The one or more processors of claim 27 , wherein workload instructions for the workload are to execute within a hardware and software domain, and the predetermined set of inefficiency patterns comprises at least one predefined inefficiency pattern that is specific to the hardware and software domain.

32 . The one or more processors of claim 27 , wherein the circuitry is further to automatically implement one or more actions aimed at improving the one or more potential inefficiencies.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2022
From: MIGACZ, SZYMON; MORKISZ, PAWEL; FIT-FLOREA, ALEX; BALA, MACIEJ; ZAKRZEWSKI, JAKUB; KRISHNAMURTHY, TRIVIKRAM; NITIN, NITIN; LYM, SANGKUG; WANG, SHANG; YU, CHENHAN; MILESI, ALEXANDRE
To: NVIDIA CORPORATION
Reel/Frame 060423/0871 →
Continuity (1)
Related Publication 20230418726A1 · Dec 28, 2023
References Cited (13)
US 10152357B1 · Espy · 2018 [cited by examiner]
US 10467132B1 · Chatterjee · 2019 [cited by examiner]
US 20040117540A1 · Hahn · 2004 [cited by examiner]
US 20080104605A1 · Steinder · 2008 [cited by examiner]
US 20110022870A1 · McGrane · 2011 [cited by examiner]
US 20140331277A1 · Frascadore · 2014 [cited by examiner]
US 20200104174A1 · Vlcek · 2020 [cited by examiner]
US 20220156114A1 · Nagpal · 2022 [cited by examiner]
US 20240078098A1 · Rister · 2024 [cited by examiner]
IEEE “IEEE Standard for Floating-Point Arithmetric”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008. [cited by applicant]
IEEE, “IEEE Standard for 802.3,” IEEE Standard for Ethernetn, IEEE Computer Society, Dec. 28, 2012, 634 pages. [cited by applicant]
Wikipedia, “IEEE 802.11,” Wikipedia the Free Encyclopedia, https://en.wikipedia.org/wiki/IEEE_802.11, most recent edit Sep. 20, 2020 [retrieved Sep. 22, 2020], 15 pages. [cited by applicant]
Wikipedia, “IEEE 802.5,” Wikepedia The Free Encyclopedia, https://en.wikipedia.org/wiki/Token_Ring, Jan. 14, 2020, 12 pages. [cited by applicant]