IP Library › Granted Patent US 12,530,224
Granted Patent B2
US 12,530,224 · App. 17/457,080 · Granted Jan 20, 2026

High throughput machine learning model runtime using CPU affinity and process niceness control

Inventor: Shi Xiaoyan (Singapore, SG)
Assignee: SAP SE
G06F9/4881G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,224
App. No.
17/457,080
Granted
Jan 20, 2026
Kind
B2
Abstract

Methods, systems, and computer-readable storage media for receiving, by a request processing engine of the ML model runtime, an inference request associated with a version model (VM), and determining, by the request processing engine that a VM-specific token and a global token are available for the inference request, the VM-specific token being available from a token pool that is specific to the VM, and in response: selecting a VM process (VMP) in a set of VMPs for execution of the inference request, the VMP being executed by a processor that is different from one or more processors executing one or more other VMPs in the set of VMPs, each VMP in the set of VMPs being specific to the VM and being designated for execution by a respective processor by a respective affinity setting, and providing the inference request to the VMP for execution.

Claims (45)

1 . A computer-implemented method for processing inference requests through a machine learning (ML) model runtime, the method being executed by one or more processors and comprising:

receiving, by a request processing engine of the ML model runtime, a first inference request associated with a first version model (VM); and

determining, by the request processing engine that a VM-specific token and a global token are available for the first inference request, the VM-specific token being available from a token pool that is specific to the first VM, and in response:

removing the VM-specific token and the global token,

selecting a first VM process (VMP) in a first set of VMPs for execution of the first inference request, the first VMP being executed by a processor that is different from one or more processors executing one or more other VMPs in the first set of VMPs, each VMP in the first set of VMPs being specific to the first VM and being designated for execution by a respective processor by a respective affinity setting, and providing the first inference request to the first VMP for execution,

wherein, based on a niceness setting and the affinity setting of each VMP in the first set of VMPs, a first processor executes the first VMP as a primary VMP of the first VM and a second processor executes a second VMP as a secondary VMP of the first VM.

2 . The method of claim 1 , wherein the first VMP is associated with the niceness setting that prioritizes the first VMP for execution by the first processor relative to the second VMP.

3 . The method of claim 2 , wherein the first VMP has a higher priority than the second VMP.

4 . The method of claim 1 , further comprising:

receiving, by the request processing engine of the ML model runtime, a second inference request associated with a second VM;

determining, by the request processing engine that one or more of a VM-specific token and a global token are unavailable for the second inference request; and

determining that a timeout has occurred for the second inference request and, in response, rejecting the second inference request.

5 . The method of claim 1 , wherein the global token is available from a global token pool.

6 . The method of claim 1 , wherein each of the request processing engine and the VMPs in the first set of VMPs is executed by a respective processor.

7 . A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for processing inference requests through a machine learning (ML) model runtime, the operations comprising:

receiving, by a request processing engine of the ML model runtime, a first inference request associated with a first version model (VM); and

determining, by the request processing engine that a VM-specific token and a global token are available for the first inference request, the VM-specific token being available from a token pool that is specific to the first VM, and in response:

removing the VM-specific token and the global token,

selecting a first VM process (VMP) in a first set of VMPs for execution of the first inference request, the first VMP being executed by a processor that is different from one or more processors executing one or more other VMPs in the first set of VMPs, each VMP in the first set of VMPs being specific to the first VM and being designated for execution by a respective processor by a respective affinity setting, and

providing the first inference request to the first VMP for execution,

wherein, based on a niceness setting and the affinity setting of each VMP in the first set of VMPs, a first processor executes the first VMP as a primary VMP of the first VM and a second processor executes a second VMP as a secondary VMP of the first VM.

8 . The non-transitory computer-readable storage medium of claim 7 , wherein the first VMP is associated with the niceness setting that prioritizes the first VMP for execution by the first processor relative to the second VMP.

9 . The non-transitory computer-readable storage medium of claim 8 , wherein the first VMP has a higher priority than the second VMP.

10 . The non-transitory computer-readable storage medium of claim 7 , wherein operations further comprise:

receiving, by the request processing engine of the ML model runtime, a second inference request associated with a second VM;

determining, by the request processing engine that one or more of a VM-specific token and a global token are unavailable for the second inference request; and

determining that a timeout has occurred for the second inference request and, in response, rejecting the second inference request.

11 . The non-transitory computer-readable storage medium of claim 7 , wherein the global token is available from a global token pool.

12 . The non-transitory computer-readable storage medium of claim 7 , wherein each of the request processing engine and the VMPs in the first set of VMPs is executed by a respective processor.

13 . A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for processing inference requests through a machine learning (ML) model runtime, the operations comprising:

receiving, by a request processing engine of the ML model runtime, a first inference request associated with a first version model (VM); and

determining, by the request processing engine that a VM-specific token and a global token are available for the first inference request, the VM-specific token being available from a token pool that is specific to the first VM, and in response:

removing the VM-specific token and the global token,

selecting a first VM process (VMP) in a first set of VMPs for execution of the first inference request, the first VMP being executed by a processor that is different from one or more processors executing one or more other VMPs in the first set of VMPs, each VMP in the first set of VMPs being specific to the first VM and being designated for execution by a respective processor by a respective affinity setting, and

providing the first inference request to the first VMP for execution,

wherein, based on a niceness setting and the affinity setting of each VMP in the first set of VMPs, a first processor executes the first VMP as a primary VMP of the first VM and a second processor executes a second VMP as a secondary VMP of the first VM.

14 . The system of claim 13 , wherein the first VMP is associated with the niceness setting that prioritizes the first VMP for execution by the first processor relative to the second VMP.

15 . The system of claim 13 , wherein the first VMP has a higher priority than the second VMP.

16 . The system of claim 13 , wherein operations further comprise:

receiving, by the request processing engine of the ML model runtime, a second inference request associated with a second VM;

determining, by the request processing engine that one or more of a VM-specific token and a global token are unavailable for the second inference request; and

determining that a timeout has occurred for the second inference request and, in response, rejecting the second inference request.

17 . The system of claim 13 , wherein the global token is available from a global token pool.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2021
From: XIAOYAN, SHI
To: SAP SE
Reel/Frame 058255/0142 →
Continuity (1)
Related Publication 20230168922A1 · Jun 1, 2023
References Cited (14)
US 7401112B1 · Matz · 2008 [cited by examiner]
US 20120054765A1 · Lee · 2012 [cited by examiner]
US 20210281662A1 · Mathur · 2021 [cited by examiner]
Beisel et al., “Cooperative multitasking for heterogeneous accelerators in the linux completely fair scheduler.” ASAP 2011-22nd IEEE International Conference on Application-specific Systems, Architectures and Processors… [cited by applicant]
Bengio et al., “Towards biologically plausible deep learning.” arXiv preprint arXiv:1502.04156, Feb. 2015, 10 pages. [cited by applicant]
Bishop, “Pattern recognition.” Machine learning 128.9, Feb. 2006, 11 pages. [cited by applicant]
Kobus et al., “Completely Fair Scheduler and its tuning.” draft on Internet, 2009, 8 pages. [cited by applicant]
Marblestone et al., “Toward an integration of deep learning and neuroscience.” Frontiers in computational neuroscience 10, 94, Sep. 2016, 41 pages. [cited by applicant]
Mitchell, “Machine learning.”, 1997, 870-877, 14 pages. [cited by applicant]
Olshausen et al., “Emergence of simple-cell receptive field properties by learning a sparse code for natural images.” Nature 381.6583, Jun. 1996, 607-609, 3 pages. [cited by applicant]
Provost et al., “Glossary of terms.” Journal of Machine Learning 30.2-3, Feb. 1998, 271-274, 4 pages. [cited by applicant]
Simon, “Too big to ignore: the business case for big data” vol. 72. John Wiley & Sons, Mar. 2013, 242 pages. [cited by applicant]
Wong et al., “Fairness and interactive performance of O (1) and CFS Linux kernel schedulers.” 2008 International Symposium on Information Technology. vol. 4. IEEE, Aug. 2008, 8 pages. [cited by applicant]
Wong et al., “Towards achieving fairness in the Linux scheduler.” ACM SIGOPS Operating Systems Review 42.5, Jul. 2008, 34-43, 10 pages. [cited by applicant]