IP Library › Granted Patent US 12,541,402
Granted Patent B2
US 12,541,402 · App. 17/659,775 · Granted Feb 3, 2026

Machine learning model layer

Inventors: Arpeet Kale (Sunnyvale, CA); Shashank Harinath (San Jose, CA)
Assignee: Salesforce, Inc.
G06F9/5044G06F9/5055G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,402
App. No.
17/659,775
Granted
Feb 3, 2026
Kind
B2
Abstract

Techniques are disclosed that pertain to facilitating the execution of machine learning (ML) models. A computer system may implement an ML model layer that permits ML models built using any of a plurality of different ML model frameworks to be submitted without a submitting entity having to define execution logic for a submitted ML model. The computer system may receive, via the ML model layer, configuration metadata for a particular ML model. The computer system may then receive a prediction request from a user to produce a prediction based on the particular ML model. The computer system may produce a prediction based on the particular ML model. As a part of producing that prediction, the computer system may select, in accordance with the received configuration metadata, one of a plurality of types of hardware resources on which to load the particular ML model.

Claims (56)

1 . A method, comprising:

implementing, by a computer system, a machine learning (ML) model layer that permits ML models built using any of a plurality of different ML model frameworks to be submitted without a submitting entity having to define execution logic for a submitted ML model;

receiving, by the computer system via the ML model layer, configuration metadata for a particular ML model, wherein the configuration metadata identifies a particular one of the plurality of different ML model frameworks that is associated with the particular ML model;

receiving, by the computer system, a first prediction request from a user to produce a prediction based on the particular ML model, wherein the first prediction request identifies a model identifier of the particular ML model; and

producing, by the computer system, a prediction based on the particular ML model, wherein the producing includes:

accessing the configuration metadata using the model identifier;

selecting, based on the particular ML model framework that is identified by the accessed configuration metadata, one of a plurality of types of hardware resources on which to load the particular ML model;

loading the particular ML model from a storage onto a hardware resource of the selected type of hardware resource; and

producing the prediction using a set of inputs from the first prediction request with the loaded, particular ML model.

2 . The method of claim 1 , wherein the method further comprises:

receiving, by the computer system, a second prediction request to produce a prediction based on the particular ML model; and

producing, by the computer system, another prediction based on the particular ML model without reloading the particular ML model on the selected type of hardware resource.

3 . The method of claim 2 , wherein the second prediction request is received from a different user than the user that provided the first prediction request.

4 . The method of claim 1 , wherein the configuration metadata specifies a maximum batch size that indicates a maximum number of prediction requests that can be issued against the particular ML model at a time.

5 . The method of claim 1 , further comprising:

maintaining, by the computer system, a set of ML models in a memory of the computer system, wherein the set of ML models includes the particular ML model, wherein the loading includes swapping the particular ML model with another ML model already loaded on the hardware resource.

6 . The method of claim 5 , wherein the swapping is performed in response to determining that a computing resource threshold associated with the hardware resource is already being consumed by ML models loaded on the hardware resource.

7 . The method of claim 1 , wherein the configuration metadata specifies an input type and an output type for the particular ML model, and wherein the producing of the prediction using the set of inputs includes pre-processing on the set of inputs to ensure that the set of inputs satisfies the input type.

8 . The method of claim 1 , wherein the configuration metadata specifies a location external to the computer system where the particular ML model is stored, and wherein the method further comprises:

after accessing the configuration metadata, the computer system accessing, based on the configuration metadata, the particular ML model from the location external to the computer system.

9 . The method of claim 8 , wherein the configuration metadata is stored at a different storage location than the particular ML model.

10 . A non-transitory computer-readable medium having program instructions stored thereon that are executable to cause a computer system to perform operations comprising:

implementing a machine learning (ML) model layer that permits ML models built using any of a plurality of different ML model frameworks to be submitted without a submitting entity having to define execution logic for a submitted ML model;

receiving, via the ML model layer, configuration metadata for a particular ML model, wherein the configuration metadata identifies a particular one of the plurality of different ML model frameworks that is associated with the particular ML model;

receiving a first prediction request to produce a prediction based on the particular ML model, wherein the first prediction request identifies a model identifier of the particular ML model; and

producing a first prediction based on the particular ML model, wherein the producing includes:

accessing the configuration metadata using the model identifier;

selecting, based on the particular ML model framework that is identified by the accessed configuration metadata, one of a plurality of types of hardware resources on which to load the particular ML model;

loading the particular ML model from a storage onto a hardware resource of the selected type of hardware resource; and

producing the prediction using a set of inputs from the first prediction request with the loaded, particular ML model.

11 . The medium of claim 10 , further comprising:

producing a second prediction based on the particular ML model without reallocating the particular ML model on the selected type of hardware resource, wherein the first prediction is produced for a first tenant of the computer system and the second prediction is produced for a second, different tenant of the computer system.

12 . The medium of claim 10 , wherein the operations further comprise:

loading a plurality of instances of the particular ML model onto hardware resources of the selected type of hardware resource; and

issuing a batch of prediction requests, including the first prediction request, against the plurality of instances.

13 . The medium of claim 10 , wherein the loading includes:

identifying, based on a replacement policy, an ML model loaded on the hardware resource; and

offloading the identified ML model from the hardware resource prior to loading the particular ML model onto the hardware resource.

14 . The medium of claim 10 , wherein the configuration metadata specifies a plurality of batch sizes indicative of respective numbers of prediction requests that can be issued against the particular ML model at a time.

15 . A system, comprising:

at least one processor;

a memory having program instructions stored thereon that are executable by the at least one processor to cause the system to perform operations comprising:

implementing a machine learning (ML) model layer that permits ML models built using any of a plurality of different ML model frameworks to be submitted without a submitting entity having to define execution logic for a submitted ML model;

receiving, via the ML model layer, configuration metadata for a particular ML model, wherein the configuration metadata identifies a particular one of the plurality of different ML model frameworks that is associated with the particular ML model;

receiving a first prediction request from a user to produce a prediction based on the particular ML model, wherein the first prediction request identifies a model identifier of the particular ML model; and

producing a prediction based on the particular ML model, wherein the producing includes:

accessing the configuration metadata using the model identifier;

selecting, based on the particular ML model framework that is identified by the accessed configuration metadata, one of a plurality of types of hardware resources on which to load the particular ML model;

loading the particular ML model from a storage onto a hardware resource of the selected type of hardware resource; and

producing the prediction using a set of inputs from the first prediction request with the loaded, particular ML model.

16 . The system of claim 15 , wherein the operations further comprise:

accessing, based on the configuration metadata, the particular ML model from a location external to the system, wherein the particular ML model and configuration metadata are stored at different storage locations.

17 . The system of claim 15 , wherein the operations further comprise:

maintaining a set of ML models in the memory of the system, wherein the set of ML models includes the particular ML model; and

wherein the loading includes swapping the particular ML model with another ML model already loaded on the hardware resource.

18 . The system of claim 15 , wherein the plurality of types of hardware resources includes at least a central processing unit and a graphics processing unit.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2022
From: KALE, ARPEET; HARINATH, SHASHANK
To: SALESFORCE.COM, INC.
Reel/Frame 059640/0344 →
Continuity (1)
Related Publication 20230333901A1 · Oct 19, 2023
References Cited (24)
US 10713594B2 · Szeto et al. · 2020 [cited by applicant]
US 11182691B1 · Zhang · 2021 [cited by examiner]
US 11614932B2 · Gumashta · 2023 [cited by examiner]
US 11720813B2 · Babu · 2023 [cited by examiner]
US 12099904B2 · Cao · 2024 [cited by examiner]
US 20170286864A1 · Fiedel · 2017 [cited by examiner]
US 20190102700A1 · Babu et al. · 2019 [cited by applicant]
US 20190188605A1 · Zavesky · 2019 [cited by examiner]
US 20200004596A1 · Sengupta · 2020 [cited by examiner]
US 20200019882A1 · Garg et al. · 2020 [cited by applicant]
US 20200110619A1 · Rajaram · 2020 [cited by examiner]
US 20200356415A1 · Goli · 2020 [cited by examiner]
US 20200380415A1 · Siracusa · 2020 [cited by examiner]
US 20210081819A1 · Polleri · 2021 [cited by examiner]
US 20210281662A1 · Mathur · 2021 [cited by examiner]
US 20220083389A1 · Poothia · 2022 [cited by examiner]
US 20220092346A1 · Jones · 2022 [cited by examiner]
US 20220215008A1 · Adibowo · 2022 [cited by examiner]
US 20220391748A1 · Nikitin · 2022 [cited by examiner]
Hunt et al.; “Chiron: Privacy-preserving Machine Learning as a Service”; arXiv:1803.05961v1 [cs.CR] Mar. 15, 2018; (Hunt_2018.pdf; pp. 1-15) (Year: 2018). [cited by examiner]
Vartak et al.; “MODELDB: Opportunities and Challenges in Managing Machine Learning Models”; Copyright 2018 IEEE; (Vartak_ 2018.pdf; pp. 16-25) (Year: 2018). [cited by examiner]
Garcia et al.; “A Cloud-Based Framework for Machine Learning Workloads and Applications”; Special Section on Scalable Deep Learning for Big Data; IEEE 2020; DOI: 10.1109/ACCESS.2020.2964386; (Garcia_2019.pdf) (Year: 202… [cited by examiner]
Sigl et al.; “Don't Fear the Reaper: A Framework for Materializing and Reusing Deep-Learning Models”; 2019 IEEE 35th International Conference on Data Engineering (ICDE); DOI 10.1109/ICDE.2019.00246; (Sigl_2019.pdf) (Yea… [cited by examiner]
Tsay et al.; “AIMMX: Artificial Intelligence Model Metadata Extractor”; 2020 IEEE/ACM 17th International Conference on Mining Software Repositories (MSR); https://doi.org/10.1145/3379597.3387448; (Tsay_2020.pdf) (Year: … [cited by examiner]