IP Library › Granted Patent US 12,725,058
Granted Patent B2
US 12,725,058 · App. 18/358,245 · Granted Sep 1, 2026

Deployment of machine learning models using large language models and few-shot learning

Inventors: Rajesh Vellore Arumugam (Singapore, SG); Anantharaman Ravi (Singapore, SG); Isaac New Yi Qing (Singapore, SG); Sundeep Gullapudi (Singapore, SG); Yi Quan Zhou (Singapore, SG)
Assignee: SAP SE
G06N5/04G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,058
App. No.
18/358,245
Granted
Sep 1, 2026
Kind
B2
Abstract

Methods, systems, and computer-readable storage media for providing, for a set of ML models, a set of training metrics determined using test data during a training phase, providing, for a production-use ML model, a set of inference metrics based on predictions generated by the production-use ML model, generating, by a prompt generator, a set of few-shot examples using the set of training metrics and the set of inference metrics, inputting, by the prompt generator, the set of few-shot examples to a LLM as prompts, transmitting, to the LLM a query, displaying, to a user, a recommendation that is received from the LLM and responsive to the query, receiving input from a user indicating a user-selected ML model responsive to the recommendation, and deploying a user-selected ML model to an inference runtime for production use.

Claims (46)

1 . A computer-implemented method for deploying machine learning (ML) models for inference in production, the method being executed by one or more processors and comprising:

providing, for a set of ML models, a set of training metrics determined using test data during a training phase of ML models in the set of ML models;

providing, for a production-use ML model, a set of inference metrics based on predictions generated by the production-use ML model;

generating, by a prompt generator, a set of few-shot examples using the set of training metrics and the set of inference metrics, the set of few-shot examples comprising two or more model identifiers, each model identifier identifying a ML model, and, for each model identifier, a sub-set of training metrics of the set of training metrics and a sub-set of inference metrics of the set of inference metrics;

inputting, by the prompt generator, the set of few-shot examples to a large language model (LLM) as prompts, the set of few-shot examples providing context to the LLM for queries associated with ML model selection;

transmitting, to the LLM a query;

displaying, to a user, a recommendation that is received from the LLM and responsive to the query;

receiving input from a user indicating a user-selected ML model responsive to the recommendation; and

deploying a user-selected ML model to an inference runtime for production use.

2 . The method of claim 1 , wherein deploying a user-selected ML model to an inference runtime for production use at least partially comprises transmitting the user-selected ML model from a ML model store to the inference runtime.

3 . The method of claim 1 , wherein the set of training metrics comprises, for each ML model in the set of ML models, a sub-set of training metrics comprising the model identifier, a code, a proposal rate, an accuracy, and a threshold.

4 . The method of claim 1 , wherein the set of inference metrics comprises sub-sets of inference metrics each comprising an auto-task accuracy, a proposal rate, and a confidence threshold.

5 . The method of claim 4 , wherein the auto-task accuracy indicates an accuracy of automatic execution of a task in response to a prediction of the production-use ML model.

6 . The method of claim 1 , wherein the query comprises a code and at least one target metric.

7 . The method of claim 1 , wherein generating, by the prompt generator, the set of few-shot examples comprises populating a prompt template.

8 . A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for deploying machine learning (ML) models for inference in production, the operations comprising:

providing, for a set of ML models, a set of training metrics determined using test data during a training phase of ML models in the set of ML models;

providing, for a production-use ML model, a set of inference metrics based on predictions generated by the production-use ML model;

generating, by a prompt generator, a set of few-shot examples using the set of training metrics and the set of inference metrics, the set of few-shot examples comprising two or more model identifiers, each model identifier identifying a ML model, and, for each model identifier, a sub-set of training metrics of the set of training metrics and a sub-set of inference metrics of the set of inference metrics;

inputting, by the prompt generator, the set of few-shot examples to a large language model (LLM) as prompts, the set of few-shot examples providing context to the LLM for queries associated with ML model selection;

transmitting, to the LLM a query;

displaying, to a user, a recommendation that is received from the LLM and responsive to the query;

receiving input from a user indicating a user-selected ML model responsive to the recommendation; and

deploying a user-selected ML model to an inference runtime for production use.

9 . The non-transitory computer-readable storage medium of claim 8 , wherein deploying a user-selected ML model to an inference runtime for production use at least partially comprises transmitting the user-selected ML model from a ML model store to the inference runtime.

10 . The non-transitory computer-readable storage medium of claim 8 , wherein the set of training metrics comprises, for each ML model in the set of ML models, a sub-set of training metrics comprising the model identifier, a code, a proposal rate, an accuracy, and a threshold.

11 . The non-transitory computer-readable storage medium of claim 8 , wherein the set of inference metrics comprises sub-sets of inference metrics each comprising an auto-task accuracy, a proposal rate, and a confidence threshold.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein the auto-task accuracy indicates an accuracy of automatic execution of a task in response to a prediction of the production-use ML model.

13 . The non-transitory computer-readable storage medium of claim 8 , wherein the query comprises a code and at least one target metric.

14 . The non-transitory computer-readable storage medium of claim 8 , wherein generating, by the prompt generator, the set of few-shot examples comprises populating a prompt template.

15 . A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for deploying machine learning (ML) models for inference in production, the operations comprising:

providing, for a set of ML models, a set of training metrics determined using test data during a training phase of ML models in the set of ML models;

providing, for a production-use ML model, a set of inference metrics based on predictions generated by the production-use ML model;

generating, by a prompt generator, a set of few-shot examples using the set of training metrics and the set of inference metrics, the set of few-shot examples comprising two or more model identifiers, each model identifier identifying a ML model, and, for each model identifier, a sub-set of training metrics of the set of training metrics and a sub-set of inference metrics of the set of inference metrics;

inputting, by the prompt generator, the set of few-shot examples to a large language model (LLM) as prompts, the set of few-shot examples providing context to the LLM for queries associated with ML model selection;

transmitting, to the LLM a query;

displaying, to a user, a recommendation that is received from the LLM and responsive to the query;

receiving input from a user indicating a user-selected ML model responsive to the recommendation; and

deploying a user-selected ML model to an inference runtime for production use.

16 . The system of claim 15 , wherein deploying a user-selected ML model to an inference runtime for production use at least partially comprises transmitting the user-selected ML model from a ML model store to the inference runtime.

17 . The system of claim 15 , wherein the set of training metrics comprises, for each ML model in the set of ML models, a sub-set of training metrics comprising the model identifier, a code, a proposal rate, an accuracy, and a threshold.

18 . The system of claim 15 , wherein the set of inference metrics comprises sub-sets of inference metrics each comprising an auto-task accuracy, a proposal rate, and a confidence threshold.

19 . The system of claim 18 , wherein the auto-task accuracy indicates an accuracy of automatic execution of a task in response to a prediction of the production-use ML model.

20 . The system of claim 15 , wherein the query comprises a code and at least one target metric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2023
From: ARUMUGAM, RAJESH VELLORE; RAVI, ANANTHARAMAN; QING, ISAAC NEW YI; ZHOU, YI QUAN; GULLAPUDI, SUNDEEP
To: SAP SE
Reel/Frame 064409/0272 →
Continuity (1)
Related Publication 20250036974A1 · Jan 30, 2025
References Cited (10)
US 12124468B1 · Zhou · 2024 [cited by examiner]
US 20230186147A1 · Sen · 2023 [cited by examiner]
US 20230326191A1 · Li · 2023 [cited by examiner]
US 20240095447A1 · Ping · 2024 [cited by examiner]
US 20240185001A1 · Nagaraju · 2024 [cited by examiner]
US 20240282298A1 · Koneru · 2024 [cited by examiner]
US 20240330279A1 · Truong · 2024 [cited by examiner]
US 20240354319A1 · Dinu · 2024 [cited by examiner]
US 20240411666A1 · Chan · 2024 [cited by examiner]
US 20240419912A1 · Somech · 2024 [cited by examiner]