IP Library › Granted Patent US 12,505,288
Granted Patent B1
US 12,505,288 · App. 18/536,917 · Granted Dec 23, 2025

Multi-machine learning system for interacting with a large language model

Inventors: Haibo Ding (Fremont, CA); Panpan Xu (Santa Clara, CA); Huan Song (San Jose, CA); James Robert Golden (Oakland, CA); Yawei Wang (Santa Clara, CA); Lin Lee Cheong (Redwood City, CA)
Assignee: Amazon Technologies, Inc.
G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,288
App. No.
18/536,917
Granted
Dec 23, 2025
Kind
B1
Abstract

Techniques disclosed may include determining a first prompt that is generated at a user interface of a user device and that is to be input to a first machine learning (ML) model. The techniques may further include determining, based at least in part on a second ML model, a classification of the first prompt. The techniques may further include generating, based at least in part on an input to a third ML model associated with the classification, a second prompt, the input being based at least in part the first prompt, an output of the third ML model comprising the second prompt. The techniques may further include performing at least one of: (i) inputting the second prompt instead of the first prompt to the first ML model, or (ii) causing the second prompt to be presented at the user interface.

Claims (85)

1 . A system comprising:

one or more storage media storing instructions; and

one or more processors configured to execute the instructions to cause the system to: receive a first prompt from a user device to be input to a large language model (LLM);

generate, by at least using a first machine learning (ML) model, a classification of the first prompt;

select, based at least in part on the classification, a second ML model from a plurality of ML models, the second ML model configured to output first LLM prompts based at least in part on a set of templates, the plurality of ML models comprising a third ML model associated with a different classification and configured to output second LLM prompts independently of any template in the set of templates;

provide a first input to the second ML model based at least in part on the first prompt and the classification;

determine a second prompt that is output by the second ML model based at least in part on the first input and a template associated with the classification;

provide the second prompt to the LLM;

receive a first response from the LLM, the first response generated using the second prompt; and

send, the first response to the user device.

2 . The system of claim 1 , wherein the execution of the instructions further causes the system to:

receive a third prompt from the user device to be input to the LLM;

generate a second classification of the third prompt by at least using the first ML model;

select the third ML model based at least in part on the second classification;

provide a second input to the third ML model based at least in part on the third prompt and the second classification;

determine a fourth prompt that is output by the third ML model based at least in part on the second input; and

provide the fourth prompt to the LLM.

3 . The system of claim 1 , wherein the execution of the instructions further causes the system to:

receive, by the second ML model, the classification;

select, using the second ML model and the classification, the template associated with the classification from the set of templates;

generate, by the second ML model, information elements based at least on the first prompt, the information elements corresponding to fields defined by the template; and

generate the second prompt using the fields defined by the template and the information elements.

4 . The system of claim 1 , wherein the execution of the instructions further causes the system to:

transmit a query to a data store, the query including the classification, the data store storing the set of templates;

receive, by the second ML model, the template associated with the classification from the data store;

receive, by the second ML model, the first prompt; and

generate, by the second ML model, information elements based at least on the first prompt, the information elements corresponding to fields defined by the template, wherein the second prompt includes the information elements.

5 . A computer-implemented method, comprising:

receiving a first prompt from a user device to be input to a large language model (LLM);

generating, by at least using a first machine learning (ML) model, a classification of the first prompt;

selecting, based at least in part on the classification, a second ML model from a plurality of ML models, the second ML model configured to output first LLM prompts based at least in part on a set of templates, the plurality of ML models comprising a third ML model associated with a different classification and configured to output second LLM prompts independently of any template in the set of templates;

providing a first input to the second ML model based at least in part on the first prompt and the classification;

determining a second prompt that is output by the second ML model based at least in part on the first input and a template associated with the classification;

providing the second prompt to the LLM;

receiving a first response from the LLM, the first response generated using the second prompt; and

sending, the first response to the user device.

6 . The computer-implemented method of claim 5 ,

wherein the template is associated with the classification.

7 . The computer-implemented method of claim 5 , wherein determining the second prompt that is output by the second ML model further comprises:

receiving, by the second ML model, the classification;

selecting, using the second ML model and the classification, the template associated with the classification from the set of templates; and

generating, by the second ML model, information elements based at least on the first prompt, the information elements corresponding to fields defined by the template, wherein the second prompt includes the information elements.

8 . The computer-implemented method of claim 5 , wherein determining the second prompt that is output by the second ML model further comprises:

transmitting a query to a data store, the query including the classification, the data store storing the set of templates;

receiving, by the second ML model, the template associated with the classification from the data store;

receiving, by the second ML model, the first prompt; and

generating, by the second ML model, information elements based at least on the first prompt, the information elements corresponding to fields defined by the template; wherein the second prompt includes the information elements.

9 . The computer-implemented method of claim 5 , wherein the second ML model is trained according to a training procedure, the training procedure comprising:

receiving a third prompt;

generating, by using the second ML model based at least in part on the third prompt, a fourth prompt;

receiving an indication that one of: (i) the third prompt or (ii) the fourth prompt were selected to be used as input to the LLM; and

updating parameters of the second ML model based at least in part on the indication.

10 . The computer-implemented method of claim 5 , wherein the second ML model is trained according to a training procedure, the training procedure comprising:

receiving a third prompt;

generating, by using the second ML model based at least in part on the third prompt, a fourth prompt;

inputting the third prompt to the LLM;

inputting the fourth prompt to the LLM;

receiving a first output from the LLM corresponding to the third prompt;

receiving a second output from the LLM corresponding to the fourth prompt;

receiving an indication that one of: (i) the first output or (ii) the second output was of greater quality; and

updating parameters of the second ML model based at least in part on the indication.

11 . One or more non-transitory computer-readable storage media storing instructions that, upon execution executable by one or more processors of a system, cause the system to perform operations comprising:

receiving a first prompt from a user device to be input to a large language model (LLM);

generating, by at least using a first machine learning (ML) model, a classification of the first prompt;

selecting, based at least in part on the classification, a second ML model from a plurality of ML models, the second ML model configured to output first LLM prompts based at least in part on a set of templates, the plurality of ML models comprising a third ML model associated with a different classification and configured to output second LLM prompts independently of any template in the set of templates;

providing a first input to the second ML model based at least in part on the first prompt and the classification;

determining a second prompt that is output by the second ML model based at least in part on the first input and a template associated with the classification;

providing the second prompt to the LLM;

receiving a first response from the LLM, the first response generated using the second prompt; and

sending, the first response to the user device.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein execution of the instructions further causes the system to perform training operations comprising:

receiving a third prompt;

generating, by using the second ML model based at least in part on the third prompt, a fourth prompt;

inputting the fourth prompt to the LLM;

receiving a second output from the LLM corresponding to the fourth prompt;

receiving an indication of a quality of the second output; and

updating parameters of the second ML model based at least in part on the indication.

13 . The non-transitory computer-readable storage medium of claim 12 , wherein the indication of the quality is determined by performing at least one of: determining a number of keywords used in the second output, determining a number of key phrases used in the second output, determining a length of the second output, determining an order of key words used in the second output, determining punctuation used in the second output, determining which keywords were used in the second output, determining which key phrases were used in the second output, determining an accuracy of the second prompt, or determining a performance metric of the second prompt.

14 . The non-transitory computer-readable storage medium of claim 12 , wherein receiving the indication further comprises:

receiving the indication from an analysis system configured to indicate which one of the third prompt or the fourth prompt has a higher prompt value metric.

15 . The non-transitory computer-readable storage medium of claim 12 , wherein the indication of the quality is determined by at least receiving user input indicating a label to be associated with the second output, the label indicating the quality.

16 . The non-transitory computer-readable storage medium of claim 11 , wherein execution of the instructions further causes the system to perform additional operations comprising:

wherein the template is associated with the classification.

17 . The non-transitory computer-readable storage medium of claim 11 , wherein the instructions are part of a plugin for program code of an application hosted by the system, and wherein the application is configured by the plugin to replace the first prompt with the second prompt as an input to the LLM.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the plugin further configures the application to present the second prompt at a user interface of the user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2023
From: DING, HAIBO; XU, PANPAN; SONG, HUAN; GOLDEN, JAMES ROBERT; WANG, YAWEI; CHEONG, LIN LEE
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 065844/0960 →
References Cited (37)
US 6266664B1 · Russell-Falla · 2001 [cited by examiner]
US 8447602B2 · Bartosik · 2013 [cited by examiner]
US 9619735B1 · Lineback · 2017 [cited by examiner]
US 10565498B1 · Zhiyanov · 2020 [cited by examiner]
US 10853579B2 · Laxman · 2020 [cited by examiner]
US 10878174B1 · Vontobel · 2020 [cited by examiner]
US 11106690B1 · Dhillon · 2021 [cited by examiner]
US 11481434B1 · Venti · 2022 [cited by examiner]
US 11532301B1 · Hajebi · 2022 [cited by examiner]
US 20030115191A1 · Copperman · 2003 [cited by examiner]
US 20120215776A1 · Guha · 2012 [cited by examiner]
US 20150278341A1 · Shen · 2015 [cited by examiner]
US 20160323398A1 · Guo · 2016 [cited by examiner]
US 20170127016A1 · Yu · 2017 [cited by examiner]
US 20170323636A1 · Xiao · 2017 [cited by examiner]
US 20180077101A1 · Desouza Sana · 2018 [cited by examiner]
US 20190108228A1 · Zeng · 2019 [cited by examiner]
US 20200118544A1 · Lee · 2020 [cited by examiner]
US 20200356653A1 · Cho · 2020 [cited by examiner]
US 20210201351A1 · Nag · 2021 [cited by examiner]
US 20220366901A1 · Rathaur · 2022 [cited by examiner]
US 20230134791A1 · Londeree · 2023 [cited by examiner]
US 20250086213A1 · Dilipkumar · 2025 [cited by examiner]
AU 2011255614A1 · 2012 [cited by examiner]
AU 2014233517A1 · 2015 [cited by examiner]
AU 2015210460A1 · 2015 [cited by examiner]
CN 110998565A · 2020 [cited by examiner]
EP 3514694B1 · 2022 [cited by examiner]
EP 4312147A2 · 2024 [cited by examiner]
WO WO2013157603A1 · 2013 [cited by examiner]
WO WO2015039165A1 · 2015 [cited by examiner]
WO WO2017160341A1 · 2017 [cited by examiner]
WO WO2021119064A1 · 2021 [cited by examiner]
WO WO2023017320A1 · 2023 [cited by examiner]
Du et al., “Template Filling with Generative Transformers”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 6-11, 2021… [cited by applicant]
Hao et al., “Optimizing Prompts for Text-to-Image Generation”, Available online at: https://arxiv.org/pdf/2212.09611, Dec. 29, 2023, pp. 1-16. [cited by applicant]
Li , “Guiding Large Language Models via Directional Stimulus Prompting”, Available online at: https://arxiv.org/pdf/2302.11520, Oct. 9, 2023, pp. 1-27. [cited by applicant]