IP Library › Granted Patent US 12,632,651
Granted Patent B2
US 12,632,651 · App. 18/477,680 · Granted May 19, 2026

Automatic language model (LM) input optimization using textual gradients

Inventors: Reid Allen Pryzant (Seattle, WA); Jerry Zheng Li (Redmond, WA); Dan Iter (Cedar Park, TX); Yin Tat Lee (Seattle, WA); Chenguang Zhu (Issaquah, WA); Nanshan Zeng (Bellevue, WA); Anup Shirgaonkar (New York, NY)
Assignee: Microsoft Technology Licensing, LLC
G06F40/20G06F40/166G06F16/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,651
App. No.
18/477,680
Granted
May 19, 2026
Kind
B2
Abstract

Systems and methods are provided for implementing automatic prompt optimization using textual gradients. In various embodiments, a feedback prompt, input into a large language model (“LLM”), is used to generate textual gradients that criticize a current prompt. The feedback prompt includes the current prompt and predictions that are incorrect compared with corresponding labels associated with minibatch data processed by the LLM using the current prompt. The textual gradients and current prompt are used in an editing prompt to the LLM to obtain a set of optimized prompts, which may be expanded using a paraphrasing prompt that is input into the LLM to generate a set of paraphrased prompts. A selection algorithm is used to select one or more optimized prompts from the set of optimized prompts and/or the set of paraphrased prompts, and the process is repeated with the selected one or more optimized prompts replacing the current prompt.

Claims (68)

1 . A system for implementing automatic prompt optimization using textual gradients, the system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the system to perform a set of operations comprising:

providing, as input to a large language model (“LLM”), a first feedback prompt requesting one or more first textual gradients, each first textual gradient including a description of one or more first flaws in an initial prompt resulting in errors in LLM predictions, the first feedback prompt including an initial prompt to be optimized and one or more first predictions that are incorrect compared with corresponding one or more labels associated with a batch of data for which the initial prompt was used to generate the one or more first predictions;

receiving, from output of the LLM in response to the first feedback prompt, the one or more first textual gradients;

providing, as input to the LLM, a first editing prompt requesting a first set of optimized prompts, based on the initial prompt and the one or more first textual gradients;

receiving, from output of the LLM, the first set of optimized prompts;

selecting one or more first optimized prompts from at least the first set of optimized prompts based at least in part on evaluation of prompt performance using a secondary LLM that is finetuned based on a curated dataset for a specific subject area;

instructing the LLM to perform a task focused on the subject area by inputting the selected one or more first optimized prompts into the LLM; and

receiving, from the LLM, results for the instructed task.

2 . The system of claim 1 , wherein the first feedback prompt further includes at least one of the batch of data and the one or more labels corresponding to the one or more first predictions that are incorrect.

3 . The system of claim 1 , wherein the set of operations further comprises:

providing, as input to the secondary LLM, the selected one or more first optimized prompts each requesting a second prediction based on the batch of data;

receiving, from output of the secondary LLM, the second prediction for each of the selected one or more first optimized prompts;

comparing each second prediction with labels contained in the batch of data that is processed by the LLM using each of the selected one or more first optimized prompts; and

based on the comparison, identifying one or more second predictions that are incorrect.

4 . The system of claim 3 , wherein the set of operations further comprises:

providing, as input to the LLM, a second feedback prompt requesting one or more second textual gradients, the second feedback prompt including the selected one or more first optimized prompts and the one or more second predictions that are incorrect compared with labels contained in the batch of data;

receiving, from output of the LLM, the one or more second textual gradients, each second textual gradient including a description of one or more second flaws in one of the selected one or more first optimized prompts;

providing, as input to the LLM, a second editing prompt requesting a second set of optimized prompts, based on the selected one or more first optimized prompts and the one or more second textual gradients;

receiving, from output of the LLM, the second set of optimized prompts; and

selecting one or more second optimized prompts from at least the second set of optimized prompts based at least in part on evaluation of prompt performance using the secondary LLM,

wherein instructing the LLM to perform the task focused on the subject area is performed by inputting the selected one or more second optimized prompts into the LLM.

5 . The system of claim 4 , wherein at least one of the first feedback prompt, the second feedback prompt, the first editing prompt, or the second editing prompt is at least one of generated by the LLM or by a second LLM.

6 . The system of claim 4 , wherein at least one of providing the first feedback prompt, providing the second feedback prompt, providing the first editing prompt, or providing the second editing prompt is performed using an application programming interface (“API”) call to the LLM.

7 . The system of claim 1 , wherein the set of operations further comprises:

providing, as input to the LLM, a first paraphrasing prompt requesting a first set of paraphrased optimized prompts, based on the first set of optimized prompts; and

receiving, from output of the LLM, the first set of paraphrased optimized prompts;

wherein selecting the one or more first optimized prompts comprises selecting from at least one of the first set of optimized prompts or the first set of paraphrased optimized prompts.

8 . The system of claim 7 , wherein selecting the one or more first optimized prompts is performed using one or more selection algorithms comprising a selection algorithm based on a scoring metric, wherein the one or more first optimized prompts are selected based on whether each optimized prompt scores above a set threshold scoring metric values.

9 . The system of claim 7 , further comprising preselecting one or more of gradients, optimized prompts, or paraphrased optimized prompts by performing at least one of:

preselection using a trained classifier;

conversion of each of the one or more of gradients, optimized prompts, or paraphrased optimized prompts into corresponding embeddings, and preselection based on distances between resultant embeddings within corresponding embedding space;

preselection using the LLM itself or another LLM; or

preselection based on a chain of thought-based preselection process.

10 . The system of claim 1 , wherein the batch of data comprises at least one of a random sample of natural language (“NL”) training data or a curated sample of the NL training data that has been labelled as difficult example training data.

11 . A computer-implemented method for implementing automatic prompt optimization using textual gradients, the method comprising:

receiving one or more textual gradients after providing a feedback prompt as input to a large language model (“LLM”), the feedback prompt including an initial prompt to be optimized and one or more predictions that are incorrect compared with corresponding one or more labels associated with a batch of data for which the initial prompt was used to generate the one or more predictions;

receiving, in response to the feedback prompt, a set of optimized prompts after providing an editing prompt as input to the LLM, the editing prompt including the initial prompt and the one or more textual gradients, each textual gradient including a description of one or more flaws in the initial prompt;

receiving a set of paraphrased optimized prompts after providing a paraphrasing prompt as input to the LLM, the paraphrasing prompt including the set of optimized prompts;

selecting one or more optimized prompts from at least one of the set of optimized prompts or the set of paraphrased optimized prompts; and

repeating, until a set condition has been met, the processes of receiving the one or more textual gradients, receiving the set of optimized prompts, receiving the set of paraphrased optimized prompts, and selecting the one or more optimized prompts, wherein, for each successive iteration, the initial prompt or a previously selected group of one or more optimized prompts is replaced with a latest selected group of one or more optimized prompts that is selected during each previous iteration.

12 . The computer-implemented method of claim 11 , wherein the set condition comprises a set number of iterations, and wherein the method further comprises:

sampling a subset of optimized prompts from at least one of the set of optimized prompts or the set of paraphrased optimized prompts; and

receiving an average score after providing the sampled subset of optimized prompts as input to the LLM to output scores corresponding to the sampled subset of optimized prompts and averaging resultant scores, the sampled subset of optimized prompts each including the batch of data;

wherein selecting the one or more optimized prompts comprises selecting based on the average score.

13 . The computer-implemented method of claim 12 , wherein selecting the one or more optimized prompts comprises one of selecting a first number of the one or more optimized prompts that have scores above the average score or selecting a remaining number of the one or more optimized prompts after removing a second number of the one or more optimized prompts that have scores below the average score.

14 . A system for implementing automatic prompt optimization using textual gradients, the system comprising:

a processing system; and

memory coupled to the processing system, the memory comprising computer executable instructions that, when executed by the processing system, causes the system to perform operations comprising:

providing, as input to a large language model (“LLM”), a feedback prompt requesting one or more textual gradients, the feedback prompt including an initial prompt to be optimized and one or more predictions that are incorrect compared with corresponding one or more labels associated with a batch of data for which the initial prompt was used to generate the one or more predictions;

receiving, from output of the LLM in response to the feedback prompt, the one or more textual gradients, each textual gradient including a description of one or more flaws in the initial prompt;

providing, as input to the LLM, an editing prompt requesting a set of optimized prompts, based on the initial prompt and the one or more textual gradients;

receiving, from output of the LLM, the set of optimized prompts;

sampling a subset of optimized prompts from at least the set of optimized prompts;

providing, as input to the LLM, the sampled subset of optimized prompts requesting an average score based on averaging resultant scores corresponding to the sampled subset of optimized prompts, the sampled subset of optimized prompts each including the batch of data;

receiving, from output of the LLM, the average score;

selecting one or more optimized prompts based on the average score; and

repeating, until a set condition has been met, the processes of providing the feedback prompt, receiving the one or more textual gradients, providing the editing prompt, receiving the set of optimized prompts, sampling the subset of optimized prompts, providing the sampled subset of optimized prompts, receiving the average score, and selecting the one or more optimized prompts, wherein, for each successive iteration, the initial prompt or a previously selected group of one or more optimized prompts is replaced with a latest selected group of one or more optimized prompts that is selected during each previous iteration.

15 . The system of claim 14 , wherein the set condition comprises one of a set number of iterations or a determined level of match between comparison of labels contained in the batch of data and one or more subsequent predictions that are generated using the selected one or more optimized prompts.

16 . The system of claim 14 , wherein the set of operations further comprises, for each iteration:

providing, as input to the LLM, a paraphrasing prompt requesting a set of paraphrased optimized prompts, based on the set of optimized prompts; and

receiving, from output of the LLM, the set of paraphrased optimized prompts;

wherein selecting the one or more optimized prompts comprises selecting from at least one of the set of optimized prompts or the set of paraphrased optimized prompts.

17 . The system of claim 16 , wherein at least one of the feedback prompt, the editing prompt, or the paraphrasing prompt is at least one of generated or optimized by the LLM.

18 . The system of claim 16 , wherein at least one of providing the feedback prompt, providing the editing prompt, providing the paraphrasing prompt, or providing the sampled subset of optimized prompts is performed using an application programming interface (“API”) call to the LLM.

19 . The system of claim 16 , wherein selecting the one or more optimized prompts is performed using one or more selection algorithms.

20 . The system of claim 19 , wherein selecting the one or more optimized prompts comprises one of selecting a first number of the one or more optimized prompts that have scores above the average score or selecting a remaining number of the one or more optimized prompts after removing a second number of the one or more optimized prompts that have scores below the average score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2023
From: PRYZANT, REID ALLEN; ZHENG LI, JERRY; ITER, DAN; LEE, YIN TAT; ZHU, CHENGUANG; ZENG, NANSHAN; SHIRGAONKAR, ANUP
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065602/0457 →
Continuity (1)
Related Publication 20250111147A1 · Apr 3, 2025
References Cited (40)
US 20220189474A1 · Sharifi · 2022 [cited by examiner]
US 20230237277A1 · Reza · 2023 [cited by applicant]
US 20250085688A1 · Carrara · 2025 [cited by examiner]
US 20250103818A1 · Bolcer · 2025 [cited by examiner]
Sordoni, Alessandro, Eric Yuan, Marc-Alexandre Côté, Matheus Pereira, Adam Trischler, Ziang Xiao, Arian Hosseini, Friederike Niedtner, and Nicolas Le Roux. “Joint prompt optimization of stacked Ilms using variational in… [cited by examiner]
International Search Report and Written Opinion received for PCT Application No. PCT/IB2024/000597, Mar. 5, 2025, 12 pages. [cited by applicant]
Pryzant, et al., “Automatic Prompt Optimization with Gradient Descent” and Beam Search, arXiv:2305.03495 v1, May 4, 2023, 11 pages. [cited by applicant]
“Auto-GPT”, Retrieved From: https://news.agpt.co/, Retrieved On: May 9, 2023, 6 Pages. [cited by applicant]
“GPT-4 Technical Report”, In Repository of arXiv:2303.08774v1, Mar. 15, 2023, 99 Pages. [cited by applicant]
Abu-Farha, et al., “From Arabic Sentiment Analysis to Sarcasm Detection: The ArSarcasm Dataset”, In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Lan… [cited by applicant]
Audibert, et al., “Best Arm Identification in Multi-Armed Bandits”, In Proceedings of 23rd Conference on Learning Theory, Jun. 27, 2010, 13 Pages. [cited by applicant]
Bird, Steven, “NLTK: The Natural Language Toolkit”, In Proceedings of COLING/ACL Interactive Presentation Sessions, Jul. 2006, pp. 69-72. [cited by applicant]
Bubeck, et al., “Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems”, In Journal of Foundations and Trends in Machine Learning, vol. 5, Issue 1, Dec. 12, 2012, 124 Pages. [cited by applicant]
Bubeck, et al., “Sparks of Artificial General Intelligence: Early Experiments with GPT-4”, In Repository of arXiv:2303.12712v1, Mar. 22, 2023, 154 Pages. [cited by applicant]
Chen, et al., “Teaching Large Language Models to Self-Debug”, In Repository of arXiv:2304.05128v1, Apr. 11, 2023, pp. 1-68. [cited by applicant]
Deng, et al., “RLPROMPT: Optimizing Discrete Text Prompts with Reinforcement Learning”, In Repository of arXiv:2205.12548v1, May 25, 2022, 18 Pages. [cited by applicant]
Gao, et al., “Making Pre-trained Language Models Better Few-shot Learners”, In Repository of arXiv:2012.15723v1, Dec. 31, 2020, 15 Pages. [cited by applicant]
Prasad, et al., “GRIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models”, In Repository of arXiv:2203.07281v1, Mar. 14, 2022, 18 Pages. [cited by applicant]
Pryzant, et al., “Automatic Prompt Optimization with “Gradient Descent” and Beam Search”, In Repository of arXiv:2305.03495v1, May 4, 2023, 11 Pages. [cited by applicant]
Qin, et al., “Learning How to Ask: Querying LMs with Mixtures of Soft Prompts”, In Repository of arXiv:2104.06599v1, Apr. 14, 2021, 11 Pages. [cited by applicant]
Reynolds, et al., “Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm”, In Proceedings of CHI Conference on Human Factors in Computing Systems, May 8, 2021, pp. 1-7. [cited by applicant]
Shin, et al., “Autoprompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts”, In Repository of arXiv:2010.15980v1, Oct. 29, 2020, 14 Pages. [cited by applicant]
Wang, Williamy. , ““Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection”, In Repository of arXiv:1705.00648v1, May 1, 2017, 5 Pages. [cited by applicant]
Wang, et al., “Self-Instruct: Aligning Language Models with Self-Generated Instructions”, In Repository of arXiv:2212.10560v1, Dec. 20, 2022, 19 Pages. [cited by applicant]
Xu, et al., “GPS: Genetic Prompt Search for Efficient Few-Shot Learning”, In Repository of arXiv:2210.17041v1, Oct. 31, 2022, 11 Pages. [cited by applicant]
Zamfirescu-Pereira, et al., “Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts”, In Proceedings of CHI Conference on Human Factors in Computing Systems, Apr. 23, 2023, 21 Pages. [cited by applicant]
Zeng, et al., “Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language”, In Repository of arXiv:2204.00598v1, Apr. 1, 2022, 20 Pages. [cited by applicant]
Zhang, et al., “Tempera: Test-time Prompt Editing via Reinforcement Learning”, In Proceedings of Eleventh International Conference on Learning Representations, May 1, 2023, pp. 1-16. [cited by applicant]
Zhou, et al., “Large Language Models are Human-Level Prompt Engineers”, In Repository of arXiv:2211.01910v1, Nov. 3, 2022, pp. 1-40. [cited by applicant]
Hambardzumyan, et al., “WARP: Word-level Adversarial ReProgramming”, In Repository of arXiv:2101.00121v1, Jan. 1, 2021, 5 Pages. [cited by applicant]
Hao, et al., “Optimizing Prompts for Text-to-Image Generation”, In Repository of arXiv:2212.09611v1, Dec. 19, 2022, pp. 1-10. [cited by applicant]
Honovich, et al., “Instruction Induction: From Few Examples to Natural Language Task Descriptions”, In Repository of arXiv:2205.10782v1, May 22, 2022, pp. 1-17. [cited by applicant]
Jiang, et al., “PromptMaker: Prompt-based Prototyping with Large Language Models”, In Proceedings of CHI Conference on Human Factors in Computing Systems, Apr. 29, 2022, 8 Pages. [cited by applicant]
Karnin, et al., “Almost Optimal Exploration in Multi-Armed Bandits”, In Proceedings of the 30th International Conference on Machine Learning, Jun. 16, 2013, 9 Pages. [cited by applicant]
Kuleshov, et al., “Algorithms for the Multi-Armed Bandit Problems”, In Repository of arXiv:1402.6028v1, Feb. 25, 2014, pp. 1-32. [cited by applicant]
Lester, et al., “The Power of Scale for Parameter-Efficient Prompt Tuning”, In Repository of arXiv:2104.08691v1, Apr. 18, 2021, 13 Pages. [cited by applicant]
Li, et al., “Competition-Level Code Generation with AlphaCode”, In Repository of arXiv:2203.07814v1, Feb. 8, 2022, 74 Pages. [cited by applicant]
Lu, et al., “Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity”, In Repository of arXiv:2104.08786v1, Apr. 18, 2021, 11 Pages. [cited by applicant]
Mollas, et al., “Ethos: An Online Hate Speech Detection Dataset”, In Repository of arXiv:2006.08328v1, Jun. 11, 2020, 8 Pages. [cited by applicant]
International Preliminary Report on Patentability (Chapter I) Received for PCT Application No. PCT/IB2024/000597, Mailed on Apr. 9, 2026, 07 Pages. [cited by applicant]