IP Library › Granted Patent US 12,724,965
Granted Patent B2
US 12,724,965 · App. 18/510,612 · Granted Sep 1, 2026

Reliable gradient-free and likelihood-free prompt tuning

Inventors: Maohao Shen (Cambridge, MA); Soumya Ghosh (Boston, MA); Prasanna Sattigeri (Acton, MA); Subhro Das (Cambridge, MA); Yuheng Bu (Philadelphia, PA); Gregory Wornell (Wellesley, MA)
Assignees: INTERNATIONAL BUSINESS MACHINES CORPORATION; Massachusetts Institute of Technology
G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,965
App. No.
18/510,612
Granted
Sep 1, 2026
Kind
B2
Abstract

Prompt embedding samples are drawn from a prior distribution and are passed into a pretrained model to receive a corresponding token label prediction for a batch of text data. Prompt embedding samples are accepted from a distribution of a first iteration; the accepted samples satisfy a condition of a distance function between a ground truth label and the corresponding token label prediction being less than a first tolerance. Embeddings are resampled from the accepted prompt embedding samples with probability proportional to weights and the resampled embeddings are perturbed via a perturbation kernel to obtain a new sample. The perturbed resampled embeddings are propagated through the pretrained model, and those that satisfy a condition are projected, where the second tolerance is decayed by one step per iteration. The projected resampled embeddings are concatenated with an embedding of a given input and inferencing is performed.

Claims (39)

1 . A method comprising:

drawing, using at least one hardware processor, prompt embedding samples from a prior distribution and passing the prompt embedding samples into a pretrained model to receive a corresponding token label prediction for a batch of text data;

accepting, using the at least one hardware processor, prompt embedding samples from a distribution of a first iteration that satisfy a condition of a distance function between a ground truth label and a corresponding token label prediction being less than a first tolerance;

resampling, using the at least one hardware processor, in a next iteration, embeddings from the accepted prompt embedding samples with probability proportional to weights and perturbing the resampled embeddings via a perturbation kernel to obtain a new sample;

propagating, using the at least one hardware processor, the perturbed resampled embeddings through the pretrained model;

projecting, to a higher dimension than a dimension of the resampled embeddings, using the at least one hardware processor, the resampled embeddings that satisfy a condition of the distance function between the ground truth label and the corresponding token label prediction being less than a second tolerance, where the second tolerance is decayed by one step per iteration;

concatenating the projected resampled embeddings with an embedding of a given input; and

performing inferencing by inputting the concatenated embeddings into the pretrained model.

2 . The method of claim 1 , wherein a number of the prompt embedding samples is designated as S and wherein the second tolerance is decayed by one step per iteration by subtracting an inverse of a total number of training data from the second tolerance.

3 . The method of claim 2 , wherein accuracy is used as the distance function.

4 . The method of claim 2 , wherein a final collection of prompt embedding samples form an approximation to a posterior.

5 . The method of claim 2 , wherein the given input is a text string and the inferencing operation determines a sentiment of the text string.

6 . A computer program product, comprising:

one or more tangible computer-readable storage media and program instructions stored on at least one of the one or more tangible computer-readable storage media, the program instructions executable by a processor, the program instructions comprising:

drawing, using at least one hardware processor, prompt embedding samples from a prior distribution and passing the prompt embedding samples into a pretrained model to receive a corresponding token label prediction for a batch of text data;

accepting, using the at least one hardware processor, prompt embedding samples from a distribution of a first iteration that satisfy a condition of a distance function between a ground truth label and the corresponding token label prediction being less than a first tolerance;

resampling, using the at least one hardware processor in a next iteration, embeddings from the accepted prompt embedding samples with probability proportional to weights and perturbing the resampled embeddings via a perturbation kernel to obtain a new sample;

propagating, using the at least one hardware processor, the perturbed resampled embeddings through the pretrained model;

projecting to a higher dimension than a dimension of the resampled embeddings, using the at least one hardware processor, the resampled embeddings that satisfy a condition of the distance function between the ground truth label and the corresponding token label prediction being less than a second tolerance, where the second tolerance is decayed by one step per iteration;

concatenating the projected resampled embeddings with an embedding of a given input; and

performing inferencing by inputting the concatenated embeddings into the pretrained model.

7 . The computer program product of claim 6 , wherein a number of the prompt embedding samples is designated as S and wherein the second tolerance is decayed by one step per iteration by subtracting an inverse of a total number of training data from the second tolerance.

8 . The computer program product of claim 7 , wherein accuracy is used as the distance function.

9 . The computer program product of claim 7 , wherein a final collection of prompt embedding samples form an approximation to a posterior.

10 . The computer program product of claim 7 , wherein the given input is a text string and the inferencing operation determines a sentiment of the text string.

11 . A system comprising:

a memory; and

at least one processor, coupled to said memory, and operative to perform operations comprising:

drawing, using at least one hardware processor, prompt embedding samples from a prior distribution and passing the prompt embedding samples into a pretrained model to receive a corresponding token label prediction for a batch of text data;

accepting, using the at least one hardware processor, prompt embedding samples from a distribution of a first iteration that satisfy a condition of a distance function between a ground truth label and the corresponding token label prediction being less than a first tolerance;

resampling, using the at least one hardware processor in a next iteration, embeddings from the accepted prompt embedding samples with probability proportional to weights and perturbing the resampled embeddings via a perturbation kernel to obtain a new sample;

propagating, using the at least one hardware processor, the perturbed resampled embeddings through the pretrained model;

projecting to a higher dimension than a dimension of the resampled embeddings, using the at least one hardware processor, the resampled embeddings that satisfy a condition of the distance function between the ground truth label and the corresponding token label prediction being less than a second tolerance, where the second tolerance is decayed by one step per iteration;

concatenating the projected resampled embeddings with an embedding of a given input; and

performing inferencing by inputting the concatenated embeddings into the pretrained model.

12 . The system of claim 11 , wherein a number of the prompt embedding samples is designated as S and wherein the second tolerance is decayed by one step per iteration by subtracting an inverse of a total number of training data from the second tolerance.

13 . The system of claim 12 , wherein accuracy is used as the distance function.

14 . The system of claim 12 , wherein a final collection of prompt embedding samples form an approximation to a posterior.

15 . The system of claim 12 , wherein the given input is a text string and the inferencing operation determines a sentiment of the text string.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2023
From: SHEN, MAOHAO; BU, YUHENG; WORNELL, GREGORY
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 065939/0750 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2023
From: GHOSH, SOUMYA; SATTIGERI, PRASANNA; DAS, SUBHRO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 065577/0989 →
Continuity (1)
Related Publication 20250156638A1 · May 15, 2025
References Cited (21)
US 20220237446A1 · Lei · 2022 [cited by examiner]
US 20230342559A1 · Bhardwaj · 2023 [cited by examiner]
US 20240005082A1 · Yin · 2024 [cited by examiner]
US 20240289550A1 · Wu · 2024 [cited by examiner]
US 20240311834A1 · Perez · 2024 [cited by examiner]
US 20250124620A1 · Arora · 2025 [cited by examiner]
CN 114661913A · 2022 [cited by applicant]
CN 114997149A · 2022 [cited by applicant]
CN 115204143A · 2022 [cited by applicant]
CN 115292484A · 2022 [cited by applicant]
CN 115311113A · 2022 [cited by applicant]
CN 115345300A · 2022 [cited by applicant]
CN 115423118A · 2022 [cited by applicant]
CN 115618006A · 2023 [cited by applicant]
CN 115795009A · 2023 [cited by applicant]
Shen M, Ghosh S, Sattigeri P, Das S, Bu Y, Wornell G. Reliable gradient-free and likelihood-free prompt tuning. arXiv preprint arXiv:2305.00593. Apr. 30, 2023. 14 pages (Grace Period Disclosure). [cited by applicant]
Lester et al., “The power of scale for parameter-efficient prompt tuning”, https://arxiv.org/pdf/2104.08691, Sep. 2, 2021, 15 pages. [cited by applicant]
Liu et al., “P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks”, https://arxiv.org/pdf/2110.07602, Mar. 20, 2022, 8 pages. [cited by applicant]
Shin et al., “AUTOPROMPT: Eliciting Knowledge from Language Models with Automatically Generated Prompts”, https://arxiv.org/pdf/2010.15980, Nov. 7, 2020, 15 pages. [cited by applicant]
Sun et al., “BBTv2: Pure Black-Box Optimization Can Be Comparable to Gradient Descent for Few-Shot Learning”, https://arxiv.org/pdf/2205.11200v1, May 23, 2022, 12 pages. [cited by applicant]
Sun et al., “Black-Box Tuning for Language-Model-as-a-Service”, https://arxiv.org/pdf/2201.03514, Jun. 27, 2022, 15 pages. [cited by applicant]