IP Library Granted Patent US 11,934,792
Granted Patent B1
US 11,934,792 · App. 18/387,657 · Granted Mar 19, 2024

Identifying prompts used for training of inference models

Inventors: Yair Adato (Kfar Ben Nun, IL); Michael Feinstein (Tel Aviv, IL); Efrat Taig (Beer Sheva, IL); Dvir Yerushalmi (Kfar Saba, IL); Ori Liberman (Netanya, IL)
Assignee: BRIA ARTIFICIAL INTELLIGENCE LTD.
G06F40/30G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,934,792
App. No.
18/387,657
Granted
Mar 19, 2024
Kind
B1
Abstract

Systems, methods and non-transitory computer readable media for identifying prompts used for training of inference models are provided. In some examples, a specific textual prompt in a natural language may be received. Further, data based on at least one parameter of an inference model may be accessed. The inference model may be a result of training a machine learning model using a plurality of training examples. Each training example of the plurality of training examples may include a respective textual content and a respective media content. The data and the specific textual prompt may be analyzed to determine a likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples. A digital signal indicative of the likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples may be generated.

Claims (54)

1. A non-transitory computer readable medium storing a software program comprising data and computer implementable instructions that when executed by at least one processor cause the at least one processor to perform operations for identifying prompts used for training of inference models, the operations comprising:

receiving a specific textual prompt in a natural language;

accessing data based on at least one parameter of an inference model, the inference model is a result of training a machine learning model using a plurality of training examples, each training example of the plurality of training examples includes a respective textual content and a respective media content;

analyzing the data and the specific textual prompt to determine a likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples; and

generating a digital signal indicative of the likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples.

2. The non-transitory computer readable medium of claim 1 , wherein the likelihood is a binary likelihood, and wherein the determining the likelihood includes determining whether the specific textual prompt is included in at least one training example of the plurality of training examples.

3. The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

using the specific textual prompt and the at least one parameter of the inference model to obtain an output of the inference model corresponding to the specific textual prompt; and

analyzing the output to determine the likelihood.

4. The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

using the specific textual prompt to generate a plurality of variations of the specific textual prompt;

for each variation of the plurality of variations, using the variation and the at least one parameter of the inference model to obtain an output of the inference model corresponding to the variation; and

analyzing the outputs to determine the likelihood.

5. The non-transitory computer readable medium of claim 4 , wherein the operations further comprise:

obtaining a plurality of directions in a mathematical space;

obtaining a specific mathematical object in the mathematical space corresponding to the specific textual prompt;

for each direction of the plurality of directions, using the specific mathematical object and the direction to determine a mathematical object in the mathematical space corresponding to the specific mathematical object and the direction; and

for each direction of the plurality of directions, generating a textual prompt corresponding to the mathematical object corresponding to the specific mathematical object and the direction, thereby generating the plurality of variations of the specific textual prompt.

6. The non-transitory computer readable medium of claim 4 , wherein the operations further comprise:

selecting a plurality of utterances, no utterance of the plurality of utterances is included in the specific textual prompt; and

for each utterance in the plurality of utterances, analyzing the specific textual prompt to generate a variation of the specific textual prompt that includes the utterance, thereby generating the plurality of variations of the specific textual prompt.

7. The non-transitory computer readable medium of claim 4 , wherein the operations further comprise:

analyzing the specific textual prompt to detect a plurality of utterances included in the specific textual prompt; and

for each utterance in the plurality of utterances, analyzing the specific textual prompt to generate a variation of the specific textual prompt that do not include the utterance, thereby generating the plurality of variations of the specific textual prompt.

8. The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

using the specific textual prompt and the at least one parameter of the inference model to obtain a gradient, corresponding to the specific textual prompt, of a function associated with the inference model; and

analyzing the gradient to determine the likelihood.

9. The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

using the specific textual prompt and the at least one parameter of the inference model to calculate a loss corresponding to the specific textual prompt and to a loss function associated with the machine learning model; and

analyzing the loss to determine the likelihood.

10. The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

analyzing the data and the specific textual prompt to determine a second likelihood that a variation version of the specific textual prompt is included in at least one training example of the plurality of training examples; and

generating a second digital signal indicative of the second likelihood.

11. The non-transitory computer readable medium of claim 1 , wherein the inference model is a generative model, and wherein for each training example of the plurality of training examples, the respective textual content is a respective input textual prompt and the respective media content is a respective desired output media content.

12. The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

analyzing the data and the specific textual prompt to determine a measure of similarity of the specific textual prompt to the textual content included in a selected training example of the plurality of training examples; and

generating a second digital signal indicative of the measure of similarity of the specific textual prompt to the selected training example of the plurality of training examples.

13. The non-transitory computer readable medium of claim 12 , wherein the measure of similarity is indicative of an amount of variation.

14. The non-transitory computer readable medium of claim 12 , wherein the second digital signal is indicative of the selected training example of the plurality of training examples.

15. The non-transitory computer readable medium of claim 12 , wherein the textual content included in the selected training example is the most similar to specific textual prompt of all the textual contents included in the plurality of training examples.

16. The non-transitory computer readable medium of claim 1 , wherein the specific textual prompt includes a first noun, wherein the textual content included in a particular training example includes a second noun, and wherein the likelihood is based on the first noun and the second noun.

17. The non-transitory computer readable medium of claim 1 , wherein the specific textual prompt includes a particular noun and a first adjective adjacent to the particular noun, wherein the textual content included in a particular training example includes the particular noun and a second adjective adjacent to the particular noun, and wherein the likelihood is based on the first adjective and the second adjective.

18. The non-transitory computer readable medium of claim 1 , wherein the specific textual prompt includes a particular verb and a first adverb adjacent to the particular verb, wherein the textual content included in a particular training example includes the particular verb and a second adverb adjacent to the particular verb, and wherein the likelihood is based on the first adverb and the second adverb.

19. A method for identifying prompts used for training of inference models, the method comprising:

receiving a specific textual prompt in a natural language;

accessing data based on at least one parameter of an inference model, the inference model is a result of training a machine learning model using a plurality of training examples, each training example of the plurality of training examples includes a respective textual content and a respective media content;

analyzing the data and the specific textual prompt to determine a likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples; and

generating a digital signal indicative of the likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples.

20. A system for identifying prompts used for training of inference models, the system comprising:

at least one processor configured to perform the operations of:

receiving a specific textual prompt in a natural language;

accessing data based on at least one parameter of an inference model, the inference model is a result of training a machine learning model using a plurality of training examples, each training example of the plurality of training examples includes a respective textual content and a respective media content;

analyzing the data and the specific textual prompt to determine a likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples; and

generating a digital signal indicative of the likelihood that the specific textual prompt is included in at least one training example of the plurality of training examples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2024
From: ADATO, YAIR; FEINSTEIN, MICHAEL; TAIG, EFRAT; YERUSHALMI, DVIR; LIBERMAN, ORI
To: BRIA ARTIFICIAL INTELLIGENCE LTD.
Reel/Frame 067265/0060 →
Continuity (3)
Continuation PCTIL2023051132 · Nov 5, 2023
Provisional Application 63525754 · Jul 10, 2023
Provisional Application 63444805 · Feb 10, 2023
Cited By (9)
US 12,271,696 US 12,307,208 US 12,431,131 US 12,488,186 US 12,541,488 US 12,626,695 US 12,688,370 US 12,705,427 US 12,724,752