IP Library › Granted Patent US 12,039,431
Granted Patent B1
US 12,039,431 · App. 18/475,588 · Granted Jul 16, 2024

Systems and methods for interacting with a multimodal machine learning model

Inventors: Noah Deutsch (San Francisco, CA); Nicholas Turley (San Francisco, CA); Benjamin Zweig (San Francisco, CA)
Assignee: OpenAI OpCo, LLC
G06N3/0455G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,039,431
App. No.
18/475,588
Granted
Jul 16, 2024
Kind
B1
Abstract

The disclosed embodiments may include a method of interacting with a multimodal machine learning model; the method may include providing a graphical user interface associated with a multimodal machine learning model. The method may further include displaying an image to a user in the graphical user interface. The method may also include receiving a textual prompt from the user and then generating input data using the image and the textual prompt. The method may further include generating an output at least in part by applying the input data to the multimodal machine learning model, the multimodal machine learning model configured using prompt engineering to identify a location in the image conditioned on the image and the textual prompt, wherein the output comprises a first location indication. The method may also include displaying, in the graphical user interface, an emphasis indicator at the indicated first location in the image.

Claims (82)

1. A method of interacting with a pre-trained multimodal machine learning model, the method comprising:

providing a graphical user interface configured to enable a user to interact with an image to generate a contextual prompt that indicates an area of emphasis in the image;

receiving the contextual prompt;

generating input data using the image and the contextual prompt;

generating a textual response to the image by applying the input data to a multimodal machine learning model configured to condition the textual response to the image on the contextual prompt; and

providing the textual response to the user,

wherein the textual response comprises a prompt suggestion and providing the textual response comprises displaying a selectable control in the graphical user interface configured to enable the user to select the prompt suggestion.

2. The method of claim 1 , wherein:

generating the input data using the image and the contextual prompt comprises:

generating an updated image based on the contextual prompt; and

generating the input data using the updated image.

3. The method of claim 1 , wherein:

generating the input data using the image and the contextual prompt comprises:

generating a segmentation mask by providing the image and the contextual prompt to a segmentation model; and

generating the input data using the image and the segmentation mask.

4. The method of claim 1 , wherein:

generating the input data using the image and the contextual prompt comprises:

generating a textual prompt or token using the contextual prompt; and

generating the input data using the image and the textual prompt or token.

5. The method of claim 4 , wherein:

the textual prompt or token indicates coordinates of a location in the image.

6. The method of claim 1 , wherein:

receiving the contextual prompt concerning the image comprises detecting a user human interface device interaction.

7. The method of claim 1 , wherein:

the graphical user interface includes an annotation tool; and

the contextual prompt concerning the image comprises an annotation generated using the annotation tool.

8. The method of claim 7 , wherein:

the annotation tool includes a loupe, a marker, or a segmentation tool.

9. The method of claim 7 , wherein:

the graphical user interface enables the user to resize an area of effect of the annotation tool.

10. The method of claim 1 , wherein:

the method further includes receiving a textual prompt from the user; the input data is further generated using the textual prompt; and

the multimodal machine learning model is configured to further condition the textual response to the image on the textual prompt.

11. The method of claim 1 , wherein:

the method further comprises, in response to selection of the control by the user:

generating second input data using the prompt suggestion and the image;

generating a second response by applying the second input data to the multimodal machine learning model; and

providing the second response to the user.

12. The method of claim 1 , wherein:

the contextual prompt indicates an object depicted in the image;

the textual response provides information about the depicted object; and

the textual response is displayed on the graphical user interface as a virtual button.

13. A system for interacting with a pre-trained multimodal machine learning model, comprising:

at least one processor; and

at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

providing a graphical user interface configured to enable a user to interact with an image to generate a contextual prompt that indicates an area of emphasis in the image;

receiving the contextual prompt;

generating input data using the image and the contextual prompt; generating a textual response to the image by applying the input data to the multimodal machine learning model configured to condition the textual response to the image on the contextual prompt; and

providing the textual response to the user,

wherein:

the textual response comprises a prompt suggestion, and

providing the textual response comprises displaying a selectable control in the graphical user interface, the selectable control being programmed to enable the user to select the prompt suggestion.

14. The system of claim 13 , wherein:

generating the input data using the image and the contextual prompt comprises:

generating an updated image based on the contextual prompt; and

generating the input data using the updated image.

15. The system of claim 13 , wherein:

generating the input data using the image and the contextual prompt comprises:

generating a segmentation mask by providing the image and the contextual prompt to a segmentation model; and

generating the input data using the image and the segmentation mask.

16. The system of claim 13 , wherein:

generating the input data using the image and the contextual prompt comprises:

generating a textual prompt or token using the contextual prompt, the textual prompt or token indicating coordinates of a location in the image; and

generating the input data using the image and the textual prompt or token.

17. The system of claim 13 , wherein:

the graphical user interface includes an annotation tool; and

the contextual prompt concerning the image comprises an annotation generated using the annotation tool.

18. The system of claim 17 , wherein:

the annotation tool includes a loupe, a marker, or a segmentation tool;

the graphical user interface enables the user to resize an area of effect of the annotation tool;

the operations further include receiving a textual prompt from the user;

the input data is further generated using the textual prompt; and

the multimodal machine learning model is configured to further condition the textual response to the image on the textual prompt.

19. The system of claim 13 , wherein:

the operations further comprise, in response to selection of the control by the user:

generating second input data using the prompt suggestion and the image;

generating a second response by applying the second input data to the multimodal machine learning model; and

providing the second response to the user.

20. The system of claim 19 , wherein:

the contextual prompt indicates an object depicted in the image;

the textual response provides information about the depicted object; and

the textual response is displayed on the graphical user interface as a virtual button.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2024
From: DEUTSCH, NOAH; TURLEY, NICHOLAS; ZWEIG, BENJAMIN
To: OPENAI OPCO LLC
Reel/Frame 066224/0358 →
Cited By (2)
US 12,613,902 US 12,748,804