IP Library Granted Patent US 12,423,821
Granted Patent B2
US 12,423,821 · App. 18/742,069 · Granted Sep 23, 2025

Systems and methods for interacting with a large language model

Inventors: Noah Deutsch (San Francisco, CA); Benjamin Zweig (San Francisco, CA)
Assignee: OpenAI OPCo, LLC
G06T7/10G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,821
App. No.
18/742,069
Granted
Sep 23, 2025
Kind
B2
Abstract

Disclosed embodiments may include a method of interacting with a multimodal machine learning model; the method may include providing a graphical user interface associated with a multimodal machine learning model. The method may further include displaying an image to a user in the graphical user interface. The method may also include receiving a textual prompt from the user and then generating input data using the image and the textual prompt. The method may further include generating an output at least in part by applying the input data to the multimodal machine learning model, the multimodal machine learning model configured using prompt engineering to identify a location in the image conditioned on the image and the textual prompt, wherein the output includes a first location indication. The method may also include displaying, in the graphical user interface, an emphasis indicator at the indicated first location in the image.

Claims (58)

1. A computer-implemented method comprising:

receiving a prompt associated with a graphical user interface (GUI);

generating input data using a snapshot of the GUI and the prompt, the input data generated in a format usable by a machine learning model, wherein generating input data comprises:

tokenizing the snapshot of the GUI and the prompt to generate a tokenized snapshot of the GUI and a tokenized prompt; and

concatenating the tokenized snapshot of the GUI and the tokenized prompt into a singular tokenized input;

generating an output by applying the input data to the machine learning model, the machine learning model being configured to identify a location in the GUI based on the prompt, the output comprising a location indication within the GUI; and

generating instructions to display a cursor at the location in the GUI.

2. The method of claim 1 , wherein the prompt is a textual prompt.

3. The method of claim 1 , wherein:

the location is associated with a GUI element, the GUI element being a selectable button; and

the method further comprises using the machine learning model to activate the cursor to select the button within the GUI.

4. The method of claim 2 , wherein the textual prompt includes an action parameter, wherein the action parameter specifies an action within the GUI.

5. The method of claim 4 , wherein the machine learning model is configured using prompt engineering to identify the action parameter.

6. The method of claim 1 , wherein:

the location is associated with a GUI element, the GUI element being a clickable element; and

the method further comprises using the machine learning model to activate the cursor to select the clickable element within the GUI.

7. The method of claim 1 , wherein generating the input data further comprises combining the snapshot of the GUI with a spatial encoding.

8. The method of claim 7 , wherein the location indication comprises coordinates within the spatial encoding or a token corresponding to coordinates within the spatial encoding.

9. The method of claim 1 , wherein:

the output comprises a second location indication within the GUI;

the method further comprises generating instructions to display a second cursor at the second location in the GUI, wherein the second location is associated with a second GUI element, the second GUI element being a selectable element; and

activating, using the machine learning model, the second cursor to select the selectable element at the second location within the GUI.

10. A system, comprising:

at least one processor; and

at least one memory containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

receiving a prompt associated with a graphical user interface (GUI);

generating input data using a snapshot of the GUI, the input data generated in a format usable by a machine learning model, wherein generating input data comprises:

tokenizing the snapshot of the GUI and the prompt to generate a tokenized snapshot of the GUI and a tokenized prompt; and

concatenating the tokenized snapshot of the GUI and the tokenized prompt into a singular tokenized input;

generating an output by applying the input data to the machine learning model, the machine learning model being configured to identify a location in the GUI based on the prompt, the output comprising a location indication; and

generating instructions to display an indicator at the location.

11. The system of claim 10 , wherein the prompt is a textual prompt.

12. The system of claim 10 , wherein:

the location is associated with a GUI element, the GUI element being a selectable button; and

the operations further comprise using the machine learning model to activate a cursor to select the button within the GUI.

13. The system of claim 11 , wherein the textual prompt includes an action parameter, wherein the action parameter specifies an action within the GUI.

14. The system of claim 13 , wherein the machine learning model is configured using prompt engineering to identify the action parameter.

15. The system of claim 10 , wherein:

the location is associated with a GUI element, the GUI element being a clickable element; and

the operations further comprise using the machine learning model to activate a cursor to select the clickable element with the GUI.

16. The system of claim 10 , wherein generating the input data further comprises combining the snapshot of the GUI with a spatial encoding.

17. The system of claim 16 , wherein the location indication comprises coordinates within the spatial encoding or a token corresponding to coordinates within the spatial encoding.

18. The system of claim 10 , wherein:

the output comprises a second location indication;

the operations further comprise generating instructions to display a cursor at the second location, wherein the second location is associated with a second GUI element, the second GUI element being a selectable button; and

activating, using the machine learning model, the cursor to select the selectable button at the second location within the GUI.

19. A machine learning system comprising:

one or more memory devices storing instructions; and

one or more processors coupled to the one or more memory devices and configured to:

receive a prompt associated with a graphical user interface (GUI);

generate input data using a snapshot of the GUI, the input data generated in a format usable by a machine learning model, wherein generating input data comprises:

tokenizing the snapshot of the GUI and the prompt to generate a tokenized snapshot of the GUI and a tokenized prompt; and

concatenating the tokenized snapshot of the GUI and the tokenized prompt into a singular tokenized input;

generating an output by applying the input data to the machine learning model, the machine learning model being configured to identify a location in the GUI based on the prompt, the output comprising a location indication; and

generating instructions to display an indicator at the location.

20. The machine learning system of claim 19 , wherein:

the location is associated with a GUI element, the GUI element being a selectable button; and

the one or more processors are configured to use the machine learning model to activate the cursor to select the button within the GUI.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2024
From: DEUTSCH, NOAH; ZWEIG, BENJAMIN
To: OPENAI OPCO LLC
Reel/Frame 067909/0954 →
Continuity (2)
Continuation 18475722 · Sep 27, 2023
Related Publication 20250104243A1 · Mar 27, 2025
References Cited (27)
US 6549929B1 · Sullivan · 2003 [cited by examiner]
US 6674905B1 · Matsugu · 2004 [cited by examiner]
US 7239740B1 · Fujieda · 2007 [cited by examiner]
US 9098313B2 · Butin · 2015 [cited by examiner]
US 9405558B2 · Butin · 2016 [cited by examiner]
US 10089742B1 · Lin · 2018 [cited by examiner]
US 11314982B2 · Price et al. · 2022 [cited by applicant]
US 11568627B2 · Price et al. · 2023 [cited by applicant]
US 11615567B2 · Harikumar · 2023 [cited by examiner]
US 11995298B1 · Rubinstein · 2024 [cited by examiner]
US 12051205B1 · Deutsch et al. · 2024 [cited by applicant]
US 20020064382A1 · Hildreth · 2002 [cited by examiner]
US 20040240708A1 · Hu · 2004 [cited by examiner]
US 20070273658A1 · Yli-Nokari · 2007 [cited by examiner]
US 20100135566A1 · Joanidopoulos · 2010 [cited by examiner]
US 20140359521A1 · Lin · 2014 [cited by examiner]
US 20170308388A1 · Gerphagnon · 2017 [cited by examiner]
US 20180268548A1 · Lin · 2018 [cited by examiner]
US 20190236394A1 · Price · 2019 [cited by examiner]
US 20200143194A1 · Hou et al. · 2020 [cited by applicant]
US 20200234084A1 · Seaton · 2020 [cited by examiner]
US 20210248748A1 · Turgutlu et al. · 2021 [cited by applicant]
US 20210383171A1 · Lee · 2021 [cited by examiner]
US 20210390700A1 · Lee · 2021 [cited by examiner]
US 20220050661A1 · Lange · 2022 [cited by examiner]
Peng Xu et al., Multimodal Learning with Transformers: A Survey, arXiv:2206.06488, May 10, 2023. [cited by applicant]
International Search Report for PCT application No. PCT /US2024/048703, dated Dec. 24, 2024, 4 pages. [cited by applicant]