IP Library Granted Patent US 12,613,892
Granted Patent B2
US 12,613,892 · App. 18/528,655 · Granted Apr 28, 2026

System and method for learning and communicating implicit stylistic preferences from historical user interaction data in text-to-image prompt engineering

Inventors: Matthew K. Hong (Los Altos, CA); Heishiro Toyoda (Los Altos, CA); Yin-Ying Chen (Los Altos, CA); Shabnam Hakimi (Los Altos, CA); Matthew Klenk (Los Altos, CA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
G06F16/3328G06F16/535G06F16/54G06T11/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,613,892
App. No.
18/528,655
Granted
Apr 28, 2026
Kind
B2
Abstract

Systems and methods are provided for implementing stylistic preferences into an image generation model with an accompanying user interface. The system can generate a plurality of images for each of a plurality of user prompts received from a first user and relate each plurality of images to the other pluralities of images. The user interface can selecting a preferred plurality of images from the pluralities of images based on input from a second user and display the pluralities of images in a node tree diagram indicating the preferred plurality of images.

Claims (52)

1 . A method comprising:

generating a plurality of images for each of a plurality of user prompts received from a first user;

relating two or more of the pluralities of images and transmitting the related pluralities of images to a second user;

selecting a preferred plurality of images from the related pluralities of images based on input from the second user;

displaying the pluralities of images in a node tree diagram indicating the preferred plurality of images;

receiving input from the second user selecting disliked nodes from the node tree diagram; and

generating a new set of images based on the selected disliked nodes and the preferred plurality of images.

2 . The method of claim 1 , further comprising:

determining an additional user prompt from the first user corresponds to the preferred plurality of images;

generating an additional plurality of images; and

adding the additional plurality of images to the node tree diagram as being related to the preferred plurality of images.

3 . The method of claim 2 , further comprising determining that the additional plurality of images is a second preferred plurality of images based on user input from the second user.

4 . The method of claim 1 , wherein the second user selects the preferred plurality of images from the node tree diagram.

5 . The method of claim 1 , wherein a machine learning model generates the pluralities of images based on the plurality of user prompts.

6 . The method of claim 1 , further comprising determining that a plurality of images is a disliked plurality of images based on user input from the second user.

7 . The method of claim 6 , further comprising attributing a positive weight to the preferred plurality of images and attributing a negative weight to the disliked plurality of images.

8 . The method of claim 7 , further comprising removing the disliked plurality of images from the node tree diagram based on the negative weight.

9 . A user interface, comprising:

a processor; and

a memory encoded with instructions, which when executed by the processor, causes the processor to:

generate a plurality of images for each of a plurality of user prompts received from a first user;

generate a node tree diagram displaying the pluralities of images based on one or more relationships between the pluralities of images and transmitting the node tree diagram to a second user;

attribute a weight to a disliked plurality of images based on receiving input from a second user selecting a node associated with the disliked plurality of images from the node tree diagram;

remove the node associated with the disliked plurality of images from the node tree diagram based on the weight; and

generate a new set of images based on the second user's selecting the node associated with the disliked plurality of images.

10 . The user interface of claim 9 , wherein a machine learning model generates the pluralities of images based on the plurality of user prompts.

11 . The user interface of claim 9 , wherein the processor is further configured to determine a preferred plurality of images based on a selection from the second user.

12 . The user interface of claim 11 , wherein the processor is further configured to:

generate additional pluralities of images based on the preferred plurality of images;

determine preferred or disliked pluralities of images based on selections from the second user;

attribute a weight to each selection;

add preferred pluralities of images to the node tree diagram; and

remove disliked pluralities of images from the node tree diagram.

13 . The user interface of claim 9 , wherein attributing the weight to the disliked plurality of images is based on the disliked plurality's relationship to other disliked pluralities of images.

14 . The user interface of claim 13 , wherein the processor is further configured to update the weight of the disliked plurality of images as the other disliked pluralities of images are added to the node tree diagram.

15 . The user interface of claim 9 , wherein removing the disliked plurality of images from the node tree diagram is based on the weight exceeding a threshold.

16 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to:

generate a plurality of images for each of a plurality of user prompts received from a first user;

relate two or more of the pluralities of images and transmit the related pluralities of images to a second user;

attribute a positive weight to a preferred plurality of images based on user selection from the second user;

generate a new set of images based on the second user's interactions with a node tree diagram and the preferred plurality of images;

receive input from the second user selecting disliked nodes from the node tree diagram; and

display all pluralities of images in a node tree diagram indicating the preferred plurality of images and the positive weight.

17 . The non-transitory machine-readable medium of claim 16 , wherein the processor is further configured to:

determine an additional user prompt from the first user corresponds to the preferred plurality of images;

generate an additional plurality of images; and

add the additional plurality of images to the node tree diagram as being related to the preferred plurality of images.

18 . The non-transitory machine-readable medium of claim 17 , wherein the processor is further configured to determine that the additional plurality of images is a second preferred plurality of images based on additional user input from the second user.

19 . The non-transitory machine-readable medium of claim 16 , wherein the second user selects the preferred plurality of images from the node tree diagram.

20 . The non-transitory machine-readable medium of claim 16 , wherein the processor is further configured to:

attribute a negative weight to a disliked plurality of images based on additional input from the second user selecting the disliked plurality of images from the pluralities of images; and

display the disliked plurality of images on the node tree diagram with an indication of the negative weight.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2026
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 074785/0424 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2023
From: HONG, MATTHEW K.; TOYODA, HEISHIRO; CHEN, YIN-YING; HAKIMI, SHABNAM; KLENK, MATTHEW
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 065756/0919 →
Continuity (1)
Related Publication 20250181612A1 · Jun 5, 2025
References Cited (18)
US 12154021B1 · Sternberg · 2024 [cited by examiner]
US 20120296897A1 · Xin-Jing · 2012 [cited by applicant]
US 20200226411A1 · Gupta · 2020 [cited by examiner]
US 20230105483A1 · Rittman · 2023 [cited by examiner]
US 20230177810A1 · Xu · 2023 [cited by applicant]
US 20240193821A1 · Denison · 2024 [cited by examiner]
US 20240320867A1 · Bean · 2024 [cited by examiner]
US 20240338860A1 · Trzyna · 2024 [cited by examiner]
CN 110532571 · 2019 [cited by applicant]
CN 110751698 · 2020 [cited by applicant]
CN 113837229 · 2021 [cited by applicant]
CN 114638905 · 2022 [cited by applicant]
CN 115018941 · 2022 [cited by applicant]
CN 115700519 · 2023 [cited by applicant]
CN 116246062 · 2023 [cited by applicant]
WO 2019052403 · 2019 [cited by applicant]
Paananen et al., “Using Text-to-Image Generation for Architectural Design Ideation,” arXiv:2304.10182v1 [cs.HC], Apr. 20, 2023, 14 pages (https://doi.org/10.48550/arXiv.2304.10182). [cited by applicant]
Soatto et al., “Taming AI Bots: Controllability of Neural States in Large Language Models,” arXiv:2305.18449v1 [cs.AI], May 29, 2023, 30 pages (https://doi.org/10.48550/arXiv.2305.18449). [cited by applicant]