IP Library Granted Patent US 12685939
Granted Patent B2
US 12685939 · App. 18/166,413 · Granted Jul 21, 2026

Cascading throughout an image dynamic user feedback responsive to the AI generated image

Inventors: Warren Benedetto (San Mateo, CA); Michael Taylor (San Bruno, CA); Jon Webb (San Mateo, CA)
Assignee: SONY INTERACTIVE ENTERTAINMENT INC.
G06F3/0482G06F3/04845G06T11/00G06T11/60G10L15/18G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12685939
App. No.
18/166,413
Granted
Jul 21, 2026
Kind
B2
Abstract

A method including generating an image using an image generation artificial intelligence system configured for implementing latent diffusion, wherein the image is decoded from a latent space representation. The method including receiving selection of a portion of the image. The method including receiving commentary of a user corresponding to the portion of the image. The method including generating a text prompt to modify the portion of the image based on the commentary. The method including encoding the portion of the image using latent diffusion based on the text prompt. The method including modifying the latent space representation of the image using latent diffusion to generate a modified image, wherein the text prompt and the encoded portion of the image are provided as input to the image generation artificial intelligence system.

Claims (86)

1 . A method, comprising:

receiving selection of a first portion of an image, wherein the image is decoded from a latent space representation of the image;

receiving user input corresponding to the first portion of the image;

identifying at least one scene physics property associated with the first portion of the image;

modifying the latent space representation of the image by:

modifying the latent space representation of the image corresponding to the first portion of the image based on the user input and the at least one scene physics property; and

responsive to modifying the latent space representation of the image corresponding to the first portion, automatically modifying the latent space representation of the image corresponding to a second portion of the image based on modifications to the latent space representation of the first portion of the image, wherein the second portion is distinct from the first portion, and wherein the modification to the latent space representation corresponding to the second portion is in accordance with the at least one scene physics property; and

decoding the modified latent space representation of the image to generate a modified image,

wherein the modified image includes one or more modifications to the second portion of the image.

2 . The method of claim 1 , wherein the receiving the user input includes:

receiving the user input formatted in a natural language; and

translating the user input into a text format.

3 . The method of claim 1 , wherein the user input includes:

receiving an indication of positive preference or negative preference associated with the first portion of the image.

4 . The method of claim 1 ,

wherein the first portion of the image is tagged using the user input to indicate the selection of the first portion of the image.

5 . The method of claim 1 , further comprising:

displaying the image in a user interface; and

providing for tagging of the first portion of the image to indicate the selection of the first portion of the image;

receiving the user input via the user interface; and

displaying the modified image in the user interface.

6 . The method of claim 1 , further comprising:

changing a scene physics property of the first portion of the image based on a text prompt; and

applying the scene physics property that is changed throughout remaining portions of the image.

7 . The method of claim 1 , further comprising:

providing a word cloud corresponding to the first portion of the image in a user interface; and

receiving at least one modification to the word cloud via the user interface,

wherein a text prompt is generated based on the word cloud that is modified.

8 . The method of claim 7 , further comprising:

receiving selection of a word in the word cloud; and

providing a drop down menu of items corresponding to the word that is selected,

wherein selection of each item in the menu of items applies a corresponding modification to the word that is selected.

9 . A non-transitory computer-readable medium storing instructions that, upon execution on a computer system, cause the computer system to perform operations comprising:

receiving selection of a first portion of an image, wherein the image is decoded from a latent space representation of the image;

receiving user input corresponding to the first portion of the image;

identifying at least one scene physics property associated with the first portion of the image;

modifying the latent space representation of the image by:

modifying the latent space representation of the image corresponding to the first portion of the image based on the user input and the at least one scene physics property; and

responsive to modifying the latent space representation of the image corresponding to the first portion, automatically modifying the latent space representation of the image corresponding to a second portion of the image based on modifications to the latent space representation of the first portion of the image, wherein the second portion is distinct from the first portion, and wherein the modification to the latent space representation corresponding to the second portion is in accordance with the at least one scene physics property; and

decoding the modified latent space representation of the image to generate a modified image,

wherein the modified image includes one or more modifications to the second portion of the image.

10 . The non-transitory computer-readable medium of claim 9 , storing further instructions that upon execution on the computer system cause the computer system to perform further operations comprising:

receiving the user input formatted in a natural language; and

translating the user input into a text format.

11 . The non-transitory computer-readable medium of claim 9 , storing further instructions that upon execution on the computer system cause the computer system to perform further operations comprising:

receiving an indication of positive preference or negative preference associated with the first portion of the image.

12 . The non-transitory computer-readable medium of claim 9 , storing further instructions that upon execution on the computer system cause the computer system to perform further operations comprising:

displaying the image in a user interface; and

providing for tagging of the first portion of the image to indicate the selection of the first portion of the image;

receiving the user input via the user interface; and

displaying the modified image in the user interface.

13 . The non-transitory computer-readable medium of claim 9 , storing further instructions that upon execution on the computer system cause the computer system to perform further operations comprising:

changing a scene physics property of the first portion of the image based on a text prompt; and

applying the scene physics property that is changed throughout remaining portions of the image.

14 . The non-transitory computer-readable medium of claim 9 , storing further instructions that upon execution on the computer system cause the computer system to perform further operations comprising:

providing a word cloud corresponding to the first portion of the image in a user interface; and

receiving at least one modification to the word cloud via the user interface,

wherein a text prompt is generated based on the word cloud that is modified.

15 . A computer system comprising:

a processor;

memory coupled to the processor and having stored therein instructions that, upon execution by the processor, configure the computer system to:

receive selection of a first portion of an image, wherein the image is decoded from a latent space representation of the image;

receive user input corresponding to the first portion of the image;

identify at least one scene physics property associated with the first portion of the image;

modify the latent space representation of the image by:

modify the latent space representation of the image corresponding to the first portion of the image based on the user input and the at least one scene physics property; and

responsive to modifying the latent space representation of the image corresponding to the first portion, automatically modify the latent space representation of the image corresponding to a second portion of the image based on modifications to the latent space representation of the first portion of the image, wherein the second portion is distinct from the first portion, and wherein the modification to the latent space representation corresponding to the second portion is in accordance with the at least one scene physics property; and

decode the modified latent space representation of the image to generate a modified image,

wherein the modified image includes one or more modifications to the second portion of the image.

16 . The computer system of claim 15 , wherein the memory stores further instructions that, upon execution by the processor, further configure the computer system to:

receive the user input formatted in a natural language; and

translate the user input into a text format.

17 . The computer system of claim 15 , wherein the memory stores further instructions that, upon execution by the processor, further configure the computer system to:

receive an indication of positive preference or negative preference associated with the first portion of the image.

18 . The computer system of claim 15 , wherein the memory stores further instructions that, upon execution by the processor, further configure the computer system to:

display the image in a user interface; and

provide for tagging of the first portion of the image to indicate the selection of the first portion of the image;

receive the user input via the user interface; and

display the modified image in the user interface.

19 . The computer system of claim 15 , wherein the memory stores further instructions that, upon execution by the processor, further configure the computer system to:

change a scene physics property of the first portion of the image based on a text prompt; and

apply the scene physics property that is changed throughout remaining portions of the image.

20 . The computer system of claim 15 , wherein the memory stores further instructions that, upon execution by the processor, further configure the computer system to:

provide a word cloud corresponding to the first portion of the image in a user interface; and

provide at least one modification to the word cloud via the user interface,

wherein a text prompt is generated based on the word cloud that is modified.