IP Library › Granted Patent US 12,530,913
Granted Patent B2
US 12,530,913 · App. 18/171,256 · Granted Jan 20, 2026

Qualifying labels automatically attributed to content in images

Inventor: Arran Green (La Mesa, CA)
Assignee: SONY INTERACTIVE ENTERTAINMENT INC.
G06V20/70G06F3/0482G06F3/167G06T7/70G06T17/00G06V10/44G06V10/764G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,913
App. No.
18/171,256
Granted
Jan 20, 2026
Kind
B2
Abstract

A method for image generation. The method including identifying a plurality of features of an image. The method including classifying each of the plurality of features using an artificial intelligence (AI) model trained to identify features in a plurality of images, wherein the plurality of features is classified as a plurality of labels, wherein the image is provided as input to the AI model. The method including receiving feedback for a label, wherein the feedback is associated with a user. The method including modifying a label based on the feedback. The method including updating the plurality of labels with the label that is modified. The method including providing as input the plurality of labels that is updated into an image generation artificial intelligence system configured for implementing latent diffusion to generate an updated image.

Claims (91)

1 . A method, comprising:

identifying a plurality of features of an image;

classifying each of the plurality of features using an artificial intelligence (AI) model trained to classify features in a plurality of images, wherein the plurality of features is classified as a plurality of labels, and wherein the image is provided as input to the AI model;

determining, based on a commentary, an object within the image to which the commentary applies;

determining one or more labels of the object to which the commentary applies;

translating the commentary into feedback for the one or more labels;

modifying the one or more labels based on the feedback;

updating the plurality of labels with the one or more labels that is modified; and

providing, as input, the plurality of labels that is updated into an image generation artificial intelligence system configured for implementing latent diffusion to generate an updated image.

2 . The method of claim 1 , further comprising:

generating the image using the image generation artificial intelligence system.

3 . The method of claim 1 , further comprising:

receiving the feedback via a user interface,

wherein the feedback is formatted in text.

4 . The method of claim 1 , further comprising:

receiving the feedback as audio, wherein the feedback is presented in natural language;

converting the audio to text; and

presenting the text via a user interface.

5 . The method of claim 1 , further comprising:

receiving identification of an object within a scene that is presented on a display;

presenting one or more labels of the object in a user interface via the display, wherein the one or more labels of the object includes the label; and

receiving identification of the label by a user via the user interface.

6 . The method of claim 5 , wherein the receiving identification of the object includes:

determining that the user is pointing to a location in physical space corresponding to a location of the object within the scene in virtual space, wherein the scene is presented on the display of a head mounted display worn by the user; and

determining that the user is pointing to the object within the scene based on the pointing.

7 . The method of claim 5 , wherein the receiving identification of the object includes:

determining that the user selects the object in the scene using a controller.

8 . The method of claim 1 , further comprising:

presenting a plurality of labels of a plurality of objects of a scene of the image in a user interface on a display, wherein the plurality of objects are presented in the user interface as a hierarchical file system of objects;

receiving selection of an object via the hierarchical file system;

presenting one or more labels of the object in the user interface via the display, wherein the one or more labels of the object includes the label; and

receiving identification of the label by the user via the user interface.

9 . The method of claim 1 , further comprising:

highlighting the object that is presented on a display.

10 . The method of claim 1 , further comprising:

determining that a user is pointing to a location in physical space corresponding to a location of an object within a scene in virtual space, wherein the scene is presented on a display of a head mounted display worn by the user;

highlighting the object in the scene;

determining that the user is selecting the object based on the pointing;

receiving commentary to modify the object from the user, wherein the commentary is presented in natural language;

determining one or more labels of the object, wherein the one or more labels of the object includes the label;

determining that the commentary applies to the label; and

translating the commentary into the feedback for the label.

11 . The method of claim 1 , further comprising;

determining a context based on the feedback for the label; and

modifying one or more of the plurality of labels based on the context,

wherein the plurality of labels that is updated includes one or more of the plurality of labels that have been modified.

12 . The method of claim 1 , further comprising:

adding a new object corresponding to the label based on the feedback;

determining a context based on the new object; and

modifying one or more of the plurality of labels based on the context,

wherein the plurality of labels that is updated includes one or more of the plurality of labels that have been modified.

13 . The method of claim 1 , further comprising:

removing an object corresponding to the label based on the feedback; and

removing the label from the plurality of labels when performing the modifying the label and when performing the updating the plurality of labels.

14 . A non-transitory computer-readable medium having instructions stored thereon, which, when executed by a processor, cause the processor to perform a method comprising:

identifying a plurality of features of an image;

classifying each of the plurality of features using an artificial intelligence (AI) model trained to classify features in a plurality of images, wherein the plurality of features is classified as a plurality of labels, wherein the image is provided as input to the AI model;

determining, based on a commentary, an object within the image to which the commentary applies;

determining one or more labels of the object to which the commentary applies;

translating the commentary into feedback for the one or more labels;

modifying the one or more labels a label based on the feedback;

updating the plurality of labels with the one or more labels that is modified; and

providing as input the plurality of labels that is updated into an image generation artificial intelligence system configured for implementing latent diffusion to generate an updated image.

15 . The non-transitory computer-readable medium of claim 14 , further comprising instructions that, when executed, cause the processor to perform the method comprising:

receiving identification of an object within a scene that is presented on a display;

presenting one or more labels of the object in a user interface via the display, wherein the one or more labels of the object includes the label; and

receiving identification of the label by a user via the user interface.

16 . The non-transitory computer-readable medium of claim 14 , further comprising instructions that, when executed, cause the processor to perform the method comprising:

presenting a plurality of labels of a plurality of objects of a scene of the image in a user interface on a display, wherein the plurality of objects are presented in the user interface as a hierarchical file system of objects;

receiving selection of an object via the hierarchical file system;

presenting one or more labels of the object in the user interface via the display, wherein the one or more labels of the object includes the label; and

receiving identification of the label by the user via the user interface.

17 . The non-transitory computer-readable medium of claim 14 , further comprising instructions that, when executed, cause the processor to perform the method comprising:

highlighting the object that is presented on a display.

18 . A computer system comprising:

a processor; and

memory coupled to the processor and having stored therein instructions that, if executed by the computer system, cause the computer system to execute a method for implementing a graphics pipeline, comprising:

identifying a plurality of features of an image;

classifying each of the plurality of features using an artificial intelligence (AI) model trained to classify features in a plurality of images, wherein the plurality of features is classified as a plurality of labels, and wherein the image is provided as input to the AI model;

determining, based on a commentary, an object within the image to which the commentary applies;

determining one or more labels of the object to which the commentary applies;

translating the commentary into feedback for the one or more labels;

modifying the one or more labels based on the feedback;

updating the plurality of labels with the one or more labels that is modified; and

providing, as input, the plurality of labels that is updated into an image generation artificial intelligence system configured for implementing latent diffusion to generate an updated image.

19 . The computer system of claim 18 , the method further comprising:

receiving identification of an object within a scene that is presented on a display;

presenting one or more labels of the object in a user interface via the display, wherein the one or more labels of the object includes the label; and

receiving identification of the label by a user via the user interface.

20 . The computer system of claim 18 , the method further comprising:

highlighting the object that is presented on a display.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: GREEN, ARRAN
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 063642/0488 →
Continuity (1)
Related Publication 20240282130A1 · Aug 22, 2024
References Cited (11)
US 10551845B1 · Kim · 2020 [cited by examiner]
US 10732783B2 · Yang · 2020 [cited by examiner]
US 11361018B2 · Gulati · 2022 [cited by examiner]
US 20170185236A1 · Yang · 2017 [cited by examiner]
US 20190163768A1 · Gulati · 2019 [cited by examiner]
US 20200272862A1 · Gardner · 2020 [cited by examiner]
US 20210232872A1 · Ries · 2021 [cited by examiner]
CN 114880441B · 2023 [cited by applicant]
ISR WO PCT/US2024/015595, dated May 27, 2024, Total 12 pages. [cited by applicant]
Robin Rombach et al.: “High-Resolution Image Synthesis with Latent Diffusion Models”, Jun. 1, 2022 (Jun. 1, 2022), Ludwig Maximillian University of Munich, XP0034194085, pp. 1-45, cited in the application the whole docu… [cited by applicant]
Ramesh Aditya et al.: DALL.E: Creating images from text:, Jan. 5, 2021 (Jan. 5, 2021), pp. 1-18, XP093134452, Retrieved from the Internet: URL:https://openai.com/research/dall-e [retrieved on Feb. 23, 2024] the whole do… [cited by applicant]