IP Library › Granted Patent US 12,354,196
Granted Patent B2
US 12,354,196 · App. 18/214,559 · Granted Jul 8, 2025

Controllable diffusion model based image gallery recommendation service

Inventors: Hong Xuan (Bellevue, WA); Li Huang (Sammamish, WA); Huangxing Li (Bellevue, WA); Xi Chen (Issaquah, WA)
Assignee: Microsoft Technology Licensing, LLC
G06T11/60G06F16/532G06F16/535G06T7/13G06T7/73G06T2200/24G06T2207/20092G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,196
App. No.
18/214,559
Granted
Jul 8, 2025
Kind
B2
Abstract

Aspects of the disclosure include methods and systems for leveraging a controllable diffusion model for dynamic image search in an image gallery recommendation service. An exemplary method can include displaying an image gallery having a plurality of gallery images and a dynamic image frame. The dynamic image frame can include a generated image and an interactive widget. The method can include receiving a user input in the interactive widget and generating, responsive to receiving the user input, an updated generated image by inputting, into a controllable diffusion model, the user input. The method can include replacing the generated image in the dynamic image frame with the updated generated image.

Claims (44)

1. A method comprising:

causing display of an image gallery comprising a plurality of gallery images, the plurality of gallery images comprising a collection of images matching an image query, the plurality of gallery images stored in an image database prior to receiving the image query, the image gallery further comprising a dynamic image frame comprising a generated image that is dynamically created in response to receiving the image query and an interactive widget for modifying the generated image, the plurality of gallery images and the generated image displayed concurrently in the image gallery;

the interactive widget including generated text corresponding to a topic of interest identified for the image query according to a topic recognition and constraint mapping applied to the image query;

receiving a user input in the interactive widget for modifying the generated image;

generating, responsive to receiving the user input, an updated generated image by inputting, into a controllable diffusion model, the user input;

replacing the generated image in the dynamic image frame with the updated generated image;

replacing, responsive to successive user inputs to modify the generated image, one or more of the plurality of gallery images that matched the image query with new gallery images selected according to learned characteristics of the updated generated image, the learned characteristics determined from the successive user inputs; and

replacing the user input with updated generated text corresponding to the updated generated image.

2. The method of claim 1 , further comprising receiving an image query in a field of the image gallery.

3. The method of claim 2 , wherein the generated image is generated by inputting, into the controllable diffusion model, the image query.

4. The method of claim 2 , wherein the plurality of gallery images and the generated image are selected according to a degree of matching to one or more features in the image query.

5. The method of claim 2 , further comprising determining one or more constraints in the image query.

6. The method of claim 5 , wherein the generated image is generated by inputting, into the controllable diffusion model, the one or more constraints.

7. The method of claim 5 , wherein the one or more constraints in the image query comprise at least one of a pose skeleton and an object boundary.

8. The method of claim 7 , wherein determining the one or more constraints comprises extracting the object boundary when a feature in the image query comprises one of a structure and a geological feature.

9. The method of claim 7 , wherein determining the one or more constraints comprises extracting the pose skeleton when a feature in the image query comprises one of a person and an animal.

10. The method of claim 1 , wherein the plurality of gallery images are sourced from an image database.

11. The method of claim 1 , wherein the interactive widget comprises a text field, and receiving the user input in the interactive widget comprises receiving a text string input into the text field.

12. The method of claim 1 , wherein the interactive widget comprises one or more of a dropdown menu, a checkbox, a slider, a color picker, a canvas interface for drawing or sketching, and a rating button.

13. The method of claim 1 , wherein the interactive widget comprises a canvas for magic wand inputs, and receiving the user input in the interactive widget comprises receiving a magic wand input graphically selecting one of a specific feature and a specific region in the generated image.

14. The method of claim 13 , wherein the interactive widget further comprises a text field, and receiving the user input in the interactive widget further comprises receiving a text string input having contextual information for the magic wand input.

15. A system having a memory, computer readable instructions, and one or more processors for executing the computer readable instructions, the computer readable instructions controlling the one or more processors to perform operations comprising:

receiving, from a client device communicatively coupled to the system, an image query;

providing, to the client device, a plurality of gallery images and a generated image according to a degree of matching to one or more features in the image query, the plurality of gallery images comprising a collection of images stored in an image database prior to receiving the image query, the generated image dynamically created in response to receiving the image query, the plurality of gallery images and the generated image displayed concurrently in an image gallery;

receiving, from an interactive widget of the client device, a user input for modifying the generated image, the interactive widget including generated text corresponding to a topic of interest identified for the image query according to a topic recognition and constraint mapping applied to the image query;

generating, responsive to the user input, an updated generated image by inputting, into a controllable diffusion model, the user input;

replacing, responsive to successive user inputs to modify the generated image, one or more of the plurality of gallery images that matched the image query with new gallery images selected according to learned characteristics of the updated generated image, the learned characteristics determined from the successive user inputs;

replacing the user input with updated generated text corresponding to the updated generated image; and

providing, to the client device, the updated generated image.

16. The system of claim 15 , wherein the generated image is generated by inputting, into the controllable diffusion model, the image query.

17. The system of claim 15 , further comprising determining one or more constraints in the image query.

18. The system of claim 17 , wherein the generated image is generated by inputting, into the controllable diffusion model, the one or more constraints.

19. A system having a memory, computer readable instructions, and one or more processors for executing the computer readable instructions, the computer readable instructions controlling the one or more processors to perform operations comprising:

receiving, from an image gallery recommendation service communicatively coupled to the system, a plurality of gallery images and a generated image, the plurality of gallery images comprising a collection of images stored in an image database prior to receiving an image query, the generated image dynamically created in response to receiving the image query;

displaying an image gallery comprising the plurality of gallery images and a dynamic image frame comprising the generated image and an interactive widget, the plurality of gallery images and the generated image displayed concurrently in the image gallery;

the interactive widget including generated text corresponding to a topic of interest identified for the image query according to a topic recognition and constraint mapping applied to the image query;

receiving a user input in the interactive widget for modifying the generated image;

transmitting the user input to the image gallery recommendation service;

receiving, from the image gallery recommendation service, an updated generated image;

replacing the generated image in the dynamic image frame with the updated generated image;

receiving, from the image gallery recommendation service, new gallery images matching the updated generated image;

replacing, responsive to successive user inputs to modify the generated image, one or more of the plurality of gallery images that matched the image query with new gallery images selected according to learned characteristics of the updated generated image, the learned characteristics determined from the successive user inputs; and

replacing the user input with updated generated text corresponding to the updated generated image.

20. The system of claim 19 , wherein the interactive widget comprises a text field, and receiving the user input in the interactive widget comprises receiving a text string input into the text field.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2023
From: HUANG, LI; XUAN, HONG; LI, HUANGXING; CHEN, XI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064227/0298 →
Continuity (1)
Related Publication 20250005822A1 · Jan 2, 2025
References Cited (27)
US 11769239B1 · Zhang · 2023 [cited by examiner]
US 11809688B1 · Parasnis · 2023 [cited by examiner]
US 11893713B1 · Zhang · 2024 [cited by examiner]
US 11908180B1 · Ho · 2024 [cited by examiner]
US 11922541B1 · Parasnis · 2024 [cited by examiner]
US 11922550B1 · Ramesh · 2024 [cited by examiner]
US 11928319B1 · Parasnis · 2024 [cited by examiner]
US 20160042252A1 · Sawhney · 2016 [cited by examiner]
US 20180253869A1 · Yumer · 2018 [cited by examiner]
US 20220236863A1 · Wilensky · 2022 [cited by examiner]
US 20230067841A1 · Saharia · 2023 [cited by examiner]
US 20230118966A1 · Liu · 2023 [cited by examiner]
US 20230177878A1 · Sekar · 2023 [cited by examiner]
US 20230214461A1 · Brooks · 2023 [cited by examiner]
US 20230267652A1 · Lupascu · 2023 [cited by examiner]
US 20230325975A1 · Zhi · 2023 [cited by examiner]
US 20230334834A1 · Bai · 2023 [cited by examiner]
US 20230377214A1 · Kansy · 2023 [cited by examiner]
US 20230377226A1 · Saharia · 2023 [cited by examiner]
US 20240037822A1 · Aberman · 2024 [cited by examiner]
US 20240062008A1 · Ghosh · 2024 [cited by examiner]
US 20240070816A1 · Jandial · 2024 [cited by examiner]
CN 115935817A · 2023 [cited by applicant]
“AI Art Generators—Ai Arts Lab”, Retrieved from the Internet URL: https://web.archive.org/web/20220926152921/https://aiartslab.com/ai-art-generators-compare-the-best-ai-image-makers/#gs.c919b2, Sep. 26, 2022, pp. 1-30. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/031517, Aug. 16, 2024, 18 pages. [cited by applicant]
Tiago Ferreira, “ImageJ User Guide—IJ 1.46r Tools”, Retrieved from the Internet URL: https://imagej.net/ij/docs/guide/146-19.html, Jun. 22, 2012, pp. 1-11. [cited by applicant]
Yang, et al., “Object matching with hierarchical skeletons”, Pattern Recognition, vol. 55, Jul. 2016, pp. 183-197. [cited by applicant]