IP Library › Granted Patent US 12,602,846
Granted Patent B2
US 12,602,846 · App. 18/236,346 · Granted Apr 14, 2026

Generating realistic machine learning-based product images for online catalogs

Inventors: Prithvishankar Srinivasan (Seattle, WA); Shih-Ting Lin (Santa Clara, CA); Min Xie (Santa Clara, CA); Shishir Kumar Prasad (Fremont, CA); Yuanzheng Zhu (Smyrna, GA); Katie Ann Forbes (Austin, TX)
Assignee: Maplebear Inc.
G06T11/60G06F16/55G06F16/583G06Q30/0643G06T11/203
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,846
App. No.
18/236,346
Granted
Apr 14, 2026
Kind
B2
Abstract

An online concierge system trains a fine-tuned generative image model for distinct categories of items based on a generative image model that takes a textual query as input and outputs and an associated image. Training of the fine-tuned generative image model is additionally based on a small set of representative images associated with the various categories, as well as textual tokens associated with the categories. Once trained, the fine-tuned generative image model can be used to generate realistic representative images for items in a database of the online concierge system that are lacking associated images. The fine-tuned model permits the generation of different variants of an item, such as different quantities or amounts, different packaging or packing density, and the like.

Claims (64)

1 . A method performed at a computer system comprising a processor and a computer-readable medium, the method for generating and managing synthetic images of items and comprising:

accessing a generative diffusion model that accepts a textual query as input and generates a synthetic image as output;

for each of a plurality of categories from a taxonomy of item categories:

generating a token corresponding to the category and not corresponding to others of the item categories;

generating a fine-tuned generative diffusion model for the categories using the generative diffusion model, sets of images corresponding to the categories, and the tokens corresponding to the categories, the fine-tuned generative diffusion model accepting a textual query as input and generating a synthetic image as output;

identifying an item in an order lacking a corresponding representative image in an database;

generating a representative image for the item using the fine-tuned generative diffusion model by:

identifying a category, from the plurality of categories arranged within a hierarchy of taxonomy, corresponding to the item in an order lacking the corresponding representative image,

identifying a token associated with that category,

generating a textual query for the category, wherein the textual query includes the token, and

specifying the textual query as input to the fine-tuned generative diffusion model;

storing the generated representative image in the database in association with the item; and

transmitting, to a picker device associated with a picker fulfilling the order through obtainment of items in the order, the order inclusive of the generated representative image for the item for display of the generated representative image on the picker device.

2 . The method of claim 1 , further comprising:

receiving a query of a customer for items from the database;

identifying items of the database corresponding to the query, the items including the item for which the representative image was generated; and

causing display of data about the items, the data including the representative image in visual association with the item.

3 . The method of claim 1 , wherein the textual query contains the token corresponding to the category of the item lacking the corresponding representative image.

4 . The method of claim 1 , wherein the textual query contains at least one of: an amount of the item, a quantity of the item, a packaging density of the item, or a packaging type of the item.

5 . The method of claim 1 , wherein the token is randomly generated.

6 . The method of claim 1 , wherein the fine-tuned generative diffusion model is generated using DREAMBOOTH.

7 . A non-transitory computer-readable storage medium containing instructions that when executed by one or more processors perform actions comprising:

accessing a generative diffusion model that accepts a textual query as input and generates a synthetic image as output;

for each of a plurality of categories from a taxonomy of item categories:

generating a token corresponding to the category and not corresponding to any of the other item categories;

generating a fine-tuned generative diffusion model for the categories using the generative diffusion model, sets of images corresponding to the categories, and the tokens corresponding to the categories, the fine-tuned generative diffusion model accepting a textual query as input and generating a synthetic image as output;

identifying an item lacking a corresponding representative image in a database;

generating a representative image for the item using the fine-tuned generative diffusion model by:

identifying a category, from the plurality of categories arranged within a hierarchy of taxonomy, corresponding to the item in an order lacking the corresponding representative image,

identifying a token associated with that category,

generating a textual query for the category, wherein the textual query includes the token, and

specifying the textual query as input to the fine-tuned generative diffusion model;

storing the generated representative image in the database in association with the item; and

transmitting, to a picker device associated with a picker fulfilling the order through obtainment of items in the order, the order inclusive of the generated representative image for the item for display of the generated representative image on the picker device.

8 . The non-transitory computer-readable storage medium of claim 7 , the actions further comprising:

receiving a query of a customer for items from the database;

identifying items of the database corresponding to the query, the items including the item for which the representative image was generated; and

causing display of data about the items, the data including the representative image in visual association with the item.

9 . The non-transitory computer-readable storage medium of claim 7 , wherein the textual query contains the token corresponding to the category of the item lacking the corresponding representative image.

10 . The non-transitory computer-readable storage medium of claim 7 , wherein the textual query contains at least one of: an amount of the item, a quantity of the item, a packaging density of the item, or a packaging type of the item.

11 . The non-transitory computer-readable storage medium of claim 7 , wherein the token is randomly generated.

12 . The non-transitory computer-readable storage medium of claim 7 , wherein the fine-tuned generative diffusion model is generated using DREAMBOOTH.

13 . A computer system comprising:

one or more computer processors; and

a computer-readable storage medium storing instructions that when executed by the one or more computer processors perform actions comprising:

accessing a generative diffusion model that accepts a textual query as input and generates a synthetic image as output;

for each of a plurality of categories from a taxonomy of item categories:

generating a token corresponding to the category and not corresponding to any of the other item categories;

generating a fine-tuned generative diffusion model for the categories using the generative diffusion model, sets of images corresponding to the categories, and the tokens corresponding to the categories, the fine-tuned generative diffusion model accepting a textual query as input and generating a synthetic image as output;

identifying an item lacking a corresponding representative image in a database;

generating a representative image for the item using the fine-tuned generative diffusion model by:

identifying a category, from the plurality of categories arranged within a hierarchy of taxonomy, corresponding to the item in an order lacking the corresponding representative image,

identifying a token associated with that category,

generating a textual query for the category, wherein the textual query includes the token, and

specifying the textual query as input to the fine-tuned generative diffusion model;

storing the generated representative image in the database in association with the item; and

transmitting, to a picker device associated with a picker fulfilling the order through obtainment of items in the order, the order inclusive of the generated representative image for them for display of the generated representative image on the picker device.

14 . The computer system of claim 13 , the actions further comprising:

receiving a query of a customer for items from the database;

identifying items of the database corresponding to the query, the items including the item for which the representative image was generated; and

causing display of data about the items, the data including the representative image in visual association with the item.

15 . The computer system of claim 13 , wherein the textual query contains the token corresponding to the category of the item lacking the corresponding representative image.

16 . The computer system of claim 13 , wherein the textual query contains at least one of: an amount of the item, a quantity of the item, a packaging density of the item, or a packaging type of the item.

17 . The computer system of claim 13 , wherein the fine-tuned diffusion model is generated using DREAMBOOTH.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2023
From: SRINIVASAN, PRITHVISHANKAR; LIN, SHIH-TING; XIE, MIN; PRASAD, SHISHIR KUMAR; ZHU, YUANZHENG; FORBES, KATIE ANN
To: MAPLEBEAR INC.
Reel/Frame 064709/0189 →
Continuity (1)
Related Publication 20250069298A1 · Feb 27, 2025
References Cited (11)
US 20180253840A1 · Tran · 2018 [cited by applicant]
US 20180357554A1 · Hazan et al. · 2018 [cited by applicant]
US 20190327328A1 · Smith et al. · 2019 [cited by applicant]
US 20220122001A1 · Choe et al. · 2022 [cited by applicant]
US 20220147838A1 · Gu et al. · 2022 [cited by applicant]
US 20220237368A1 · Tran · 2022 [cited by applicant]
US 20230377099A1 · Kreis · 2023 [cited by examiner]
US 20240264723A1 · Davidson · 2024 [cited by examiner]
US 20250005281A1 · Yuan · 2025 [cited by examiner]
US 20250014246A1 · Difonzo · 2025 [cited by examiner]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US2024/032619, Sep. 10, 2024, 12 pages. [cited by applicant]