IP Library Granted Patent US 11,714,849
Granted Patent B2
US 11,714,849 · App. 17/894,090 · Granted Aug 1, 2023

Image generation system and method

Inventors: Huiling Zhou (Hangzhou, CN); Jinbao Xue (Hangzhou, CN); Zhikang Li (Hangzhou, CN); Jie Liu (Hangzhou, CN); Shuai Bai (Beijing, CN); Chang Zhou (Hangzhou, CN); Hongxia Yang (Hangzhou, CN); Jingren Zhou (West Lafayette, IN)
Assignee: Alibaba Damo (Hangzhou) Technology Co., Ltd.
G06F16/583G06F16/5866G06F18/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,714,849
App. No.
17/894,090
Filed
Aug 23, 2022
Granted
Aug 1, 2023
Kind
B2
Art Unit
2665
USPC
382/100
Abstract

Embodiments of this application provide an image generation system and method. In an exemplary manufacturing industry scenario, a style requirement of a product category in a manufacturing industry is automatically captured according to user behavior data and product description information associated with the product category. Based on these data, a style description text may be generated and converted to product images by using a text prediction-based image generation model. The product images are further screened by using an image-text matching model, to obtain a product image with high quality. This process covers from style description text mining to text-to-image prediction to image quality evaluation. It provides an automation product image generation capability for the manufacturing industry, shorten a cycle of designing and producing the product image in the manufacturing industry, and improve production efficiency of the product image.

Claims (62)

1. An image generation method, comprising:

generating, according to user behavior data associated with a specified object category and object description information of the specified object category in a manufacturing industry, a style description text for the specified object category, wherein the style description text reflects a style requirement of the specified object category;

inputting the style description text into a text prediction-based first image generation model for image generation, to obtain a plurality of initial object images, wherein the first image generation model comprises an encoder-decoder structure implemented based on a vector quantization generative adversarial network (VQGAN) and a sparse attention mechanism, wherein the inputting the style description text into the text prediction-based first image generation model to obtain the plurality of initial object images comprises:

inputting a text sequence corresponding to the style description text into the first image generation model, and generating a plurality of image sequences according to the text sequence based on a codebook, wherein the codebook represents a quantization text representation of the image sequences; and

respectively performing image reconstruction on the plurality of image sequences, to obtain the plurality of initial object images; and

inputting the plurality of initial object images and the style description text into a second image-text matching model for matching, to obtain at least one candidate object image of which a matching degree meets a threshold requirement.

2. The method according to claim 1 , wherein the generating a plurality of image sequences according to the text sequence based on a codebook comprises:

inputting the text sequence into an encoder of the first image generation model, and encoding the text sequence, to obtain a first image feature; and

inputting the first image feature into a decoder of the first image generation model, and respectively decoding the first image feature based on the codebook, to obtain the plurality of image sequences.

3. The method according to claim 2 , wherein the respectively decoding the first image feature based on the codebook, to obtain the plurality of image sequences comprises:

decoding the first image feature based on the codebook and by using a sparse attention mechanism in the decoder of the first image generation model, to obtain the plurality of image sequences.

4. The method according to claim 3 , wherein a length of each of the image sequences is greater than or equal to 4096.

5. The method according to claim 2 , wherein the respectively performing image reconstruction on the plurality of image sequences, to obtain the plurality of initial object images comprises:

inputting the plurality of image sequences into an image sequence model, and respectively performing, by a decoder of the image sequence model, image reconstruction on the plurality of image sequences, to obtain the plurality of initial object images, wherein

the image sequence model adopts an encoder-decoder structure, and the image sequence model is a neural network model obtained by training the codebook.

6. The method according to claim 5 , further comprising:

acquiring a plurality of cross-field original sample images; and

performing model training by using a vector quantization generative adversarial network and using the plurality of original sample images, to obtain the image sequence model and the codebook.

7. The method according to claim 6 , wherein the plurality of original sample images comprise a first sample group and a second sample group, and the performing model training by using a vector quantization generative adversarial network and using the plurality of original sample images, to obtain the image sequence model and the codebook comprises:

performing non-adversarial training on an initial model by using the original sample images in the first sample group, to obtain an image sequence model in an intermediate state; and

performing adversarial training on the image sequence model in the intermediate state by using the vector quantization generative adversarial network and using the original sample images in the second sample group, to obtain the image sequence model and the codebook.

8. The method according to claim 1 , wherein the inputting the plurality of initial object images and the style description text into a second image-text matching model for matching, to obtain at least one candidate object image of which a matching degree meets a threshold requirement comprises:

inputting the plurality of initial object images and the style description text into the second image-text matching model, wherein the second image-text matching model is configured to respectively perform feature encoding on the plurality of initial object images and the style description text and map the plurality of initial object images and the style description text to a same semantic space, to obtain a plurality of second image features and a text feature; and

selecting, according to matching degrees between the plurality of second image features and the text feature, at least one initial object image of which a matching degree is greater than a threshold from the plurality of initial object images as the candidate object image.

9. The method according to claim 1 , wherein the generating, according to user behavior data associated with a specified object category and object description information of the specified object category in a first manufacturing industry, a style description text for the specified object category comprises:

performing text mining on the user behavior data associated with the specified object category, to obtain an object property and a category description in which a user is interested;

performing text mining on description information of a new object that appears within a latest time period in the specified object category, to obtain an object property and a category description of the new object;

obtaining category property data from a category-property-value knowledge system of the first manufacturing industry according to the object property and the category description in which the user is interested and the object property and the category description of the new object, wherein the category property data comprises at least a commodity style property; and

generating the style description text according to the category property data.

10. The method according to claim 1 , further comprising:

displaying the at least one candidate object image to an evaluation system, obtaining a selected target object image in response to selection of the evaluation system, and using the target object image for a subsequent manufacturing link.

11. A system for image generation, comprising one or more processors and one or more non-transitory computer-readable memories coupled to the one or more processors and configured with instructions executable by the one or more processors to cause the system to perform operations comprising:

generating, according to user behavior data associated with a specified object category and object description information of the specified object category in a manufacturing industry, a style description text for the specified object category, wherein the style description text reflects a style requirement of the specified object category;

inputting the style description text into a text prediction-based first image generation model for image generation, to obtain a plurality of initial object images, wherein the first image generation model comprises an encoder-decoder structure implemented based on a vector quantization generative adversarial network (VQGAN) and a sparse attention mechanism; and

inputting the plurality of initial object images and the style description text into a second image-text matching model for matching, to obtain at least one candidate object image of which a matching degree meets a threshold requirement, wherein the inputting the plurality of initial object images and the style description text into the second image-text matching model to obtain at least one candidate object image comprises:

inputting the plurality of initial object images and the style description text into the second image-text matching model, wherein the second image-text matching model is configured to respectively perform feature encoding on the plurality of initial object images and the style description text and map the plurality of initial object images and the style description text to a same semantic space, to obtain a plurality of second image features and a text feature; and

selecting, according to matching degrees between the plurality of second image features and the text feature, at least one initial object image of which a matching degree is greater than a threshold from the plurality of initial object images as the candidate object image.

12. The system of claim 11 , wherein the inputting the style description text into the text prediction-based first image generation model to obtain the plurality of initial object images comprises:

inputting a text sequence corresponding to the style description text into the first image generation model, and generating a plurality of image sequences according to the text sequence based on a codebook, wherein the codebook represents a quantization text representation of the image sequences; and

respectively performing image reconstruction on the plurality of image sequences, to obtain the plurality of initial object images.

13. The system of claim 12 , wherein the generating a plurality of image sequences according to the text sequence based on a codebook comprises:

inputting the text sequence into an encoder of the first image generation model, and encoding the text sequence, to obtain a first image feature; and

inputting the first image feature into a decoder of the first image generation model, and respectively decoding the first image feature based on the codebook, to obtain the plurality of image sequences.

14. The system of claim 13 , wherein the respectively decoding the first image feature based on the codebook, to obtain the plurality of image sequences comprises:

decoding the first image feature based on the codebook and by using a sparse attention mechanism in the decoder of the first image generation model, to obtain the plurality of image sequences.

15. The system of claim 13 , wherein the respectively performing image reconstruction on the plurality of image sequences, to obtain the plurality of initial object images comprises:

inputting the plurality of image sequences into an image sequence model, and respectively performing, by a decoder of the image sequence model, image reconstruction on the plurality of image sequences, to obtain the plurality of initial object images, wherein

the image sequence model adopts an encoder-decoder structure, and the image sequence model is a neural network model obtained by training the codebook.

16. The system of claim 15 , wherein the operations further comprise:

acquiring a plurality of cross-field original sample images; and

performing model training by using a vector quantization generative adversarial network and using the plurality of original sample images, to obtain the image sequence model and the codebook.

17. The system of claim 16 , wherein the plurality of original sample images comprise a first sample group and a second sample group, and the performing model training by using a vector quantization generative adversarial network and using the plurality of original sample images, to obtain the image sequence model and the codebook comprises:

performing non-adversarial training on an initial model by using the original sample images in the first sample group, to obtain an image sequence model in an intermediate state; and

performing adversarial training on the image sequence model in the intermediate state by using the vector quantization generative adversarial network and using the original sample images in the second sample group, to obtain the image sequence model and the codebook.

18. The system of claim 11 , wherein the operations further comprise:

displaying the at least one candidate object image to an evaluation system, obtaining a selected target object image in response to selection of the evaluation system, and using the target object image for a subsequent manufacturing link.

19. A computer-implemented method, comprising:

generating, according to user behavior data associated with a specified object category and object description information of the specified object category in a manufacturing industry, a style description text for the specified object category, wherein the style description text reflects a style requirement of the specified object category;

inputting the style description text into a text prediction-based first image generation model for image generation, to obtain a plurality of initial object images, wherein the first image generation model comprises an encoder-decoder structure implemented based on a vector quantization generative adversarial network (VQGAN) and a sparse attention mechanism;

inputting the plurality of initial object images and the style description text into a second image-text matching model for matching, to obtain at least one candidate object image of which a matching degree meets a threshold requirement; and

displaying the at least one candidate object image to an evaluation system, obtaining a selected target object image in response to selection of the evaluation system, and using the target object image for a subsequent manufacturing link.

20. The computer-implemented method of claim 19 , wherein the manufacturing industry is a clothing industry, a printing industry, an articles-for-daily-use industry, a furniture industry, an appliance industry, or a passenger car industry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2022
From: ZHOU, HUILING; XUE, JINBAO; LI, ZHIKANG; LIU, JIE; BAI, SHUAI; ZHOU, CHANG; YANG, HONGXIA; ZHOU, JINGREN
To: ALIBABA DAMO (HANGZHOU) TECHNOLOGY CO., LTD.
Reel/Frame 060876/0431 →
Priority Claims (1)
CN 202111015905.2 · Aug 31, 2021 · national
Continuity (1)
Related Publication 20230068103A1 · Mar 2, 2023
Cited By (2)
US 12,646,347 US 12,731,381