Model training based on synthetic data
In model training based on synthetic data, synthetic images are generated by providing respective text prompts into a text-to-image generation model. Respective training labels associated with the synthetic images are also generated based on the used text prompts. A target model, which is configured to perform an image classification task, is trained based at least in part on the synthetic images and the associated training labels. Through this solution, a large scale of synthetic images can be automatically obtained and applicable for training a model for image classification, to improve the model performance with data-scare setting or in the case of model pre-training where the training data amount matters.
1 . A method comprising:
generating a plurality of synthetic images by providing a plurality of text prompts into a text-to-image generation model;
generating respective training labels to be associated with the plurality of synthetic images based on the plurality of text prompts; and
performing training of a target model based at least in part on the plurality of synthetic images and the associated training labels, the target model being configured to perform an image classification task.
2 . The method of claim 1 , wherein performing the training of the target model comprises:
determining respective quality scores of the plurality of synthetic images;
filtering the plurality of synthetic images based on the determined respective quality scores, to discard at least one of the plurality of synthetic images; and
performing the training of the target model based at least in part on the filtered synthetic images and the associated training labels.
3 . The method of claim 2 , wherein determining the respective quality scores of the plurality of synthetic images comprises:
providing the plurality of synthetic images into a pre-trained image classification model, respectively, to obtain a plurality of classification confidences generated by the pre-trained image classification model for the plurality of synthetic images; and
determining the respective quality scores of the plurality of synthetic images based on the plurality of classification confidences.
4 . The method of claim 1 , wherein performing the training of the target model comprises:
performing the training of the target model using a soft cross-entropy (SCE) loss function.
5 . The method of claim 1 , wherein the image classification task is configured to distinguish between a plurality of classes with a plurality of class names, and the associated training labels at least indicate the plurality of class names.
6 . The method of claim 5 , further comprising:
generating the plurality of text prompts through at least one of the following:
filling at least one of the plurality of class names into a text template, to obtain at least one text prompt, or
providing at least one of the plurality of class names into a trained word-to-sentence model, to obtain at least one text prompt generated by the word-to-sentence model.
7 . The method of claim 5 , wherein the target model comprises a pre-trained model and the image classification task is a downstream task for the pre-trained model with zero-shot setting for the plurality of classes, and wherein performing the training of the target model comprises:
fine-tuning the pre-trained model based on the plurality of synthetic images and the associated training labels.
8 . The method of claim 5 , wherein the target model comprises a pre-trained model and the image classification task is a downstream task for the pre-trained model with few-shot learning for the plurality of classes, each class having a number of non-synthetic images with associated training labels, and wherein performing the training of the target model comprises:
fine-tuning the pre-trained model based on the plurality of synthetic images and the associated training labels, and the non-synthetic the associated training labels for the plurality of classes.
9 . The method of claim 8 , wherein fine-tuning the pre-trained model comprises:
filtering the plurality of synthetic images by:
for a first synthetic image associated with a first class, determining a first feature similarity between the first synthetic image and a first non-synthetic image in the first class, and a second feature similarity between the first synthetic image and a second non-synthetic image in a second class, and
in accordance with a determination that the second feature similarity is higher than the first feature similarity, discarding the first synthetic image; and
fine-tuning the pre-trained model based on the filtered synthetic images and the associated training labels, and the non-synthetic images with the associated training labels.
10 . The method of claim 8 , wherein the text-to-image generation model comprises a diffusion model, and generating the plurality of synthetic images comprises:
generating the plurality of synthetic images by further providing the non-synthetic images for the plurality of classes into the diffusion model as guidance information.
11 . The method of claim 8 , wherein fine-tuning the pre-trained model comprises:
fine-tuning the pre-trained model by:
performing a phase-wise training procedure on the pre-trained model, the non-synthetic images being used for training in a first phase of the phase-wise training procedure, and the plurality of synthetic images being used for training in a second phase of the phase-wise training procedure, or
performing a mix-training procedure on the pre-trained model by mixing the non-synthetic images and the plurality of synthetic images for training.
12 . The method of claim 8 , wherein the pre-trained model comprises at least one batch normalization (BN) layer, and wherein fine-tuning the pre-trained model comprises:
fine-tuning the pre-trained model with the at least one BN layer fixed.
13 . The method of claim 1 , wherein performing the training of the target model comprises:
pre-training the target model based at least in part on the plurality of synthetic images and the associated training labels, and
wherein the image classification task is a downstream task after the pre-training.
14 . A system, comprising:
at least one processor; and
at least one memory communicatively coupled to the at least one processor and comprising computer-readable instructions that upon execution by the at least one processor cause the at least one processor to perform acts comprising:
generating a plurality of synthetic images by providing a plurality of text prompts into a text-to-image generation model;
generating respective training labels to be associated with the plurality of synthetic images based on the plurality of text prompts; and
performing training of a target model based at least in part on the plurality of synthetic images and the associated training labels, the target model being configured to perform an image classification task.
15 . The system of claim 14 , wherein performing the training of the target model comprises:
determining respective quality scores of the plurality of synthetic images;
filtering the plurality of synthetic images based on the determined respective quality scores, to discard at least one of the plurality of synthetic images; and
performing the training of the target model based at least in part on the filtered synthetic images and the associated training labels.
16 . The system of claim 15 , wherein determining the respective quality scores of the plurality of synthetic images comprises:
providing the plurality of synthetic images into a pre-trained image classification model, respectively, to obtain a plurality of classification confidences generated by the pre-trained image classification model for the plurality of synthetic images; and
determining the respective quality scores of the plurality of synthetic images based on the plurality of classification confidences.
17 . The system of claim 14 , wherein the image classification task is configured to distinguish between a plurality of classes with a plurality of class names, and the associated training labels at least indicate the plurality of class names.
18 . The system of claim 17 , further comprising:
generating the plurality of text prompts through at least one of the following:
filling at least one of the plurality of class names into a text template, to obtain at least one text prompt, or
providing at least one of the plurality of class names into a trained word-to-sentence model, to obtain at least one text prompt generated by the word-to-sentence model.
19 . The system of claim 17 , wherein the target model comprises a pre-trained model and the image classification task is a downstream task for the pre-trained model with zero-shot setting for the plurality of classes, and wherein performing the training of the target model comprises:
fine-tuning the pre-trained model based on the plurality of synthetic images and the associated training labels; or
wherein the target model comprises a pre-trained model and the image classification task is a downstream task for the pre-trained model with few-shot learning for the plurality of classes, each class having a number of non-synthetic images with associated training labels, and wherein performing the training of the target model comprises:
fine-tuning the pre-trained model based on the plurality of synthetic images and the associated training labels, and the non-synthetic images the associated training labels for the plurality of classes.
20 . A non-transitory computer-readable storage medium, storing computer-readable instructions that upon execution by a device cause the device to perform acts comprising:
generating a plurality of synthetic images by providing a plurality of text prompts into a text-to-image generation model;
generating respective training labels to be associated with the plurality of synthetic images based on the plurality of text prompts; and
performing training of a target model based at least in part on the plurality of synthetic images and the associated training labels, the target model being configured to perform an image classification task.