IP Library Granted Patent US 11,978,245
Granted Patent B2
US 11,978,245 · App. 16/050,329 · Granted May 7, 2024

Method and apparatus for generating image

Inventors: Tao He (Beijing, CN); Gang Zhang (Beijing, CN); Jingtuo Liu (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G06V10/82G06F18/214G06F18/2413G06N3/044G06N3/045G06N3/047G06N3/08G06N3/084G06N7/00G06N20/00G06T7/251G06T7/97G06T11/00G06V10/764G06V40/165G06V40/167G06V40/168G06V40/172G06N7/01G06N20/10G06T2207/20081G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,978,245
App. No.
16/050,329
Granted
May 7, 2024
Kind
B2
Abstract

The present disclosure discloses a method and apparatus for generating an image. A specific embodiment of the method comprises: acquiring at least two frames of facial images extracted from a target video; and inputting the at least two frames of facial images into a pre-trained generative model to generate a single facial image. The generative model updates a model parameter using a loss function in a training process, and the loss function is determined based on a probability of the single facial generative image being a real facial image and a similarity between the single facial generative image and a standard facial image. According to this embodiment, authenticity of the single facial image generated by the generative model may be enhanced, and then a quality of a facial image obtained based on the video is improved.

Claims (53)

1. A method for generating an image, comprising:

acquiring at least two frames of facial images of a same person extracted from a target video; and

inputting the at least two frames of facial images of the same person into one pre-trained generative model to generate a single facial image, wherein the generative model is obtained through:

inputting a single facial generative image outputted by an initial generative model into a pre-trained discriminative model to generate a probability of the single facial generative image being a real facial image;

determining a loss function of the initial generative model based on the probability and a similarity between the single facial generative image and a standard facial image, the similarity including a Jaccard similarity coefficient between the single facial generative image and the standard facial image, or a Euclidean distance between the single facial generative image and the standard facial image, the standard facial image and the single facial generative image containing facial information of the same person; and

updating a model parameter of the initial generative model using the loss function determined based on the probability and the similarity to obtain the generative model,

wherein determining the loss function of the initial generative model comprises:

extracting respectively feature information of the single facial generative image and feature information of the standard facial image using a pre-trained recognition model, and calculating a Euclidean distance between the feature information of the single facial generative image and the feature information of the standard facial image; and

obtaining the loss function of the initial generative model according to the probability and the Euclidean distance, and

wherein the method is performed by at least one processor.

2. The method according to claim 1 , wherein the initial generative model is trained and obtained by:

using at least two frames of initial training facial sample images extracted from an initial training video as an input, and using a preset initial training facial image as an output by using a machine learning method, the at least two frames of initial training facial sample images and the initial training facial image containing the facial information of the given person.

3. The method according to claim 1 , wherein the discriminative model is trained and obtained by:

using a first sample image as an input and use annotation information of the first sample image as an output by using a machine learning method, the first sample image comprising a positive sample image with annotation information and a negative sample image with annotation information, wherein the negative sample image is an image outputted by the generative model.

4. The method according to claim 1 , wherein the recognition model is trained and obtained by:

using a second sample image as an input and feature information of the second sample image as an output by using a machine learning method.

5. The method according to claim 1 , wherein the generative model is a Long-Short Term Memory Model.

6. An apparatus for generating an image, comprising:

at least one processor; and

a memory storing instructions, the instructions when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

acquiring at least two frames of facial images of a same person extracted from a target video;

inputting the at least two frames of facial images of the same person into one pre-trained generative model to generate a single facial image; and

training the generative model,

wherein the generative model is obtained through:

inputting a single facial generative image outputted by an initial generative model into a pre-trained discriminative model to generate a probability of the single facial generative image being a real facial image;

determining a loss function of the initial generative model based on the probability and a similarity between the single facial generative image and a standard facial image, the similarity including a Jaccard similarity coefficient between the single facial generative image and the standard facial image, or a Euclidean distance between the single facial generative image and the standard facial image, the standard facial image and the single facial generative image containing facial information of the same person; and

updating a model parameter of the initial generative model using the loss function determined based on the probability and the similarity to obtain the generative model,

wherein determining the loss function of the initial generative model comprises:

extracting respectively feature information of the single facial generative image and feature information of the standard facial image using a pre-trained recognition model, and calculating a Euclidean distance between the feature information of the single facial generative image and the feature information of the standard facial image; and

obtaining the loss function of the initial generative model based on the probability and the Euclidean distance.

7. The apparatus according to claim 6 , wherein the initial generative model is trained and obtained by:

using at least two frames of initial training facial sample images extracted from an initial training video as an input, and using a preset initial training facial image as an output by using a machine learning method to train and obtain the initial generative model, the at least two frames of initial training facial sample images and the initial training facial image containing the facial information of the same person.

8. The apparatus according to claim 6 , wherein the discriminative model is trained and obtained by:

using a first sample image as an input and using annotation information of the first sample image as an output by using the machine learning method to train and obtain the discriminative model, the first sample image comprising a positive sample image with annotation information and a negative sample image with annotation information, wherein the negative sample image is an image outputted by the generative model.

9. The apparatus according to claim 6 , wherein the recognition model is trained and obtained by:

using a second sample image as an input and feature information of the second sample image as an output by using the machine learning method to train and obtain the recognition model.

10. The apparatus according to claim 6 , wherein the generative model is a Long-Short Term Memory Model.

11. A non-transitory computer medium storing a computer program, wherein the computer program, when executed by a processor, causes the processor to perform operations, the operations comprising:

acquiring at least two frames of facial images of a same person extracted from a target video; and

inputting the at least two frames of facial images of the same person into one pre-trained generative model to generate a single facial image, wherein the generative model is obtained through:

inputting a single facial generative image outputted by an initial generative model into a pre-trained discriminative model to generate a probability of the single facial generative image being a real facial image;

determining a loss function of the initial generative model based on the probability and a similarity between the single facial generative image and a standard facial image, the similarity including a Jaccard similarity coefficient between the single facial generative image and the standard facial image, or a Euclidean distance between the single facial generative image and the standard facial image, the standard facial image and the single facial generative image containing facial information of the same person; and

updating a model parameter of the initial generative model using the loss function determined based on the probability and the similarity to obtain the generative model,

wherein determining the loss function of the initial generative model comprises:

extracting respectively feature information of the single facial generative image and feature information of the standard facial image using a pre-trained recognition model, and calculating a Euclidean distance between the feature information of the single facial generative image and the feature information of the standard facial image; and

obtaining the loss function of the initial generative model based on the probability and the Euclidean distance.

12. The non-transitory computer medium according to claim 11 , wherein the initial generative model is trained and obtained by:

using at least two frames of initial training facial sample images extracted from an initial training video as an input, and using a preset initial training facial image as an output by using a machine learning method to train and obtain the initial generative model, the at least two frames of initial training facial sample images and the initial training facial image containing the facial information of the same person.

13. The non-transitory computer medium according to claim 11 , wherein the discriminative model is trained and obtained by:

using a first sample image as an input and using annotation information of the first sample image as an output by using the machine learning method to train and obtain the discriminative model, the first sample image comprising a positive sample image with annotation information and a negative sample image with annotation information, wherein the negative sample image is an image outputted by the generative model.

14. The non-transitory computer medium according to claim 13 , wherein the recognition model is trained and obtained by:

using a second sample image as an input and feature information of the second sample image as an output by using the machine learning method to train and obtain the recognition model.

15. The non-transitory computer medium according to claim 11 , wherein the generative model is a Long-Short Term Memory Model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2020
From: HE, TAO; ZHANG, GANG; LIU, JINGTUO
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 052938/0036 →
Priority Claims (1)
CN 201710806070.X · Sep 8, 2017 · national
Continuity (1)
Related Publication 20190080148A1 · Mar 14, 2019