IP Library › Granted Patent US 11,836,836
Granted Patent B2
US 11,836,836 · App. 17/527,068 · Granted Dec 5, 2023

Methods and apparatuses for generating model and generating 3D animation, devices and storage mediums

Inventor: Shaoxiong Yang (Beijing, CN)
Assignee: Beijing Baidu Netcom Science Technology Co., Ltd.
G06T13/20G06N3/045G06N3/08G06T17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,836,836
App. No.
17/527,068
Granted
Dec 5, 2023
Kind
B2
Abstract

Methods and apparatuses for generating a model and generating a 3D animation, devices, and storage mediums are provided. The method for generating a model may include: acquiring a preset sample set; acquiring pre-established generative adversarial nets, the generative adversarial nets including a generator and a discriminator; and performing training steps as follows: selecting a sample from the sample set; extracting a sample audio feature from the sample audio of the sample; inputting the sample audio feature into the generator to obtain a pseudo 3D mesh vertex sequence of the sample; inputting the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence of the sample into the discriminator to discriminate authenticity of 3D mesh vertices; and in response to determining that the generative adversarial nets meet a training completion condition, obtaining a trained generator as a model for generating a 3D animation.

Claims (68)

1. A method for generating a model, the method comprising:

acquiring a preset sample set, the sample set comprising at least one sample, and each of the at least one sample comprising a sample audio and a real 3D mesh vertex sequence;

acquiring pre-established generative adversarial nets, the generative adversarial nets comprising a generator and a discriminator, wherein the discriminator comprises at least one of: a 3D mesh vertex frame discriminator, a 3D mesh vertex sequence discriminator, or an audio and 3D mesh vertex sequence synchronization discriminator; and

performing training steps as follows:

selecting a sample from the sample set;

extracting a sample audio feature from the sample audio of the sample;

inputting the sample audio feature into the generator to obtain a pseudo 3D mesh vertex sequence of the sample;

forming a first matrix by splicing the pseudo 3D mesh vertex sequence and the sample audio;

forming a second matrix by splicing the real 3D mesh vertex sequence and the sample audio;

inputting the first matrix and the second matrix into the audio and 3D mesh vertex sequence synchronization discriminator to discriminate whether the audio and 3D mesh vertex sequences including the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequences are synchronized; and

in response to determining that the generative adversarial nets meet a training completion condition, obtaining a trained generator as a model for generating a 3D animation.

2. The method according to claim 1 , wherein the method further comprises:

adjusting a relevant parameter in the generative adversarial nets to make a loss value converge, in response to determining that the generative adversarial nets do not meet the training completion condition, and continue performing the training steps based on the adjusted generative adversarial nets.

3. The method according to claim 1 , wherein the discriminator comprises at least one of:

a 3D mesh vertex frame discriminator, or a 3D mesh vertex sequence discriminator.

4. The method according to claim 3 , wherein, in response to determining that the discriminator comprises the 3D mesh vertex frame discriminator, the inputting the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence of the sample into the discriminator to discriminate authenticity of 3D mesh vertices, comprises:

inputting each 3D mesh vertex frame in the pseudo 3D mesh vertex sequence and each 3D mesh vertex frame in the real 3D mesh vertex sequence into the 3D mesh vertex frame discriminator to discriminate authenticity of a single 3D mesh vertex frame.

5. The method according to claim 3 , wherein, in response to determining that the discriminator comprises the 3D mesh vertex sequence discriminator, the inputting the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence of the sample into the discriminator to discriminate authenticity of 3D mesh vertices, comprises:

inputting the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence into the 3D mesh vertex sequence discriminator to discriminate authenticity of the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence.

6. The method according to claim 1 , wherein the generator comprises an audio encoding module and an expression decoding module.

7. A method for generating a 3D animation, the method comprising:

extracting an audio feature from an audio;

inputting the audio feature into a generator of generative adversarial nets generated in the method according to claim 1 , to generate a 3D mesh vertex sequence; and

rendering the 3D mesh vertex sequence to obtain a 3D animation.

8. An electronic device, comprising:

at least one processor; and

a memory, communicatively connected with the at least one processor;

the memory storing instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, causing the at least one processor to perform operations, the operations comprising:

extracting an audio feature from an audio;

inputting the audio feature into a generator of generative adversarial nets generated in the method according to claim 1 , to generate a 3D mesh vertex sequence; and

rendering the 3D mesh vertex sequence to obtain a 3D animation.

9. A non-transitory computer readable storage medium, storing computer instructions, the computer instructions being used to cause a computer to perform operations, the operations comprising:

extracting an audio feature from an audio;

inputting the audio feature into a generator of generative adversarial nets generated in the method according to claim 1 , to generate a 3D mesh vertex sequence; and

rendering the 3D mesh vertex sequence to obtain a 3D animation.

10. An electronic device, comprising:

at least one processor; and

a memory, communicatively connected with the at least one processor;

the memory storing instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, causing the at least one processor to perform operations, the operations comprising:

acquiring a preset sample set, the sample set comprising at least one sample, and each of the at least one sample comprising a sample audio and a real 3D mesh vertex sequence;

acquiring pre-established generative adversarial nets, the generative adversarial nets comprising a generator and a discriminator, wherein the discriminator comprises at least one of: a 3D mesh vertex frame discriminator, a 3D mesh vertex sequence discriminator, or an audio and 3D mesh vertex sequence synchronization discriminator; and

performing training steps as follows:

selecting a sample from the sample set;

extracting a sample audio feature from the sample audio of the sample;

inputting the sample audio feature into the generator to obtain a pseudo 3D mesh vertex sequence of the sample;

forming a first matrix by splicing the pseudo 3D mesh vertex sequence and the sample audio;

forming a second matrix by splicing the real 3D mesh vertex sequence and the sample audio;

inputting the first matrix and the second matrix into the audio and 3D mesh vertex sequence synchronization discriminator to discriminate whether the audio and 3D mesh vertex sequences including the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence are synchronized; and

in response to determining that the generative adversarial nets meet a training completion condition, obtaining a trained generator as a model for generating a 3D animation.

11. The electronic device according to claim 10 , wherein the operations further comprise:

adjusting a relevant parameter in the generative adversarial nets to make a loss value converge, in response to determining that the generative adversarial nets do not meet the training completion condition, and continue performing the training steps based on the adjusted generative adversarial nets.

12. The electronic device according to claim 10 , wherein the discriminator comprises at least one of: a 3D mesh vertex frame discriminator, or a 3D mesh vertex sequence discriminator.

13. The electronic device according to claim 12 , wherein, in response to determining that the discriminator comprises the 3D mesh vertex frame discriminator, the inputting the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence of the sample into the discriminator to discriminate authenticity of 3D mesh vertices, comprises:

inputting each 3D mesh vertex frame in the pseudo 3D mesh vertex sequence and each 3D mesh vertex frame in the real 3D mesh vertex sequence into the 3D mesh vertex frame discriminator to discriminate authenticity of a single 3D mesh vertex frame.

14. The electronic device according to claim 12 , wherein, in response to determining that the discriminator comprises the 3D mesh vertex sequence discriminator, the inputting the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence of the sample into the discriminator to discriminate authenticity of 3D mesh vertices, comprises:

inputting the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence into the 3D mesh vertex sequence discriminator to discriminate authenticity of the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence.

15. The electronic device according to claim 9 , wherein the generator comprises an audio encoding module and an expression decoding module.

16. A non-transitory computer readable storage medium, storing computer instructions, the computer instructions being used to cause a computer to perform operations, the operations comprising:

acquiring a preset sample set, the sample set comprising at least one sample, and each of the at least one sample comprising a sample audio and a real 3D mesh vertex sequence;

acquiring pre-established generative adversarial nets, the generative adversarial nets comprising a generator and a discriminator, wherein the discriminator comprises at least one of: a 3D mesh vertex frame discriminator, a 3D mesh vertex sequence discriminator, or an audio and 3D mesh vertex sequence synchronization discriminator; and

performing training steps as follows:

selecting a sample from the sample set;

extracting a sample audio feature from the sample audio of the sample;

inputting the sample audio feature into the generator to obtain a pseudo 3D mesh vertex sequence of the sample;

forming a first matrix by splicing the pseudo 3D mesh vertex sequence and the sample audio;

forming a second matrix by splicing the real 3D mesh vertex sequence and the sample audio;

inputting the first matrix and the second matrix into the audio and 3D mesh vertex sequence synchronization discriminator to discriminate whether the audio and 3D mesh vertex sequences including the pseudo 3D mesh vertex sequence and the real 3D mesh vertex sequence are synchronized; and

in response to determining that the generative adversarial nets meet a training completion condition, obtaining a trained generator as a model for generating a 3D animation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2022
From: YANG, SHAOXIONG
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 060656/0060 →
Priority Claims (1)
CN 202011485571.0 · Dec 16, 2020 · national
Continuity (1)
Related Publication 20220076470A1 · Mar 10, 2022