IP Library › Granted Patent US 11,948,236
Granted Patent B2
US 11,948,236 · App. 17/527,473 · Granted Apr 2, 2024

Method and apparatus for generating animation, electronic device, and computer readable medium

Inventors: Shaoxiong Yang (Beijing, CN); Yang Zhao (Beijing, CN); Chen Zhao (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06T13/40G06T13/205G06V20/46G06V40/165G06V40/171G06V40/174G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,948,236
App. No.
17/527,473
Granted
Apr 2, 2024
Kind
B2
Abstract

The present disclosure discloses a method and apparatus for generating animation. An implementation of the method may include: processing a to-be-processed material to generate a normalized text; analyzing the normalized text to generate a Chinese pinyin sequence of the normalized text; generating a reference audio based on the to-be-processed material; and obtaining a animation of facial expressions corresponding to the timing sequence of the reference audio based on the Chinese pinyin sequence and the reference audio.

Claims (61)

1. A method for generating animation, comprising:

processing to-be-processed material to generate a normalized text;

analyzing the normalized text to generate a Chinese pinyin sequence of the normalized text;

generating a reference audio based on the to-be-processed material;

aligning the reference audio and respective Chinese pinyins in the Chinese pinyin sequence according to a timing sequence of the reference audio, to obtain time stamps of the respective Chinese pinyins;

searching in a pinyin-expression coefficient dictionary to obtain an expression coefficient sequence corresponding to each of the Chinese pinyins, the pinyin-expression coefficient dictionary being used to represent a corresponding relationship between a Chinese pinyin and an expression coefficient sequence;

stitching, based on the time stamps, expression coefficient sequences corresponding to the Chinese pinyins in the Chinese pinyin sequence, to obtain an expression coefficient sequence corresponding to the timing sequence of the reference audio; and

obtaining animation of facial expressions corresponding to the timing sequence of the reference audio based on the expression coefficient sequence corresponding to the timing sequence of the reference audio.

2. The method according to claim 1 , wherein the pinyin-expression coefficient dictionary is obtained by annotating through:

recording a video of a voice actor reading each Chinese pinyin, to obtain pinyin videos in one-to-one corresponding relationship with the Chinese pinyins;

performing a face key point detection on each video frame in each of the pinyin videos; and

calculating an expression coefficient based on detected face key points, to obtain the pinyin-expression coefficient dictionary including expression coefficient sequences in one-to-one corresponding relationship with the Chinese pinyins.

3. The method according to claim 1 , wherein the obtaining the animation of the facial expressions based on the expression coefficient sequence corresponding to the timing sequence of the reference audio comprises:

performing a weighted summation on the expression coefficient sequence corresponding to the timing sequence of the reference audio and an expression base, to obtain a sequence of three-dimensional face models;

rendering the sequence of three-dimensional face models to obtain a sequence of video frame pictures; and

synthesizing the sequence of video frame pictures to obtain the animation of facial expressions.

4. The method according to claim 1 , wherein the to-be-processed material includes an audio to be processed, and the processing the to-be-processed material to generate the normalized text comprises:

performing speech recognition processing on the audio to be processed to generate a Chinese text; and

performing text normalization processing on the Chinese text to generate the normalized text.

5. A non-transitory computer readable storage medium, storing computer instructions, wherein the computer instructions are used to cause a processor to perform the method according to claim 1 .

6. A method for generating animation, comprising:

processing to-be-processed material to generate a normalized text;

analyzing the normalized text to generate a Chinese pinyin sequence of the normalized text;

generating a reference audio based on the to-be-processed material;

aligning the reference audio and the Chinese pinyins in the Chinese pinyin sequence according to a timing sequence of the reference audio, to obtain time stamps of the Chinese pinyins;

searching in a pinyin-grid dictionary to obtain a three-dimensional face grid sequence corresponding to the each of the Chinese pinyins, the pinyin-grid dictionary being used to represent a corresponding relationship between a Chinese pinyin and a three-dimensional face grid sequence;

stitching, based on the time stamps, three-dimensional face grid sequences corresponding to the Chinese pinyins in the Chinese pinyin sequence, to obtain a three-dimensional face grid sequence corresponding to the timing sequence of the reference audio;

obtaining the animation of facial expressions based on the three-dimensional face grid sequence corresponding to the timing sequence of the reference audio.

7. The method according to claim 6 , wherein the obtaining the animation of facial expressions based on the three-dimensional face grid sequence corresponding to the timing sequence of the reference audio comprises:

rendering the three-dimensional face grid sequence corresponding to the timing sequence of the reference audios, to obtain a sequence of video frame pictures; and

synthesizing the sequence of video frame pictures to obtain the animation of facial expressions.

8. An electronic device, comprising:

at least one processor; and

a storage device connected with the at least one processor,

wherein the storage device stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform the method according to claim 6 .

9. The device according to claim 8 , wherein the obtaining the animation of facial expressions based on the three-dimensional face grid sequence corresponding to the timing sequence of the reference audio comprises:

rendering the three-dimensional face grid sequence corresponding to the timing sequence of the reference audios, to obtain a sequence of video frame pictures; and

synthesizing the sequence of video frame pictures to obtain the animation of facial expressions.

10. A non-transitory computer readable storage medium, storing computer instructions, wherein the computer instructions are used to cause a processor to perform the method according to claim 6 .

11. An electronic device, comprising:

at least one processor; and

a storage device connected with the at least one processor,

wherein the storage device stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

processing to-be-processed material to generate a normalized text;

analyzing the normalized text to generate a Chinese pinyin sequence of the normalized text;

generating a reference audio based on the to-be-processed material;

aligning the reference audio and respective Chinese pinyins in the Chinese pinyin sequence according to a timing sequence of the reference audio, to obtain time stamps of the respective Chinese pinyins;

searching in a pinyin-expression coefficient dictionary to obtain an expression coefficient sequence corresponding to each of the Chinese pinyins, the pinyin-expression coefficient dictionary being used to represent a corresponding relationship between a Chinese pinyin and an expression coefficient sequence;

stitching, based on the time stamps, expression coefficient sequences corresponding to the Chinese pinyins in the Chinese pinyin sequence, to obtain an expression coefficient sequence corresponding to the timing sequence of the reference audio; and

obtaining animation of facial expressions corresponding to the timing sequence of the reference audio based on the expression coefficient sequence corresponding to the timing sequence of the reference audio.

12. The device according to claim 11 , wherein the pinyin-expression coefficient dictionary is obtained by annotating through:

recording a video of a voice actor reading each Chinese pinyin, to obtain pinyin videos in one-to-one corresponding relationship with the Chinese pinyins;

performing a face key point detection on each video frame in each of the pinyin videos; and

calculating an expression coefficient based on detected face key points, to obtain the pinyin-expression coefficient dictionary including expression coefficient sequences in one-to-one corresponding relationship with the Chinese pinyins.

13. The device according to claim 11 , wherein the obtaining the animation of the facial expressions based on the expression coefficient sequence corresponding to the timing sequence of the reference audio comprises:

performing a weighted summation on the expression coefficient sequence corresponding to the timing sequence of the reference audio and an expression base, to obtain a sequence of three-dimensional face models;

rendering the sequence of three-dimensional face models to obtain a sequence of video frame pictures; and

synthesizing the sequence of video frame pictures to obtain the animation of facial expressions.

14. The device according to claim 11 , wherein the to-be-processed material includes an audio to be processed, and the processing the to-be-processed material to generate the normalized text comprises:

performing speech recognition processing on the audio to be processed to generate a Chinese text; and

performing text normalization processing on the Chinese text to generate the normalized text.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2022
From: YANG, SHAOXIONG; ZHAO, CHEN
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 060959/0883 →
EMPLOYMENT CONTRACT Recorded Aug 31, 2022
From: ZHAO, YANG
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061363/0571 →
Priority Claims (1)
CN 202011430467.1 · Dec 9, 2020 · national
Continuity (1)
Related Publication 20220180584A1 · Jun 9, 2022
Cited By (1)
US 12,646,256