IP Library › Granted Patent US 11,562,732
Granted Patent B2
US 11,562,732 · App. 17/024,662 · Granted Jan 24, 2023

Method and apparatus for predicting mouth-shape feature, and electronic device

Inventors: Yuqiang Liu (Beijing, CN); Tao Sun (Beijing, CN); Wenfu Wang (Beijing, CN); Guanbo Bao (Beijing, CN); Zhe Peng (Beijing, CN); Lei Jia (Beijing, CN)
Assignee: Baidu Online Network Technology (Beijing) Co., Ltd.
G10L15/02G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,732
App. No.
17/024,662
Granted
Jan 24, 2023
Kind
B2
Abstract

A method and apparatus for predicting a mouth-shape feature, and an electronic device are provided. A specific implementation of the method comprises: recognizing a phonetic posterior gram (PPG) of a phonetic feature; and performing a prediction on the PPG by using a neural network model, to predict a mouth-shape feature of the phonetic feature, the neural network model being obtained by training with training samples and an input thereof including a PPG and an output thereof including a mouth-shape feature, and the training samples including a PPG training sample and a mouth-shape feature training sample.

Claims (51)

1. A method for predicting a mouth-shape feature, comprising:

recognizing a phonetic posterior gram (PPG) of a phonetic feature of a speech; and

performing a prediction on the PPG by using a neural network model, to predict a mouth-shape feature of the phonetic feature, the neural network model being obtained by training with training samples and an input thereof including a PPG and an output thereof including a mouth-shape feature, and the training samples including a PPG training sample and a mouth-shape feature training sample,

wherein the PPG training sample comprises:

PPGs of target phonetic features, a target phonetic feature being a phonetic feature having complete semantics which is obtained by slicing according to the semantics of the speech; and

the mouth-shape feature training sample comprises:

mouth-shape features corresponding to the PPGs of the target phonetic features,

wherein the neural network model is a multi-branch neural network model, and the mouth-shape feature of the phonetic feature includes at least two of:

a regression mouth-shape point, a mouth-shape thumbnail, a blend shape coefficient, or three-dimensional morphable models (3DMM) expression coefficient,

wherein the multi-branch neural network model has a plurality of branch networks, and each branch network predicts a mouth-shape feature of the regression mouth-shape point, the mouth-shape thumbnail, the blend shape coefficient, and the 3DMM expression coefficient.

2. The method according to claim 1 , wherein a frequency of a target phonetic feature matches a frequency of a mouth-shape feature corresponding to a PPG of the target phonetic feature.

3. The method according to claim 1 , wherein the multi-branch neural network model is a recurrent neural network (RNN) model having an autoregressive mechanism, and a process of training the RNN model includes:

performing the training by using a mouth-shape feature training sample of a frame preceding a current frame as an input, by using a PPG training sample of the current frame as a condition constraint, and a mouth-shape feature training sample of the current frame as a target.

4. The method according to claim 1 , further comprising:

performing predictions on PPGs of pieces of pieces of real speech data using the neural network model, to obtain mouth-shape features of the pieces of real speech data; and

constructing a mouth-shape feature index library based on the mouth-shape features of the pieces of real speech data, the mouth-shape feature index library being used for synthesizing a mouth shape of a virtual image.

5. An electronic device, comprising:

at least one processor; and

a memory, communicated with the at least one processor,

wherein the memory stores an instruction executable by the at least one processor, and the instruction, when executed by the at least one processor, causes the at least one processor to perform operations, the operations comprise:

recognizing a phonetic posterior gram (PPG) of a phonetic feature of a speech; and

performing a prediction on the PPG by using a neural network model, to predict a mouth-shape feature of the phonetic feature, the neural network model being obtained by training with training samples and an input thereof including a PPG and an output thereof including a mouth-shape feature, and the training samples including a PPG training sample and a mouth-shape feature training sample,

wherein the PPG training sample comprises:

PPGs of target phonetic features, a target phonetic feature being a phonetic feature having complete semantics which is obtained by slicing according to the semantics of the speech; and

the mouth-shape feature training sample comprises:

mouth-shape features corresponding to the PPGs of the target phonetic features,

wherein the neural network model is a multi-branch neural network model, and the mouth-shape feature of the phonetic feature includes at least two of:

a regression mouth-shape point, a mouth-shape thumbnail, a blend shape coefficient, or three-dimensional morphable models (3DMM) expression coefficient,

wherein the multi-branch neural network model has a plurality of branch networks, and each branch network predicts a mouth-shape feature of the regression mouth-shape point, the mouth-shape thumbnail, the blend shape coefficient, and the 3DMM expression coefficient.

6. The electronic device according to claim 5 , wherein a frequency of a target phonetic feature matches a frequency of a mouth-shape feature corresponding to a PPG of the target phonetic feature.

7. The electronic device according to claim 5 , wherein the multi-branch neural network model is a recurrent neural network (RNN) model having an autoregressive mechanism, and a process of training the RNN model includes:

performing the training by using a mouth-shape feature training sample of a frame preceding a current frame as an input, by using a PPG training sample of the current frame as a condition constraint, and a mouth-shape feature training sample of the current frame as a target.

8. The electronic device according to claim 5 , wherein the operations further comprise:

performing predictions on PPGs of pieces of real speech data using the neural network model, to obtain mouth-shape features of the pieces of real speech data; and

constructing a mouth-shape feature index library based on the mouth-shape features of the pieces of real speech data, the mouth-shape feature index library being used for synthesizing a mouth shape of a virtual image.

9. A non-transitory computer readable storage medium, storing a computer instruction, wherein the computer instruction, when executed by a processor, causes the processor to perform operations, the operations comprise:

recognizing a phonetic posterior gram (PPG) of a phonetic feature of a speech; and

performing a prediction on the PPG by using a neural network model, to predict a mouth-shape feature of the phonetic feature, the neural network model being obtained by training with training samples and an input thereof including a PPG and an output thereof including a mouth-shape feature, and the training samples including a PPG training sample and a mouth-shape feature training sample,

wherein the PPG training sample comprises:

PPGs of target phonetic features, a target phonetic feature being a phonetic feature having complete semantics which is obtained by slicing according to the semantics of the speech; and

the mouth-shape feature training sample comprises:

mouth-shape features corresponding to the PPGs of the target phonetic features,

wherein the neural network model is a multi-branch neural network model, and the mouth-shape feature of the phonetic feature includes at least two of:

a regression mouth-shape point, a mouth-shape thumbnail, a blend shape coefficient, or three-dimensional morphable models (3DMM) expression coefficient,

wherein the multi-branch neural network model has a plurality of branch networks, and each branch network predicts a mouth-shape feature of the regression mouth-shape point, the mouth-shape thumbnail, the blend shape coefficient, and the 3DMM expression coefficient.

10. The medium according to claim 9 , wherein a frequency of a target phonetic feature matches a frequency of a mouth-shape feature corresponding to a PPG of the target phonetic feature.

11. The medium according to claim 9 , wherein the multi-branch neural network model is a recurrent neural network (RNN) model having an autoregressive mechanism, and a process of training the RNN model includes:

performing the training by using a mouth-shape feature training sample of a frame preceding a current frame as an input, by using a PPG training sample of the current frame as a condition constraint, and a mouth-shape feature training sample of the current frame as a target.

12. The medium according to claim 9 , wherein the operations further comprise:

performing predictions on PPGs of pieces of pieces of real speech data using the neural network model, to obtain mouth-shape features of the pieces of real speech data; and

constructing a mouth-shape feature index library based on the mouth-shape features of the pieces of real speech data, the mouth-shape feature index library being used for synthesizing a mouth shape of a virtual image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2020
From: LIU, YUQIANG; SUN, TAO; WANG, WENFU; BAO, GUANBO; PENG, ZHE; JIA, LEI
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 053809/0768 →
Priority Claims (1)
CN 202010091799.5 · Feb 13, 2020 · national
Continuity (1)
Related Publication 20210256962A1 · Aug 19, 2021