IP Library › Granted Patent US 10,410,621
Granted Patent B2
US 10,410,621 · App. 15/758,280 · Granted Sep 10, 2019

Training method for multiple personalized acoustic models, and voice synthesis method and device

Inventor: Xiulin Li (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G10L13/02G10L13/08G10L15/02G10L15/04G10L15/063G10L15/142G10L15/183G10L15/1807G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,410,621
App. No.
15/758,280
Filed
Mar 7, 2018
Granted
Sep 10, 2019
Kind
B2
Examiner
GAY, SONIA L
Art Unit
2657
USPC
704/266
Abstract

A training method for multiple personalized acoustic models, and a voice synthesis method and device, for voice synthesis. The method comprises: training a reference acoustic model, based on first acoustic feature data of training voice data and first text annotation data corresponding to the training voice data (S 11 ); acquiring voice data of a target user (S 12 ); training a first target user acoustic model according to the reference acoustic model and the voice data (S 13 ); generating second acoustic feature data of the first text annotation data, according to the first target user acoustic model and the first text annotation data (S 14 ); and training a second target user acoustic model, based on the first text annotation data and the second acoustic feature data (S 15 ).

Claims (41)

1. A method for training personalized multiple acoustic models for speech synthesis, comprising:

training a reference acoustic model based on first acoustic feature data of training speech data and first text annotation data corresponding to the training speech data;

obtaining speech data of a target user;

training a first target user acoustic model according to the reference acoustic model and the speech data;

generating second acoustic feature data of the first text annotation data according to the first target user acoustic model and the first text annotation data; and

training a second target user acoustic model based on the first text annotation data and the second acoustic feature data.

2. The method according to claim 1 , wherein training a first target user acoustic model according to the reference acoustic model and the speech data comprises:

performing an acoustic feature extraction on the speech data to obtain third acoustic feature data of the speech data;

performing speech annotation on the speech data to obtain second text annotation data of the speech data; and

training the first target user acoustic model according to the reference acoustic model, the third acoustic feature data, and the second text annotation data.

3. The method according to claim 2 , wherein training the first target user acoustic model according to the reference acoustic model, the third acoustic feature data, and the second text annotation data comprises:

obtaining a neural network structure of the reference acoustic model; and

training the first target user acoustic model according to the third acoustic feature data, the second text annotation data, and the neural network structure of the reference acoustic model.

4. The method according to claim 1 , wherein training a second target user acoustic model based on the first text annotation data and the second acoustic feature data, comprises:

performing training on the first text annotation data and the second acoustic feature data based on a hidden Markov model, and building the second target user acoustic model according to a result of the training.

5. A method for speech synthesis using a second target user acoustic model, wherein the second target user acoustic model is obtained by training a reference acoustic model based on first acoustic feature data of training speech data and first text annotation data corresponding to the training speech data; obtaining speech data of a target user; training the first target user acoustic model according to the reference acoustic model and the speech data, generating second acoustic feature data of the first text annotation data according to the first target user acoustic model and the first text annotation data; and training the second target user acoustic model based on the first text annotation data and the second acoustic feature data;

the method for speech synthesis comprises:

obtaining a text to be synthesized, and performing word segmentation on the text to be synthesized;

performing part-of-speech tagging on the text to be synthesized after the word segmentation, and performing a prosody prediction on the text to be synthesized after the part-of-speech tagging via a prosody prediction model, to generate prosodic features of the text to be synthesized;

performing phonetic notation on the text to be synthesized according to a result of the word segmentation, a result of the part-of-speech tagging, and the prosodic features, to generate a result of phonetic notation of the text to be synthesized;

inputting the result of phonetic notation, the prosodic features, and context features of the text to be synthesized to the second target user acoustic model, and performing an acoustic prediction on the text to be synthesized via the second target user acoustic model, to generate an acoustic parameter sequence of the text to be synthesized; and

generating a speech synthesis result of the text to be synthesized according to the acoustic parameter sequence.

6. The method according to claim 1 , wherein obtaining speech data of a target user comprises:

designing recording text according to phone coverage and prosody coverage;

providing the recording text to the target user; and

receiving the speech data of the recording text read by the target user.

7. The method according to claim 2 , wherein training a first target user acoustic model according to the reference acoustic model and the speech data further comprises:

performing pre-processing on the speech data of the target user, wherein the pre-processing comprises data de-noising, data detection, data filtration and segmentation.

8. The method according to claim 2 , wherein training the first target user acoustic model according to the reference acoustic model, the third acoustic feature data, and the second text annotation data comprises:

obtaining a neural network structure of the reference acoustic model;

updating parameters in the neural network structure of the reference acoustic model by performing an iterative operation via a neural network adaptive technology according to the third acoustic feature data, the second text annotation data and the neural network structure of the reference acoustic model, to obtain the first target user acoustic model.

9. The method according to claim 1 , wherein training a second target user acoustic model based on the first text annotation data and the second acoustic feature data comprises:

performing training on the first text annotation data and the second acoustic feature data based on a hidden Markov model; and

building the second target user acoustic model according to a result of the training.

10. The method according to claim 2 , wherein training the first target user acoustic model according to the reference acoustic model, the third acoustic feature data, and the second text annotation data comprises:

obtaining a neural network structure of the reference acoustic model; and

training the first target user acoustic model according to the third acoustic feature data, the second text annotation data, and the neural network structure of the reference acoustic model.

11. The method according to claim 2 , wherein training a second target user acoustic model based on the first text annotation data and the second acoustic feature data, comprises:

performing training on the first text annotation data and the second acoustic feature data based on a hidden Markov model, and building the second target user acoustic model according to a result of the training.

12. The method according to claim 3 , wherein training a second target user acoustic model based on the first text annotation data and the second acoustic feature data, comprises:

performing training on the first text annotation data and the second acoustic feature data based on a hidden Markov model, and building the second target user acoustic model according to a result of the training.

Priority Claims (1)
CN 2015 1 0684475 · Oct 20, 2015 · national
Continuity (1)
Related Publication 20180254034A1 · Sep 6, 2018