IP Library › Granted Patent US 11,417,314
Granted Patent B2
US 11,417,314 · App. 16/797,267 · Granted Aug 16, 2022

Speech synthesis method, speech synthesis device, and electronic apparatus

Inventors: Chenxi Sun (Beijing, CN); Tao Sun (Beijing, CN); Xiaolin Zhu (Beijing, CN); Wenfu Wang (Beijing, CN)
Assignee: Baidu Online Network Technology (Beijing) Co., Ltd.
G10L13/047G06N3/08G10L13/02G10L13/08G10L13/10G10L13/06G10L13/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,417,314
App. No.
16/797,267
Granted
Aug 16, 2022
Kind
B2
Abstract

A speech synthesis method, a speech synthesis device, and an electronic apparatus are provided, which relate to a field of speech synthesis. Specific implementation solution is the following: inputting text information into an encoder of an acoustic model, to output a text feature of a current time step; splicing the text feature of the current time step with a spectral feature of a previous time step to obtain a spliced feature of the current time step, and inputting the spliced feature of the current time step into an decoder of the acoustic model to obtain a spectral feature of the current time step; and inputting the spectral feature of the current time step into a neural network vocoder, to output speech.

Claims (55)

1. A speech synthesis method, comprising:

inputting text information into an encoder of an acoustic model, to output a text feature of a current time step;

splicing the text feature of the current time step with a spectral feature of a previous time step to obtain a spliced feature of the current time step, and inputting the spliced feature of the current time step into a decoder of the acoustic model to obtain a spectral feature of the current time step; and

inputting the spectral feature of the current time step into a neural network vocoder, to output speech;

wherein the splicing the text feature of the current time step with a spectral feature of a previous time step to obtain a spliced feature of the current time step, and inputting the spliced feature of the current time step into a decoder of the acoustic model to obtain a spectral feature of the current time step comprises:

inputting the spliced feature of the previous time step into at least one gated recurrent unit and a fully connected layer in the decoder, to output a first spectral feature of the previous time step;

inputting the first spectral feature of the previous time step into another fully connected layer, to obtain a second spectral feature of the previous time step;

splicing the text feature of the current time step with the second spectral feature of the previous time step, to obtain the spliced feature of the current time step; and

inputting the spliced feature of the current time step into the decoder of the acoustic model, to obtain a first spectral feature of the current time step.

2. The speech synthesis method according to claim 1 , wherein the inputting text information into an encoder of an acoustic model, to output a text feature of a current time step comprises:

passing the text information through at least one fully connected layer and a gated recurrent unit in the encoder, to output the text feature of the current time step.

3. The speech synthesis method according to claim 1 , wherein the inputting the spectral feature of the current time step into a neural network vocoder, to output speech comprises:

inputting the first spectral feature of the current time step into at least one convolutional neural network, to obtain a second spectral feature of the current time step; and

inputting the first spectral feature of the current time step or the second spectral feature of the current time step into the neural network vocoder, to output the speech.

4. The speech synthesis method according to claim 3 , further comprising:

calculating a first loss according to the first spectral feature of the current time step and a true spectral feature;

calculating a second loss according to the second spectral feature of the current time step and the true spectral feature; and

training the acoustic model by taking the first loss and the second loss as starting points of a gradient back propagation.

5. A speech synthesis device, comprising:

one or more processors; and

a storage device configured to store one or more programs, wherein

the one or more programs, when executed by the one or more processors, cause the one or more processors to:

input text information into an encoder of an acoustic model, to output a text feature of a current time step;

splice the text feature of the current time step with a spectral feature of a previous time step to obtain a spliced feature of the current time step, and input the spliced feature of the current time step into a decoder of the acoustic model to obtain a spectral feature of the current time step; and

input the spectral feature of the current time step into a neural network vocoder, to output speech;

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

input the spliced feature of the previous time step into at least one gated recurrent unit and a fully connected layer in the decoder, to output a first spectral feature of the previous time step;

input the first spectral feature of the previous time step into another fully connected layer, to obtain a second spectral feature of the previous time step;

splice the text feature of the current time step with the second spectral feature of the previous time step, to obtain the spliced feature of the current time step; and

input the spliced feature of the current time step into the decoder of the acoustic model, to obtain a first spectral feature of the current time step.

6. The speech synthesis device according to claim 5 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

pass the text information through at least one fully connected layer and a gated recurrent unit in the encoder, to output the text feature of the current time step.

7. The speech synthesis device according to claim 5 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

input the first spectral feature of the current time step into at least one convolutional neural network, to obtain a second spectral feature of the current time step; and

input the first spectral feature of the current time step or the second spectral feature of the current time step into the neural network vocoder, to output the speech.

8. The speech synthesis device according to claim 7 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

calculate a first loss according to the first spectral feature of the current time step and a true spectral feature; calculate a second loss according to the second spectral feature of the current time step and the true spectral feature; and train the acoustic model by taking the first loss and the second loss as starting points of a gradient back propagation.

9. A non-transitory computer-readable storage medium comprising computer executable instructions stored thereon, wherein the executable instructions, when executed by a computer, causes the computer to:

input text information into an encoder of an acoustic model, to output a text feature of a current time step;

splice the text feature of the current time step with a spectral feature of a previous time step to obtain a spliced feature of the current time step, and input the spliced feature of the current time step into a decoder of the acoustic model to obtain a spectral feature of the current time step; and

input the spectral feature of the current time step into a neural network vocoder, to output speech;

wherein the executable instructions, when executed by the computer, causes the computer further to:

input the spliced feature of the previous time step into at least one gated recurrent unit and a fully connected layer in the decoder, to output a first spectral feature of the previous time step;

input the first spectral feature of the previous time step into another fully connected layer, to obtain a second spectral feature of the previous time step;

splice the text feature of the current time step with the second spectral feature of the previous time step, to obtain the spliced feature of the current time step; and

input the spliced feature of the current time step into the decoder of the acoustic model, to obtain a first spectral feature of the current time step.

10. The non-transitory computer-readable storage medium according to claim 9 , wherein the executable instructions, when executed by the computer, causes the computer further to:

pass the text information through at least one fully connected layer and a gated recurrent unit in the encoder, to output the text feature of the current time step.

11. The non-transitory computer-readable storage medium according to claim 9 , wherein the executable instructions, when executed by the computer, causes the computer further to:

input the first spectral feature of the current time step into at least one convolutional neural network, to obtain a second spectral feature of the current time step; and

input the first spectral feature of the current time step or the second spectral feature of the current time step into the neural network vocoder, to output the speech.

12. The non-transitory computer-readable storage medium according to claim 11 , wherein the executable instructions, when executed by the computer, causes the computer further to:

calculate a first loss according to the first spectral feature of the current time step and a true spectral feature;

calculate a second loss according to the second spectral feature of the current time step and the true spectral feature; and

train the acoustic model by taking the first loss and the second loss as starting points of a gradient back propagation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2020
From: SUN, CHENXI; SUN, TAO; ZHU, XIAOLIN; WANG, WENFU
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051924/0868 →
Priority Claims (1)
CN 201910888456.9 · Sep 19, 2019 · national
Continuity (1)
Related Publication 20210090550A1 · Mar 25, 2021
Cited By (1)
US 12,387,710