IP Library › Granted Patent US 11,769,482
Granted Patent B2
US 11,769,482 · App. 17/489,616 · Granted Sep 26, 2023

Method and apparatus of synthesizing speech, method and apparatus of training speech synthesis model, electronic device, and storage medium

Inventors: Wenfu Wang (Beijing, CN); Tao Sun (Beijing, CN); Xilei Wang (Beijing, CN); Junteng Zhang (Beijing, CN); Zhengkun Gao (Beijing, CN); Lei Jia (Beijing, CN)
Assignee: Beijing Baidu Netcom Science Technology Co., Ltd.
G10L13/10G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,769,482
App. No.
17/489,616
Granted
Sep 26, 2023
Kind
B2
Abstract

The present disclosure provides a method and apparatus of synthesizing a speech, a method and apparatus of training a speech synthesis model, an electronic device, and a storage medium. The method of synthesizing a speech includes acquiring a style information of a speech to be synthesized, a tone information of the speech to be synthesized, and a content information of a text to be processed; generating an acoustic feature information of the text to be processed, by using a pre-trained speech synthesis model, based on the style information, the tone information, and the content information of the text to be processed; and synthesizing the speech for the text to be processed, based on the acoustic feature information of the text to be processed.

Claims (36)

1. A method of synthesizing a speech, comprising:

acquiring a style information of a speech to be synthesized, a tone information of the speech to be synthesized, and a content information of a text to be processed;

generating an acoustic feature information of the text to be processed, by using a pre-trained speech synthesis model, based on the style information, the tone information, and the content information of the text to be processed; and

synthesizing the speech for the text to be processed, based on the acoustic feature information of the text to be processed,

wherein the acquiring the style information of the speech to be synthesized comprises:

acquiring a description information of an input style of a user; and determining a style identifier, from a preset style table, corresponding to the input style according to the description information of the input style, as the style information of the speech to be synthesized.

2. The method according to claim 1 , wherein the generating an acoustic feature information of the text to be processed, by using a pre-trained speech synthesis model, based on the style information, the tone information, and the content information of the text to be processed comprising:

encoding the content information of the text to be processed, by using a content encoder in the speech synthesis model, so as to obtain a content encoded feature;

encoding the content information of the text to be processed and the style information by using a style encoder in the speech synthesis model, so as to obtain a style encoded feature;

encoding the tone information by using a tone encoder in the speech synthesis model, so as to obtain a tone encoded feature; and

decoding by using a decoder in the speech synthesis model based on the content encoded feature, the style encoded feature, and the tone encoded feature, so as to generate the acoustic feature information of the text to be processed.

3. An electronic device, comprising:

at least one processor; and

a memory in communication with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to implement the method according to claim 1 .

4. A non-transitory computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions, when executed, cause a computer to implement the method according to claim 1 .

5. The method according to claim 1 , wherein the acquiring the style information of the speech to be synthesized further comprises:

acquiring an audio information described in an input style; and extracting a tone information of the input style from the audio information, as the style information of the speech to be synthesized.

6. A method of training a speech synthesis model, comprising:

acquiring a plurality of training data, wherein each of the plurality of training data contains a training style information of a speech to be synthesized, a training tone information of the speech to be synthesized, a content information of a training text, a style feature information using a training style corresponding to the training style information to describe the content information of the training text, and a target acoustic feature information using the training style corresponding to the training style information and a training tone corresponding to the training tone information to describe the content information of the training text; and

training the speech synthesis model by using the plurality of training data.

7. The method according to claim 6 , wherein the training the speech synthesis model by using the plurality of training data comprises:

encoding the content information of the training text, the training style information and the training tone information in each of the plurality of training data by using a content encoder, a style encoder, and a tone encoder in the speech synthesis model, respectively, so as to obtain a training content encoded feature, a training style encoded feature, and a training tone encoded feature sequentially;

extracting a target training style encoded feature by using a style extractor in the speech synthesis model, based on the content information of the training text and the style feature information using the training style corresponding to the training style information to describe the content information of the training text;

decoding by using a decoder in the speech synthesis model based on the training content encoded feature, the target training style encoded feature, and the training tone encoded feature, so as to generate a predicted acoustic feature information of the training text;

constructing a comprehensive loss function based on the training style encoded feature, the target training style encoded feature, the predicted acoustic feature information, and the target acoustic feature information; and

adjusting parameters of the content encoder, the style encoder, the tone encoder, the style extractor, and the decoder in response to the comprehensive loss function not converging, so that the comprehensive loss function tends to converge.

8. The method according to claim 7 , wherein constructing the comprehensive loss function based on the training style encoded feature, the target training style encoded feature, the predicted acoustic feature information, and the target acoustic feature information comprises:

constructing a style loss function based on the training style encoded feature and the target training style encoded feature;

constructing a reconstruction loss function based on the predicted acoustic feature information and the target acoustic feature information; and

generating the comprehensive loss function based on the style loss function and the reconstruction loss function.

9. An electronic device, comprising:

at least one processor; and

a memory in communication with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to implement the method according to claim 6 .

10. A non-transitory computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions, when executed, cause a computer to implement the method according to claim 6 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: WANG, WENFU; SUN, TAO; WANG, XILEI; ZHANG, JUNTENG; GAO, ZHENGKUN; JIA, LEI
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 057650/0596 →
Priority Claims (1)
CN 202011253104.5 · Nov 11, 2020 · national
Continuity (1)
Related Publication 20220020356A1 · Jan 20, 2022