IP Library › Granted Patent US 12,198,674
Granted Patent B2
US 12,198,674 · App. 17/629,483 · Granted Jan 14, 2025

Speech synthesis method and apparatus, and storage medium

Inventors: Zhizheng Wu (Beijing, CN); Wei Song (Beijing, CN)
Assignees: BEIJING JINGDONG SHANGKE INFORMATION TECHNOLOGY CO., LTD.; BEIJING JINGDONG CENTURY TRADING CO., LTD.
G10L13/047G10L13/06G10L25/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,674
App. No.
17/629,483
Granted
Jan 14, 2025
Kind
B2
Abstract

Disclosed are a speech synthesis method and apparatus, and a storage medium. The method comprises: acquiring a symbol sequency of a statement to be synthesized, wherein the statement to be synthesized comprises a recorded statement characterizing a target object and a query result statement for the target object; encoding the symbol sequence by using a pre-set encoding model, in order to obtain a feature vector set; acquiring recording acoustic features corresponding to the recorded statement; predicting, according to a pre-set decoding model, the feature vector set, a pre-set attention model and the recording acoustic features, acoustic features corresponding to the statement to be synthesized, in order to obtain predicted acoustic features corresponding to the statement to be synthesized, wherein the pre-set attention model is a model that uses the feature vector set to generate a context vector used for decoding, and the predicted acoustic features are composed of at least one associated acoustic feature; and performing feature conversion and synthesis on the predicted acoustic features to obtain a speech corresponding to the sentence to be synthesized.

Claims (97)

1. A speech synthesis method performed by a speech synthesis apparatus, comprising:

acquiring a symbol sequence of a statement to be synthesized, wherein the statement to be synthesized includes a recorded statement that represents a target object and a query result statement for the target object;

performing encoding processing on the symbol sequence to obtain a feature vector set by using a preset encoding model;

acquiring a recording acoustic feature corresponding to the recorded statement;

performing prediction on an acoustic feature corresponding to the statement to be synthesized based on a preset decoding model, the feature vector set, a preset attention model, and the recording acoustic feature to obtain predicted acoustic features corresponding to the statement to be synthesized, wherein the preset attention model is a model for generating a context vector for decoding by using the feature vector set, and the predicted acoustic features consist of at least one related acoustic feature;

performing feature conversion and synthesis on the predicted acoustic features to obtain speech corresponding to the statement to be synthesized;

wherein the performing prediction on the acoustic feature corresponding to the statement to be synthesized based on the preset decoding model, the feature vector set, the preset attention model, and the recording acoustic feature to obtain the predicted acoustic features corresponding to the statement to be synthesized comprises:

when i is equal to 1, acquiring an initial acoustic feature at an ith decoding time, predicting a 1st acoustic feature corresponding to the statement to be synthesized based on the initial acoustic feature, the preset decoding model, the feature vector set, and the preset attention model, where the i is an integer greater than 0;

in a case where the i is greater than 1, when the ith decoding time is a decoding time of the recorded statement, acquiring a jth frame of recording acoustic feature from the recording acoustic feature, taking the jth frame of recording acoustic feature from the recording acoustic feature as an (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, and predicting an ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model, where the i is the integer greater than 0;

when the ith decoding time is a decoding time of the query result statement, taking one frame of acoustic feature of an (i−1)th acoustic feature corresponding to the statement to be synthesized as the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, and predicting the ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model;

continuing to perform a prediction process of an (i+1)th decoding time until a decoding of the statement to be synthesized is ended to obtain an nth acoustic feature corresponding to the statement to be synthesized, wherein n is a total number of frames of decoding times of the statement to be synthesized and is an integer greater than 1; and

taking the ith acoustic feature to the nth acoustic feature corresponding to the statement to be synthesized as the predicted acoustic features; and

transmitting the speech to a playback module and playing the speech by the playback module.

2. The method of claim 1 , wherein the preset decoding model comprises a first Recurrent Neural Network (RNN) and a second RNN; and the predicting the ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model comprises:

performing a first non-linear transformation on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized to obtain an intermediate feature vector;

performing a first matrix operation and a second non-linear transformation on the intermediate feature vector to obtain an ith intermediate hidden variable by using the first RNN;

performing a context vector calculation on the feature vector set and the ith intermediate hidden variable to obtain an ith context vector by using the preset attention model;

performing a second matrix operation and a third non-linear transformation on the ith context vector and the ith intermediate hidden variable to obtain an ith hidden variable by using the second RNN; and

performing a linear transformation on the ith hidden variable to obtain the ith acoustic feature corresponding to the statement to be synthesized according to a preset number of frames.

3. The method of claim 2 , wherein the feature vector set comprises a feature vector corresponding to each symbol in the symbol sequence; the performing the context vector calculation on the feature vector set and the ith intermediate hidden variable to obtain the ith context vector by using the preset attention model comprises:

performing the context vector calculation on the feature vector corresponding to each symbol in the symbol sequence and the ith intermediate hidden variable to obtain an ith group of attention values by using the preset attention model; and

performing weighted summation on the feature vector set to obtain the ith context vector according to the ith group of attention values.

4. The method of claim 3 , wherein after predicting the ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model, and before continuing to perform the prediction process of the (i+1)th decoding time, the method further comprises:

determining an ith target symbol corresponding to a maximum attention value from the ith group of attention values; and

at least one of

when the ith target symbol is a non-terminal symbol of the recorded statement, determining the (i+1)th decoding time as the decoding time of the recorded statement;

when the ith target symbol is a non-terminal symbol of the query result statement, determining the (i+1)th decoding time as the decoding time of the query result statement;

when the ith target symbol is the non-terminal symbol of the recorded statement, and a terminal symbol of the recorded statement is not a terminal symbol of the statement to be synthesized, determining the (i+1)th decoding time as the decoding time of the query result statement;

when the ith target symbol is a terminal symbol of the query result statement, and the terminal symbol of the query result statement is not the terminal symbol of the statement to be synthesized, determining the (i+1)th decoding time as the decoding time of the recorded statement; or

when the ith target symbol is the terminal symbol of the statement to be synthesized, determining the (i+1)th decoding time as a decoding time of the statement to be synthesized.

5. The method of claim 1 , wherein the performing encoding processing on the symbol sequence to obtain the feature vector set by using the preset encoding model comprises:

performing vector conversion on the symbol sequence to obtain an initial feature vector set by using the preset encoding model; and

performing non-linear transformation and feature extraction on the initial feature vector to obtain the feature vector set.

6. The method of claim 1 , wherein the performing feature conversion and synthesis on the predicted acoustic features to obtain the speech corresponding to the statement to be synthesized comprises:

performing the feature conversion on the predicted acoustic features to obtain a linear spectrum; and

performing reconstruction and synthesis on the linear spectrum to obtain the speech.

7. The method according to claim 1 , wherein the symbol sequence is a letter sequence or a phoneme sequence.

8. The method according to claim 1 , wherein before acquiring the symbol sequence of the statement to be synthesized, the method further comprises:

acquiring a sample symbol sequence corresponding to each of at least one sample synthesis statement, wherein each of the at least one sample synthesis statement represents a sample object and a reference query result for the sample object;

acquiring an initial speech synthesis model, the initial acoustic feature, and a sample acoustic feature corresponding to the at least one sample synthesis statement, wherein the initial speech synthesis model is a model configured for encoding processing, and prediction; and

training the initial speech synthesis model by using the sample symbol sequence, the initial acoustic feature, and the sample acoustic feature to obtain the preset encoding model, the preset decoding model, and the preset attention model.

9. A speech synthesis apparatus, comprising:

a memory for storing instructions executable by a processor;

the processor configured to execute the instructions to perform operations of:

acquiring a symbol sequence of a statement to be synthesized, wherein the statement to be synthesized includes a recorded statement that represents a target object and a query result statement for the target object;

performing encoding processing on the symbol sequence to obtain a feature vector set by using a preset encoding model;

acquiring a recording acoustic feature corresponding to the recorded statement;

performing prediction on an acoustic feature corresponding to the statement to be synthesized based on a preset decoding model, the feature vector set, a preset attention model, and the recording acoustic feature to obtain predicted acoustic features corresponding to the statement to be synthesized, the preset attention model being a model for generating a context vector for decoding by using the feature vector set, and the predicted acoustic features consisting of at least one related acoustic feature;

performing feature conversion and synthesis on the predicted acoustic features to obtain speech corresponding to the statement to be synthesized;

wherein the performing prediction on the acoustic feature corresponding to the statement to be synthesized based on the preset decoding model, the feature vector set, the preset attention model, and the recording acoustic feature to obtain the predicted acoustic features corresponding to the statement to be synthesized comprises:

when i is equal to 1, acquiring an initial acoustic feature at an ith decoding time, predicting a 1st acoustic feature corresponding to the statement to be synthesized based on the initial acoustic feature, the preset decoding model, the feature vector set, and the preset attention model, where the i is an integer greater than 0;

in a case where the i is greater than 1, when the ith decoding time is a decoding time of the recorded statement, acquiring a jth frame of recording acoustic feature from the recording acoustic feature, taking the jth frame of recording acoustic feature from the recording acoustic feature as an (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, and predicting an ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model, where the i is the integer greater than 0;

when the ith decoding time is a decoding time of the query result statement, taking one frame of acoustic feature of an (i−1)th acoustic feature corresponding to the statement to be synthesized as the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, and predicting the ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model;

continuing to perform a prediction process of an (i+1)th decoding time until a decoding of the statement to be synthesized is ended to obtain an nth acoustic feature corresponding to the statement to be synthesized, wherein n is a total number of frames of decoding times of the statement to be synthesized, and is an integer greater than 1; and

taking the ith acoustic feature to the nth acoustic feature corresponding to the statement to be synthesized as the predicted acoustic features; and

transmitting the speech to a playback module and playing the speech by the playback module.

10. The apparatus of claim 9 , wherein the preset decoding model comprises a first Recurrent Neural Network (RNN) and a second RNN; and the predicting the ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model comprises:

performing a first non-linear transformation on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized to obtain an intermediate feature vector;

performing a first matrix operation and a second non-linear transformation on the intermediate feature vector to obtain an ith intermediate hidden variable by using the first RNN;

performing a context vector calculation on the feature vector set and the ith intermediate hidden variable to obtain an ith context vector by using the preset attention model;

performing a second matrix operation and a third non-linear transformation on the ith context vector and the ith intermediate hidden variable to obtain an ith hidden variable by using the second RNN; and

performing a linear transformation on the ith hidden variable to obtain the ith acoustic feature corresponding to the statement to be synthesized according to a preset number of frames.

11. The apparatus of claim 10 , wherein the feature vector set comprises a feature vector corresponding to each symbol in the symbol sequence; the performing the context vector calculation on the feature vector set and the ith intermediate hidden variable to obtain the ith context vector by using the preset attention model comprises:

performing the context vector calculation on the feature vector corresponding to each symbol in the symbol sequence and the ith intermediate hidden variable to obtain an ith group of attention values by using the preset attention model; and

performing weighted summation on the feature vector set to obtain the ith context vector according to the ith group of attention values.

12. The apparatus of claim 11 , wherein the processor is further configured to execute the instructions to perform the operations of:

after predicting the ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model, and before continuing to perform the prediction process of the (i+1)th decoding time, determining an ith target symbol corresponding to a maximum attention value from the ith group of attention values; and

at least one of:

when the ith target symbol is a non-terminal symbol of the recorded statement, determining the (i+1)th decoding time as the decoding time of the recorded statement;

when the ith target symbol is a non-terminal symbol of the query result statement, determining the (i+1)th decoding time as the decoding time of the query result statement;

when the ith target symbol is the non-terminal symbol of the recorded statement, and a terminal symbol of the recorded statement is not a terminal symbol of the statement to be synthesized, determining the (i+1)th decoding time as the decoding time of the query result statement;

when the ith target symbol is a terminal symbol of the query result statement, and the terminal symbol of the query result statement is not the terminal symbol of the statement to be synthesized, determining the (i+1)th decoding time as the decoding time of the recorded statement; or

when the ith target symbol is the terminal symbol of the statement to be synthesized, determining the (i+1)th decoding time as a decoding time of the statement to be synthesized.

13. The apparatus of claim 9 , wherein the performing encoding processing on the symbol sequence to obtain the feature vector set by using the preset encoding model comprises:

performing vector conversion on the symbol sequence to obtain an initial feature vector set by using the preset encoding model; and

performing non-linear transformation and feature extraction on the initial feature vector to obtain the feature vector set.

14. The apparatus of claim 9 , wherein the performing feature conversion and synthesis on the predicted acoustic features to obtain the speech corresponding to the statement to be synthesized comprises:

performing the feature conversion on the predicted acoustic features to obtain a linear spectrum; and

performing reconstruction and synthesis on the linear spectrum to obtain the speech.

15. The apparatus according to claim 9 , wherein the symbol sequence is a letter sequence or a phoneme sequence.

16. The apparatus according to claim 9 , wherein before acquiring the symbol sequence of the statement to be synthesized, the operations further comprise:

acquiring a sample symbol sequence corresponding to each of at least one sample synthesis statement, wherein each of the at least one sample synthesis statement represents a sample object and a reference query result for the sample object;

acquiring an initial speech synthesis model, the initial acoustic feature, and a sample acoustic feature corresponding to the at least one sample synthesis statement, wherein the initial speech synthesis model is a model configured for encoding processing, and prediction; and

training the initial speech synthesis model by using the sample symbol sequence, the initial acoustic feature, and the sample acoustic feature to obtain the preset encoding model, the preset decoding model, and the preset attention model.

17. A non-transitory computer-readable storage medium, which is located in a speech synthesis apparatus and has stored a program that when executed by at least one processor in the speech synthesis apparatus, causes the at least one processor to execute a speech synthesis method, the speech synthesis method comprising:

acquiring a symbol sequence of a statement to be synthesized, wherein the statement to be synthesized includes a recorded statement that represents a target object and a query result statement for the target object;

performing encoding processing on the symbol sequence to obtain a feature vector set by using a preset encoding model;

acquiring a recording acoustic feature corresponding to the recorded statement;

performing prediction on an acoustic feature corresponding to the statement to be synthesized based on a preset decoding model, the feature vector set, a preset attention model, and the recording acoustic feature to obtain predicted acoustic features corresponding to the statement to be synthesized, wherein the preset attention model is a model for generating a context vector for decoding by using the feature vector set, and the predicted acoustic features consist of at least one related acoustic feature;

performing feature conversion and synthesis on the predicted acoustic features to obtain speech corresponding to the statement to be synthesized;

wherein the performing prediction on the acoustic feature corresponding to the statement to be synthesized based on the preset decoding model, the feature vector set, the preset attention model, and the recording acoustic feature to obtain the predicted acoustic features corresponding to the statement to be synthesized comprises:

when i is equal to 1, acquiring an initial acoustic feature at an ith decoding time, predicting a 1st acoustic feature corresponding to the statement to be synthesized based on the initial acoustic feature, the preset decoding model, the feature vector set, and the preset attention model, where the i is an integer greater than 0;

in a case where the i is greater than 1, when the ith decoding time is a decoding time of the recorded statement, acquiring a jth frame of recording acoustic feature from the recording acoustic feature, taking the jth frame of recording acoustic feature from the recording acoustic feature as an (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, and predicting an ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model, where the i is the integer greater than 0;

when the ith decoding time is a decoding time of the query result statement, taking one frame of acoustic feature of an (i−1)th acoustic feature corresponding to the statement to be synthesized as the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, and predicting the ith acoustic feature corresponding to the statement to be synthesized based on the (i−1)th frame of acoustic feature corresponding to the statement to be synthesized, the preset decoding model, the feature vector set, and the preset attention model;

continuing to perform a prediction process of an (i+1)th decoding time until a decoding of the statement to be synthesized is ended to obtain an nth acoustic feature corresponding to the statement to be synthesized, wherein n is a total number of frames of decoding times of the statement to be synthesized, and is an integer greater than 1; and

taking the ith acoustic feature to the nth acoustic feature corresponding to the statement to be synthesized as the predicted acoustic features; and

transmitting the speech to a playback module and playing the speech by the playback module.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2022
From: WU, ZHIZHENG; SONG, WEI
To: BEIJING JINGDONG SHANGKE INFORMATION TECHNOLOGY CO., LTD.; BEIJING JINGDONG CENTURY TRADING CO., LTD.
Reel/Frame 059261/0651 →
Priority Claims (1)
CN 201910878228.3 · Sep 17, 2019 · national
Continuity (1)
Related Publication 20220270587A1 · Aug 25, 2022
References Cited (28)
US 6795808B1 · Strubbe · 2004 [cited by examiner]
US 11922924B2 · Yang · 2024 [cited by examiner]
US 20160365087A1 · Freud · 2016 [cited by applicant]
US 20190065486A1 · Lin et al. · 2019 [cited by applicant]
US 20190122651A1 · Arik et al. · 2019 [cited by applicant]
US 20190180732A1 · Ping et al. · 2019 [cited by applicant]
CN 1945691A · 2007 [cited by applicant]
CN 105261355A · 2016 [cited by applicant]
CN 105355193A · 2016 [cited by applicant]
CN 107871494A · 2018 [cited by applicant]
CN 109036375A · 2018 [cited by applicant]
CN 109697974A · 2019 [cited by applicant]
CN 109767755A · 2019 [cited by applicant]
CN 109979429A · 2019 [cited by applicant]
CN 110033755A · 2019 [cited by applicant]
EP 1256932A2 · 2002 [cited by applicant]
JP 2003295880A · 2003 [cited by applicant]
JP 2006133559A · 2006 [cited by applicant]
JP 2008107454A · 2008 [cited by applicant]
JP 2019120841A · 2019 [cited by applicant]
WO 2019040173A1 · 2019 [cited by applicant]
WO 2019139430A1 · 2019 [cited by applicant]
Lee, et al. “Voice Imitating Text-to-Speech Neural Networks”, arXiv:1806.00927v1, Jun. 4, 2018 (Year: 2018). [cited by examiner]
“Hybrid Unit Selection Speech Synthesis System Target Cost Construction”, 2018, Cai Wenbin, Wei Yunlong and Xu Haihua, «Computer Engineering and Application, 54(24), http://cea.ceaj.org/CN/10.3778/j. 6 pages with Englis… [cited by applicant]
“Tacotron: A Fully End-to-End Text-to-Speech Synthesis Model”, Mar. 2017, Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zonghengy Yang, Ying Xiao, Zhifent Chen, Samy Bengio, Quoc … [cited by applicant]
“Application of Speech Synthesis Based on Tacotron Model”, Dec. 2018, Source: Penguin—Deep Learning Daily Digest, reprinted from the Internet at: https://cloud.tencent.com/developer/news/376583, 8 pgs. [cited by applicant]
International Search Report in the international application No. PCT/CN2020/079930, mailed on Jun. 23, 2020, 3 pgs. [cited by applicant]
English translation of the Written Opinion of the International Search Authority in the international application No. PCT/CN2020/079930, mailed on Jun. 23, 2020, 5 pgs. [cited by applicant]
Cited By (1)
US 12,626,686