IP Library › Granted Patent US 11,488,578
Granted Patent B2
US 11,488,578 · App. 17/205,121 · Granted Nov 1, 2022

Method and apparatus for training speech spectrum generation model, and electronic device

Inventors: Zhijie Chen (Beijing, CN); Tao Sun (Beijing, CN); Lei Jia (Beijing, CN)
Assignee: Beijing Baidu Netcom Science and Technology Co., Ltd.
G10L13/047G10L25/18G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,488,578
App. No.
17/205,121
Granted
Nov 1, 2022
Kind
B2
Abstract

The present application discloses a method and an apparatus for training a speech spectrum generation model, as well as an electronic device, and relates to the technical field of speech synthesis and deep learning. A specific implementation is as follows: inputting a first text sequence into the speech spectrum generation model to generate an analog spectrum sequence corresponding to the first text sequence, and obtain a first loss value of the analog spectrum sequence according to a preset loss function; inputting the analog spectrum sequence corresponding to the first text sequence into an adversarial loss function model, which is a generative adversarial network model, to obtain a second loss value of the analog spectrum sequence; and training the speech spectrum generation model based on the first loss value and the second loss value.

Claims (62)

1. A method for training a speech spectrum generation model, the method comprising:

inputting a first text sequence as a real spectrum sequence into the speech spectrum generation model to generate an analog spectrum sequence corresponding to the first text sequence, and to obtain a first loss value of the analog spectrum sequence according to a preset loss function, wherein the first loss value represents a loss in terms of intelligibility of the analog spectrum sequence relative to the real spectrum sequence;

inputting the analog spectrum sequence corresponding to the first text sequence into an adversarial loss function model, which is a generative adversarial network model, to obtain a second loss value of the analog spectrum sequence, wherein the second loss value represents a loss in terms of clarity of the analog spectrum sequence relative to the real spectrum sequence; and

feeding the first loss value and the second loss value back to the speech spectrum generation model according to a preset ratio that is determined depending on characterisitics of speakers in different sound banks, and training the speech spectrum generation model based on the first loss value and the second loss value fed back.

2. The method according to claim 1 , wherein prior to inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence, the method further comprises:

obtaining a real spectrum sequence corresponding to a second text sequence, and an analog spectrum sequence corresponding to the second text sequence, which is generated by the speech spectrum generation model; and

training the adversarial loss function model based on the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence, and

wherein inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence comprises inputting the analog spectrum sequence corresponding to the first text sequence into the trained adversarial loss function model to obtain the second loss value.

3. The method according to claim 2 , wherein training the adversarial loss function model based on the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence comprises:

inputting the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence into the adversarial loss function model separately to obtain a third loss value; and

training the adversarial loss function model based on the third loss value,

wherein the third loss value represents a loss of the analog spectrum sequence corresponding to the second text sequence relative to the real spectrum sequence corresponding to the second text sequence.

4. The method according to claim 1 , wherein inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence comprises:

inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain an original loss value;

down-sampling the analog spectrum sequence corresponding to the first text sequence N times to obtain each down-sampled analog spectrum sequence;

inputting each down-sampled analog spectrum sequence into the adversarial loss function model separately to obtain a loss value corresponding to said down-sampled analog spectrum sequence; and

obtaining the second loss value based on the loss values corresponding to all of the down-sampled analog spectrum sequences and the original loss value.

5. The method according to claim 1 , wherein the adversarial loss function model adopts a deep convolutional neural network model.

6. The method according to claim 1 , wherein the speech spectrum generation model comprises a Tacotron model.

7. The method according to claim 1 , wherein the speech spectrum generation model comprises a Text To Speech (TTS) model.

8. An electronic device, comprising:

at least one processor; and

a memory communicatively coupled to the at least one processor,

wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform a method for training a speech spectrum generation model, the method comprising:

inputting a first text sequence as a real spectrum sequence into the speech spectrum generation model to generate an analog spectrum sequence corresponding to the first text sequence, and to obtain a first loss value of the analog spectrum sequence according to a preset loss function, wherein the first loss value represents a loss in terms of intelligibility of the analog spectrum sequence relative to the real spectrum sequence;

inputting the analog spectrum sequence corresponding to the first text sequence into an adversarial loss function model, which is a generative adversarial network model, to obtain a second loss value of the analog spectrum sequence, wherein the second loss value represents a loss in terms of clarity of the analog spectrum sequence relative to the real spectrum sequence; and

feeding the first loss value and the second loss value back to the speech spectrum generation model according to a preset ratio that is determined depending on characterisitics of speakers in different sound banks, and training the speech spectrum generation model based on the first loss value and the second loss value fed back.

9. The electronic device according to claim 8 , wherein prior to inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence, the method further comprises:

obtaining a real spectrum sequence corresponding to a second text sequence, and an analog spectrum sequence corresponding to the second text sequence, which is generated by the speech spectrum generation model; and

training the adversarial loss function model based on the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence, and

wherein inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence comprises inputting the analog spectrum sequence corresponding to the first text sequence into the trained adversarial loss function model to obtain the second loss value.

10. The electronic device according to claim 9 , wherein training the adversarial loss function model based on the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence comprises:

inputting the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence into the adversarial loss function model separately to obtain a third loss value; and

training the adversarial loss function model based on the third loss value,

wherein the third loss value represents a loss of the analog spectrum sequence corresponding to the second text sequence relative to the real spectrum sequence corresponding to the second text sequence.

11. The electronic device according to claim 8 , wherein inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence comprises:

inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain an original loss value;

down-sampling the analog spectrum sequence corresponding to the first text sequence N times to obtain each down-sampled analog spectrum sequence;

inputting each down-sampled analog spectrum sequence into the adversarial loss function model separately to obtain a loss value corresponding to said down-sampled analog spectrum sequence; and

obtaining the second loss value based on the loss values corresponding to all of the down-sampled analog spectrum sequences and the original loss value.

12. The electronic device according to claim 8 , wherein the adversarial loss function model adopts a deep convolutional neural network model.

13. The electronic device according to claim 8 , wherein the speech spectrum generation model comprises a Tacotron model.

14. The electronic device according to claim 8 , wherein the speech spectrum generation model comprises a Text To Speech (TTS) model.

15. A non-transitory computer-readable storage medium having computer instructions stored thereon, which are used to realize a method for training a speech spectrum generation model, the method comprising:

inputting a first text sequence as a real spectrum sequence into the speech spectrum generation model to generate an analog spectrum sequence corresponding to the first text sequence, and to obtain a first loss value of the analog spectrum sequence according to a preset loss function, wherein the first loss value represents a loss in terms of intelligibility of the analog spectrum sequence relative to the real spectrum sequence;

inputting the analog spectrum sequence corresponding to the first text sequence into an adversarial loss function model, which is a generative adversarial network model, to obtain a second loss value of the analog spectrum sequence, wherein the second loss value represents a loss in terms of clarity of the analog spectrum sequence relative to the real spectrum sequence; and

feeding the first loss value and the second loss value back to the speech spectrum generation model according to a preset ratio that is determined depending on characterisitics of speakers in different sound banks, and training the speech spectrum generation model based on the first loss value and the second loss value fed back.

16. The non-transitory computer-readable storage medium according to claim 15 , wherein prior to inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence, the method further comprises:

obtaining a real spectrum sequence corresponding to a second text sequence, and an analog spectrum sequence corresponding to the second text sequence, which is generated by the speech spectrum generation model; and

training the adversarial loss function model based on the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence, and

wherein inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence comprises inputting the analog spectrum sequence corresponding to the first text sequence into the trained adversarial loss function model to obtain the second loss value.

17. The non-transitory computer-readable storage medium according to claim 16 , wherein training the adversarial loss function model based on the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence comprises:

inputting the real spectrum sequence corresponding to the second text sequence and the analog spectrum sequence corresponding to the second text sequence into the adversarial loss function model separately to obtain a third loss value; and

training the adversarial loss function model based on the third loss value,

wherein the third loss value represents a loss of the analog spectrum sequence corresponding to the second text sequence relative to the real spectrum sequence corresponding to the second text sequence.

18. The non-transitory computer-readable storage medium according to claim 15 , wherein inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain the second loss value of the analog spectrum sequence comprises:

inputting the analog spectrum sequence corresponding to the first text sequence into the adversarial loss function model to obtain an original loss value;

down-sampling the analog spectrum sequence corresponding to the first text sequence N times to obtain each down-sampled analog spectrum sequence;

inputting each down-sampled analog spectrum sequence into the adversarial loss function model separately to obtain a loss value corresponding to said down-sampled analog spectrum sequence; and

obtaining the second loss value based on the loss values corresponding to all of the down-sampled analog spectrum sequences and the original loss value.

19. The non-transitory computer-readable storage medium according to claim 15 , wherein the adversarial loss function model adopts a deep convolutional neural network model.

20. The non-transitory computer-readable storage medium according to claim 15 , wherein the speech spectrum generation model comprises a Tacotron model or a Text To Speech (TTS) model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2021
From: CHEN, ZHIJIE; SUN, TAO; JIA, LEI
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 056830/0200 →
Priority Claims (1)
CN 202010858104.1 · Aug 24, 2020 · national
Continuity (1)
Related Publication 20210201887A1 · Jul 1, 2021