IP Library › Granted Patent US 11,798,527
Granted Patent B2
US 11,798,527 · App. 17/445,385 · Granted Oct 24, 2023

Systems and methods for synthesizing speech

Inventors: Peng Zhang (Hangzhou, CN); Xinhui Hu (Hangzhou, CN); Xinkang Xu (Hangzhou, CN); Jian Lu (Hangzhou, CN)
Assignee: ZHEJIANG TONGHU ASHUN INTELLIGENT TECHNOLOGY CO., LTD.
G10L13/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,798,527
App. No.
17/445,385
Granted
Oct 24, 2023
Kind
B2
Abstract

The present disclosure discloses a method for synthesizing a speech. The method includes generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer; and training the speech synthesis model when an evaluation index meets a preset condition, wherein the evaluation index includes one or more quality indexes determined based on at least a part of the text and at least a part of the speech.

Claims (53)

1. A method, that is implemented on a computing device having at least one processor and at least one storage medium including a set of instructions for synthesizing a speech, comprising:

generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer; and

training the speech synthesis model when an evaluation index meets a preset condition, wherein the evaluation index includes one or more quality indexes determined based on at least a part of the text and at least a part of the speech,

wherein the evaluation index includes a first weight matrix; and the method further comprises:

generating the first weight matrix based on the text with the speech synthesis model, wherein elements in the first weight matrix are configured to represent a probability that speech frames of the speech is aligned with characters of the text.

2. The method of claim 1 , wherein:

the speech synthesis model includes end-to-end models based on attention mechanisms.

3. The method of claim 1 , further including:

obtaining a first effect score of the speech synthesis model based on one or more of a total count of the speech frames, a total count of the characters, and the first weight matrix.

4. The method of claim 1 , further including:

determining an importance index of each weight in the first weight matrix;

generating a second weight matrix based on the first weight matrix.

5. The method of claim 1 , wherein the evaluation index includes a second effect score of the speech synthesis model, and the method further comprises:

generating the second effect score of the speech synthesis model based on at least one of a duration of the speech and a correct ending position of a sentence corresponding to the speech.

6. The method of claim 5 , further including:

designating the sentence corresponding to the speech that does not end at the correct ending position as an abnormal sentence, and determining the correct ending position of the abnormal sentence.

7. The method of claim 6 , wherein the determining the correct ending position of the abnormal sentence includes:

obtaining a recognized result based on the speech;

determining the correct ending position of the abnormal sentence based on the recognized result.

8. The method of claim 1 , wherein the training the speech synthesis model when the evaluation index meets the preset condition includes:

training the speech synthesis model based on an abnormal training database, wherein the abnormal training database is configured to store the abnormal sentence and the correct ending position of the abnormal sentence.

9. The method of claim 8 , wherein the training includes:

generating an embedding feature based on the abnormal sentence in the abnormal training database by the embedding layer;

training the position layer based on the embedding feature, wherein the position layer takes the embedding feature as a training sample and the correct ending position of the abnormal sentence as a label, and the position layer is configured to update the speech synthesis model.

10. A system for synthesizing a speech, comprising:

at least one storage medium including a set of instructions; and

at least one processor in communication with the at least one storage medium; wherein when executing the set of instructions, the at least one processor is configured to direct the system to perform operations including:

generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer; and

training the speech synthesis model when an evaluation index meets a preset condition, wherein the evaluation index includes one or more quality indexes determined based on at least a part of the text and at least a part of the speech,

wherein the evaluation index includes a first weight matrix; and the at least one processor is further configured to direct the system to perform:

generating the first weight matrix based on the text with the speech synthesis model, wherein elements in the first weight matrix are configured to represent a probability that speech frames of the speech is aligned with characters of the text.

11. The system of claim 10 , further including:

obtaining a first effect score of the speech synthesis model based on one or more of a total count of the speech frames, a total count of the characters, and the first weight matrix.

12. The system of claim 10 further including:

determining an importance index of each weight in the first weight matrix;

generating a second weight matrix based on the first weight matrix.

13. The system of claim 10 , wherein the evaluation index includes a second effect score of the speech synthesis model, and the method further comprises:

generating the second effect score of the speech synthesis model based on at least one of a duration of the speech and a correct ending position of a sentence corresponding to the speech.

14. The system of claim 13 , further including:

designating the sentence corresponding to the speech that does not end at the correct ending position as an abnormal sentence, and determining the correct ending position of the abnormal sentence.

15. The system of claim 14 , wherein the determining the correct ending position of the abnormal sentence includes:

obtaining a recognized result based on the speech;

determining the correct ending position of the abnormal sentence based on the recognized result.

16. The system of claim 10 , wherein the training the speech synthesis model when the evaluation index meets the preset condition includes:

training the speech synthesis model based on an abnormal training database, wherein the abnormal training database is configured to store the abnormal sentence and the correct ending position of the abnormal sentence.

17. The system of claim 16 , wherein the training includes:

generating an embedding feature based on the abnormal sentence in the abnormal training database by the embedding layer;

training the position layer based on the embedding feature, wherein the position layer takes the embedding feature as a training sample and the correct ending position of the abnormal sentence as a label, and the position layer is configured to update the speech synthesis model.

18. A non-transitory computer-readable storage medium, comprising instructions that, when executed by at least one processor, direct the at least processor to perform a method for synthesizing a speech, the method comprising:

generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer; and

training the speech synthesis model when an evaluation index meets a preset condition, wherein the evaluation index includes one or more quality indexes determined based on at least a part of the text and at least a part of the speech,

wherein the evaluation index includes a first weight matrix; and the method further comprises:

generating the first weight matrix based on the text with the speech synthesis model, wherein elements in the first weight matrix are configured to represent a probability that speech frames of the speech is aligned with characters of the text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2021
From: ZHANG, PENG; HU, XINHUI; XU, XINKANG; LU, JIAN
To: ZHEJIANG TONGHUASHUN INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 057291/0445 →
Priority Claims (2)
CN 202010835266.3 · Aug 19, 2020 · national
CN 202011148521.3 · Oct 23, 2020 · national
Continuity (1)
Related Publication 20220059072A1 · Feb 24, 2022