IP Library › Granted Patent US 12,148,415
Granted Patent B2
US 12,148,415 · App. 18/465,143 · Granted Nov 19, 2024

Systems and methods for synthesizing speech

Inventors: Peng Zhang (Hangzhou, CN); Xinhui Hu (Hangzhou, CN); Xinkang Xu (Hangzhou, CN); Jian Lu (Hangzhou, CN)
Assignee: ZHEJIANG TONGHUASHUN INTELLIGENT TECHNOLOGY CO., LTD.
G10L13/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,148,415
App. No.
18/465,143
Granted
Nov 19, 2024
Kind
B2
Abstract

The present disclosure discloses a method for synthesizing a speech. The method includes generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer; and training the speech synthesis model when an evaluation index meets a preset condition, wherein the evaluation index includes one or more quality indexes determined based on at least a part of the text and at least a part of the speech.

Claims (54)

1. A method, that is implemented on a computing device having at least one processor and at least one storage medium including a set of instructions for synthesizing a speech, comprising:

generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model is configured to output the speech corresponding to the text and a stop token indicating where the speech should stop;

obtaining an evaluation index, wherein the evaluation index includes a second effect score of the speech synthesis model, and the second effect score is generated based on at least one of a duration of the speech and a correct ending position of a sentence corresponding to the speech; and

training the speech synthesis model when the evaluation index meets a preset condition.

2. The method of claim 1 , wherein the stop token is used to determine the end of the sentence corresponding to the speech, and the second effect score is used to evaluate the accuracy of the stop token predicted with the speech synthesis model.

3. The method of claim 2 , wherein obtaining the evaluation index further includes:

designating the sentence corresponding to the speech that has a duration greater than or equal to a first target threshold as an abnormal sentence.

4. The method of claim 2 , wherein obtaining the evaluation index further includes:

obtaining a recognized result based on the speech and determining a correct ending position of the sentence corresponding to the speech based on the recognized result;

determining a correct ending position of an abnormal sentence based on the recognized result; and

designating the sentence corresponding to the speech that does not end at the correct ending position as the abnormal sentence.

5. The method of claim 4 , wherein the recognized result includes valid phoneme and invalid phoneme, and the determining the correct ending position of the sentence corresponding to the speech based on the recognized result includes:

comparing the recognized result with the text corresponding to the speech; and

determining the correct ending position of the sentence corresponding to the speech based on the comparison result.

6. The method of claim 1 , wherein the second effect score is represented by a count of abnormal sentence, and the preset condition includes the count of the abnormal sentence is greater than or equal to a second target threshold.

7. The method of claim 1 , wherein the evaluation index further includes a first effect score of the speech synthesis model, and the method further including:

obtaining the first effect score of the speech synthesis model based on one or more of a total count of the speech frames, a total count of the characters, and a first weight matrix;

wherein the first weight matrix is generated based on the text with the speech synthesis model, and elements in the first weight matrix are configured to represent a probability that speech frame of the speech is aligned with characters of the text.

8. The method of claim 1 , wherein:

the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer.

9. The method of claim 8 , wherein the training the speech synthesis model when the evaluation index meets the preset condition includes:

training the speech synthesis model based on an abnormal training database, wherein the abnormal training database is configured to store an abnormal sentence and the correct ending position of the abnormal sentence.

10. The method of claim 9 , wherein the training includes:

generating an embedding feature based on the abnormal sentence in the abnormal training database by the embedding layer; and

training the position layer based on the embedding feature, wherein the position layer takes the embedding feature as a training sample and the correct ending position of the abnormal sentence as a label, and the position layer is configured to update the speech synthesis model.

11. A system for synthesizing a speech, comprising:

at least one storage medium including a set of instructions; and

at least one processor in communication with the at least one storage medium;

wherein when executing the set of instructions, the at least one processor is configured to direct the system to perform operations including:

generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model is configured to output the speech corresponding to the text and a stop token indicating where the speech should stop;

obtaining an evaluation index, wherein the evaluation index includes a second effect score of the speech synthesis model, and the second effect score is generated based on at least one of a duration of the speech and a correct ending position of a sentence corresponding to the speech; and

training the speech synthesis model when the evaluation index meets a preset condition.

12. The system of claim 11 , wherein the stop token is used to determine the end of the sentence corresponding to the speech, and the second effect score is used to evaluate the accuracy of the stop token predicted with the speech synthesis model.

13. The system of claim 12 , wherein obtaining the evaluation index further includes:

designating the sentence corresponding to the speech that has a duration greater than or equal to a first target threshold as an abnormal sentence.

14. The system of claim 12 , wherein obtaining the evaluation index further includes:

obtaining a recognized result based on the speech and determining a correct ending position of the sentence corresponding to the speech based on the recognized result;

determining a correct ending position of an abnormal sentence based on the recognized result; and

designating the sentence corresponding to the speech that does not end at the correct ending position as the abnormal sentence.

15. The system of claim 14 , wherein the recognized result includes valid phoneme and invalid phoneme, and the determining the correct ending position of the sentence corresponding to the speech based on the recognized result includes:

comparing the recognized result with the text corresponding to the speech; and

determining the correct ending position of the sentence corresponding to the speech based on the comparison result.

16. The system of claim 11 , wherein the second effect score is represented by a count of abnormal sentence, and the preset condition includes the count of the abnormal sentence is greater than or equal to a second target threshold.

17. The system of claim 11 , wherein:

the speech synthesis model includes an embedding layer, a speech synthesis layer, and a position layer.

18. The system of claim 17 , wherein the training the speech synthesis model when the evaluation index meets the preset condition includes:

training the speech synthesis model based on an abnormal training database, wherein the abnormal training database is configured to store an abnormal sentence and the correct ending position of the abnormal sentence.

19. The system of claim 18 , wherein the training includes:

generating an embedding feature based on the abnormal sentence in the abnormal training database by the embedding layer; and

training the position layer based on the embedding feature, wherein the position layer takes the embedding feature as a training sample and the correct ending position of the abnormal sentence as a label, and the position layer is configured to update the speech synthesis model.

20. A non-transitory computer-readable storage medium, comprising instructions that, when executed by at least one processor, direct the at least processor to perform a method for synthesizing a speech, the method comprising:

generating the speech based on a text with a speech synthesis model, wherein the speech synthesis model is configured to output the speech corresponding to the text and a stop token indicating where the speech should stop;

obtaining an evaluation index, wherein the evaluation index includes a second effect score of the speech synthesis model, and the second effect score is generated based on at least one of a duration of the speech and a correct ending position of a sentence corresponding to the speech; and

training the speech synthesis model when the evaluation index meets a preset condition.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: ZHANG, PENG; HU, XINHUI; XU, XINKANG; LU, JIAN
To: ZHEJIANG TONGHUASHUN INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 064990/0752 →
Priority Claims (2)
CN 202010835266.3 · Aug 19, 2020 · national
CN 202011148521.3 · Oct 23, 2020 · national
Continuity (2)
Continuation 17445385 · Aug 18, 2021
Related Publication 20230419948A1 · Dec 28, 2023