IP Library Granted Patent US 12,118,979
Granted Patent B2
US 12,118,979 · App. 18/346,694 · Granted Oct 15, 2024

Text-to-speech synthesis system and method

Inventors: Piero Perucci (Zurich, CH); Martin Reber (Zurich, CH); Vijeta Avijeet (Zurich, CH)
Assignee: Telepathy Labs, Inc.
G10L13/047G06N3/08G10L13/08G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,118,979
App. No.
18/346,694
Granted
Oct 15, 2024
Kind
B2
Abstract

A method, computer program product, and computer system for text-to-speech synthesis is disclosed. Synthetic speech data for an input text may be generated. The synthetic speech data may be compared to recorded reference speech data corresponding to the input text. Based on, at least in part, the comparison of the synthetic speech data to the recorded reference speech data, at least one feature indicative of at least one difference between the synthetic speech data and the recorded reference speech data may be extracted. A speech gap filling model may be generated based on, at least in part, the at least one feature extracted. A speech output may be generated based on, at least in part, the speech gap filling model.

Claims (51)

1. A computing system including one or more processors and one or more memories configured to perform operations comprising:

comparing synthetic speech data for an input text to recorded reference speech data corresponding to the input text;

extracting at least one feature indicative of at least one difference between the synthetic speech data and the recorded reference speech data based on, at least in part, the comparison of the synthetic speech data to the recorded reference speech data;

generating a speech gap filling model based on, at least in part, the at least one feature extracted;

generating a speech output based on, at least in part, the speech gap filling model;

comparing the speech output generated for a second input text to recorded reference speech data corresponding to the second input text; and

extracting an updated at least one feature indicative of at least one difference between the speech output generated for the second input text and the recorded reference speech data corresponding to the second input text based on, at least in part, the comparison of the speech output for the second input text to the recorded reference speech data corresponding to the second input text.

2. The computing system of claim 1 , wherein generating the speech output comprises:

generating an interim set of parameters;

processing the interim set of parameters based on, at least in part, the speech gap filling model to generate a final set of parameters; and

generating the speech output based on, at least in part, the final set of parameters.

3. The computing system of claim 1 , wherein the synthetic speech data generated is based on, at least in part, at least one of a parametric acoustic model and a linguistic model pre-configured for a speaker.

4. The computing system of claim 1 , wherein the synthetic speech data generated is further based on, at least in part, the recorded reference speech data pre-recorded by a speaker.

5. The computing system of claim 1 further comprising aligning the synthetic speech data and the recorded reference speech data preceding the comparison.

6. The computing system of claim 5 , wherein aligning the synthetic speech data and the recorded reference speech data comprises implementing one or more of pitch shifting, time normalization, and time alignment between the synthetic speech data and the recorded reference speech data.

7. The computing system of claim 1 further comprising training a neural network based on, at least in part, the at least one feature to generate the speech gap filling model.

8. The computing system of claim 1 further comprising updating the speech gap filling model based on, at least in part, the updated at least one feature.

9. A computer-implemented method, comprising:

comparing synthetic speech data for an input text to recorded reference speech data corresponding to the input text;

extracting at least one feature indicative of at least one difference between the synthetic speech data and the recorded reference speech data based on, at least in part, the comparison of the synthetic speech data to the recorded reference speech data;

generating a speech gap filling model based on, at least in part, the at least one feature extracted; and

generating a speech output based on, at least in part, the speech gap filling model;

comparing the speech output generated for a second input text to recorded reference speech data corresponding to the second input text; and

extracting an updated at least one feature indicative of at least one difference between the speech output generated for the second input text and the recorded reference speech data corresponding to the second input text based on, at least in part, the comparison of the speech output for the second input text to the recorded reference speech data corresponding to the second input text.

10. The computer-implemented method of claim 9 , wherein generating the speech output comprises:

generating an interim set of parameters;

processing the interim set of parameters based on, at least in part, the speech gap filling model to generate a final set of parameters; and

generating the speech output based on, at least in part, the final set of parameters.

11. The computer-implemented method of claim 9 , wherein the synthetic speech data generated is based on, at least in part, at least one of a parametric acoustic model and a linguistic model pre-configured for a speaker.

12. The computer-implemented method of claim 9 , wherein the synthetic speech data generated is further based on, at least in part, the recorded reference speech data pre-recorded by a speaker.

13. The computer-implemented method of claim 9 further comprising aligning the synthetic speech data and the recorded reference speech data preceding the comparison.

14. The computer-implemented method of claim 13 , wherein aligning the synthetic speech data and the recorded reference speech data comprises implementing one or more of pitch shifting, time normalization, and time alignment between the synthetic speech data and the recorded reference speech data.

15. The computer-implemented method of claim 9 further comprising training a neural network based on, at least in part, the at least one feature to generate the speech gap filling model.

16. The computer-implemented method of claim 9 further comprising updating the speech gap filling model based on, at least in part, the updated at least one feature.

17. A computer program product residing on a computer readable storage medium having a plurality of instructions stored thereon which, when executed across one or more processors, causes at least a portion of the one or more processors to perform operations comprising:

comparing synthetic speech data for an input text to recorded reference speech data corresponding to the input text;

extracting at least one feature indicative of at least one difference between the synthetic speech data and the recorded reference speech data based on, at least in part, the comparison of the synthetic speech data to the recorded reference speech data;

generating a speech gap filling model based on, at least in part, the at least one feature extracted;

generating a speech output based on, at least in part, the speech gap filling model;

comparing the speech output generated for a second input text to recorded reference speech data corresponding to the second input text; and

extracting an updated at least one feature indicative of at least one difference between the speech output generated for the second input text and the recorded reference speech data corresponding to the second input text based on, at least in part, the comparison of the speech output for the second input text to the recorded reference speech data corresponding to the second input text.

18. The computer program product of claim 17 , wherein generating the speech output comprises:

generating an interim set of parameters;

processing the interim set of parameters based on, at least in part, the speech gap filling model to generate a final set of parameters; and

generating the speech output based on, at least in part, the final set of parameters.

19. The computer program product of claim 17 , wherein the synthetic speech data generated is based on, at least in part, at least one of a parametric acoustic model and a linguistic model pre-configured for a speaker.

20. The computer program product of claim 17 , wherein the synthetic speech data generated is further based on, at least in part, the recorded reference speech data pre-recorded by a speaker.

21. The computer program product of claim 17 further comprising aligning the synthetic speech data and the recorded reference speech data preceding the comparison.

22. The computer program product of claim 21 , wherein aligning the synthetic speech data and the recorded reference speech data comprises implementing one or more of pitch shifting, time normalization, and time alignment between the synthetic speech data and the recorded reference speech data.

23. The computer program product of claim 17 further comprising training a neural network based on, at least in part, the at least one feature to generate the speech gap filling model.

24. The computer program product of claim 17 further comprising updating the speech gap filling model based on, at least in part, the updated at least one feature.

Continuity (4)
Continuation 17880007 · Aug 3, 2022
Continuation 17041822
Provisional Application 62649312 · Mar 28, 2018
Related Publication 20230368775A1 · Nov 16, 2023