IP Library Granted Patent US 11,929,059
Granted Patent B2
US 11,929,059 · App. 17/004,460 · Granted Mar 12, 2024

Method, device, and computer readable storage medium for text-to-speech synthesis using machine learning on basis of sequential prosody feature

Inventors: Taesu Kim (Suwon-si Gyeonggi-do, KR); Younggun Lee (Seoul, KR)
Assignee: NEOSAPIENCE, INC.
G10L13/10G06N3/08G06N20/00G10L13/027G10L13/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,929,059
App. No.
17/004,460
Granted
Mar 12, 2024
Kind
B2
Abstract

The present disclosure relates to a text-to-speech synthesis method using machine learning based on a sequential prosody feature. The text-to-speech synthesis method includes receiving input text, receiving a sequential prosody feature, and generating output speech data for the input text reflecting the received sequential prosody feature by inputting the input text and the received sequential prosody feature to an artificial neural network text-to-speech synthesis model.

Claims (35)

1. A text-to-speech synthesis method using machine learning based on a sequential prosody feature, comprising:

receiving an input text;

receiving a sequential prosody feature; and

generate output speech data for the input text reflecting the received sequential prosody feature by inputting the input text and the received sequential prosody feature to an artificial neural network text-to-speech synthesis model,

wherein receiving the sequential prosody feature includes receiving a plurality of embedding vectors representing the sequential prosody feature,

wherein the artificial neural network text-to-speech synthesis model includes an encoder and a decoder,

wherein the method further includes inputting the received plurality of embedding vectors to an attention module to generate a plurality of converted embedding vectors corresponding to respective parts of the input text provided to the encoder, wherein lengths of the plurality of converted embedding vectors varies with a length of the input text, and

wherein generating the output speech data for the input text includes:

inputting the generated plurality of converted embedding vectors to the encoder of the artificial neural network text-to-speech synthesis model, and

generating output speech data for the input text reflecting the plurality of converted embedding vectors.

2. The text-to-speech synthesis method of claim 1 , wherein the artificial neural network text-to-speech synthesis model is generated by performing machine learning based on a plurality of learning texts and data representing learning speeches corresponding to the plurality of learning texts, and

wherein the data representing the learning speeches includes sequential prosody features of the learning speeches.

3. The text-to-speech synthesis method of claim 1 , wherein the sequential prosody feature includes, in chronological order, prosody information corresponding to at least one unit from among frame, character, phoneme, syllable, or word, and

wherein the prosody information includes at least one of information on a volume of sound, information on a pitch of the sound, information on a length of sound, information on a pause duration of sound, or information on a style of the sound.

4. The text-to-speech synthesis method of claim 3 ,

wherein each of the plurality of embedding vectors corresponds to the prosody information included in the chronological order.

5. The text-to-speech synthesis method of claim 1 , wherein generating the output speech data for the input text further includes:

inputting the received plurality of embedding vectors to the decoder of the artificial neural network text-to-speech synthesis model.

6. The text-to-speech synthesis method of claim 4 , further comprising receiving an articulatory feature of a speaker,

wherein generating the output speech data for the input text includes generating output speech data for the input text, which simulates speech of the speaker and reflects a plurality of embedding vectors representing the sequential prosody feature.

7. The text-to-speech synthesis method of claim 6 , wherein receiving the articulatory feature of the speaker includes receiving a sequential prosody feature of the speaker,

wherein receiving the plurality of embedding vectors includes normalizing the received plurality of embedding vectors based on the sequential prosody feature of the speaker, and

wherein generating the output speech data for the input text includes generating the output speech data for the input text, which simulates the speech of the speaker and reflects the normalized plurality of embedding vectors.

8. The text-to-speech synthesis method of claim 7 , wherein normalizing the received plurality of embedding vectors includes:

calculating an average value of embedding vectors representing the sequential prosody feature of the speaker at each time-step, and

subtracting the received plurality of embedding vectors by the average value of the embedding vectors calculated at each time-step.

9. The text-to-speech synthesis method of claim 1 , wherein receiving the sequential prosody feature includes receiving prosody information on at least a part of the input text through a user interface, and

wherein generating the output speech data for the input text reflecting the received sequential prosody feature includes generating output speech data for the input text reflecting the prosody information on at least the part of the input text.

10. The text-to-speech synthesis method of claim 9 , wherein the prosody information on at least the part of the input text is input through a tag provided in a speech synthesis markup language.

11. The text-to-speech synthesis method of claim 1 , further comprising:

receiving prosody information on at least a part of the input text through a user interface; and

changing the received sequential prosody feature based on the received prosody information on at least a part of the input text,

wherein generating the output speech data for the input text reflecting the received sequential prosody feature includes generating output speech data for the input text reflecting the changed sequential prosody feature.

12. The text-to-speech synthesis method of claim 11 , wherein the prosody information on the at least the part of the input text, which is used to change the received sequential prosody feature, is input through a tag provided in a speech synthesis markup language.

13. A non-transitory computer-readable storage medium having a program recorded thereon, the program comprising instructions of performing operations of the text-to-speech synthesis method using machine learning based on the sequential prosody feature of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2020
From: KIM, TAESU; LEE, YOUNGGUN
To: NEOSAPIENCE, INC.
Reel/Frame 053615/0906 →
Priority Claims (2)
KR 10-2018-0090134 · Aug 2, 2018 · national
KR 10-2019-0094065 · Aug 1, 2019 · national
Continuity (2)
Continuation PCTKR2019009659 · Aug 2, 2019
Related Publication 20200394998A1 · Dec 17, 2020