IP Library Granted Patent US 11,769,483
Granted Patent B2
US 11,769,483 · App. 17/533,459 · Granted Sep 26, 2023

Multilingual text-to-speech synthesis

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,769,483
App. No.
17/533,459
Filed
Nov 23, 2021
Granted
Sep 26, 2023
Kind
B2
Art Unit
2656
USPC
704/259
Abstract

A multilingual text-to-speech synthesis method and system are disclosed. The method includes receiving an articulatory feature of a speaker regarding a first language, receiving an input text of a second language, and generating output speech data for the input text of the second language that simulates the speaker's speech by inputting the input text of the second language and the articulatory feature of the speaker regarding the first language to a single artificial neural network multilingual text-to-speech synthesis model. The single artificial neural network multilingual text-to-speech synthesis model is generated by learning similarity information between phonemes of the first language and phonemes of the second language based on a first learning data of the first language and a second learning data of the second language.

Claims (33)

1. A method for multilingual text-to-speech synthesis, comprising:

receiving an articulatory feature of a speaker regarding a first language;

receiving an input text of a second language; and

generating output speech data for the input text of the second language that simulates the speaker's speech by inputting the input text of the second language and the articulatory feature of the speaker regarding the first language to a single artificial neural network multilingual text-to-speech synthesis model,

wherein the single artificial neural network multilingual text-to-speech synthesis model is generated by learning similarity information between phonemes of the first language and phonemes of the second language based on a first learning data of the first language and a second learning data of the second language,

the first learning data does not include a language identifier for the first language, and

the second learning data does not include a language identifier for the second language.

2. The method of claim 1 , wherein the first learning data of the first language includes a learning text of the first language and learning speech data of the first language corresponding to the learning text of the first language, and

the second learning data of the second language includes a learning text of the second language and learning speech data of the second language corresponding to the learning text of the second language.

3. The method of claim 1 , further comprising:

receiving an emotion feature of the speaker in the first language; and

generating output speech data for the input text of the second language that simulates the speaker's speech and emotion by inputting the input text of the second language, the articulatory feature of the speaker regarding the first language, and the emotion feature to the single artificial neural network multilingual text-to-speech synthesis model.

4. The method of claim 1 , further comprising:

receiving a prosody feature of the speaker in the first language; and

generating output speech data for the input text of the second language that simulates the speaker's speech and prosody by inputting the input text of the second language, the articulatory feature of the speaker regarding the first language, and the prosody feature to the single artificial neural network multilingual text-to-speech synthesis model.

5. The method of claim 4 , wherein the prosody feature includes at least one of information on utterance speed, information on accentuation, information on voice pitch, or information on pause duration.

6. The method of claim 1 , wherein receiving the articulatory feature includes:

receiving an input speech of the first language; and

extracting a feature vector from the input speech of the first language to generate the articulatory feature of the speaker regarding the first language.

7. The method of claim 6 , wherein receiving the input text of the second language includes:

converting the input speech of the first language into an input text of the first language; and

converting the input text of the first language into an input text of the second language.

8. A non-transitory computer-recordable storage medium having recorded thereon a program comprising instructions for performing each step according to the method of claim 1 .

9. A system for multilingual text-to-speech synthesis, comprising:

a memory; and

at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory,

wherein the at least one computer-readable program includes instructions for:

receiving an articulatory feature of a speaker regarding a first language;

receiving an input text of a second language; and

generating output speech data for the input text of the second language that simulates the speaker's speech by inputting the input text of the second language and the articulatory feature of the speaker regarding the first language to a single artificial neural network multilingual text-to-speech synthesis model,

wherein the single artificial neural network multilingual text-to-speech synthesis model is generated by learning similarity information between phonemes of the first language and phonemes of the second language based on a first learning data of the first language and a second learning data of the second language, and

the first learning data does not include a language identifier for the first language, and

the second learning data does not include a language identifier for the second language.