IP Library Granted Patent US 12676140
Granted Patent B2
US 12676140 · App. 18/770,736 · Granted Jul 7, 2026

Speech translation method and system using multilingual text-to-speech synthesis model

Inventors: Taesu Kim (Suwon-si, KR); Younggun Lee (Seoul, KR)
Assignee: NEOSAPIENCE, INC.
G10L13/10G06F40/40G06N3/04G06N3/044G06N3/045G06N3/08G10L13/033G10L13/047G10L13/086G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12676140
App. No.
18/770,736
Granted
Jul 7, 2026
Kind
B2
Abstract

A speech translation method using a multilingual text-to-speech synthesis model includes receiving input speech data of the first language and an articulatory feature of a speaker regarding the first language, converting the input speech data of the first language into a text of the first language, converting the text of the first language into a text of the second language, and generating output speech data for the text of the second language that simulates the speaker's speech by inputting the text of the second language and the articulatory feature of the speaker to a single artificial neural network text-to-speech synthesis model.

Claims (37)

1 . A speech translation method using a multilingual text-to-speech synthesis model, comprising:

receiving input speech data of first language and an articulatory feature of a speaker regarding the first language;

converting the input speech data of the first language into a text of the first language;

converting the text of the first language into a text of second language; and

generating output speech data for the text of the second language that simulates speech of the speaker by inputting the text of the second language and the articulatory feature of the speaker to a single artificial neural network text-to-speech synthesis model, wherein the single artificial neural network text-to-speech synthesis model is trained by inputting a plurality of learning texts of the first language, learning speech data of the first language corresponding to the plurality of learning texts of the first language, a speaker vector associated with the learning speech data of the first language, a plurality of learning texts of the second language, learning speech data of the second language corresponding to the plurality of learning texts of the second language, and a speaker vector associated with the learning speech data of the second language to the single artificial neural network text-to-speech synthesis model, the second language being different from the first language, and

wherein the articulatory feature of the speaker regarding the first language is generated by extracting a feature vector from the input speech data uttered by the speaker in the first language, and

wherein each of the plurality of learning texts of the first language and the plurality of learning texts of the second language includes a plurality of text embedding vectors corresponding to text divided by units of a syllable, a character, or a phoneme.

2 . The speech translation method according to claim 1 , further comprising generating an emotion feature of the speaker regarding the first language from the input speech data of the first language,

wherein the generating the output speech data for the text of the second language that simulates the speech of the speaker includes generating the output speech data for the text of the second language that simulates the speech of the speaker by inputting the text of the second language, the articulatory feature, and the emotion feature of the speaker regarding the first language to the single artificial neural network text-to-speech synthesis model.

3 . The speech translation method according to claim 2 , wherein the emotion feature includes information on emotions inherent in a content uttered by the speaker.

4 . The speech translation method according to claim 1 , further comprising generating a prosody feature of the speaker regarding the first language from the input speech data of the first language,

wherein the generating the output speech data for the text of the second language that simulates the speech of the speaker includes generating the output speech data for the text of the second language that simulates the speech of the speaker by inputting the text of the second language, the articulatory feature, and the prosody feature of the speaker regarding the first language to the single artificial neural network text-to-speech synthesis model.

5 . The speech translation method according to claim 4 , wherein the prosody feature includes at least one of information on utterance speed, information on accentuation, information on voice pitch, and information on pause duration.

6 . The speech translation method according to claim 1 , wherein the articulatory feature of the speaker regarding the first language includes a speaker ID or a speaker embedding vector.

7 . A non-transitory computer readable storage medium having recorded thereon a program comprising instructions for performing the steps of the method according to claim 1 .

8 . A video translation method using a multilingual text-to-speech synthesis model, comprising:

receiving video data including input speech data of first language, a text of the first language corresponding to the input speech data of the first language, and an articulatory feature of a speaker regarding the first language;

deleting the input speech data of the first language from the video data;

converting the text of the first language into a text of second language;

generating output speech data for the text of the second language that simulates speech of the speaker by inputting the text of the second language and the articulatory feature of the speaker regarding the first language to a single artificial neural network text-to-speech synthesis model; and

combining the output speech data with the video data,

wherein the single artificial neural network text-to-speech synthesis model is trained by inputting a plurality of learning texts of the first language, learning speech data of the first language corresponding to the plurality of learning texts of the first language, a speaker vector associated with the learning speech data of the first language, a learning text of the second language and learning speech data of the second language corresponding to the learning text of the second language, and a speaker vector associated with the learning speech data of the second language to the single artificial neural network text-to-speech synthesis model, the second language being different from the first language, and

wherein the articulatory feature of the speaker regarding the first language is generated by extracting a feature vector from the input speech data uttered by the speaker in the first language, and

wherein each of the plurality of learning texts of the first language and the plurality of learning texts of the second language includes a plurality of text embedding vectors corresponding to text divided by units of a syllable, a character, or a phoneme.

9 . The video translation method according to claim 8 , further comprising generating an emotion feature of the speaker regarding the first language from the input speech data of the first language,

wherein the generating the output speech data for the text of the second language that simulates the speech of the speaker includes generating the output speech data for the text of the second language that simulates the speech of the speaker by inputting the text of the second language, the articulatory feature, and the emotion feature of the speaker regarding the first language to the single artificial neural network text-to-speech synthesis model.

10 . The video translation method according to claim 8 , further comprising generating a prosody feature of the speaker regarding the first language from the input speech data of the first language,

wherein the generating the output speech data for the text of the second language that simulates the speech of the speaker includes generating the output speech data for the text of the second language that simulates the speech of the speaker by inputting the text of the second language, the articulatory feature, and the prosody feature of the speaker regarding the first language to the single artificial neural network text-to-speech synthesis model.

11 . A speech translation system using a multilingual text-to-speech synthesis model, comprising:

a memory; and

at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one computer-readable program includes instructions for:

receiving input speech data of first language and an articulatory feature of a speaker regarding the first language;

converting the input speech data of the first language into a text of the first language;

converting the text of the first language into a text of second language; and

generating output speech data for the text of the second language that simulates speech of the speaker by inputting the text of the second language and the articulatory feature of the speaker to a single artificial neural network text-to-speech synthesis model, wherein the single artificial neural network text-to-speech synthesis model is trained by inputting a learning text of the first language, learning speech data of the first language corresponding to the learning text of the first language, a speaker vector associated with the learning speech data of the first language, a learning text of the second language, learning speech data of the second language corresponding to the learning text of the second language, and a speaker vector associated with the learning speech data of the second language to the single artificial neural network text-to-speech synthesis model, the second language being different from the first language, and

wherein the articulatory feature of the speaker regarding the first language is generated by extracting a feature vector from the input speech data uttered by the speaker in the first language, and

wherein each of the plurality of learning texts of the first language and the plurality of learning texts of the second language includes a plurality of text embedding vectors corresponding to text divided by units of a syllable, a character, or a phoneme.