IP Library Granted Patent US 11,393,447
Granted Patent B2
US 11,393,447 · App. 16/498,360 · Granted Jul 19, 2022

Speech synthesizer using artificial intelligence, method of operating speech synthesizer and computer-readable recording medium

Inventor: Jonghoon Chae (Seoul, KR)
Assignee: LG ELECTRONICS INC.
G10L13/02G06N3/08G10L13/08G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,393,447
App. No.
16/498,360
Granted
Jul 19, 2022
Kind
B2
Abstract

A speech synthesizer using artificial intelligence includes a memory, a communication unit for receiving utterance information of words uttered by a user, and a processor for acquiring a plurality of utterance intonation phrase (IP) ratios respectively corresponding to a plurality of words uttered by the user based on the utterance information, acquiring a plurality of non-utterance IP ratios respectively corresponding to a plurality of unuttered words based on the utterance information and the plurality of utterance IP ratios, and generating a personalized synthesized speech model based on the plurality of utterance IP ratios and the plurality of non-utterance IP ratios. A plurality of classes indicating reading break of a word includes first to third classes. A minor class has a smallest count among the first to third classes. Each of the utterance and non-utterance IP ratios is a ratio in which a word is classified as the minor class.

Claims (41)

1. A speech synthesizer using artificial intelligence, comprising:

a memory;

a communication unit configured to receive utterance information of words uttered by a user from a terminal; and

a processor configured to acquire a plurality of utterance intonation phrase (IP) ratios respectively corresponding to a plurality of words uttered by the user based on the utterance information, acquire a plurality of non-utterance IP ratios respectively corresponding to a plurality of unuttered words based on the utterance information and the plurality of utterance IP ratios, and generate a personalized synthesized speech model based on the plurality of utterance IP ratios and the plurality of non-utterance IP ratios, wherein the personalized synthesized speech model is a model for outputting a synthesized speech, to which reading break of words uttered by the user is applied, and is an artificial neural network based model learned by a deep learning algorithm or a machine learning algorithm,

wherein a plurality of classes indicating reading break of a word includes a first class corresponding to first reading break, a second class corresponding to second reading break greater than the first reading break and a third class corresponding to third reading break greater than the second reading break,

wherein a minor class has a smallest count among the first to third classes, and

wherein each of the utterance IP ratios and the non-utterance IP ratios is a ratio in which a word is classified as the minor class.

2. The speech synthesizer according to claim 1 , wherein the utterance information comprises reading break of each uttered word, a part of speech of each uttered word, and a position of each uttered word in a sentence.

3. The speech synthesizer according to claim 2 ,

wherein the processor is further configured to acquire an utterance IP ratio of each uttered word using a number of times that the user reads each uttered word with break corresponding to the minor class.

4. The speech synthesizer according to claim 3 ,

wherein the memory stores an IP ratio model for determining an IP ratio of an unuttered word, and

wherein the processor is further configured to:

determine a probability in which the unuttered word is classified as the minor class, using each unuttered word, a property of each unuttered word, an IP ratio of an uttered word and labeling data, and

determine a probability of being classified as a determined IP class as a non-utterance IP ratio of the unuttered word.

5. The speech synthesizer according to claim 1 , wherein the personalized synthesized speech model is a model for inferring a probability that each word is classified as the minor class, using text data corresponding to the plurality of uttered words and the plurality of unuttered words, an IP ratio of each word, and a labeled probability of being classified as the minor class, which is labeled with each word.

6. The speech synthesizer according to claim 5 , wherein the processor is further configured to transmit the personalized synthesized speech model to the terminal through the communication unit.

7. The speech synthesizer according to claim 1 , wherein, when the utterance information of the unuttered words is received, the processor is further configured to acquire the plurality of non-utterance IP ratios using the received utterance information of the unuttered words and to retrain the personalized synthesized speech model using the plurality of acquired unuttered IP ratios.

8. A method of operating a speech synthesizer using artificial intelligence, the method comprising:

receiving, via a communication unit of the speech synthesizer, utterance information of words uttered by a user from a terminal;

acquiring, using a processor of the speech synthesizer, a plurality of utterance intonation phrase (IP) ratios respectively corresponding to a plurality of words uttered by the user based on the utterance information;

acquiring, using the processor of the speech synthesizer, a plurality of non-utterance IP ratios respectively corresponding to a plurality of unuttered words based on the utterance information and the plurality of utterance IP ratios,

generating, using the processor of the speech synthesizer, a personalized synthesized speech model based on the plurality of utterance IP ratios and the plurality of non-utterance IP ratios, wherein the personalized synthesized speech model is a model for outputting a synthesized speech, to which reading break of words uttered by the user is applied, and is an artificial neural network based model learned by a deep learning algorithm or a machine learning algorithm, and

outputting, using the processor of the speech synthesizer, a synthesized speech based on an utterance style of the user using the generated personalized synthesized speech model,

wherein a plurality of classes indicating reading break of a word includes a first class corresponding to first reading break, a second class corresponding to second reading break greater than the first break and a third class corresponding to third reading break greater than the second break,

wherein a minor class has a smallest count among the first to third classes, and

wherein each of the utterance IP ratios and the non-utterance IP ratios is a ratio in which a word is classified as the minor class.

9. The method according to claim 8 , wherein the utterance information comprises reading break of each uttered word, a part of speech of each uttered word, and a position of each uttered word in a sentence.

10. The method according to claim 9 , wherein acquiring the plurality of utterance IP ratios further comprises acquiring an utterance IP ratio of each uttered word using a number of times that the user reads each uttered word with break corresponding to the minor class.

11. The method according to claim 10 , further comprising:

determining, via the processor of the speech synthesizer, a probability in which the unuttered word is classified as the minor class, using each unuttered word, a property of each unuttered word, an IP ratio of an uttered word and labeling data, and

determining, via the processor of the speech synthesizer, a probability of being classified as a determined IP class as a non-utterance IP ratio of the unuttered word.

12. The method according to claim 8 wherein the personalized synthesized speech model is a model for inferring a probability that each word is classified as the minor class, using text data corresponding to the plurality of uttered words and the plurality of unuttered words, an IP ratio of each word, and a labeled probability of being classified as the minor class, which is labeled with each word.

13. A non-transitory computer-readable recording medium for performing a method of operating a speech synthesizer using artificial intelligence, the method comprising:

receiving utterance information of words uttered by a user from a terminal;

acquiring a plurality of utterance intonation phrase (IP) ratios respectively corresponding to a plurality of words uttered by the user based on the utterance information;

acquiring a plurality of non-utterance IP ratios respectively corresponding to a plurality of unuttered words based on the utterance information and the plurality of utterance IP ratios, and

generating a personalized synthesized speech model based on the plurality of utterance IP ratios and the plurality of non-utterance IP ratios, wherein the personalized synthesized speech model is a model for outputting a synthesized speech, to which reading break of words uttered by the user is applied, and is an artificial neural network based model learned by a deep learning algorithm or a machine learning algorithm,

wherein a plurality of classes indicating reading break of a word includes a first class corresponding to first reading break, a second class corresponding to second reading break greater than the first break and a third class corresponding to third reading break greater than the second break,

wherein a minor class has a smallest count among the first to third classes, and

wherein each of the utterance IP ratios and the non-utterance IP ratios is a ratio in which a word is classified as the minor class.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2019
From: CHAE, JONGHOON
To: LG ELECTRONICS INC.
Reel/Frame 050508/0259 →
Continuity (1)
Related Publication 20210358473A1 · Nov 18, 2021
Cited By (1)
US 12,620,386