IP Library › Granted Patent US 11,404,045
Granted Patent B2
US 11,404,045 · App. 17/007,793 · Granted Aug 2, 2022

Speech synthesis method and apparatus

Inventors: Seungdo Choi (Suwon-si, KR); Kyoungbo Min (Suwon-si, KR); Sangjun Park (Suwon-si, KR); Kihyun Choo (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L13/08G10L13/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,404,045
App. No.
17/007,793
Granted
Aug 2, 2022
Kind
B2
Abstract

A speech synthesis method performed by an electronic apparatus to synthesize speech from text and includes: obtaining text input to the electronic apparatus; obtaining a text representation by encoding the text using a text encoder of the electronic apparatus; obtaining an audio representation of a first audio frame set from an audio encoder of the electronic apparatus, based on the text representation; obtaining an audio representation of a second audio frame set based on the text representation and the audio representation of the first audio frame set; obtaining an audio feature of the second audio frame set by decoding the audio representation of the second audio frame set; and synthesizing speech based on an audio feature of the first audio frame set and the audio feature of the second audio frame set.

Claims (52)

1. A method, performed by an electronic apparatus, of synthesizing speech from text, the method comprising:

obtaining text input to the electronic apparatus;

obtaining a text representation of the text by encoding the text using a text encoder of the electronic apparatus;

obtaining a first audio representation of a first audio frame set of the text from an audio encoder of the electronic apparatus, based on the text representation;

obtaining a first audio feature of the first audio frame set by decoding the first audio representation of the first audio frame set;

obtaining a second audio representation of a second audio frame set of the text based on the text representation and the first audio representation of the first audio frame set;

obtaining a second audio feature of the second audio frame set by decoding the second audio representation of the second audio frame set;

generating feedback information by combining audio feature information of at least one audio frame of the second audio frame set with compression information about the at least one audio frame of the second audio frame set; and

synthesizing speech corresponding to the text based on the first audio feature of the first audio frame set and the second audio feature of the second audio frame set.

2. The method of claim 1 , wherein the second audio frame set includes at least one audio frame succeeding a last audio frame of the first audio frame set.

3. The method of claim 1 , wherein the generating of the feedback information comprises:

obtaining the audio feature information of the at least one audio frame of the second audio frame set; and

obtaining the compression information about the at least one audio frame of the second audio frame set.

4. The method of claim 1 , wherein the feedback information is used to obtain an audio feature of a third audio frame set succeeding the second audio frame set.

5. The method of claim 1 , wherein the compression information includes at least one of a first magnitude of an amplitude value of an audio signal corresponding to the at least one audio frame, a second magnitude of a root means square (RMS) of the amplitude value of the audio signal, or a third magnitude of a peak value of the audio signal.

6. The method of claim 1 , wherein the obtaining of the second audio representation comprises:

obtaining attention information for identifying a portion of the text representation requiring attention, based on at least part of the text representation and the first audio representation of the first audio frame set; and

obtaining the second audio representation of the second audio frame set based on the text representation and the attention information.

7. An electronic apparatus for synthesizing speech from text, the electronic apparatus comprising:

a memory storing instructions; and

at least one processor communicatively coupled to the memory, wherein the at least one processor is configured to execute the instructions to:

obtain text input to the electronic apparatus;

obtain a text representation of the text by encoding the text;

obtain a first audio representation of a first audio frame set of the text based on the text representation;

obtain a first audio feature of the first audio frame set by decoding the first audio representation of the first audio frame set;

obtain a second audio representation of a second audio frame set of the text based on the text representation and the first audio representation of the first audio frame set;

obtain a second audio feature of the second audio frame set by decoding the second audio representation of the second audio frame set;

generate feedback information by combining audio feature information of at least one audio frame of the second audio frame set with compression information about the at least one audio frame of the second audio frame set; and

synthesize speech corresponding to the text based on the first audio feature of the first audio frame set and the second audio feature of the second audio frame set.

8. The electronic apparatus of claim 7 , wherein the second audio frame set includes at least one audio frame succeeding a last audio frame of the first audio frame set.

9. The electronic apparatus of claim 7 , wherein the at least one processor is further configured to generate the feedback information based on the second audio feature of the second audio frame set by obtaining the audio feature information of the at least one audio frame of the second audio frame set, and obtaining the compression information about the at least one audio frame of the second audio frame set.

10. The electronic apparatus of claim 7 , wherein the feedback information is used to obtain an audio feature of a third audio frame set succeeding the second audio frame set.

11. The electronic apparatus of claim 7 , wherein the compression information includes at least one of a first magnitude of an amplitude value of an audio signal corresponding to the at least one audio frame, a second magnitude of a root means square (RMS) of the amplitude value of the audio signal, or a third magnitude of a peak value of the audio signal.

12. The electronic apparatus of claim 7 , wherein the at least one processor is further configured to obtain the second audio representation by obtaining attention information for identifying a portion of the text representation requiring attention, based on at least part of the text representation and the first audio representation of the first audio frame set, and obtain the second audio representation of the second audio frame set based on the text representation and the attention information.

13. A non-transitory computer-readable recording medium having recorded thereon a program for executing, on an electronic apparatus, a method of synthesizing speech from text, the method comprising:

obtaining text input to the electronic apparatus;

obtaining a text representation of the text by encoding the text using a text encoder of the electronic apparatus;

obtaining a first audio representation of a first audio frame set of the text from an audio encoder of the electronic apparatus, based on the text representation;

obtaining a first audio feature of the first audio frame set by decoding the first audio representation of the first audio frame set;

obtaining a second audio representation of a second audio frame set of the text based on the text representation and the first audio representation of the first audio frame set;

obtaining a second audio feature of the second audio frame set by decoding the second audio representation of the second audio frame set;

generate feedback information by combining audio feature information of at least one audio frame of the second audio frame set with compression information about the at least one audio frame of the second audio frame set; and

synthesizing speech corresponding to the text based on the first audio feature of the first audio frame set and the second audio feature of the second audio frame set.

14. A method, performed by an electronic apparatus, of synthesizing speech from text, the method comprising:

obtaining a first audio feature of a first audio frame set by decoding a text representation of a text input;

obtaining a second audio feature of a second audio frame set by decoding the text representation and a combination of a third audio feature of at least one audio frame of the first audio frame set and compression information of the at least one audio frame of the first audio frame set; and

synthesizing speech corresponding to the text input based on the first audio feature of the first audio frame set and the second audio feature of the second audio frame set.

15. The method of claim 14 , wherein the second audio frame set includes at least one audio frame succeeding a last audio frame of the first audio frame set.

16. The method of claim 14 , wherein the compression information includes at least one of a first magnitude of an amplitude value of an audio signal corresponding to the at least one audio frame of the first audio frame set, a second magnitude of a root means square (RMS) of the amplitude value of the audio signal, or a third magnitude of a peak value of the audio signal.

17. The method of claim 14 , wherein the obtaining of the second audio feature comprises:

obtaining attention information for identifying a portion of the text representation requiring attention, based on at least part of the text representation and a first audio representation of the first audio frame set; and

obtaining a second audio representation of the second audio frame set based on the text representation and the attention information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2020
From: CHOI, SEUNGDO; MIN, KYOUNGBO; PARK, SANGJUN; CHOO, KIHYUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 053646/0414 →
Priority Claims (1)
KR 10-2020-0009391 · Jan 23, 2020 · national
Continuity (2)
Provisional Application 62894203 · Aug 30, 2019
Related Publication 20210065678A1 · Mar 4, 2021
Cited By (1)
US 12,272,349