IP Library › Granted Patent US 11,580,955
Granted Patent B1
US 11,580,955 · App. 17/218,740 · Granted Feb 14, 2023

Synthetic speech processing

Inventors: Yixiong Meng (San Jose, CA); Roberto Barra Chicote (Cambridge, GB); Grzegorz Beringer (Gdansk, PL); Zeya Chen (Lynnwood, WA); Jie Liang (Bellevue, WA); James Garnet Droppo (Carnation, WA); Chia-Hao Chang (Shoreline, WA); Oguz Hasan Elibol (Sunnyvale, CA)
Assignee: Amazon Technologies, Inc.
G10L13/08G10L13/027G10L13/0335G10L13/047G10L15/063G10L19/008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,955
App. No.
17/218,740
Granted
Feb 14, 2023
Kind
B1
Abstract

A speech-processing system receives input data representing text. A first encoder processes segments of the text to determine embedding data representing the text, and a second encoder processes corresponding audio data to determine prosodic data corresponding to the text. The embedding and prosodic data is processed to create output data including a representation of speech corresponding to the text and prosody.

Claims (59)

1. A computer-implemented method for generating synthesizing speech, the method comprising:

receiving first data representing text;

processing the first data to determine phoneme data representing sounds to be played corresponding to the text;

determining first audio data representing speech corresponding to the first data, the first audio data corresponding to a first prosody;

processing, using a phoneme encoder, the phoneme data to determine phoneme embedding data;

processing the first audio data, using a reference encoder trained to generate a second prosody corresponding to the text, to determine probability data representing a probability distribution;

determining second embedding data corresponding to the second prosody by selecting a value of the probability data;

determining third embedding data corresponding to at least one vocal characteristic;

processing, using an attention component, the phoneme embedding data, the second embedding data, and the third embedding data to determine attended embedding data; and

processing, using a decoder, the attended embedding data to determine second audio data, the second audio data including a representation of speech corresponding to the first data and exhibiting the second prosody.

2. The computer-implemented method of claim 1 , further comprising:

determining, using a text-to-speech system, that the first data corresponds to the first prosody; and

processing the first data, using the text-to-speech system, to determine the first audio data.

3. A computer-implemented method comprising:

processing, using a first encoder, first data to determine first embedding data corresponding to speech to be synthesized;

processing, using a second encoder trained to output variations in prosody of speech, first audio data corresponding to a first prosody to determine probability data representing variations in prosody;

determining second embedding data corresponding to a second prosody by selecting a value of the probability data, the value corresponding to the second prosody;

determining third embedding data corresponding to at least one vocal characteristic;

processing, using an attention component, the first embedding data, the second embedding data, and the third embedding data to determine second data; and

processing, using a decoder, the second data to determine second audio data, the second audio data including a representation of speech corresponding to the first data exhibiting the second prosody.

4. The computer-implemented method of claim 3 , further comprising:

determining that the first data corresponds to the second prosody; and

based at least in part on determining that the first data corresponds to the second prosody, selecting the first audio data.

5. The computer-implemented method of claim 3 , further comprising:

processing the first data with a text-to-speech system to determine the first audio data.

6. The computer-implemented method of claim 3 , wherein:

determining the probability data comprises determining a mean of a probability distribution and a variance of the probability distribution.

7. The computer-implemented method of claim 6 , wherein:

selecting the value comprises sampling the probability distribution by processing the mean and the variance.

8. The computer-implemented method of claim 3 ,

wherein the second encoder is configured to determine variations in prosody corresponding to a probability distribution.

9. The computer-implemented method of claim 8 ,

wherein the variations in prosody represent a variation in speed of speech and a variation in pitch of an utterance.

10. The computer-implemented method of claim 9 ,

wherein the variations in prosody represent a number of sets of variations in the speed and in the pitch of the utterance.

11. The computer-implemented method of claim 3 , further comprising:

determining that the third embedding data corresponds to the first data.

12. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

process, using a first encoder, first data to determine first embedding data corresponding to speech to be synthesized;

process, using a second encoder trained to predict variations in prosody of speech, first audio data corresponding to a first prosody to determine probability data representing of the variations in the prosody;

determine second embedding data corresponding to a second prosody by selecting a value of the probability data, the value corresponding to the second prosody;

determine third embedding data corresponding to at least one vocal characteristic;

process, using an attention component, the first embedding data, the second embedding data, and the third embedding data to determine second data; and

process, using a decoder, the second data to determine second audio data, the second audio data including a representation of speech corresponding to the first data and exhibiting the second prosody.

13. The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

process the first data with a text-to-speech system to determine the first audio data.

14. The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

determine that the first data corresponds to the second prosody; and

based at least in part on determining that the first data corresponds to the second prosody, select the first audio data.

15. The system of claim 12 , wherein the instructions that cause the system to determine the probability data comprise instructions to determine a mean corresponding to a probability distribution and a variance corresponding to the probability distribution.

16. The system of claim 15 , wherein the instructions that cause the system to select the value comprise instructions to:

sample the probability distribution by processing the mean and the variance.

17. The system of claim 12 , wherein the second encoder is configured to determine variations in prosody corresponding to a probability distribution.

18. The system of claim 17 , wherein the variations in prosody represent a variation in speed of speech and a variation in pitch of an utterance.

19. The system of claim 18 , wherein the variations in prosody represent a number of sets of variations in the speed and in the pitch of the utterance.

20. The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

determining that the third embedding data corresponds to the first data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2022
From: MENG, YIXIONG; BARRA CHICOTE, ROBERTO; BERINGER, GRZEGORZ; CHEN, ZEYA; LIANG, JIE; DROPPO, JAMES GARNET; CHANG, CHIA-HAO; ELIBOL, OGUZ HASAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 061845/0841 →
Cited By (3)
US 12,243,511 US 12,387,619 US 12,688,863