IP Library › Granted Patent US 11,367,431
Granted Patent B2
US 11,367,431 · App. 16/818,542 · Granted Jun 21, 2022

Synthetic speech processing

Inventors: Antonio Bonafonte (Cambridge, GB); Panagiotis Agis Oikonomou Filandras (Greater London, GB); Bartosz Perz (London, GB); Arent van Korlaar (London, GB); Ioannis Douratsos (Cambridge, GB); Jonas Felix Ananda Rohnke (Bassingbourn, GB); Elena Sokolova (London, GB); Andrew Paul Breen (Norwich, GB); Nikhil Sharma (Woodinville, WA)
Assignee: Amazon Technologies, Inc.
G10L13/10G10L13/047G10L15/16G10L15/1815G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,367,431
App. No.
16/818,542
Granted
Jun 21, 2022
Kind
B2
Abstract

A speech-processing system receives both text data and natural-understanding data (e.g., a domain, intent, and/or entity) related to a command represented in the text data. The system uses the natural-understanding data to vary vocal characteristics in determining spectrogram data corresponding to the text data based on the natural-understanding data.

Claims (77)

1. A computer-implemented method for generating speech, the method comprising:

receiving, from a user device, first audio data representing a command;

processing the first audio data using an automatic speech-recognition (ASR) component to determine first text data representing the speech;

processing the first text data using a natural-language understanding (NLU) component to determine natural-understanding data comprising a representation of an entity in the first text data;

processing the natural-understanding data using a dialog manager component to determine second text data representing a response to the first audio data, wherein the response includes a reference to the entity;

processing the second text data with a linguistic encoder of a text-to-speech (TTS) component to determine first encoded data representing words in the command;

processing the second text data with a second encoder of the TTS component to determine second encoded data corresponding to the natural-understanding data;

processing the first encoded data, the second encoded data, and the natural-understanding data with an attention network of the TTS component to determine weighted encoded data, the weighted encoded data corresponding to a variation in synthetic speech that places an emphasis on a name of the entity; and

processing the weighted encoded data with a speech decoder of the TTS component to determine second audio data, the second audio data corresponding to a the variation in the synthetic speech.

2. The computer-implemented method of claim 1 , further comprising:

processing second NLU data using the dialog manager component to determine third text data representing a second response to a second command;

processing the third text data and second natural-understanding data with a rephrasing component to determine fourth text data, the fourth text data including a representation of the entity and at least a first word unrepresented in the third text data; and

processing the fourth text data with the TTS component to determine third audio data.

3. The computer-implemented method of claim 1 , further comprising:

processing third text data using the NLU component to determine second natural-understanding data comprising an intent to repeat a word represented in the first audio data;

processing the second natural-understanding data using the dialog manager component to determine fourth text data representing a response to the third text data;

processing the second natural-understanding data with the attention network to determine second weighted encoded data, the second weighted encoded data corresponding to the emphasis on the word; and

processing the second weighted encoded data with the speech decoder to determine third audio data, the third audio data corresponding to a second vocal characteristic associated with the word.

4. The computer-implemented method of claim 1 , further comprising:

determining a domain associated with the natural-understanding data;

determining that first data stored in a computer memory indicates that a style of speech is associated with the domain;

determining second data representing the style of speech,

wherein the weighted encoded data is further based at least in part on the second data.

5. A computer-implemented method comprising:

receiving first input data corresponding to a response to a command;

receiving second input data comprising a machine representation of the command;

processing the first input data with a first model to determine first encoded data representing words;

processing the first input data with a second model to determine second encoded data corresponding to the second input data;

processing the first encoded data using the second encoded data and the second input data to determine third encoded data; and

processing the third encoded data with a third model to determine audio data, the audio data corresponding to a variation in synthesized speech associated with the second input data.

6. The computer-implemented method of claim 5 , further comprising:

processing the audio data using a vocoder to determine output audio data; and

causing output of the output audio data.

7. The computer-implemented method of claim 5 , further comprising:

receiving third input data corresponding to a second response to a second command;

processing the third input data with a fourth model to determine fourth input data different from the third input data, the fourth input data corresponding to the second input data; and

processing the fourth input data with the first model, the second model, and the third model to determine second audio data.

8. The computer-implemented method of claim 7 , further comprising:

prior to processing the third input data and, determining that the response corresponds to the second response and that the command corresponds to the second command.

9. The computer-implemented method of claim 5 , further comprising:

determining a style of speech associated with a domain associated with the response;

wherein the third encoded data is further based at least in part on the style of speech.

10. The computer-implemented method of claim 5 , further comprising:

determining a score representing a degree of the variation; and

determining that the score is less than a threshold.

11. The computer-implemented method of claim 5 , wherein processing the first input data with the second model further comprises:

processing an intermediate output of the second model with at least one recurrent layer.

12. The computer-implemented method of claim 5 , further comprising:

processing the second input data and fourth encoded data with the third model to determine second audio data, the second audio data corresponding to a second variation in the synthesized speech associated with fourth input data.

13. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive first input data corresponding to a response to a command;

receive second input data comprising a machine representation of the command;

process the first input data with a first model to determine first encoded data representing words;

process the first input data with a second model to determine second encoded data corresponding to the second input data;

process the first encoded data using the second encoded data and the second input data with an attention network to determine third encoded data; and

process the third encoded data with a third model to determine audio data, the audio data corresponding to a variation in synthesized speech associated with the second input data.

14. The system of claim 13 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

process the audio data using a vocoder to determine output audio data; and

cause output of the output audio data.

15. The system of claim 13 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

receive third input data corresponding to a second response to a second command;

process the third input data with a fourth model to determine fourth input data different from the third input data; and

process the fourth input data with the first model, the second model, and the third model to determine second audio data.

16. The system of claim 15 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

prior to processing the third input data and, determine that the response corresponds to the second response and that the command corresponds to the second command.

17. The system of claim 13 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

determine a style of speech associated with a domain associated with the response;

wherein the third encoded data is further based at least in part on the style of speech.

18. The system of claim 13 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

determine a score representing a degree of the variation; and

determine that the score is less than a threshold.

19. The system of claim 13 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

process an intermediate output of the second model with at least one recurrent layer.

20. The system of claim 13 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

process the second input data and fourth encoded data with the third model to determine second audio data, the second audio data corresponding to a second variation in the synthesized speech associated with fourth input data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2020
From: BONAFONTE, ANTONIO; OIKONOMOU FILANDRAS, PANAGIOTIS AGIS; PERZ, BARTOSZ; VAN KORLAAR, ARENT; DOURATSOS, IOANNIS; ROHNKE, JONAS FELIX ANANDA; SOKOLOVA, ELENA; BREEN, ANDREW PAUL; SHARMA, NIKHIL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 052111/0243 →
Continuity (1)
Related Publication 20210287656A1 · Sep 16, 2021
Cited By (1)
US 12,586,582