IP Library Granted Patent US 11,562,739
Granted Patent B2
US 11,562,739 · App. 16/786,629 · Granted Jan 24, 2023

Content output management based on speech quality

Inventors: Andrew Smith (Seattle, WA); Christopher Schindler (Bainbridge Island, WA); Karthik Ramakrishnan (Bellevue, WA); Rohit Prasad (Lexington, MA); Michael George (Seattle, WA); Rafal Kuklinski (Otomin, PL)
Assignee: Amazon Technologies, Inc.
G10L15/20G10L13/033G10L13/10G10L15/1807
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,739
App. No.
16/786,629
Granted
Jan 24, 2023
Kind
B2
Abstract

Techniques for ensuring content output to a user conforms to a quality of the user's speech, even when a speechlet or skill ignores the speech's quality, are described. When a system receives speech, the system determines an indicator of the speech's quality (e.g., whispered, shouted, fast, slow, etc.) and persists the indicator in memory. When the system receives output content from a speechlet or skill, the system checks whether the output content is in conformity with the speech quality indicator. If the content conforms to the speech quality indicator, the system may cause the content to be output to the user without further manipulation. But, if the content does not conform to the speech quality indicator, the system may manipulate the content to render it in conformity with the speech quality indicator and output the manipulated content to the user.

Claims (76)

1. A computer-implemented method, comprising:

receiving, from a first device, first audio data representing first speech;

determining a user identifier associated with the first audio data;

determining a trained model associated with the user identifier, the trained model configured to determine an emotion corresponding to speech associated with the user identifier;

processing, using the trained model, the first audio data to determine a first indicator representing a first emotion corresponding to the first speech;

performing speech processing on the first audio data to generate natural language understanding (NLU) results data comprising at least an intent corresponding to the first speech;

determining a first component configured to process the NLU results data;

sending, to the first component, the NLU results data and the first indicator;

determining, by the first component, first content responsive to the first speech;

determining, by the first component, second content responsive to the first speech;

determining, by the first component and based at least in part on the first indicator, that the first content is to be output instead of the second content;

receiving the first content from the first component; and

causing output of the first content.

2. The computer-implemented method of claim 1 , further comprising:

generating metadata representing text-to-speech (TTS) processing is to be performed based at least in part on the first emotion; and

performing, based at least in part on the metadata, TTS processing on the first content to generate second audio data,

wherein causing output of the first content comprises causing output of audio corresponding to the second audio data.

3. The computer-implemented method of claim 2 , wherein the second audio data matches the first emotion.

4. The computer-implemented method of claim 1 , further comprising:

determining the first indicator based at least in part on processing image data corresponding to the first speech.

5. The computer-implemented method of claim 4 , wherein the image data represents a gesture of a user.

6. The computer-implemented method of claim 4 , wherein the image data represents a face of a user.

7. The computer-implemented method of claim 1 , wherein the first content corresponds to second audio data and the computer-implemented method further comprises:

generating, based at least in part on the first indicator, metadata representing the second audio data is to be output at a first volume level;

sending the second audio data to the first device; and

sending the metadata to the first device, the metadata causing the first device to output audio, corresponding to the second audio data, at the first volume level.

8. The computer-implemented method of claim 1 , wherein causing output of the first content comprises causing a second device, different from the first device, to output the first content.

9. A system comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive, from a first device, first audio data representing first speech;

determine a user identifier associated with the first audio data;

determine a trained model associated with the user identifier, the trained model configured to determine a speech quality corresponding to speech associated with the user identifier;

process, using the trained model, the first audio data to determine a first indicator representing a first speech quality corresponding to the first speech;

perform speech processing on the first audio data to generate natural language understanding (NLU) results data comprising at least an intent corresponding to the first speech;

determine a first component configured to process the NLU results data;

send, to the first component, the NLU results data and the first indicator;

determine, by the first component, first content responsive to the first speech;

determine, by the first component, second content responsive to the first speech;

determine, by the first component and based at least in part on the first indicator, that the first content is to be output instead of the second content;

receive the first content from the first component; and

cause output of the first content.

10. The system of claim 9 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

generate metadata representing text-to-speech (TTS) processing is to be performed based at least in part on the first speech quality; and

perform, based at least in part on the metadata, TTS processing on the first content to generate second audio data,

wherein the instructions that cause output of the first content further cause the system to cause output of audio corresponding to the second audio data.

11. The system of claim 9 , wherein the first speech quality corresponds to a first emotion corresponding to the first speech.

12. The system of claim 10 , wherein the second audio data matches the first speech quality.

13. The system of claim 9 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine the first indicator based at least in part on processing image data corresponding to the first speech.

14. The system of claim 13 , wherein the image data represents a gesture of a user.

15. The system of claim 13 , wherein the image data represents a face of a user.

16. The system of claim 9 , wherein the first content corresponds to second audio data and the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

generate, based at least in part on the first indicator, metadata representing the second audio data is to be output at a first volume level;

send the second audio data to the first device; and

send the metadata to the first device, the metadata causing the first device to output audio, corresponding to the second audio data, at the first volume level.

17. The system of claim 9 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

cause a second device, different from the first device, to output the first content.

18. The computer-implemented method of claim 1 , further comprising:

determining user profile data associated with the user identifier;

determining the user profile data indicates the first emotion is enabled to be used to perform processing related to the first audio data; and

after determining the user profile data indicates the first emotion is enabled to be used to perform processing related to the first audio data, causing the first content to be output.

19. The system of claim 9 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine user profile data associated with the user identifier;

determine the user profile data indicates the first speech quality is enabled to be used to perform processing related to the first audio data; and

after determining the user profile data indicates the first speech quality is enabled to be used to perform processing related to the first audio data, cause the first content to be output.

20. The computer-implemented method of claim 1 , wherein causing output of the first content comprises causing the first device to output the first content in a first manner corresponding to the first emotion and the computer-implemented method further comprises:

while causing output of the first content, receiving, from the first device, second audio data representing second speech;

processing, using the trained model, the second audio data to determine a second indicator representing a second emotion corresponding to the second speech; and

in response to the second indicator, causing the first device to output the first content in a second manner corresponding to the second emotion, the second manner being different from the first manner.

21. The computer-implemented method of claim 20 , further comprising:

determining synthesized speech corresponding to a response to the second speech; and

causing the first device to output the synthesized speech while the first device is outputting the first content in the second manner.

22. The computer-implemented method of claim 21 , wherein:

the synthesized speech is whispered synthesized speech; and

the second manner of outputting the first content involves outputting the first content at a lower volume than outputting the first content in the first manner.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2020
From: SMITH, ANDREW; SCHINDLER, CHRISTOPHER; RAMAKRISHNAN, KARTHIK; PRASAD, ROHIT; GEORGE, MICHAEL; KUKLINSKI, RAFAL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 051774/0137 →
Continuity (2)
Continuation 15933676 · Mar 23, 2018
Related Publication 20200251104A1 · Aug 6, 2020