IP Library › Granted Patent US 11,488,591
Granted Patent B1
US 11,488,591 · App. 16/510,060 · Granted Nov 1, 2022

Altering audio to improve automatic speech recognition

Inventors: Gregory Michael Hart (Mercer Island, WA); William Spencer Worley, III (Half Moon Bay, CA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L15/20G10L17/00G11B27/005H03G3/32H03G5/02H04R3/12G10L15/26G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,488,591
App. No.
16/510,060
Granted
Nov 1, 2022
Kind
B1
Abstract

Techniques for altering audio being output by a voice-controlled device, or another device, to enable more accurate automatic speech recognition (ASR) by the voice-controlled device. For instance, a voice-controlled device may output audio within an environment using a speaker of the device. While outputting the audio, a microphone of the device may capture sound within the environment and may generate an audio signal based on the captured sound. The device may then analyze the audio signal to identify speech of a user within the signal, with the speech indicating that the user is going to provide a subsequent command to the device. Thereafter, the device may alter the output of the audio (e.g., attenuate the audio, pause the audio, switch from stereo to mono, etc.) to facilitate speech recognition of the user's subsequent command.

Claims (64)

1. A device comprising:

at least one speaker;

at least one microphone;

one or more processors; and

computer-readable media storing computer-executable instructions that, when executed on the one or more processors, cause the device to perform operations, the operations comprising:

causing the at least one speaker to output first content;

receiving a first input audio signal generated by the at least one microphone based at least in part on first sound from a user, the first sound captured by the at least one microphone;

determining predefined audio within the first input audio signal, the predefined audio comprising one or more words indicating that the user is going to provide a subsequent command to the device;

altering output of the first content by the at least one speaker for a first period of time based at least in part on determining the predefined audio within the first input audio signal;

receiving a second input audio signal generated by the at least one microphone based at least in part on second sound captured by the at least one microphone during at least a portion of the first period of time;

determining a voice command in the second input audio signal; and

causing, based at least in part on the voice command, the at least one speaker to output second content different from the first content for a second period of time that is after the first period of time.

2. The device of claim 1 , the operations further comprising:

determining an identify of the user; and

determining a user profile associated with the user.

3. The device of claim 2 , wherein the altering the output of the first audio content is based at least in part on the user profile.

4. The device of claim 2 , further comprising:

a camera,

wherein the determining the identity of the user is based at least in part on image data captured by the camera.

5. The device of claim 1 , wherein altering the output of the first content comprises lowering a volume at which the at least one speaker outputs the first content during the first period of time.

6. The device of claim 1 , wherein altering the output of the first content comprises stopping output of the first content for the first period of time.

7. The device of claim 1 , wherein altering the output of the first content comprises switching from outputting the first content in stereo to outputting the first content in mono for the first period of time.

8. The device of claim 1 , further comprising:

a switch configurable in a first position that couples the at least one speaker to a power source and a second position that decouples the at least one speaker from the power source,

wherein the operations further comprise configuring, based at least in part on determining the predefined audio, the switch in the second position.

9. The device of claim 1 , the operations further comprising:

determining a type of the first content;

wherein the altering the output of the first content comprises altering the output in a first manner based at least in part on the first content being a first type, and

wherein the altering the output of the first content comprises altering the output in a second manner based at least in part on the first content being a second type.

10. The device of claim 1 , the operations further comprising:

determining an audible response to the verbal command,

wherein the second content includes the audible response to the verbal command.

11. A device comprising:

at least one speaker;

at least one microphone;

one or more processors; and

computer-readable media storing computer-executable instructions that, when executed on the one or more processors, cause the device to perform operations, the operations comprising:

causing the at least one speaker to output first content;

receiving a first input audio signal generated by the at least one microphone based at least in part on first sound from a user, the first sound captured by the at least one microphone;

determining predefined audio within the first input audio signal, the predefined audio comprising one or more words indicating that the user is going to provide a subsequent command to the device;

determining a user profile associated with the user;

altering, based at least in part on determining the predefined audio, output of the first content by the at least one speaker for a first period of time;

receiving a second input audio signal generated by the at least one microphone based at least in part on second sound captured by the at least one microphone during at least a portion of the first period of time;

determining a voice command in the second input audio signal; and

causing, based at least in part on the voice command, the at least one speaker to output second content different from the first content for a second period of time that is after the first period of time.

12. The device of claim 11 , wherein the determining the user profile comprises comparing at least one of at least a portion of the first input audio signal or at least a portion of the voice command to a voice print associated with the user profile.

13. The device of claim 11 , wherein the altering the output is based at least in part on the user profile.

14. The device of claim 11 , further comprising:

a camera;

wherein the determining the user profile is based at least at least in part on image data captured by the camera.

15. The device of claim 14 , the operations further comprising:

performing facial recognition techniques on the image data to identify the user,

wherein the user profile is associated with the user.

16. The device of claim 11 , the operations further comprising:

determining the second content based at least in part on the user profile.

17. The device of claim 11 , wherein altering the output of the first content comprises at least one of lowering a volume at which the at least one speaker outputs the first content during the first period of time or stopping output of the first content for the first time.

18. The device of claim 11 , wherein altering the output of the first content comprises switching from outputting the first content in stereo to outputting the first content in mono for the first period of time.

19. The device of claim 11 , the operations further comprising:

determining a type of the first content;

wherein the altering the output of the first content comprises altering the output in a first manner based at least in part on the first content being a first type, and

wherein the altering the output of the first content comprises altering the output in a second manner based at least in part on the first content being a second type.

20. The device of claim 19 , the operations further comprising:

determining an audible response to the verbal command,

wherein the second content includes the audible response to the verbal command.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2019
From: RAWLES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 049737/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2019
From: HART, GREGORY M.; WORLEY, WILLIAM SPENCER, III
To: RAWLES LLC
Reel/Frame 049737/0781 →
Continuity (3)
Continuation 15918608 · Mar 12, 2018
Continuation 14994926 · Jan 13, 2016
Continuation 13627890 · Sep 26, 2012
Cited By (1)
US 12,490,003