IP Library › Granted Patent US 10,304,475
Granted Patent B1
US 10,304,475 · App. 15/676,273 · Granted May 28, 2019

Trigger word based beam selection

Inventors: Rui Wang (Sunnyvale, CA); Amit Singh Chhetri (Santa Clara, CA); Xiaoxue Li (San Jose, CA); Trausti Thor Kristjansson (San Jose, CA); Philip Ryan Hilmes (Sunnyvale, CA)
Assignee: Amazon Technologies, Inc.
G10L21/0216G10L15/16G10L15/22G10L15/26G10L2015/088G10L2015/223G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,304,475
App. No.
15/676,273
Granted
May 28, 2019
Kind
B1
Abstract

An audio capture device that incorporates a beamformer and beam-specific trigger word detection. Audio data from each beam is processed by a low power trigger word detector, such as a neural network or other trained model to detect if audio data (such as an audio frame or feature vector corresponding thereto) likely includes part of a trigger word. The beam that either most strongly represents a trigger word portion or represents a trigger word portion most early in time may be selected for further processing such as speech processing or confirmation by a more robust power intensive trigger word detector.

Claims (110)

1. A computer-implemented method comprising:

receiving input audio data corresponding to input audio captured by a microphone array;

performing beamforming on the input audio data to determine first beamformed audio data corresponding to a first direction and second beamformed audio data corresponding to a second direction;

processing the first beamformed audio data to determine a first plurality of feature vectors corresponding to a first time period;

processing the first plurality of feature vectors using a first neural network to determine a first score, the first score corresponding to a likelihood that at least a first portion of a wakeword is represented in the first beamformed audio data corresponding to first time period;

processing the second beamformed audio data to determine a second plurality of feature vectors corresponding to a second time period;

processing the second plurality of feature vectors using a second neural network to determine a second score, the second score corresponding to a likelihood that at least a second portion of the wakeword is represented in the second beamformed audio data corresponding to the second time period;

determining, based on the first score exceeding a threshold, that the first portion of the wakeword is represented in the first beamformed audio data;

determining, based on the second score exceeding the threshold, that the second portion of the wakeword is represented in the second beamformed audio data;

determining that the first portion of the wakeword represented in the first beamformed audio data corresponds to the first time period;

determining that the second portion of the wakeword represented in the second beamformed audio data corresponds to the second time period;

selecting the first beamformed audio data in response to the first time period being prior to the second time period; and

sending the first beamformed audio data for further processing.

2. The computer-implemented method of claim 1 , further comprising:

sending the first beamformed audio data to a wakeword component;

operating the wakeword component to compare the first beamformed audio data to a stored audio signature corresponding to the wakeword;

determining, using the wakeword component, that the first beamformed audio data comprises the wakeword; and

sending the first beamformed audio data to a remote device for speech processing.

3. A computer-implemented method comprising:

receiving input audio data corresponding to input audio captured by a microphone array;

performing beamforming on the input audio data to determine first beamformed audio data corresponding to a first direction and second beamformed audio data corresponding to a second direction;

determining at least one first feature vector corresponding to a first portion of the first beamformed audio data and at least one second feature vector corresponding to a first portion of the second beamformed audio data;

using a first trained model to process the at least one first feature vector to determine a first score, the first score corresponding to a likelihood that at least a first portion of a wakeword is represented in the first portion of the first beamformed audio data;

using a second trained model to process the at least one second feature vector to determine a second score, the second score corresponding to a likelihood that at least a second portion of the wakeword is represented in the first portion of the second beamformed audio data;

determining, based on at least the first score exceeding a threshold, that at least the first portion of the wakeword is represented in the first portion of the first beamformed audio data;

determining, based on at least the second score exceeding the threshold, that at least the second portion of the wakeword is represented in the first portion of the second beamformed audio data; and

selecting, based at least on the first score and the second score, at least a second portion of the first beamformed audio data for further processing by a speech processing component configured to identify a command to perform an action.

4. The computer-implemented method of claim 3 , further comprising:

determining that at least the first portion of the wakeword represented in the first portion of the first beamformed audio data corresponds to a first time period;

determining a plurality of scores, including the second score, wherein each of the plurality of scores corresponds to audio data captured within a time window after the first time period; and

determining the first score is greater than each of the plurality of scores.

5. The computer-implemented method of claim 3 , wherein:

the first portion of the first beamformed audio data corresponds to a first time period;

the first portion of the second beamformed audio data corresponds to the first time period; and

the method further comprises determining that the first score is greater than the second score.

6. The computer-implemented method of claim 3 , wherein the first portion of the first beamformed audio data and the first portion of the second beamformed audio data correspond to a first time period and the method further comprises:

determining a third portion of the first beamformed audio data and a second portion of the second beamformed audio data, wherein the third portion of the first beamformed audio data and the second portion of the second beamformed audio data correspond to a second time period after the first time period;

using the first trained model to process the third portion of the first beamformed audio data to determine a third score;

using the second trained model to process the second portion of the second beamformed audio data to determine a fourth score;

determining that the fourth score is greater than the third score;

determining a difference between the fourth score and the third score;

determining that the difference does not exceed a threshold; and

selecting, based at least in part on the difference not exceeding the threshold, at least a fourth portion of the first beamformed audio data for further processing.

7. A computer-implemented method, comprising:

receiving input audio data corresponding to input audio captured by a microphone array;

determining first audio data corresponding to a first direction and second audio data corresponding to a second direction;

using a first trained model to process the first audio data to determine a first score, the first score corresponding to a likelihood that at least a first portion of a wakeword is represented in the first audio data;

using a second trained model to process the second audio data to determine a second score, the second score corresponding to a likelihood that at least a second portion of the wakeword is represented in the second audio data;

determining, based on at least the first score exceeding a threshold, that at least the first portion of the wakeword is represented in the first audio data;

determining, based on at least the second score exceeding the threshold, that at least the second portion of the wakeword is represented in the second audio data;

determining that at least the first portion of the wakeword represented in the first audio data corresponds to a first time period;

determining that at least the second portion of the wakeword represented in the second audio data corresponds to a second time period; and

selecting, based at least in part on the first score, the second score, and the first time period being before the second time period, further audio data corresponding to the first direction for further processing.

8. The computer-implemented method of claim 3 , further comprising:

sending at least the first portion of the first beamformed audio data to a further wakeword detection component; and

by the further wakeword detection component, comparing the first portion of the first beamformed audio data to a stored audio signature corresponding to the wakeword to determine that the first portion of the first beamformed audio data represents the wakeword.

9. The computer-implemented method of claim 8 , wherein comparing the first portion of the first beamformed audio data to the stored audio signature with the further wakeword detection component uses more computing power than using the first trained model to process the at least one first feature vector to determine the first score.

10. The computer-implemented method of claim 3 , further comprising:

sending the at least one first feature vector to a further wakeword detection component.

11. The computer-implemented method of claim 3 , further comprising:

sending the second portion of the first beamformed audio data to the speech processing component.

12. A device comprising:

at least one processor;

at least one microphone array comprising a plurality of microphones; and

at least one memory including instructions operable to be executed by the at least one processor to configure the device to:

receive input audio data corresponding to input audio captured by the at least one microphone array,

perform beamforming on the input audio data to determine first beamformed audio data corresponding to a first direction and second beamformed audio data corresponding to a second direction,

determine at least one first feature vector corresponding to a first portion of the first beamformed audio data and at least one second feature vector corresponding to a first portion of the second beamformed audio data,

use a first trained model to process the at least one first feature vector to determine a first score, the first score corresponding to a likelihood that at least a first portion of a wakeword is represented in the first portion of the first beamformed audio data,

use a second trained model to process the at least one second feature vector to determine a second score, the second score corresponding to a likelihood that at least a second portion of the wakeword is represented in the first portion of the second beamformed audio data,

determine, based on at least the first score exceeding a threshold, that at least the first portion of the wakeword is represented in the first portion of the first beamformed audio data,

determine, based on at least the second score exceeding the threshold, that at least the second portion of the wakeword is represented in the first portion of the second beamformed audio data, and

select, based at least on the first score and the second score, at least a second portion of the first beamformed audio data for further processing by a speech processing component configured to identify a command to perform an action.

13. The device of claim 12 , wherein the memory includes additional instructions operable to be executed by the at least one processor to further configure the device to:

determine that at least the first portion of the wakeword represented in the first portion of the first beamformed audio data corresponds to a first time period;

determine a plurality of scores, including the second score, wherein each of the plurality of scores corresponds to audio data captured within a time window after the first time period; and

determine the first score is greater than each of the plurality of scores.

14. The device of claim 12 , wherein:

the first portion of the first beamformed audio data corresponds to a first time period;

the first portion of the second beamformed audio data corresponds to the first time period; and

the memory includes additional instructions operable to be executed by the at least one processor to further configure the device to determine that the first score is greater than the second score.

15. The device of claim 12 , wherein the first portion of the first beamformed audio data and the first portion of the second beamformed audio data correspond to a first time period and the memory includes additional instructions operable to be executed by the at least one processor to further configure the device to:

determine a third portion of the first beamformed audio data and a second portion of the second beamformed audio data, wherein the third portion of the first beamformed audio data and the second portion of the second beamformed audio data correspond to a second time period after the first time period;

use the first trained model to process the third portion of the first beamformed audio data to determine a third score;

use the second trained model to process the second portion of the second beamformed audio data to determine a fourth score;

determine that the fourth score is greater than the third score;

determine a difference between the fourth score and the third score;

determine that the difference does not exceed a threshold; and

select, based at least in part on the difference not exceeding the threshold, at least a fourth portion of the first beamformed audio data for further processing.

16. A device, comprising:

at least one processor;

at least one microphone array comprising a plurality of microphones; and

at least one memory including instructions operable to be executed by the at least one processor to configure the device to:

receive input audio data corresponding to input audio captured by the at least one microphone array,

determine first audio data corresponding to a first direction and second audio data corresponding to a second direction,

use a first trained model to process the first audio data to determine a first score, the first score corresponding to a likelihood that at least a first portion of a wakeword is represented in the first audio data,

use a second trained model to process the second audio data to determine a second score, the second score corresponding to a likelihood that at least a second portion of the wakeword is represented in the second audio data,

determine, based on at least the first score exceeding a threshold, that at least the first portion of the wakeword is represented in the first audio data,

determine, based on at least the second score exceeding the threshold, that at least the second portion of the wakeword is represented in the second audio data,

determine that at least the first portion of the wakeword represented in the first audio data corresponds to a first time period;

determine that at least the second portion of the wakeword represented in the second audio data corresponds to a second time period, and

select, based at least in part on the first score, the second score, and the first time period being before the second time period, further audio data corresponding to the first direction for further processing.

17. The device of claim 12 , wherein the memory includes additional instructions operable to be executed by the at least one processor to further configure the device to:

send at least the first portion of the first beamformed audio data to a further wakeword detection component; and

by the further wakeword detection component, compare the first portion of the first beamformed audio data to a stored audio signature corresponding to the wakeword to determine that the first portion of the first beamformed audio data represents the wakeword.

18. The device of claim 17 , wherein execution of the additional instructions to compare the first portion of the first beamformed audio data to the stored audio signature with the further wakeword detection component uses more computing power than execution of the instructions to use the first trained model to process the at least one first feature vector to determine the first score.

19. The device of claim 12 , wherein the memory includes additional instructions operable to be executed by the at least one processor to further configure the device to:

send the at least one first feature vector to a further wakeword detection component.

20. The device of claim 12 , wherein the memory includes additional instructions operable to be executed by the at least one processor to further configure the device to:

send the second portion of the first beamformed audio data to the speech processing component.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2017
From: WANG, RUI; CHHETRI, AMIT SINGH; LI, XIAOXUE; KRISTJANSSON, TRAUSTI THOR; HILMES, PHILIP RYAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 043284/0476 →
Cited By (51)
US 12,189,744 US 12,192,713 US 12,211,490 US 12,212,945 US 12,217,748 US 12,217,765 US 12,230,291 US 12,236,932 US 12,277,368 US 12,279,096 US 12,279,100 US 12,283,269 US 12,288,558 US 12,322,390 US 12,327,549 US 12,327,556 US 12,327,571 US 12,360,734 US 12,361,958 US 12,375,052 US 12,375,855 US 12,387,716 US 12,424,220 US 12,425,782 US 12,438,977 US 12,444,418 US 12,470,880 US 12,494,217 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,525,250 US 12,574,691 US 12,578,779 US 12,579,978 US 12,581,250 US 12,610,200 US 12,626,717 US 12,634,642 US 12,640,148 US 12,664,984 US 12,699,543 US 12,711,962 US 12,713,188 US 12,732,547 US 12,738,262 US 12,744,035 US 12,748,566 US 12,750,623