IP Library › Granted Patent US 12,190,875
Granted Patent B1
US 12,190,875 · App. 17/490,572 · Granted Jan 7, 2025

Preemptive wakeword detection

Inventors: Eli Joshua Fidler (Toronto, CA); Aaron Challenner (Cambridge, MA); Zoe Adams (Orange County, CA); Sree Hari Krishnan Parthasarathi (Toronto, CA); Gengshen Fu (Sharon, MA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L15/02G10L15/08G10L15/187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,875
App. No.
17/490,572
Granted
Jan 7, 2025
Kind
B1
Abstract

Systems and methods for preemptive wakeword detection are disclosed. For example, a first part of a wakeword is detected from audio data representing a user utterance. When this occurs, on-device speech processing is initiated prior to when the entire wakeword is detected. When the entire wakeword is detected, results from the on-device speech processing and/or the audio data is sent to a speech processing system to determine a responsive action to be performed by the device. When the entire wakeword is not detected, on-device processing is canceled and the device refrains from sending the audio data to the speech processing system.

Claims (105)

1. A device, comprising:

one or more processors; and

non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving first audio data representing a first user utterance captured by a microphone of the device;

determining, from a comparison of a first audio signature of the first audio data to a stored audio signature corresponding to a first part of a wakeword, that a portion of the first audio data corresponds to the first part of the wakeword, the wakeword, when detected, causing the device to send the first audio data to a speech processing system, the first part of the wakeword being one or more starting syllables of the wakeword;

performing, in response to the portion of the first audio data corresponding to the first part of the wakeword, automatic speech recognition (ASR) on the first audio data such that first ASR data is generated; and

determining that an entirety of the wakeword is detected in the first audio data.

2. The device of claim 1 , the operations further comprising:

storing first data indicating a first confidence threshold for detecting the first part of the wakeword;

receiving second data indicating when both the first part of the wakeword is detected and the entirety of the wakeword is detected in sample audio data; and

generating, based at least in part on the second data, third data indicating a second confidence threshold for detecting the first part of the wakeword, the second confidence threshold differing from the first confidence threshold.

3. The device of claim 1 , the operations further comprising:

receiving second audio data representing a second user utterance;

detecting the first part of the wakeword in the second audio data;

initiating speech processing on the second audio data in response to detecting the first part of the wakeword in the second audio data such that second data representing a first part of the second user utterance is generated;

determining that the entirety of the wakeword is absent from the second data; and

causing the speech processing to cease in response to the entirety of the wakeword being absent from the second data.

4. The device of claim 1 , the operations further comprising:

in response to determining that the portion of the first audio data corresponds to the first part of the wakeword, selecting a first speech processing component associated with the wakeword to perform speech processing on the first audio data, wherein performing the speech processing is performed by the first speech processing component;

receiving second audio data representing a second user utterance;

detecting a first portion of a keyword other than the wakeword from the second audio data;

selecting a second speech processing component associated with the keyword in response to detecting the first portion of the keyword; and

performing, using the second speech processing component, speech processing on the second audio data.

5. A method, comprising:

receiving first audio data representing a first user utterance;

determining, based at least in part on a first audio signature of the first audio data corresponding to a stored audio signature of a first part of a first keyword, that at least a portion of the first audio data corresponds to the first part of the first keyword;

initiating generation of first data indicating words included in the first user utterance;

detecting an entirety of the first keyword from the first audio data; and

causing generation of the first data to proceed based at least in part on the entirety of the first keyword being detected.

6. The method of claim 5 , further comprising:

receiving second audio data representing a second user utterance;

detecting, at a wakeword engine of a device that received the first audio data, the first part of the first keyword in the second audio data;

initiating, on the device, speech processing on the second audio data based at least in part on detecting the first part of the first keyword in the second audio data;

determining, at the wakeword engine, that the entirety of the first keyword is undetected; and

causing the speech processing to cease based at least in part on the entirety of the first keyword being undetected.

7. The method of claim 5 , further comprising:

receiving second audio data representing a second user utterance;

detecting the first part of the first keyword in the second audio data;

performing speech processing on the second audio data based at least in part on detecting the first part of the first keyword in the second audio data such that second data representing the second user utterance is generated;

determining, from the second data, that the entirety of the first keyword is absent from the second data; and

causing the speech processing to cease based at least in part on the entirety of the first keyword being absent from the second data.

8. The method of claim 5 , further comprising:

based at least in part on determining that the portion of the first audio data corresponds to the first part of the first keyword, selecting a first speech processing component associated with the first keyword to perform speech processing on the first audio data, wherein performing the speech processing is performed by the first speech processing component;

receiving second audio data representing a second user utterance;

detecting a first portion of a second keyword from the second audio data;

selecting a second speech processing component associated with the second keyword based at least in part on detecting the first portion of the second keyword; and

performing, using the second speech processing component, speech processing on the second audio data.

9. The method of claim 5 , further comprising:

storing second data indicating a first confidence threshold for detecting the first part of the first keyword;

receiving third data indicating when both the first part of the first keyword is detected and the entirety of the first keyword is detected in sample audio data; and

generating, based at least in part on the third data, fourth data indicating a second confidence threshold for detecting the first part of the first keyword, the second confidence threshold differing from the first confidence threshold.

10. The method of claim 5 , wherein:

determining that the at least the portion of the first audio data corresponds to the first part of the first keyword is performed utilizing a wakeword model of a wakeword engine disposed on a device that received the first audio data; and

detecting the entirety of the first keyword is performed by the wakeword engine based at least in part on the wakeword model determining that the at least the portion of the first audio data corresponds to the first keyword.

11. The method of claim 5 , wherein:

determining that the at least the portion of the first audio data corresponds to the first part of the first keyword is performed utilizing a first wakeword model of a wakeword engine disposed on a device that received the first audio data and configured to detect the first part of the first keyword; and

detecting the entirety of the first keyword is performed by the wakeword engine utilizing a second wakeword model configured to detect the entirety of the first keyword.

12. The method of claim 5 , further comprising:

receiving second audio data representing a second user utterance;

detecting, prior to sending the second audio data to a wakeword engine of a device that received the first audio data, a feature of the second audio data that corresponds to the first part of the first keyword;

initiating, on the device, speech processing on the second audio data based at least in part on detecting the feature;

detecting, at the wakeword engine, the entirety of the wakeword; and

causing the speech processing to continue based at least in part on the entirety of the first keyword being detected at the wakeword engine.

13. A device, comprising:

one or more processors; and

non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving first audio data representing a first user utterance;

determining, based at least in part on a first audio signature of the first audio data corresponding to a stored audio signature of a first part of a first keyword, that at least a portion of the first audio data corresponds to the first part of the first keyword;

initiating generation of first data indicating words included in the first user utterance;

detecting an entirety of the first keyword from the first audio data; and

causing generation of the first data to proceed based at least in part on the entirety of the first keyword being detected.

14. The device of claim 13 , the operations further comprising:

receiving second audio data representing a second user utterance;

detecting, at a wakeword engine of a device that received the first audio data, the first part of the first keyword in the second audio data;

initiating, on the device, speech processing on the second audio data based at least in part on detecting the first part of the first keyword in the second audio data;

determining, at the wakeword engine, that the entirety of the first keyword is undetected; and

causing the speech processing to cease based at least in part on the entirety of the first keyword being undetected.

15. The device of claim 13 , the operations further comprising:

receiving second audio data representing a second user utterance;

detecting the first part of the first keyword in the second audio data;

performing speech processing on the second audio data based at least in part on detecting the first part of the first keyword in the second audio data such that second data representing the second user utterance is generated;

determining, from the second data, that the entirety of the first keyword is absent from the second data; and

causing the speech processing to cease based at least in part on the entirety of the first keyword being absent from the second data.

16. The device of claim 13 , the operations further comprising:

based at least in part on determining that the portion of the first audio data corresponds to the first part of the first keyword, selecting a first speech processing component associated with the first keyword to perform speech processing on the first audio data, wherein performing the speech processing is performed by the first speech processing component;

receiving second audio data representing a second user utterance;

detecting a first portion of a second keyword from the second audio data;

selecting a second speech processing component associated with the second keyword based at least in part on detecting the first portion of the second keyword; and

performing, using the second speech processing component, speech processing on the second audio data.

17. The device of claim 13 , the operations further comprising:

storing second data indicating a first confidence threshold for detecting the first part of the first keyword;

receiving third data indicating when both the first part of the first keyword is detected and the entirety of the first keyword is detected in sample audio data; and

generating, based at least in part on the third data, fourth data indicating a second confidence threshold for detecting the first part of the first keyword, the second confidence threshold differing from the first confidence threshold.

18. The device of claim 13 , wherein:

determining that the at least the portion of the first audio data corresponds to the first part of the first keyword is performed utilizing a wakeword model of a wakeword engine disposed on a device that received the first audio data; and

detecting the entirety of the first keyword is performed by the wakeword engine based at least in part on the wakeword model determining that the at least the portion of the first audio data corresponds to the first keyword.

19. The device of claim 13 , wherein:

determining that the at least the portion of the first audio data corresponds to the first part of the first keyword is performed utilizing a first wakeword model of a wakeword engine disposed on a device that received the first audio data and configured to detect the first part of the first keyword; and

detecting the entirety of the first keyword is performed by the wakeword engine utilizing a second wakeword model configured to detect the entirety of the first keyword.

20. The device of claim 13 , the operations further comprising:

receiving second audio data representing a second user utterance;

detecting, prior to sending the second audio data to a wakeword engine of a device that received the first audio data, a feature of the second audio data that corresponds to the first part of the first keyword;

initiating, on the device, speech processing on the second audio data based at least in part on detecting the feature;

detecting, at the wakeword engine, the entirety of the wakeword; and

causing the speech processing to continue based at least in part on the entirety of the first keyword being detected at the wakeword engine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2024
From: FIDLER, ELI JOSHUA; CHALLENNER, AARON; ADAMS, ZOE; PARTHASARATHI, SREE HARI KRISHNAN; FU, GENGSHEN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 068472/0729 →
References Cited (15)
Arik, et al., “Convolutional recurrent neural networks for small-footprint keyword spotting,” CoRR, vol. abs/1703.05390, 2017, 5 pages. [cited by applicant]
Rich Caruana, “Multitask learning,” Machine Learning, vol. 28, 35 pages. [cited by applicant]
Ceolini, et al., “Event-driven pipeline for low-latency low-compute keyword spotting and speaker verification system,” in ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP… [cited by applicant]
Chen, et al. “Small-footprint keyword spotting using deep neural networks,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, 5 pages. [cited by applicant]
Du, et al., “Low-latency convolutional recurrent neural network for keyword spotting,” in 2018 Joint 10th International Conference on Soft Computing and Intelligent Systems (SCIS) and 19th International Symposium on Adv… [cited by applicant]
Kumar, et al., “On convolutional lstm modeling for joint wake-word detection and text dependent speaker verification,” in Interspeech, 2018, 5 pages. [cited by applicant]
Panchapagesan, et al. “Multi-task learning and weighted crossentropy for dnn-based keyword spotting,” in Interspeech, 2016, 5 pages. [cited by applicant]
Sainath, et al. “Convolutional neural networks for small-footprint keyword spotting,” in Interspeech, 2015, 5 pages. [cited by applicant]
Seo, et al., “Temporal convolution for realtime keyword spotting on mobile devices,” in Interspeech, 2019. 5 pages. [cited by applicant]
Sigtia, et al., “Progressive voice trigger detection Accuracy vs latency,” in ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing 2021, 5 pages. [cited by applicant]
Sun, et al., “Max pooling loss training of long short-term memory networks for small-footprint keyword spotting,” in Spoken Language Technology Workshop (SLT), 2016 IEEE, 15 pages. [cited by applicant]
Tang, et al. “Deep residual learning for small-footprint keyword spotting,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5 pages. [cited by applicant]
Tucker, et al. “Model compression applied to small-footprint keyword spotting,” in Interspeech, 2016. 5 pages. [cited by applicant]
Yamamoto, et al., “Small-footprint magic word detection method using convolutional Istm neural network,” in Interspeech, 2019, 5 pages. [cited by applicant]
Zhang, et al., “Autokws: Keyword spotting with differentiable architecture search,” ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 3 pages. [cited by applicant]
Cited By (2)
US 12,412,578 US 12,548,557