IP Library Granted Patent US 10,192,546
Granted Patent B1
US 10,192,546 · App. 14/672,277 · Granted Jan 29, 2019

Pre-wakeword speech processing

Inventors: Kurt Wesley Piersol (San Jose, CA); Gabriel Beddingfield (Fremont, CA)
Assignee: AMAZON TECHNOLOGIES, INC.
G10L15/08G10L17/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,192,546
App. No.
14/672,277
Granted
Jan 29, 2019
Kind
B1
Abstract

A system for capturing and processing portions of a spoken utterance command that may occur before a wakeword. The system buffers incoming audio and indicates locations in the audio where the utterance changes, for example when a long pause is detected. When the system detects a wakeword within a particular utterance, the system determines the most recent utterance change location prior to the wakeword and sends the audio from that location to the end of the command utterance to a server for further speech processing.

Claims (45)

1. A computer-implemented method, comprising:

receiving audio;

storing, in non-transitory memory, audio data representing the audio;

determining a first location in the audio data that includes a first amount of non-speech audio data;

determining a wakeword at a second location in the audio data, the audio data including non-wakeword speech between the first location and the second location;

determining a third location in the audio data that includes a second amount of non-speech audio data, the third location being after the second location in the audio data; and

selecting, for speech processing, a portion of the audio data starting with the first location and ending with the third location, the portion of the audio data comprising at least the non-wakeword speech.

2. The computer-implemented method of claim 1 , further comprising:

sending, to at least one remote device, the portion of the audio data;

receiving, from the at least one remote device, output data; and

presenting output content corresponding to the output data.

3. The computer-implemented method of claim 1 , wherein:

the audio is received using a microphone; and

the non-transitory memory is associated with the microphone.

4. The computer-implemented method of claim 1 , wherein the first amount is configured based at least in part on a language of the non-wakeword speech.

5. The computer-implemented method of claim 1 , wherein the first amount is configured based at least in part on an identity of a user that spoke the non-wakeword speech.

6. The computer-implemented method of claim 1 , further comprising:

determining a change in at least one of a tone, speed, pitch, source direction, frequency, volume, prosody, or energy of the non-wakeword speech,

wherein the first location is associated with the change.

7. The computer-implemented method of claim 1 , further comprising comprising:

determining a confidence score associated with the first location.

8. The computer-implemented method of claim 1 , wherein the audio data includes second non-wakeword speech between the second location and the third location.

9. The computer-implemented method of claim 8 , wherein the portion of the audio data comprises the non-wakeword speech and the second non-wakeword speech.

10. A computing device, comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the computing device to:

receive audio;

store, in non-transitory memory, audio data representing at least some of the audio;

determine a first location in the audio data that includes a first amount of non-speech audio data;

determine a wakeword at a second location in the audio data, the audio data including non-wakeword speech between the first location and the second location;

determine a third location in the audio data that includes a second number amount of non-speech audio data, the third location being after the second location in the audio data; and

determine, for speech processing, a portion of the audio data starting with the first location and ending with the third location, the portion of the audio data comprising at least the non-wakeword speech.

11. The computing device of claim 10 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the computing device to:

send, to at least one remote device, the portion of the audio data;

receive, from the at least one remote device, output data; and

present output content corresponding to the output data.

12. The computing device of claim 10 , wherein:

the audio is received using a microphone; and

the non-transitory memory is associated with the microphone.

13. The computing device of claim 10 , wherein the first amount is configured based at least in part on a language of the non-wakeword speech.

14. The computing device of claim 10 , wherein the first amount is configured based at least in part on an identity of a user that spoke the non-wakeword speech.

15. The computing device of claim 10 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the computing device to:

determine a change in at least one of a tone, speed, pitch, prosody, or energy of the non-wakeword speech, wherein the first location is associated with the change.

16. The computing device of claim 10 , wherein the audio data includes second non-wakeword speech between the second location and the third location.

17. The computing device of claim 16 , wherein the portion of the audio data comprises the non-wakeword speech and the second non-wakeword speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2018
From: BEDDINGFIELD, GABRIEL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 047798/0523 →
Cited By (86)
US 12,190,069 US 12,190,702 US 12,190,879 US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,210,841 US 12,210,843 US 12,211,490 US 12,211,502 US 12,216,894 US 12,217,009 US 12,217,010 US 12,217,748 US 12,219,314 US 12,223,282 US 12,223,285 US 12,223,286 US 12,223,287 US 12,230,291 US 12,236,199 US 12,236,932 US 12,236,952 US 12,242,812 US 12,242,813 US 12,242,814 US 12,249,321 US 12,254,277 US 12,254,278 US 12,254,887 US 12,260,181 US 12,260,182 US 12,260,234 US 12,277,954 US 12,283,269 US 12,293,763 US 12,301,635 US 12,314,660 US 12,321,697 US 12,327,549 US 12,327,556 US 12,333,404 US 12,340,180 US 12,353,827 US 12,360,734 US 12,361,943 US 12,367,879 US 12,375,855 US 12,386,434 US 12,386,491 US 12,387,716 US 12,393,777 US 12,400,085 US 12,406,146 US 12,412,567 US 12,424,220 US 12,430,503 US 12,430,504 US 12,430,505 US 12,431,125 US 12,431,128 US 12,456,008 US 12,475,916 US 12,477,470 US 12,499,320 US 12,505,832 US 12,513,479 US 12,518,107 US 12,518,756 US 12,524,619 US 12,525,234 US 12,554,935 US 12,556,890 US 12,562,159 US 12,562,167 US 12,567,408 US 12,579,978 US 12,585,883 US 12,596,881 US 12,603,098 US 12,608,171 US 12,613,730 US 12,619,452 US 12,699,543 US 12,711,962