IP Library Granted Patent US 10,074,371
Granted Patent B1
US 10,074,371 · App. 15/458,628 · Granted Sep 11, 2018

Voice control of remote device by disabling wakeword detection

Inventors: Peng Wang (Kirkland, WA); Pathivada Rajsekhar Naidu (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L15/1815G10L15/30G10L2015/088G10L2015/223H04M7/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,074,371
App. No.
15/458,628
Granted
Sep 11, 2018
Kind
B1
Abstract

A system configured to enable remote control to allow a first user to provide assistance to a second user. The system may receive a command from the second user granting remote control to the first user, enabling the first user to initiate a voice command on behalf of the second user. In some examples, the system may enable the remote control by treating a voice command originating from the first user as though it originated from the second user instead. For example, the system may receive the voice command from a first device associated with the first user but may route the voice command as though it was received by a second device associated with the second user. To enable this functionality, during a remote control session the first device may disable wakeword detection so that the voice command is correctly routed to the second device.

Claims (134)

1. A computer-implemented method, comprising:

generating, by a first speech-controlled device in a first environment, first audio data using one or more microphones;

sending at least a portion of the first audio data to a second speech-controlled device in a second environment physically remote from the first environment, the first audio data corresponding to a communication from the first speech-controlled device to the second speech-controlled device;

detecting that a wakeword is represented in the first audio data;

generating second audio data including at least a portion of the first audio data following the wakeword, the second audio data corresponding to a first command to disable wakeword detection;

sending the second audio data to a remote server;

receiving, from the remote server in response to the first command, an instruction to disable wakeword detection;

disabling the wakeword detection;

generating, while the wakeword detection is disabled, third audio data corresponding to the communication from the first speech-controlled device to the second speech-controlled device and including a second command to perform a second action; and

sending at least a portion of the third audio data to the second speech-controlled device.

2. The computer-implemented method of claim 1 , further comprising:

detecting, using a wakeword detection component, that the wakeword is represented in the third audio data;

determining that the wakeword detection is disabled;

determining to ignore the wakeword;

detecting, using the wakeword detection component, that a second wakeword is represented in the third audio data, the second wakeword corresponding to a command to enable the wakeword detection; and

enabling the wakeword detection.

3. The computer-implemented method of claim 1 , further comprising, by the first speech-controlled device prior to receiving the instruction to disable the wakeword detection:

detecting, using the wakeword detection, that a second wakeword is represented in the first audio data;

generating fourth audio data including at least a portion of the first audio data following the second wakeword; and

sending the fourth audio data to the remote server, the fourth audio data corresponding to a third command to perform a third action.

4. The computer-implemented method of claim 1 , further comprising:

receiving fourth audio data originating from the second speech-controlled device;

outputting first audio corresponding to a first portion of the fourth audio data, the first audio having a first volume level;

generating fifth audio data using the one or more microphones;

detecting that the wakeword is represented in the fifth audio data;

outputting second audio corresponding to a second portion of the fourth audio data, the second audio having a second volume level that is less than the first volume level;

receiving a notification that the second speech-controlled device is granted remote control of the first speech-controlled device;

receiving sixth audio data originating from the second speech-controlled device;

outputting third audio corresponding to a first portion of the sixth audio data, the third audio having the first volume level;

generating, using the one or more microphones, seventh audio data corresponding to at least a portion of the third audio;

detecting that the wakeword is represented in the seventh audio data; and

outputting fourth audio corresponding to a second portion of the sixth audio data, the fourth audio having the first volume level.

5. A computer-implemented method, comprising:

generating, by a first device in a first environment, first audio data;

sending at least a portion of the first audio data to a second device in a second environment physically remote from the first environment, the first audio data corresponding to a communication from the first device to the second device;

detecting that a wakeword is represented in the first audio data;

generating second audio data including at least a portion of the first audio data following the wakeword, the second audio data corresponding to a first command to disable wakeword detection;

sending at least the second audio data to a remote server;

receiving, from the remote server in response to the first command, an instruction to disable wakeword detection;

disabling the wakeword detection;

generating, while the wakeword detection is disabled, third audio data corresponding to the communication from the first device to the second device and including a second command to perform an action; and

sending at least a portion of the third audio data to the second device.

6. The computer-implemented method of claim 5 , further comprising, by the first device prior to receiving the instruction to disable the wakeword detection:

detecting that the wakeword is represented in the first audio data;

generating fourth audio data including at least a portion of the first audio data following the wakeword; and

sending the fourth audio data to the remote server, the fourth audio data corresponding to a third command to perform a second action associated with a first user profile corresponding to the first device.

7. The computer-implemented method of claim 6 , further comprising, by the first device prior to receiving the instruction to disable the wakeword detection:

detecting that a second wakeword is represented in the first audio data, the second wakeword being different from the first wakeword;

generating fifth audio data including at least a portion of the first audio data following the second wakeword; and

sending the fifth audio data to the remote server, the fifth audio data corresponding to a fourth command to perform a third action associated with a second user profile corresponding to the second device.

8. The computer-implemented method of claim 5 , further comprising:

receiving, from the remote server, a second instruction to enable the wakeword detection; and

enabling the wakeword detection.

9. The computer-implemented method of claim 5 , further comprising:

detecting, using a wakeword detection component, that the wakeword is represented in the third audio data;

determining that the wakeword detection is disabled; and

determining to ignore the wakeword.

10. The computer-implemented method of claim 5 , further comprising:

detecting, using a wakeword detection component, that a second wakeword is represented in the third audio data, the second wakeword different than the first wakeword and corresponding to a third command to enable the wakeword detection; and

enabling the wakeword detection.

11. The computer-implemented method of claim 5 , further comprising, by the first device prior to receiving the instruction to disable the wakeword detection:

detecting that wakeword is represented in the first audio data;

generating fourth audio data including at least a portion of the first audio data following the wakeword; and

sending the fourth audio data to the remote server, the fourth audio data corresponding to a third command to perform a second action associated with a user profile corresponding to the second device.

12. The computer-implemented method of claim 5 , further comprising:

receiving fourth audio data originating from the second device;

outputting first audio corresponding to a first portion of the fourth audio data, the first audio having a first volume level;

generating fifth audio data;

detecting that the wakeword is represented in the fifth audio data;

outputting second audio corresponding to a second portion of the fourth audio data, the second audio having a second volume level that is less than the first volume level;

receiving a notification that the second device is granted remote control of the user profile;

receiving sixth audio data originating from the second device;

outputting third audio corresponding to a first portion of the sixth audio data, the third audio having the first volume level;

generating seventh audio data corresponding to at least a portion of the sixth audio;

detecting that the wakeword is represented in the seventh audio data; and

outputting fourth audio corresponding to a second portion of the sixth audio data, the fourth audio having the first volume level.

13. The computer-implemented method of claim 5 , wherein generating the third audio data further comprises:

capturing speech using at least one microphone associated with the first device, wherein the speech is generated by a user in the first environment.

14. The computer-implemented method of claim 5 , wherein generating the third audio data further comprises:

generating, by the first device while the wakeword detection is disabled, the third audio data, wherein the third audio data corresponds to speech generated in the first environment.

15. A system comprising:

at least one processor; and

memory including instructions operable to be executed by the at least one processor to to cause the system to:

receive, by a first device in a first environment, from a remote server, an instruction to begin a remote control session between the first device and a second device in a second environment physically remote from the first environment;

disable wakeword detection in response to receiving the instruction to begin the remote control session;

generate, by the first device, first audio data;

send at least a portion of the first audio data to the second device, the first audio data corresponding to a communication from the first device to the second device;

generate, while the wakeword detection is disabled, second audio data corresponding to the communication from the first device to the second device and including a command to perform an action; and

send at least a portion of the second audio data to the second device.

16. The system of claim 15 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the first device to, prior to receiving the instruction to begin the remote control session:

generate third audio data;

detect that a wakeword is represented in the third audio data;

generate fourth audio data including at least a portion of the third audio data following the wakeword; and

send the fourth audio data to the remote server, the fourth audio data corresponding to a second command to perform a second action associated with a first user profile corresponding to the first device.

17. The system of claim 16 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the first device to, prior to receiving the instruction to begin the remote control session:

detect that a second wakeword is represented in the third audio data;

generate fifth audio data including at least a portion of the third audio data following the second wakeword; and

send the fifth audio data to the remote server, the fifth audio data corresponding to a third command to perform a third action associated with a second user profile corresponding to the second device.

18. The system of claim 15 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the first device to:

receive, from the remote server, a second instruction to end the remote control session; and

enable the wakeword detection.

19. The system of claim 15 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the first device to:

detect, using a wakeword detection component, that a wakeword is represented in the second audio data;

determine that the wakeword detection is disabled; and

determine to ignore the wakeword.

20. The system of claim 19 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the first device to:

detect, using the wakeword detection component, that a second wakeword is represented in the second audio data, the second wakeword corresponding to a command to enable the wakeword detection; and

enable the wakeword detection.

21. The system of claim 15 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the first device to, prior to receiving the instruction to begin the remote control session:

generate third audio data;

detect that a wakeword is represented in the third audio data;

generate fourth audio data including at least a portion of the third audio data following the wakeword; and

send the fourth audio data to the remote server, the fourth audio data corresponding to a second command to perform a second action associated with a user profile corresponding to the second device.

22. The system of claim 15 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the first device to:

receive third audio data originating from the second device;

output first audio corresponding to a first portion of the third audio data, the first audio having a first volume level;

generate fourth audio data;

detect that a wakeword is represented in the fourth audio data;

output second audio corresponding to a second portion of the third audio data, the second audio having a second volume level that is less than the first volume level;

receive a notification that the second device is granted remote control of the user profile;

receive fifth audio data originating from the second device;

output third audio corresponding to a first portion of the fifth audio data, the third audio having the first volume level;

generate sixth audio data corresponding to at least a portion of the fifth audio;

detect that the wakeword is represented in the sixth audio data; and

output fourth audio corresponding to a second portion of the fifth audio data, the fourth audio having the first volume level.

23. A computer-implemented method, comprising:

generating, by a first device in a first environment, first audio data;

sending at least a portion of the first audio data to a second device in a second environment physically remote from the first environment, the first audio data corresponding to a communication from the first device to the second device;

receiving, from a remote server, an instruction to disable wakeword detection;

disabling the wakeword detection;

generating, while the wakeword detection is disabled, second audio data corresponding to the communication from the first device to the second device and including a command to perform an action;

sending at least a portion of the second audio data to the second device;

receiving, from the remote server, a second instruction to enable the wakeword detection; and

enabling the wakeword detection.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2017
From: WENG, PENG; NAIDU, PATHIVADA RAJSEKHAR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 041572/0917 →
Cited By (71)
US 12,190,702 US 12,192,713 US 12,197,712 US 12,197,817 US 12,199,985 US 12,200,297 US 12,204,932 US 12,211,490 US 12,211,502 US 12,216,894 US 12,217,748 US 12,219,314 US 12,223,282 US 12,230,291 US 12,236,932 US 12,236,952 US 12,243,532 US 12,254,887 US 12,260,234 US 12,277,954 US 12,279,096 US 12,283,269 US 12,288,558 US 12,293,763 US 12,301,635 US 12,315,510 US 12,322,390 US 12,327,549 US 12,327,556 US 12,333,404 US 12,360,734 US 12,361,943 US 12,367,879 US 12,374,334 US 12,375,052 US 12,375,850 US 12,379,895 US 12,386,434 US 12,386,491 US 12,387,716 US 12,424,220 US 12,431,128 US 12,438,977 US 12,444,418 US 12,450,025 US 12,462,802 US 12,464,302 US 12,477,470 US 12,495,258 US 12,498,899 US 12,501,229 US 12,505,832 US 12,513,479 US 12,518,755 US 12,518,756 US 12,548,564 US 12,556,890 US 12,573,400 US 12,574,697 US 12,579,978 US 12,580,785 US 12,608,171 US 12,613,730 US 12,619,452 US 12,626,703 US 12,652,508 US 12,659,682 US 12,666,217 US 12,699,543 US 12,711,962 US 12,718,803