IP Library › Granted Patent US 11,513,766
Granted Patent B2
US 11,513,766 · App. 16/895,869 · Granted Nov 29, 2022

Device arbitration by multiple speech processing systems

Inventor: Stanislaw Ignacy Pasko (Zawonia, PL)
Assignee: Amazon Technologies, Inc.
G06F3/167G10L15/22G10L15/32G10L25/60G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,513,766
App. No.
16/895,869
Granted
Nov 29, 2022
Kind
B2
Abstract

A device can perform device arbitration, even when the device is unable to communicate with a remote system over a wide area network (e.g., the Internet). Upon detecting a wakeword in an utterance, the device can wait a period of time for data to arrive at the device, which, if received, indicates to the device that another speech interface device in the environment detected an utterance. If the device receives data prior to the period of time lapsing, the device can determine the earliest-occurring wakeword based on multiple wakeword occurrence times, and may designate whichever device that detected the wakeword first as the designated device to perform an action with respect to the user speech. To account for differences in sound capture latency between speech interface devices, a pre-calculated time offset value can be applied to wakeword occurrence time(s) during device arbitration.

Claims (51)

1. A first device comprising:

one or more processors; and

memory storing computer-executable instructions that, when executed by the one or more processors, cause the first device to:

receive contextual data from a second device that received an utterance that represents user speech;

receive, from the second device, audio data generated by the second device based at least in part on the user speech;

determine, based at least in part on the contextual data, that the second device was used more recently than the first device;

determine, based at least in part on determining that the first second device was used more recently than the first device, the second device is to perform an action with respect to the user speech;

process the audio data using a speech processing component of the first device to determine a speech processing result; and

send directive data to the second device based at least in part on the speech processing result, the directive data causing the second device to perform the action.

2. The first device of claim 1 , wherein determining that the second device was used more recently than the first device comprises determining that the second device responded to a previous user input.

3. The first device of claim 1 , wherein determining that the second device was used more recently than the first device comprises determining that the second device outputted content more recently than the first device.

4. The first device of claim 1 , further comprising a first microphone, wherein the utterance is received by the first device using the first microphone, and wherein the second device has a second microphone that received the utterance, the second microphone different than the first microphone.

5. The first device of claim 1 , wherein the audio data is second audio data, and wherein the computer-executable instructions, when executed by the one or more processors, further cause the first device to:

determine a first signal strength value associated with first audio data generated by the first device based at least in part on the user speech; and

determine, based at least in part on energy data received from the second device, a second signal strength value associated with the second audio data generated by the second device based at least in part on the user speech,

wherein determining the second device is to perform the action is further based on at least one of the first signal strength value or the second signal strength value.

6. The first device of claim 1 , wherein the contextual data indicates an amount of time that has transpired since a time at which the second device was last used.

7. The first device of claim 1 , wherein the directive data causes the second device to output content via an output device of the second device.

8. A method comprising:

receiving contextual data at a first device from a second device that received an utterance that represents user speech;

receiving, by the first device and from the second device, audio data generated by the second device based at least in part on the user speech;

determining, by the first device and based at least in part on the contextual data, that the second device was used more recently than the first device;

determining, based at least in part on the determining that the second device was used more recently than the first device, the second device is to perform an action with respect to the user speech;

processing the audio data using a speech processing component of the first device to determine a speech processing result; and

sending, by the first device, directive data to the second device based at least in part on the speech processing result, the directive data causing the second device to perform the action.

9. The method of claim 8 , wherein the determining that the second device was used more recently than the first device comprises determining that the second device responded to a previous user input.

10. The method of claim 8 , wherein the determining that the second device was used more recently than the first device comprises determining that the second device outputted content more recently than the first device.

11. The method of claim 8 , further comprising receiving the utterance by the first device using a first microphone of the first device, and wherein the second device has a second microphone that received the utterance, the second microphone different than the first microphone.

12. The method of claim 8 , wherein the audio data is second audio data, the method further comprising:

determining, by the first device, a first signal strength value associated with first audio data generated by the first device based at least in part on the user speech; and

determining, by the first device and based at least in part on energy data received from the second device, a second signal strength value associated with the second audio data generated by the second device based at least in part on the user speech,

wherein the determining the second device is to perform the action is further based on at least one of the first signal strength value or the second signal strength value.

13. The method of claim 8 , wherein the contextual data indicates an amount of time that has transpired since a time at which the second device was last used.

14. A first device comprising:

one or more processors; and

memory storing computer-executable instructions that, when executed by the one or more processors, cause the first device to:

receive contextual data from a second device that received an utterance that represents user speech;

receive, from the second device, audio data generated by the second device based at least in part on the user speech;

determine, based at least in part on the contextual data, that the second device was used more recently than the first device;

determine, based at least in part on determining that the second device was used more recently than the first device, the first device is to refrain from performing an action with respect to the user speech;

process the audio data using a speech processing component of the first device to determine a speech processing result; and

send directive data to the second device based at least in part on the speech processing result, the directive data causing the second device to perform the action.

15. The first device of claim 14 , wherein determining that the second device was used more recently than the first device comprises determining that the second device responded to a previous user input.

16. The first device of claim 14 , wherein determining that the second device was used more recently than the first device comprises determining that the second device outputted content more recently than the first device.

17. The first device of claim 14 , further comprising a first microphone, wherein the utterance is received by the first device using the first microphone, and wherein the second device has a second microphone that received the utterance, the second microphone different than the first microphone.

18. The first device of claim 14 , wherein the audio data is second audio data, and wherein the computer-executable instructions, when executed by the one or more processors, further cause the first device to:

determine a first signal strength value associated with first audio data generated by the first device based at least in part on the user speech; and

determine, based at least in part on energy data received from the second device, a second signal strength value associated with the second audio data generated by the second device based at least in part on the user speech,

wherein determining the first device is to refrain from performing the action is further based on at least one of the first signal strength value or the second signal strength value.

19. The first device of claim 14 , wherein the contextual data indicates an amount of time that has transpired since a time at which the second device was last used.

20. The first device of claim 14 , wherein the directive data causes the second device to continue capturing the user speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2020
From: PASKO, STANISLAW IGNACY
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 052869/0441 →
Continuity (2)
Continuation 15948519 · Apr 9, 2018
Related Publication 20200301661A1 · Sep 24, 2020