IP Library Granted Patent US 12,125,489
Granted Patent B1
US 12,125,489 · App. 18/219,411 · Granted Oct 22, 2024

Speech recognition using multiple voice-enabled devices

Inventors: Sahil Puri (Ajax, CA); Zhengran Li (San Jose, CA); Zhan Xu (Los Gatos, CA); Oliver Sinsik Chiu (Etobicoke, CA); Sembhayya Gollakota (Sunnyvale, CA); Bruno Dufour (Mississauga, CA)
Assignee: Amazon Technologies, Inc.
G10L15/34G06F40/20G10L15/22G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,125,489
App. No.
18/219,411
Granted
Oct 22, 2024
Kind
B1
Abstract

Techniques for using multiple voice-enabled devices in a user environment to reduce the latency for obtaining responses to user utterances from a remote system. The voice-enabled devices may each establish connections with the remote system to have the remote system perform supplemental speech processing for utterances the devices are unable to process locally. One voice-enabled device may have a higher-latency connection to the remote system, and another voice-enabled device may have a lower-latency connection to the remote system. The lower-latency device may send an utterance to the remote system before the higher-latency device is able, and the remote system may begin processing the utterance faster than if the lower-latency device sent the utterance. The remote system may then provide a response for the utterance to the higher-latency device in less time than if the remote system had to wait for the utterance from the higher-latency device.

Claims (36)

1. A method comprising:

receiving, at a first device configured with a first speech processing component, audio data representing a voice command;

receiving first data indicating that a first role of a second device is to process voice commands on behalf of the first device;

receiving second data indicating that a second role of the first device is to perform actions responsive to the voice commands;

determining, utilizing the first speech processing component and based at least in part on the first data and the second data, that a second speech processing component is to respond to the voice command;

sending a request for the second device to utilize the second speech processing component to generate a response to the voice command;

receiving third data representing the response from the second device; and

causing an action to be performed by the first device utilizing the third data.

2. The method of claim 1 , further comprising outputting audio requesting that the voice command be provided again to the second device, wherein sending the request is based at least in part on outputting the audio.

3. The method of claim 1 , further comprising sending a command to the second device, the command configured to cause the second speech processing component of the second device to be enabled to process the audio data.

4. The method of claim 1 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a device type associated with the second device.

5. The method of claim 1 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a communication protocol utilized by the first device and the second device.

6. The method of claim 1 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on first natural language understanding capabilities of the first device and second natural language understanding capabilities of the second device.

7. The method of claim 1 , further comprising receiving an indication that the second device is associated with controlling operations of a vehicle, wherein determining that the second speech processing component is to respond to the voice command is based at least in part on the second device being associated with controlling operations of the vehicle.

8. The method of claim 1 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a voice command type of the voice command.

9. The method of claim 1 , further comprising determining that the first speech processing component is unable to respond to the voice command, wherein determining that the second speech processing component is to respond to the voice command is based at least in part on the first speech processing component being unable to response to the voice command.

10. The method of claim 1 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a user profile associated with the voice command.

11. A system comprising:

one or more processors; and

non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving, at a first device configured with a first speech processing component, audio data representing a voice command;

receiving first data indicating that a first role of a second device is to process voice commands on behalf of the first device;

receiving second data indicating that a second role of the first device is to perform actions responsive to the voice commands;

determining, utilizing the first speech processing component and based at least in part on the first data and the second data, that a second speech processing component is to respond to the voice command;

sending a request for the second device to utilize the second speech processing component to generate a response to the voice command;

receiving third data representing the response from the second device; and

causing an action to be performed by the first device utilizing the third data.

12. The system of claim 11 , the operations further comprising outputting audio requesting that the voice command be provided again to the second device, wherein sending the request is based at least in part on outputting the audio.

13. The system of claim 11 , the operations further comprising sending a command to the second device, the command configured to cause the second speech processing component of the second device to be enabled to process the audio data.

14. The system of claim 11 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a device type associated with the second device.

15. The system of claim 11 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a communication protocol utilized by the first device and the second device.

16. The system of claim 11 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on first natural language understanding capabilities of the first device and second natural language understanding capabilities of the second device.

17. The system of claim 11 , the operations further comprising receiving an indication that the second device is associated with controlling operations of a vehicle, wherein determining that the second speech processing component is to respond to the voice command is based at least in part on the second device being associated with controlling operations of the vehicle.

18. The system of claim 11 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a voice command type of the voice command.

19. The system of claim 11 , the operations further comprising determining that the first speech processing component is unable to respond to the voice command, wherein determining that the second speech processing component is to respond to the voice command is based at least in part on the first speech processing component being unable to response to the voice command.

20. The system of claim 11 , wherein determining that the second speech processing component is to respond to the voice command is based at least in part on a user profile associated with the voice command.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: PURI, SAHIL; LI, ZHENGRAN; XU, ZHAN; CHIU, OLIVER SINSIK; GOLLAKOTA, SEMBHAYYA; DUFOUR, BRUNO
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064981/0914 →
Continuity (1)
Continuation 17078954 · Oct 23, 2020
Cited By (1)
US 12,647,488