IP Library › Granted Patent US 11,199,906
Granted Patent B1
US 11,199,906 · App. 14/018,331 · Granted Dec 14, 2021

Global user input management

Inventors: Ryan Halley Curtis (Allston, MA); Andrew Dean Christian (Lincoln, MA)
Assignee: AMAZON TECHNOLOGIES, INC.
G06F3/017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,199,906
App. No.
14/018,331
Granted
Dec 14, 2021
Kind
B1
Abstract

Systems and approaches enable concurrent interaction with multiple user applications in a multi-tasking environment. User input, such as voice commands, head movement, hand or finger gestures, device motion, can be received to a centralized component of a system. State information for each user application can be determined, and the centralized component can send a recognized command or gesture to the appropriate user application(s) based on the state information and/or rules for propagating user input. Additionally, users can configure the input modalities of each user application to customize interaction with systems.

Claims (59)

1. A computing system, comprising:

one or more processors;

one or more microphones;

one or more cameras; and

memory including instructions that, when executed by the one or more processors, cause the computing system to:

execute a first application of the computing system;

execute a second application of the computing system during a first period of time in which the first application is also executed;

capture audio data during the first period of time using the one or more microphones;

capture image data during the first period of time using the one or more cameras;

process the audio data to identify a first keyword;

process the image data to identify a first gesture;

determine that the first keyword corresponds to the first application, and the first gesture corresponds to the second application;

send, based on the first keyword corresponding to the first application, a first command to the first application; and

send, based on the first gesture corresponding to the second application, a second command to the second application.

2. The computing system of claim 1 , further comprising further instructions that, when executed by the one or more processors, further cause the computing system to:

receive, from the first application, a registration of the first keyword.

3. The computing system of claim 2 , further comprising further instructions that, when executed by the one or more processors, further cause the computing system to:

prioritize the first application for receiving the first command over the second application receiving the second command.

4. A computer-implemented method, comprising:

associating one or more first keywords with a first application;

associating one or more second gestures with a second application;

executing the first application on a computing device during a first period of time in which the second application is also executed on the computing device;

receiving audio input data captured during the first period of time by one or more audio input components of the computing device;

receiving image data captured during the first period of time by one or more cameras of the computing device;

processing the audio input data to identify a first keyword;

processing the image data to identify a first gesture;

determining that the first keyword corresponds to the first application, and the first gesture corresponds to the second application;

sending, based on the first keyword corresponding to the first application, a first command to the first application; and

sending, based on the first gesture corresponding to the second application, a second command to the second application.

5. The computer-implemented method of claim 4 , wherein the image data corresponds to lip movement, the method further comprising:

analyzing the audio input data and the image data corresponding to lip movement to enhance recognition of the audio input data.

6. The computer-implemented method of claim 4 , further comprising:

determining that the first application has focus; and

determining that the second application does not have focus.

7. The computer-implemented method of claim 4 , further comprising:

receiving, from the first application, a registration of the one or more first keywords.

8. The computer-implemented method of claim 7 , further comprising:

prioritizing the first application for receiving the first command over the second application receiving the second command.

9. The computer-implemented method of claim 8 , further comprising:

setting a prioritization of the first application over the second application based at least in part upon a category of the first application, a time a user last directly interacted with the first application, a percentage of a display screen corresponding to the first application, or a frequency of usage of the first application.

10. The computer-implemented method of claim 4 , further comprising:

capturing second input data using a second input component of the computing device; and

processing the second input data to increase a confidence level associated with identifying the first keyword.

11. A non-transitory computer-readable storage medium storing instructions, the instructions when executed by a processor causing a computing device to:

associate one or more first keywords with a first application;

associate one or more second gestures with a second application;

execute the first application on the computing device during a first period of time in which the second application is also executed on the computing device;

receive audio input data captured during the first period of time by one or more audio input components of the computing device;

receive image data captured during the first period of time by one or more cameras of the computing device;

process the audio input data to identify a first keyword;

process the image data to identify a first gesture;

determine that the first keyword corresponds to the first application, and the first gesture corresponds to the second application;

send, based on the first keyword corresponding to the first application, a first command to the first application; and

send, based on the first gesture corresponding to the second application, a second command to the second application.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the image data corresponds to lip movement, further comprising further instructions that, when executed by the processor, further cause the computing device to:

analyze the audio input data and the image data corresponding to lip movement to enhance recognition of the audio input data.

13. The non-transitory computer-readable storage medium of claim 11 , further comprising further instructions that, when executed by the processor, further cause the computing device to:

determine that the first application has focus; and

determine that the second application does not have focus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2013
From: CURTIS, RYAN HALLEY; CHRISTIAN, ANDREW DEAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 031668/0122 →
Cited By (2)
US 12,267,454 US 12,738,053