IP Library Granted Patent US 12694878
Granted Patent B2
US 12694878 · App. 18/203,170 · Granted Jul 28, 2026

Methods and systems for combined voice and gesture control

Inventors: Dhananjay Lal (Englewood, CO); Reda Harb (Issaquah, WA)
Assignee: Adeia Guides Inc.
G10L17/22G06F3/017G06V40/171G06V40/20G06V40/70G10L17/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694878
App. No.
18/203,170
Granted
Jul 28, 2026
Kind
B2
Abstract

Systems and methods are provided for enabling the combined voice and gesture control of a computing device. A voice command is received at a computing device from a user, and a gesture being performed by the user concurrently to the voice command being received is identified. A location of the user performing the gesture is identified in an area proximate to the user, and a direction vector extending from the user is identified based on the gesture. A target computing device in the area proximate to the user is identified from a plurality of computing devices based on the location of the user and direction vector, and the command is transmitted to the target computing device.

Claims (116)

1 . A method comprising:

receiving, at a first computing device, a voice command from a user;

identifying, at the first computing device, a gesture being performed by the user concurrently with receiving the voice command;

identifying, at the first computing device, a location of the user performing the gesture in an area proximate to the user, wherein the location of the user is identified relative to the first computing device;

identifying, based on, at least in part, the gesture, a direction vector extending from the user;

identifying, at the first computing device and based on, at least in part, the location of the user relative to the first computing device and the direction vector, a target second computing device and a target third computing device in the area proximate to the user from a plurality of computing devices, wherein the first computing device, the target second computing device and the target third computing device are different; p 1 determining that the target second computing device and the target third computing device are within a threshold distance from the user;

determining a first angle subtended by the target second computing device and the direction vector;

determining a second angle subtended by the target third computing device and the direction vector;

identifying that the first angle is less than the second angle; and

selecting the target second computing device based on, at least in part, the target second computing device being within the threshold distance from the user and the first angle being less than the second angle;

transmitting, from the first computing device to the target second computing device, a command, wherein the command is based on, at least in part, the received voice command.

2 . The method of claim 1 , wherein:

the method further comprises determining, at the first computing device and based on, at least in part, the received voice command and the gesture, a plurality of candidate computing devices; and

identifying the target second computing device from the plurality of computing devices further comprises identifying the target second computing device from the plurality of candidate computing devices.

3 . The method of claim 2 , wherein:

the method further comprises accessing, at the first computing device, a spatial map indicating locations of each of the plurality of candidate computing devices; and

identifying the location of the user performing the gesture further comprises identifying the location of the user on the spatial map relative to at least a subset of the locations of each of the plurality of candidate computing devices.

4 . The method of claim 3 , wherein the candidate computing devices are members of a group and the method further comprises:

adding, at the first computing device, an additional candidate computing device to the group;

generating, at the first computing device and based on, at least in part, the additional candidate computing device being added to the group, a request to update the spatial map;

receiving, at the first computing device, a location of the additional candidate computing device; and

updating, at the first computing device, the spatial map with the location of the additional candidate computing device.

5 . The method of claim 1 , wherein:

the method further comprises:

identifying, at the first computing device and based on, at least in part, a captured image, a plurality of users local to the first computing device; and

identifying, at the first computing device and via image processing, the user from whom the voice command was received; and

identifying the gesture further comprises identifying the gesture being performed by the identified user from whom the voice command was received.

6 . The method of claim 5 , wherein:

the method further comprises:

identifying, at the first computing device and via speech recognition and a user profile, a user associated with the voice command; and

accessing, at the first computing device and via the user profile, an image associated with the user; and

identifying, at the first computing device, the user from whom the voice command was received further comprises identifying the user via image processing based on, at least in part, the accessed image associated with the user.

7 . The method of claim 1 , wherein:

the method further comprises:

identifying, at the first computing device and based on, at least in part, a captured video, a plurality of users local to the first computing device;

identifying, at the first computing device and via video processing, lip movements of the plurality of users; and

identifying, at the first computing device and based on, at least in part, the lip movements of the plurality of users, the user from whom the voice command was received; and

identifying the gesture further comprises identifying the gesture being performed by the identified user from whom the voice command was received.

8 . The method of claim 1 , wherein:

the gesture is a first gesture; and

the method further comprises:

receiving, from a capture device, a capture of a plurality of users local to the first computing device;

identifying, based on, at least in part, the capture, the user from whom the voice command was received;

transmitting a plurality of commands to the capture device, wherein the plurality of commands comprises commands to move the capture device to keep the user from whom the command was received in the received capture;

outputting a query; and

identifying, via the capture, a response to the query comprising a second gesture being performed by the user.

9 . The method of claim 2 , wherein:

the method further comprises:

identifying, at the first computing device, a group of computing devices associated with the first computing device that receives the voice command;

identifying, at the first computing device and based on, at least in part, the command, a type of computing device; and

accessing, at the first computing device and to identify a command history, a user profile; and

identifying the plurality of candidate computing devices further comprises filtering the initial plurality of identified candidate computing devices based on at least one of:

the group of computing devices associated with the first computing device that receives the voice command;

the identified type of computing device;

the computing devices capable of performing the command; and

the identified command history.

10 . The method of claim 1 , wherein identifying the direction vector extending from the user further comprises identifying the direction vector based on at least one of: a direction of a user's finger, a direction of a user's hand, a direction of a user's head movement and/or a direction of a user's gaze.

11 . A system comprising:

input/output circuitry configured to:

receive, at a first computing device, a voice command from a user;

processing circuitry configured to:

identify, at the first computing device, a gesture being performed by the user concurrently with receiving the voice command;

identify, at the first computing device, a location of the user performing the gesture in an area proximate to the user, wherein the location of the user is identified relative to the first computing device;

identify, based on, at least in part, the gesture, a direction vector extending from the user;

identify, at the first computing device and based on, at least in part, the location of the user relative to the first computing device and the direction vector, a target second computing device and a target third computing device in the area proximate to the user from a plurality of computing devices, wherein the first computing device, the target second computing device and the target third computing device are different;

determine that the target second computing device and the target third computing device are within a threshold distance from the user;

determine a first angle subtended by the target second computing device and the direction vector;

determine a second angle subtended by the target third computing device and the direction vector;

identify that the first angle is less than the second angle; and

select the target second computing device based on, at least in part, the target second computing device being within the threshold distance from the user and the first angle being less than the second angle;

transmit, from the first computing device to the target second computing device, a command, wherein the command is based on, at least in part, the received voice command.

12 . The system of claim 11 , wherein:

the processing circuitry is further configured to determine, at the first computing device and based on, at least in part, the received voice command and the gesture, a plurality of candidate computing devices; and

the processing circuitry configured to identify the target second computing device from the plurality of computing devices is further configured to identify the target second computing device from the plurality of candidate computing devices.

13 . The system of claim 12 , wherein:

the processing circuitry is further configured to access, at the first computing device, a spatial map indicating locations of each of the plurality of candidate computing devices; and

the processing circuitry configured to identify the location of the user performing the gesture is further configured to identify the location of the user on the spatial map relative to at least a subset of the locations of each of the plurality of candidate computing devices.

14 . The system of claim 13 , wherein the candidate computing devices are members of a group and the processing circuitry is further configured to:

add, at the first computing device, an additional candidate computing device to the group; generate, at the first computing device and based on, at least in part, the additional candidate computing device being added to the group, a request to update the spatial map;

receive, at the first computing device, a location of the additional candidate computing device; and

update, at the first computing device, the spatial map with the location of the additional candidate computing device.

15 . The system of claim 11 , wherein:

the processing circuitry is further configured to:

identify, at the first computing device and based on, at least in part, a captured image, a plurality of users local to the first computing device; and

identify, at the first computing device and via image processing, the user from whom the voice command was received; and

the processing circuitry configured to identify the gesture is further configured to identify the gesture being performed by the identified user from whom the voice command was received.

16 . The system of claim 15 , wherein:

the processing circuitry is further configured to:

identify, at the first computing device and via speech recognition and a user profile, a user associated with the voice command; and

access, at the first computing device and via the user profile, an image associated with the user; and

the processing circuitry configured to identify, at the first computing device, the user from whom the voice command was received is further configured to identify the user via image processing based on, at least in part, the accessed image associated with the user.

17 . The system of claim 11 , wherein:

the processing circuitry is further configured to:

identify, at the first computing device based on, at least in part, a captured video, a plurality of users local to the first computing device;

identify, at the first computing device and via video processing, lip movements of the plurality of users; and

identify, at the first computing device and based on, at least in part, the lip movements of the plurality of users, the user from whom the voice command was received; and

the processing circuitry configured to identify the gesture is further configured to identify the gesture being performed by the identified user from whom the voice command was received.

18 . The system of claim 11 , wherein:

the gesture is a first gesture; and

the processing circuitry is further configured to:

receive, from a capture device, a capture of a plurality of users local to the first computing device;

identify, based on, at least in part, the capture, the user from whom the voice command was received;

transmit a plurality of commands to the capture device, wherein the plurality of commands comprises commands to move the capture device to keep the user from whom the command was received in the received capture;

output a query; and

identify, via the capture, a response to the query comprising a second gesture being performed by the user.

19 . The system of claim 12 , wherein:

the processing circuitry is further configured to:

identify, at the first computing device, a group of computing devices associated with the first computing device that receives the voice command;

identify, at the first computing device and based on, at least in part, the command, a type of computing device; and

access, at the first computing device and to identify a command history, a user profile; and

the processing circuitry configured to identify the plurality of candidate computing devices is further configured to filter the initial plurality of identified candidate computing devices based on at least one of:

the group of computing devices associated with the first computing device that receives the voice command;

the identified type of computing device;

the computing devices capable of performing the command; and

the identified command history.

20 . The system of claim 11 , wherein the processing circuitry configured to identify the direction vector extending from the user further comprises processing circuitry configured to identify the direction vector based on at least one of: a direction of a user's finger, a direction of a user's hand, a direction of a user's head movement and/or a direction of a user's gaze.