IP Library › Granted Patent US 11,543,888
Granted Patent B2
US 11,543,888 · App. 16/946,532 · Granted Jan 3, 2023

Intent detection with a computing device

Inventors: Archana Kannan (Sunnyvale, CA); Roza Chojnacka (Jersey City, NY); Jamieson Kerns (Santa Monica, CA); Xiyang Luo (Mountain View, CA); Meltem Oktem (Sunnlyvale, CA); Nada Elassal (Santa Clara, CA)
Assignee: GOOGLE LLC
G06F3/017G06F3/16G06K9/6256G06T7/20G06T7/73G06T11/20G06V30/274G06V40/28G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,543,888
App. No.
16/946,532
Filed
Jun 25, 2020
Granted
Jan 3, 2023
Kind
B2
Examiner
SADIO, INSA
Art Unit
2628
USPC
345/156
Abstract

A method can perform a process with a method including capturing an image, determining an environment that a user is operating a computing device, detecting a hand gesture based on an object in the image, determining, using a machine learned model, an intent of a user based on the hand gesture and the environment, and executing a task based at least on the determined intent.

Claims (65)

1. A method, comprising:

capturing an image;

determining an environment in which a user is operating a computing device;

detecting a hand gesture based on an object in the image;

selecting a machine learned model configured to:

generate a bounding box,

determine a feature based on a portion of the object within the bounding box, and

identify the object based on the feature;

determining, using the machine learned model, an intent of the user based on the portion of the object within the bounding box and the environment; and

executing a task based at least on the intent of the user.

2. The method of claim 1 , wherein determining the intent of the user further includes:

translating an interaction of the user with a real-world, and

using the interaction and the hand gesture to determine the intent of the user.

3. The method of claim 1 , wherein the machine learned model is based on a computer vision model.

4. The method of claim 1 , wherein

a first machine learned model and a second machine learned model are used to determine the intent of the user, the method further comprising:

continuous tracking of a hand associated with the hand gesture using the second machine learned model.

5. The method of claim 1 , wherein the image is captured using a single non-depth sensing camera of the computing device.

6. The method of claim 1 , wherein the task is based on use of a computer assistant.

7. The method of claim 1 , wherein the task includes at least one of a visual and audible output.

8. The method of claim 1 , wherein

the machine learned model is trained using a plurality of images including at least one hand gesture,

the machine learned model is trained using a plurality of ground-truth images of hand gestures,

a loss function is used to confirm a match between a hand gesture and a ground-truth image of a hand gesture, and

the detecting of the hand gesture based on the object in the image includes matching the object to the hand gesture matched to the ground-truth image of the hand gesture.

9. The method of claim 1 , wherein

the machine learned model is trained using a plurality of images each including at least one object, and

the at least one object has an associated ground-truth box.

10. A system comprising:

a memory storing a set of instructions; and

a processor configured to execute the set of instructions to cause the system to:

capture an image;

determine an environment in which a user is operating a computing device;

detect a hand gesture based on an object in the image;

select a machine learned model configured to:

generate a bounding box,

determine a feature based on a portion of the object within the bounding box, and

identify the object based on the feature;

determine, using the machine learned model, an intent of the user based on the portion of the object within the bounding box and the environment; and

execute a task based at least on the intent of the user.

11. The system of claim 10 , wherein determining the intent of the user further includes:

translating an interaction of the user with a real-world, and

using the interaction and the hand gesture to determine the intent of the user.

12. The system of claim 10 , wherein the machine learned model is based on a computer vision model.

13. The system of claim 10 , wherein

a first machine learned model and a second machine learned model are used to determine the intent of the user; the set of instructions are executed by the processor to further cause the system:

continuously track of the hand using the second machine learned model.

14. The system of claim 10 , wherein the image is captured using a single non-depth sensing camera of the computing device.

15. The system of claim 10 , wherein the task is executed using a computer assistant.

16. The system of claim 10 , wherein the task includes at least one of a visual and audible output.

17. The system of claim 10 , wherein

the machine learned model is trained using a plurality of images including at least one hand gesture,

the machine learned model is trained using a plurality of ground-truth images of hand gestures,

a loss function is used to confirm a match between a hand gesture and a ground-truth image of a hand gesture, and

the detecting of the hand gesture based on the object in the image includes matching the object to the hand gesture matched to the ground-truth image of the hand gesture.

18. A non-transitory computer readable storage medium containing instructions that when executed by a processor of a computer system cause the processor to perform steps comprising:

capturing an image;

determining an environment in which a user is operating a computing device;

detecting a hand gesture based on an object in the image;

selecting a machine learned model configured to:

generate a bounding box,

determine a feature based on a portion of the object within the bounding box, and

identify the object based on the feature;

determining, using the machine learned model, an intent of the user based on the portion of the object within the bounding box and the environment; and

executing a task based at least on the intent of the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2020
From: KANNAN, ARCHANA; CHOJNACKA, ROZA; KERNS, JAMIESON; LUO, XIYANG; OKTEM, MELTEM; ELASSAL, NADA
To: GOOGLE LLC
Reel/Frame 053210/0939 →
Continuity (2)
Provisional Application 62867389 · Jun 27, 2019
Related Publication 20200409469A1 · Dec 31, 2020