IP Library › Granted Patent US 11,960,793
Granted Patent B2
US 11,960,793 · App. 18/148,582 · Granted Apr 16, 2024

Intent detection with a computing device

Inventors: Archana Kannan (Sunnyvale, CA); Roza Chojnacka (Jersey City, NY); Jamieson Kerns (Santa Monica, CA); Xiyang Luo (Mountain View, CA); Meltem Oktem (San Jose, CA); Nada Elassal (Santa Clara, CA)
Assignee: GOOGLE LLC
G06F3/167G06F3/017G06F3/16G06F18/214G06T7/20G06T7/73G06T11/20G06V10/82G06V20/20G06V30/19173G06V30/274G06V40/113G06V40/28G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,960,793
App. No.
18/148,582
Granted
Apr 16, 2024
Kind
B2
Abstract

A method can perform a process with a method including capturing an image, determining an environment that a user is operating a computing device, detecting a hand gesture based on an object in the image, determining, using a machine learned model, an intent of a user based on the hand gesture and the environment, and executing a task based at least on the determined intent.

Claims (67)

1. A method, comprising:

determining an environment in which a user is operating a computing device;

detecting a verbal command using the computing device; and

determining, using a machine learned model, an intent of the user based on the verbal command and the environment;

executing a first task based on the intent of the user, the first task including:

capturing an image; and

detecting an object in the image; and

executing a second task based on the verbal command and a feature of the object.

2. The method of claim 1 , wherein

the machine learned model is a first machine learned model, and

the detecting of the object in the image includes:

using a second machine learned model to:

generate a bounding box; and

determine the feature of the object based on a portion of the object within the bounding box.

3. The method of claim 2 , wherein the second machine learned model is based on a computer vision model.

4. The method of claim 2 , wherein

the second machine learned model is trained using a plurality of images each including at least one object, and

the at least one object has an associated ground-truth box.

5. The method of claim 1 , wherein determining the intent of the user further includes:

translating an interaction of the user with a real-world, and

using the interaction and the verbal command to determine the intent of the user.

6. The method of claim 1 , wherein the second task is based on use of a computer assistant.

7. The method of claim 1 , wherein the second task includes at least one of a visual and audible output.

8. The method of claim 1 , wherein the image is captured using a single non-depth sensing camera of the computing device.

9. The method of claim 1 , wherein

a first machine learned model and a second machine learned model are used to determine the intent of the user, the method further comprising:

continuous tracking of a hand associated with a hand gesture using the second machine learned model.

10. A system comprising:

a memory storing a set of instructions; and

a processor configured to execute the set of instructions to cause the system to:

determining an environment in which a user is operating a computing device;

detecting a verbal command using the computing device;

determining, using a machine learned model, an intent of the user based on the verbal command and the environment;

executing a first task based on the intent of the user, the first task including:

capturing an image; and

detecting an object in the image; and

executing a second task based on the verbal command and a feature of the object.

11. The system of claim 10 , wherein

the machine learned model is a first machine learned model, and

the detecting of the object in the image includes:

using a second machine learned model to:

generate a bounding box;

determine the feature based on a portion of the object within the bounding box.

12. The system of claim 11 , wherein the second machine learned model is based on a computer vision model.

13. The system of claim 11 , wherein

the second machine learned model is trained using a plurality of images each including at least one object, and

the at least one object has an associated ground-truth box.

14. The system of claim 10 , wherein determining the intent of the user further includes:

translating an interaction of the user with a real-world, and

using the interaction and the verbal command to determine the intent of the user.

15. The system of claim 10 , wherein the second task is based on use of a computer assistant.

16. The system of claim 10 , wherein the second task includes at least one of a visual and audible output.

17. The system of claim 10 , wherein the image is captured using a single non-depth sensing camera of the computing device.

18. The system of claim 10 , wherein

a first machine learned model and a second machine learned model are used to determine the intent of the user, the set of instructions further cause the system to:

continuous tracking of a hand associated with a hand gesture using the second machine learned model.

19. A non-transitory computer readable storage medium containing instructions that when executed by a processor of a computer system cause the processor to perform steps comprising:

determining an environment in which a user is operating a computing device;

detecting a verbal command using the computing device;

determining, using a machine learned model, an intent of the user based on the verbal command and the environment;

executing a first task based on the intent of the user, the first task including:

capturing an image; and

detecting an object in the image; and

executing a second task based on the verbal command and a feature of the object.

20. The non-transitory computer readable storage medium of claim 19 , wherein

a first machine learned model and a second machine learned model are used to determine the intent of the user, the steps further comprising:

continuous tracking of a hand associated with a hand gesture using the second machine learned model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2023
From: KANNAN, ARCHANA; CHOJNACKA, ROZA; KERNS, JAMIESON; LUO, XIYANG; OKTEM, MELTEM; ELASSAL, NADA
To: GOOGLE LLC
Reel/Frame 062774/0729 →
Continuity (3)
Continuation 16946532 · Jun 25, 2020
Provisional Application 62867389 · Jun 27, 2019
Related Publication 20230289134A1 · Sep 14, 2023