IP Library Granted Patent US 11,152,001
Granted Patent B2
US 11,152,001 · App. 16/722,964 · Granted Oct 19, 2021

Vision-based presence-aware voice-enabled device

Inventors: Boyan Ivanov Bonev (Santa Clara, CA); Pascale El Kallassi (Menlo Park, CA); Patrick A. Worfolk (San Jose, CA)
Assignee: SYNAPTICS INCORPORATED
G10L15/22G06K9/00362G10L15/08G10L15/25G10L15/30G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,152,001
App. No.
16/722,964
Granted
Oct 19, 2021
Kind
B2
Abstract

A method and apparatus for voice search. A voice search system for a voice-enabled device captures one or more images of a scene, detects a user in the one or more images, determines whether the position of the user satisfies an attention-based trigger condition for initiating a voice search operation, and selectively transmits a voice query to a network resource based at least in part on the determination. The voice query may include audio recorded from the scene and/or the one or more images captured of the scene. The voice search system may further determine whether the trigger condition is satisfied as a result of a false trigger and disable the voice-enabled device from transmitting the voice query to the network resource based at least in part on the trigger condition being satisfied as the result of a false trigger.

Claims (55)

1. A method of performing voice searches by a voice-enabled device, comprising:

capturing one or more images of a scene;

detecting a user in the one or more images;

determining whether a position of the user satisfies an attention-based trigger condition for initiating a voice search operation; and

selectively performing the voice search operation based at least in part on whether the position of the user satisfies the attention-based trigger condition, wherein the voice search operation includes:

generating a voice query that includes audio recorded from the scene; and

outputting a response based at least in part on one or more results of the voice query.

2. The method of claim 1 , wherein the trigger condition is satisfied when the user is facing, looking at, or attending to the voice-enabled device.

3. The method of claim 1 , wherein the voice query further includes the one or more images captured of the scene.

4. The method of claim 1 , wherein the voice search operation further comprises:

transmitting the voice query to a network resource when the trigger condition is satisfied.

5. The method of claim 4 , wherein the transmitting comprises:

recording the audio from the scene upon detecting that the trigger condition is satisfied.

6. The method of claim 1 , wherein the determining of whether the position of the user satisfies the attention-based trigger condition comprises:

listening for a trigger word; and

determining whether the position of the user satisfies the attention-based trigger condition after detecting the trigger word.

7. The method of claim 4 , further comprising:

determining that the trigger condition is satisfied as a result of a false trigger; and

disabling the voice-enabled device from transmitting the voice query to the network resource based at least in part on the trigger condition being satisfied as the result of a false trigger.

8. The method of claim 7 , wherein the disabling includes disabling the voice-enabled device from recording the audio from the scene.

9. The method of claim 7 , wherein the false trigger is determined based on at least one of an amount of lip movement by the user captured in the one or more images, a volume level of the audio recorded from the scene, a directionality of the audio recorded from the scene, or a response from a network service indicating that the voice query resulted in an unsuccessful voice search operation.

10. The method of claim 7 , wherein the disabling comprises:

determining a first location of the user in the scene based on the one or more images; and

updating a false-trigger probability distribution upon determining that the trigger condition is satisfied, the false-trigger probability distribution indicating a likelihood of a false trigger occurring at the first location, wherein the voice query is not transmitted to the network resource when the likelihood of a false trigger occurring at the first location exceeds a threshold probability.

11. The method of claim 10 , wherein the updating comprises:

updating the false-trigger probability distribution to indicate an increase in the likelihood of a false trigger occurring at the first location when the trigger condition is satisfied as a result of a false trigger; and

updating the false-trigger probability distribution to indicate a decrease in the likelihood of a false trigger occurring at the first location when the trigger condition is satisfied not as a result of a false trigger.

12. A voice-enabled device, comprising:

processing circuitry; and

memory storing instructions that, when executed by the processing circuitry, causes the voice-enabled device to:

capture one or more image of a scene;

detect a user in the one or more images;

determine whether a position of the user satisfies an attention-based trigger condition for initiating a voice search operation; and

selectively perform the voice search operation based at least in part on whether the position of the user satisfies the attention-based trigger condition, wherein execution of the instructions for performing the voice search operation further causes the voice-enabled device to:

generate a voice query that includes audio recorded from the scene or the one or more images captured of the scene; and

output a response based at least in part on one or more results of the voice query.

13. The voice-enabled device of claim 12 , wherein the trigger condition is satisfied when the user is facing, looking at, or attending to the voice-enabled device.

14. The voice-enabled device of claim 12 , wherein execution of the instructions for performing the voice search operation further causes the voice-enabled device to:

transmit the voice query to a network resource when the trigger condition is satisfied.

15. The voice-enabled device of claim 14 , wherein execution of the instructions for transmitting the voice query to the network resource causes the voice-enabled device to:

record the audio from the scene upon detecting that the trigger condition is satisfied.

16. The voice-enabled device of claim 12 , wherein execution of the instructions for determining whether the position of the user satisfies the trigger condition causes the voice-enabled device to:

listen for a trigger word; and

determine whether the position of the user satisfies the attention-based trigger condition after detecting the trigger word.

17. The voice-enabled device of claim 14 , wherein execution of the instructions further causes the voice-enabled device to:

determine that the trigger condition is satisfied as a result of a false trigger, wherein the false trigger is determined based on at least one of an amount of lip movement by the user captured in the one or more images, a volume level of the audio recorded from the scene, a directionality of the audio recorded from the scene, or a response from a network service indicating that the voice query resulted in an unsuccessful voice search operation; and

disable the voice-enabled device from transmitting the voice query to the network resource based at least in part on the trigger condition being satisfied as the result of a false trigger.

18. The voice-enabled device of claim 17 , wherein execution of the instructions for disabling the voice-enabled device from transmitting the voice query to the network resource causes the voice-enabled device to:

disable the voice-enabled device from recording the audio from the scene.

19. The voice-enabled device of claim 17 , wherein execution of the instructions for disabling the voice-enabled device from transmitting the voice query to the network resource causes the voice-enabled device to:

determine a first location of the user in the scene based on the one or more images; and

update a false-trigger probability distribution upon determining that the trigger condition is satisfied, the false-trigger probability distribution indicating a likelihood of a false trigger occurring at the first location, wherein the voice query is not transmitted to the network resource when the likelihood of a false trigger occurring at the first location exceeds a threshold probability.

20. The voice-enabled device of claim 19 , wherein execution of the instructions for updating the false-trigger probability distribution causes the voice-enabled device to:

update the false-trigger probability distribution to indicate an increase in the likelihood of a false trigger occurring at the first location when the trigger condition is satisfied as a result of a false trigger; and

update the false-trigger probability distribution to indicate a decrease in the likelihood of a false trigger occurring at the first location when the trigger condition is satisfied not as a result of a false trigger.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2020
From: IVANOV BONEV, BOYAN; EL KALLASSI, PASCALE; WORFOLK, PATRICK A.
To: SYNAPTICS INCORPORATED
Reel/Frame 053332/0934 →
SECURITY INTEREST Recorded Feb 14, 2020
From: SYNAPTICS INCORPORATED
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 051936/0103 →