IP Library Granted Patent US 9,747,900
Granted Patent B2
US 9,747,900 · App. 14/164,354 · Granted Aug 29, 2017

Method and apparatus for using image data to aid voice recognition

Inventors: Robert A Zurek (Antioch, IL); Adrian M Schuster (West Olive, MI); Fu-Lin Shau (Lake Zurich, IL); Jincheng Wu (Naperville, IL)
Assignee: Google Technology Holdings LLC
G10L15/22G06F3/013B60N2/002G06K9/00335G06K9/00597G06K9/00832G10L15/20G10L15/24G10L15/25G10L21/0208G10L25/78G10L2015/227G10L2021/02166H04R2430/20H04R2460/07H04R2499/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,747,900
App. No.
14/164,354
Granted
Aug 29, 2017
Kind
B2
Abstract

A device performs a method for using image data to aid voice recognition. The method includes the device capturing image data of a vicinity of the device and adjusting, based on the image data, a set of parameters for voice recognition performed by the device. The set of parameters for the device performing voice recognition include, but are not limited to: a trigger threshold of a trigger for voice recognition; a set of beamforming parameters; a database for voice recognition; and/or an algorithm for voice recognition, wherein the algorithm can include using noise suppression or using acoustic beamforming.

Claims (62)

1. A computer-implemented method comprising:

obtaining one or more images that are generated by one or more cameras of a mobile device;

analyzing one or more of the images;

identifying one or more features of an environment in which the mobile device is operating based on the analysis of one or more of the images;

determining, based on the identified one or more features, a trigger threshold against which respective values of voice commands are compared, each value of a voice command indicating a likelihood that received audio data corresponds to the voice command;

after determining the trigger threshold against which respective values voice commands are compared, receiving particular audio data;

determining that a value of a particular voice command associated with the particular audio data satisfies the trigger threshold; and

in response to determining that the value of the particular voice command associated with the particular audio data satisfies the trigger threshold, performing the particular voice command.

2. The method of claim 1 , wherein the one or more features includes a number of persons in the environment in which the mobile device is operating.

3. The method of claim 1 , wherein the one or more cameras comprise a front camera of the mobile device and a rear camera of the mobile device.

4. The method of claim 1 , wherein determining a trigger threshold against which respective values measures of voice commands are compared comprises determining a phoneme matching threshold against which a number of matching phonemes between the voice commands and reference commands are compared.

5. The method of claim 1 , wherein determining a trigger threshold against which respective values measures of voice commands are compared comprises determining a probability of the mobile device receiving audio from a source other than a speaker of the voice commands.

6. The method of claim 1 , comprising:

determining that the mobile device is inside a motor vehicle; and

based on determining that the mobile device is inside the motor vehicle, reducing the trigger threshold.

7. The method of claim 1 , comprising:

determining that one person is in the environment in which the mobile device is operating; and

based on determining that one person is in the environment in which the mobile device is operating, reducing the trigger threshold.

8. The method of claim 1 , comprising:

determining that more than one person is in the environment in which the mobile device is operating; and

based on determining that more than one person is in the environment in which the mobile device is operating, increasing the trigger threshold.

9. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining one or more images that are generated by one or more cameras of a mobile device;

analyzing one or more of the images;

identifying one or more features of an environment in which the mobile device is operating based on the analysis of one or more of the images;

determining, based on the identified one or more features, a trigger threshold against which respective values of voice commands are compared, each value of a voice command indicating a likelihood that received audio data corresponds to the voice command;

after determining the trigger threshold against which respective values of voice commands are compared, receiving particular audio data;

determining that a value of a particular voice command associated with the particular audio data satisfies the trigger threshold; and

in response to determining that the value of the particular voice command associated with the particular audio data satisfies the trigger threshold, performing the particular voice command.

10. The system of claim 9 , wherein the one or more features includes a number of persons in the environment in which the mobile device is operating.

11. The system of claim 9 , wherein the one or more cameras comprise a front camera of the mobile device and a rear camera of the mobile device.

12. The system of claim 9 , wherein determining a trigger threshold against which respective values measures of voice commands are compared comprises determining a phoneme matching threshold against which a number of matching phonemes between the voice commands and reference commands are compared.

13. The system of claim 9 , wherein determining a trigger threshold against which respective values measures of voice commands are compared comprises determining a probability of the mobile device receiving audio from a source other than a speaker of the voice commands.

14. The system of claim 9 , wherein the operations further comprise:

determining that the mobile device is inside a motor vehicle; and

based on determining that the mobile device is inside the motor vehicle, reducing the trigger threshold.

15. The system of claim 9 , wherein the operations further comprise:

determining that one person is in the environment in which the mobile device is operating; and

based on determining that one person is in the environment in which the mobile device is operating, reducing the trigger threshold.

16. The system of claim 9 , wherein the operations further comprise:

determining that more than one person is in the environment in which the mobile device is operating; and

based on determining that more than one person is in the environment in which the mobile device is operating, increasing the trigger threshold.

17. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

obtaining one or more images that are generated by one or more cameras of a mobile device;

analyzing one or more of the images;

identifying one or more features of an environment in which the mobile device is operating based on the analysis of one or more of the images;

determining, based on the identified one or more features, a trigger threshold against which respective values of voice commands are compared, each value of a voice command indicating a likelihood that received audio data corresponds to the voice command;

after determining the trigger threshold against which respective values of voice commands are compared, receiving particular audio data;

determining that a value of a particular voice command associated with the particular audio data satisfies the trigger threshold; and

in response to determining that the value of the particular voice command associated with the particular audio data satisfies the trigger threshold, performing the particular voice command.

18. The medium of claim 17 , wherein the one or more features includes a number of persons in the environment in which the mobile device is operating.

19. The medium of claim 17 , wherein the one or more cameras comprise a front camera of the mobile device and a rear camera of the mobile device.

20. The medium of claim 17 , wherein determining a trigger threshold against which respective values of voice commands are compared comprises determining a phoneme matching threshold against which a number of matching phonemes between the voice commands and reference commands are compared.

21. The medium of claim 17 , wherein the operations further comprise:

determining that the mobile device is inside a motor vehicle; and

based on determining that the mobile device is inside the motor vehicle, reducing the trigger threshold.

22. The medium of claim 17 , wherein the operations further comprise:

determining that one person is in the environment in which the mobile device is operating; and

based on determining that one person is in the environment in which the mobile device is operating, reducing the trigger threshold.

23. The method of claim 1 , comprising:

determining the confidence value associated with the particular voice command based on a degree to which phonemes of the particular voice command match phonemes stored as reference data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2014
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 034244/0014 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2014
From: ZUREK, ROBERT A; SCHUSTER, ADRIAN M; SHAU, FU-LIN; WU, JINCHENG
To: MOTOROLA MOBILITY LLC
Reel/Frame 032623/0205 →
Continuity (2)
Provisional Application 61827048 · May 24, 2013
Related Publication 20140350924A1 · Nov 27, 2014