IP Library Granted Patent US 10,923,124
Granted Patent B2
US 10,923,124 · App. 16/416,427 · Granted Feb 16, 2021

Method and apparatus for using image data to aid voice recognition

Inventors: Robert A. Zurek (Antioch, IL); Adrian M. Schuster (West Olive, MI); Fu-Lin Shau (Lake Zurich, IL); Jincheng Wu (Naperville, IL)
G10L15/22G06F3/013G06K9/00255G06K9/00335G06K9/00604G06K9/00832G10L15/20G10L15/25G10L15/26G10L21/0208B60N2/002G06K9/00597G10L15/24G10L25/78G10L2015/223G10L2015/227G10L2021/02166H04R2430/20H04R2460/07H04R2499/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,923,124
App. No.
16/416,427
Granted
Feb 16, 2021
Kind
B2
Abstract

A device performs a method for using image data to aid voice recognition. The method includes the device capturing ( 302 ) image data of a vicinity of the device and adjusting ( 304 ), based on the image data, a set of parameters for voice recognition performed by the device ( 102 ). The set of parameters for the device performing voice recognition include, but are not limited to: a trigger threshold of a trigger for voice recognition; a set of beamforming parameters; a database for voice recognition; and/or an algorithm for voice recognition. The algorithm may include using noise suppression or using acoustic beamforming.

Claims (45)

1. A computer-implemented method comprising:

receiving, by one or more computing devices, audio data of an utterance;

obtaining, by the one or more computing devices, image data while receiving the audio data of the utterance;

based on the image data obtained while receiving the audio data, determining, by the one or more computing devices, to perform speech recognition on a first portion of the audio data without performing speech recognition on a second portion of the audio data;

determining, by the one or more computing devices, a number of people included in the image data;

based on the number of people included in the image data, determining, by the one or more computing devices, a level of accuracy for speech recognition; and

performing, by the one or more computing devices, speech recognition on the first portion of the audio data according to the level of accuracy for speech recognition without performing speech recognition on the second portion of the audio data.

2. The method of claim 1 , further comprising:

determining, by the one or more computing devices, that the image data includes a representation of a person,

wherein determining to perform speech recognition on the first portion of the audio data without performing speech recognition on the second portion of the audio data is based on determining that the image data includes a representation of a person.

3. The method of claim 2 , wherein determining that the image data includes a representation of a person comprises determining that the image data includes a representation of both eyes, a nose, and a mouth of a person.

4. The method of claim 2 , wherein determining that the image data includes a representation of a person comprises determining that the image data includes a representation of an authorized user of the one or more computing devices.

5. The method of claim 1 , further comprising:

determining, by the one or more computing devices, that the image data includes a representation of a person who spoke the utterance,

wherein determining to perform speech recognition on the first portion of the audio data without performing speech recognition on the second portion of the audio data is based on determining that the image data includes a representation of the person who spoke the utterance.

6. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving audio data of an utterance;

obtaining image data while receiving the audio data of the utterance;

based on the image data obtained while receiving the audio data, determining to perform speech recognition on a first portion of the audio data without performing speech recognition on a second portion of the audio data;

determining a number of people included in the image data;

based on the number of people included in the image data, determining a level of accuracy for speech recognition; and

performing speech recognition on the first portion of the audio data according to the level of accuracy for speech recognition without performing speech recognition on the second portion of the audio data.

7. The system of claim 6 , wherein the operations further comprise:

determining that the image data includes a representation of a person,

wherein determining to perform speech recognition on the first portion of the audio data without performing speech recognition on the second portion of the audio data is based on determining that the image data includes a representation of a person.

8. The system of claim 7 , wherein determining that the image data includes a representation of a person comprises determining that the image data includes a representation of both eyes, a nose, and a mouth of a person.

9. The system of claim 7 , wherein determining that the image data includes a representation of a person comprises determining that the image data includes a representation of an authorized user of one or more computing devices.

10. The system of claim 6 , wherein the operations further comprise:

determining that the image data includes a representation of a person who spoke the utterance,

wherein determining to perform speech recognition on the first portion of the audio data without performing speech recognition on the second portion of the audio data is based on determining that the image data includes a representation of the person who spoke the utterance.

11. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving audio data of an utterance;

obtaining image data while receiving the audio data of the utterance;

based on the image data obtained while receiving the audio data, determining to perform speech recognition on a first portion of the audio data without performing speech recognition on a second portion of the audio data;

determining a number of people included in the image data;

based on the number of people included in the image data, determining a level of accuracy for speech recognition; and

performing speech recognition on the first portion of the audio data according to the level of accuracy for speech recognition without performing speech recognition on the second portion of the audio data.

12. The medium of claim 11 , wherein the operations further comprise:

determining that the image data includes a representation of a person,

wherein determining to perform speech recognition on the first portion of the audio data without performing speech recognition on the second portion of the audio data is based on determining that the image data includes a representation of a person.

13. The medium of claim 12 , wherein determining that the image data includes a representation of a person comprises determining that the image data includes a representation of both eyes, a nose, and a mouth of a person.

14. The medium of claim 11 , wherein the operations further comprise:

determining that the image data includes a representation of a person who spoke the utterance,

wherein determining to perform speech recognition on the first portion of the audio data without performing speech recognition on the second portion of the audio data is based on determining that the image data includes a representation of the person who spoke the utterance.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2019
From: ZUREK, ROBERT A.; SCHUSTER, ADRIAN M.; SHAU, FU-LIN; WU, JINCHENG
To: MOTOROLA MOBILITY LLC
Reel/Frame 049226/0276 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2019
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 049228/0001 →
Continuity (4)
Continuation 15464704 · Mar 21, 2017
Continuation 14164354 · Jan 27, 2014
Provisional Application 61827048 · May 24, 2013
Related Publication 20190341044A1 · Nov 7, 2019
Cited By (1)
US 12,334,074