IP Library Granted Patent US 11,942,087
Granted Patent B2
US 11,942,087 · App. 17/147,991 · Granted Mar 26, 2024

Method and apparatus for using image data to aid voice recognition

Inventors: Robert A. Zurek (Antioch, IL); Adrian M. Schuster (West Olive, MI); Fu-Lin Shau (Lake Zurich, IL); Jincheng Wu (Naperville, IL)
Assignee: Google Technology Holdings LLC
G10L15/22G06F3/013G06V20/59G06V40/166G06V40/19G06V40/20G10L15/20G10L15/25G10L15/26G10L21/0208B60N2/002G06V40/18G10L2015/223G10L2015/227G10L15/24G10L2021/02166G10L25/78H04R2430/20H04R2460/07H04R2499/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,942,087
App. No.
17/147,991
Granted
Mar 26, 2024
Kind
B2
Abstract

A device performs a method for using image data to aid voice recognition. The method includes the device capturing image data of a vicinity of the device and adjusting, based on the image data, a set of parameters for voice recognition performed by the device. The set of parameters for the device performing voice recognition include, but are not limited to: a trigger threshold of a trigger for voice recognition; a set of beamforming parameters; a database for voice recognition; and/or an algorithm for voice recognition. The algorithm may include using noise suppression or using acoustic beamforming.

Claims (46)

1. A computer-implemented method executed on data processing hardware of a computing device that causes the data processing hardware to perform operations comprising:

receiving a first acoustic signal comprising first voice data corresponding to an utterance spoken by a user of the computing device;

determining, by processing the first voice data, that the utterance spoken by the user comprises a specific verbal command directed toward the computing device that instructs the computing device to process subsequent speech spoken by the user; and

in response to determining that the utterance spoken by the user comprises the specific verbal command:

receiving image data of a vicinity of the computing device, the image data captured by a camera of the computing device;

determining a direction of the user relative to the computing device based on the image data captured by the camera;

adjusting, based on the direction of the user relative to the computing device, a microphone beamform; and

receiving, using the microphone beamform, a second acoustic signal comprising second voice data corresponding to the subsequent speech spoken by the user.

2. The computer-implemented method of claim 1 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to better isolate and capture the second voice data in the second acoustic signal.

3. The computer-implemented method of claim 1 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to reduce an amount of noise captured in the second acoustic signal from acoustic sources other than the user.

4. The computer-implemented method of claim 1 , wherein the operations further comprise:

determining that the user is gazing toward the computing device based on the image data captured by the camera,

wherein adjusting the microphone beamform is further based on the determining that the user is gazing toward the computing device.

5. The computer-implemented method of claim 1 , wherein the operations further comprise:

determining that the image data captured by the camera includes a representation of a person,

wherein determining the direction of the user is based on the image data and the determination that the image data captured by the camera includes the representation of the person.

6. The computer-implemented method of claim 5 , wherein determining that the image data captured by the camera includes the representation of the person comprises determining that the image data includes a representation of both eyes, a nose, and a mouth of a person.

7. The computer-implemented method of claim 5 , wherein determining that the image data captured by the camera includes the representation of the person comprises determining that the image data includes a representation of an authorized user of the computing device.

8. The computer-implemented method of claim 1 , wherein the operations further comprise, after adjusting the microphone beamform, performing speech recognition on the second voice data corresponding to the subsequent speech spoken by the user.

9. The computer-implemented method of claim 1 , wherein receiving the image data of the vicinity of the computing device comprises receiving the image data of the vicinity of the computing device while receiving the first acoustic signal comprising the first voice data corresponding to the utterance spoken by the user.

10. The computer-implemented method of claim 1 , wherein the first acoustic signal further comprises noises captured from acoustic sources other than the user.

11. A computing device associated with a user, the computing device comprising:

a plurality of microphones;

a camera;

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a first acoustic signal comprising first voice data corresponding to an utterance spoken by a user of the computing device;

determining, by processing the first voice data, that the utterance spoken by the user comprises a specific verbal command directed toward the computing device that instructs the computing device to process subsequent speech spoken by the user; and

in response to determining that the utterance spoken by the user comprises the specific verbal command:

receiving image data of a vicinity of the computing device, the image data captured by a camera of the computing device;

determining a direction of the user relative to the computing device based on the image data captured by the camera;

adjusting, based on the direction of the user relative to the computing device, a microphone beamform; and

receiving, using the microphone beamform, a second acoustic signal comprising second voice data corresponding to the subsequent speech spoken by the user.

12. The computing device of claim 11 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to better isolate and capture the second voice data in the second acoustic signal.

13. The computing device of claim 11 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to reduce an amount of noise captured in the second acoustic signal from acoustic sources other than the user.

14. The computing device of claim 11 , wherein the operations further comprise:

determining that the user is gazing toward the computing device based on the image data captured by the camera,

wherein adjusting the microphone beamform is further based on the determining that the user is gazing toward the computing device.

15. The computing device of claim 11 , wherein the operations further comprise:

determining that the image data captured by the camera includes a representation of a person,

wherein determining the direction of the user is based on the image data and the determination that the image data captured by the camera includes the representation of the person.

16. The computing device of claim 15 , wherein determining that the image data captured by the camera includes the representation of the person comprises determining that the image data includes a representation of both eyes, a nose, and a mouth of a person.

17. The computing device of claim 15 , wherein determining that the image data captured by the camera includes the representation of the person comprises determining that the image data includes a representation of an authorized user of the computing device.

18. The computing device of claim 11 , wherein the operations further comprise, after adjusting the microphone beamform, performing speech recognition on the second voice data corresponding to the subsequent speech spoken by the user.

19. The computing device of claim 11 , wherein receiving the image data of the vicinity of the computing device comprises receiving the image data of the vicinity of the computing device while receiving the first acoustic signal comprising the first voice data corresponding to the utterance spoken by the user.

20. The computing device of claim 11 , wherein the first acoustic signal further comprises noises captured from acoustic sources other than the user.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2021
From: ZUREK, ROBERT A.; SCHUSTER, ADRIAN M.; SHAU, FU-LIN; WU, JINCHENG
To: MOTOROLA MOBILITY LLC
Reel/Frame 054907/0488 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2021
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 054979/0019 →