IP Library Granted Patent US 12,334,074
Granted Patent B2
US 12,334,074 · App. 18/606,066 · Granted Jun 17, 2025

Method and apparatus for using image data to aid voice recognition

Inventors: Robert A. Zurek (Antioch, IL); Adrian M. Schuster (West Olive, MI); Fu-Lin Shau (Lake Zurich, IL); Jincheng Wu (Naperville, IL)
Assignee: Google Technology Holdings LLC
G10L15/22G06F3/013G06V20/59G06V40/166G06V40/19G06V40/20G10L15/20G10L15/25G10L15/26G10L21/0208G06V40/18G10L2015/223G10L2015/227G10L15/24G10L2021/02166G10L25/78H04R2430/20H04R2460/07H04R2499/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,074
App. No.
18/606,066
Granted
Jun 17, 2025
Kind
B2
Abstract

A device performs a method for using image data to aid voice recognition. The method includes the device capturing image data of a vicinity of the device and adjusting, based on the image data, a set of parameters for voice recognition performed by the device. The set of parameters for the device performing voice recognition include, but are not limited to: a trigger threshold of a trigger for voice recognition; a set of beamforming parameters; a database for voice recognition; and/or an algorithm for voice recognition. The algorithm may include using noise suppression or using acoustic beamforming.

Claims (62)

1. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

obtaining image data comprising a representation of a first user and a second user;

obtaining audio data comprising:

a first voice data corresponding to the first user speaking; and

a second voice data corresponding to the second user speaking;

associating, based on the image data, the first voice data to a first voice-recognition database of the first user speaking and the second voice data to a second voice-recognition database of the second user speaking;

generating, using speech-to-text conversion, a transcription of the audio data;

annotating, based on the first voice data associated with the first voice-recognition database, a first portion of the transcription corresponding to the first voice data with a first annotation identifying the first user; and

annotating, based on the second voice data associated with the second voice-recognition database, a second portion of the transcription corresponding to the second voice data with a second annotation identifying the second user.

2. The computer-implemented method of claim 1 , wherein the operations further comprise:

determining a direction of the first user relative to a computing device that captured the image data; and

adjusting, based on the direction of the first user relative to the computing device that captured the image data, a microphone beamform.

3. The computer-implemented method of claim 2 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to better isolate and capture the first voice data.

4. The computer-implemented method of claim 1 , wherein the operations further comprise:

determining a direction of the second user relative to a computing device that captured the image data; and

adjusting, based on the direction of the second user relative to the computing device that captured the image data, a microphone beamform.

5. The computer-implemented method of claim 4 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to better isolate and capture the second voice data.

6. The computer-implemented method of claim 1 , wherein the operations further comprise:

determining that the first user is gazing toward a computing device that captured the image data; and

based on determining that the first user is gazing toward the computing device that captured the image data, adjusting a microphone beamform to better isolate and capture the first voice data.

7. The computer-implemented method of claim 1 , wherein the operations further comprise:

determining that the second user is gazing toward a computing device that captured the image data; and

based on determining that the second user is gazing toward the computing device that captured the image data, adjusting a microphone beamform to better isolate and capture the second voice data.

8. The computer-implemented method of claim 1 , wherein the audio data further comprises noises captured from acoustic sources other than the first user and the second user.

9. The computer-implemented method of claim 1 , wherein:

the first annotation comprises a name of the first user; and

the second annotation comprises a name of the second user.

10. The computer-implemented method of claim 1 , wherein:

the first annotation comprises a title of the first user; and

the second annotation comprises a title of the second user.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining image data comprising a representation of a first user and a second user;

obtaining audio data comprising:

a first voice data corresponding to the first user speaking; and

a second voice data corresponding to the second user speaking;

associating, based on the image data, the first voice data to a first voice-recognition database of the first user speaking and the second voice data to a second voice-recognition database of the second user speaking;

generating, using speech-to-text conversion, a transcription of the audio data;

annotating, based on the first voice data associated with the first voice- recognition database, a first portion of the transcription corresponding to the first voice data with a first annotation identifying the first user; and

annotating, based on the second voice data associated with the second voice-recognition database, a second portion of the transcription corresponding to the second voice data with a second annotation identifying the second user.

12. The system of claim 11 , wherein the operations further comprise:

determining a direction of the first user relative to a computing device that captured the image data; and

adjusting, based on the direction of the first user relative to the computing device that captured the image data, a microphone beamform.

13. The system of claim 12 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to better isolate and capture the first voice data.

14. The system of claim 11 , wherein the operations further comprise:

determining a direction of the second user relative to a computing device that captured the image data; and

adjusting, based on the direction of the second user relative to the computing device that captured the image data, a microphone beamform.

15. The system of claim 14 , wherein adjusting the microphone beamform comprises adjusting the microphone beamform to better isolate and capture the second voice data.

16. The system of claim 11 , wherein the operations further comprise:

determining that the first user is gazing toward a computing device that captured the image data; and

based on determining that the first user is gazing toward the computing device that captured the image data, adjusting a microphone beamform to better isolate and capture the first voice data.

17. The system of claim 11 , wherein the operations further comprise:

determining that the second user is gazing toward a computing device that captured the image data; and

based on determining that the second user is gazing toward the computing device that captured the image data, adjusting a microphone beamform to better isolate and capture the second voice data.

18. The system of claim 11 , wherein the audio data further comprises noises captured from acoustic sources other than the first user and the second user.

19. The system of claim 11 , wherein:

the first annotation comprises a name of the first user; and

the second annotation comprises a name of the second user.

20. The system of claim 11 , wherein:

the first annotation comprises a title of the first user; and

the second annotation comprises a title of the second user.

Continuity (6)
Continuation 17147991 · Jan 13, 2021
Continuation 16416427 · May 20, 2019
Continuation 15464704 · Mar 21, 2017
Continuation 14164354 · Jan 27, 2014
Provisional Application 61827048 · May 24, 2013
Related Publication 20240221745A1 · Jul 4, 2024
References Cited (91)
US 4951079A · Hoshino · 1990 [cited by examiner]
US 5983186A · Miyazawa · 1999 [cited by examiner]
US 6021278A · Bernardi · 2000 [cited by examiner]
US 6101338A · Bernardi · 2000 [cited by examiner]
US 7043429B2 · Chang · 2006 [cited by examiner]
US 7778632B2 · Kurlander · 2010 [cited by examiner]
US 7962331B2 · Miller · 2011 [cited by examiner]
US 8073590B1 · Zilka · 2011 [cited by examiner]
US 8078397B1 · Zilka · 2011 [cited by examiner]
US 8131458B1 · Zilka · 2012 [cited by examiner]
US 8249309B2 · Kurzweil · 2012 [cited by examiner]
US 8265862B1 · Zilka · 2012 [cited by examiner]
US 8285791B2 · Ratcliff · 2012 [cited by examiner]
US 8589160B2 · Weeks et al. · 2013 [cited by applicant]
US 8612211B1 · Shires · 2013 [cited by examiner]
US 8660847B2 · Soemo · 2014 [cited by examiner]
US 8666750B2 · Buck · 2014 [cited by examiner]
US 8700392B1 · Hart · 2014 [cited by examiner]
US 9094576B1 · Karakotsios · 2015 [cited by examiner]
US 9420227B1 · Shires · 2016 [cited by examiner]
US 9479736B1 · Karakotsios · 2016 [cited by examiner]
US 9747900B2 · Zurek · 2017 [cited by examiner]
US 10311868B2 · Zurek · 2019 [cited by examiner]
US 10482904B1 · Hardie et al. · 2019 [cited by applicant]
US 10923124B2 · Zurek · 2021 [cited by examiner]
US 11133027B1 · Hardie et al. · 2021 [cited by applicant]
US 11148296B2 · Breazeal · 2021 [cited by applicant]
US 11195531B1 · Adams et al. · 2021 [cited by applicant]
US 11562731B2 · Thomson · 2023 [cited by examiner]
US 11942087B2 · Zurek · 2024 [cited by examiner]
US 20020105575A1 · Hinde · 2002 [cited by examiner]
US 20030018475A1 · Basu · 2003 [cited by examiner]
US 20040267521A1 · Cutler · 2004 [cited by examiner]
US 20040267536A1 · Hershey · 2004 [cited by examiner]
US 20050086051A1 · Brulle-Drews · 2005 [cited by examiner]
US 20050102133A1 · Rees · 2005 [cited by examiner]
US 20050128311A1 · Rees · 2005 [cited by examiner]
US 20050272415A1 · McConnell · 2005 [cited by examiner]
US 20060100876A1 · Nishizaki · 2006 [cited by examiner]
US 20070244700A1 · Kahn · 2007 [cited by examiner]
US 20080037837A1 · Noguchi · 2008 [cited by examiner]
US 20080059147A1 · Afify · 2008 [cited by examiner]
US 20080262849A1 · Buck · 2008 [cited by examiner]
US 20080289002A1 · Portele · 2008 [cited by examiner]
US 20090018831A1 · Morita · 2009 [cited by examiner]
US 20090037171A1 · McFarland · 2009 [cited by examiner]
US 20090043576A1 · Miller · 2009 [cited by examiner]
US 20100103242A1 · Linaker · 2010 [cited by examiner]
US 20100241432A1 · Michaelis · 2010 [cited by examiner]
US 20100268534A1 · Kishan Thambiratnam · 2010 [cited by examiner]
US 20100328316A1 · Stroila · 2010 [cited by examiner]
US 20110043652A1 · King · 2011 [cited by examiner]
US 20110092249A1 · Evanitsky · 2011 [cited by examiner]
US 20110150270A1 · Carpenter · 2011 [cited by examiner]
US 20110184735A1 · Flaks · 2011 [cited by examiner]
US 20110224979A1 · Raux · 2011 [cited by examiner]
US 20110257971A1 · Morrison · 2011 [cited by examiner]
US 20120014567A1 · Allegra · 2012 [cited by examiner]
US 20120089397A1 · Arai · 2012 [cited by examiner]
US 20120143605A1 · Thorsen · 2012 [cited by examiner]
US 20120173224A1 · Anisimovich · 2012 [cited by examiner]
US 20120215539A1 · Juneja · 2012 [cited by examiner]
US 20120236025A1 · Jacobsen · 2012 [cited by examiner]
US 20120327177A1 · Kee · 2012 [cited by examiner]
US 20130046537A1 · Weeks · 2013 [cited by examiner]
US 20130053007A1 · Cosman · 2013 [cited by examiner]
US 20130060571A1 · Soemo · 2013 [cited by examiner]
US 20130073583A1 · Licata · 2013 [cited by examiner]
US 20130090931A1 · Ghovanloo · 2013 [cited by examiner]
US 20130096918A1 · Harada · 2013 [cited by examiner]
US 20130124207A1 · Sarin · 2013 [cited by examiner]
US 20130173701A1 · Goyal · 2013 [cited by examiner]
US 20130339024A1 · Kojima · 2013 [cited by examiner]
US 20140350924A1 · Zurek · 2014 [cited by examiner]
US 20150206535A1 · Iwai · 2015 [cited by examiner]
US 20160231818A1 · Zhang et al. · 2016 [cited by applicant]
US 20170193996A1 · Zurek · 2017 [cited by examiner]
US 20190341044A1 · Zurek · 2019 [cited by examiner]
US 20210134293A1 · Zurek · 2021 [cited by examiner]
US 20240221745A1 · Zurek · 2024 [cited by examiner]
EP 1443498A1 · 2004 [cited by applicant]
EP 1884421B1 · 2008 [cited by applicant]
EP 2065871A1 · 2009 [cited by applicant]
WO 2002015560A2 · 2002 [cited by applicant]
WO WO2014189666A1 · 2014 [cited by examiner]
Sergei Preobrazhensky, Optimizing Acoustic Array Beamforming to Aid a Speech Recognition System, https://kb.osu.edu/dspace/bitstream/handle/1811/52868/ECE_683H_HONORS_THESIS.pdf?sequence=1, May 2012, all pages. [cited by applicant]
Bub et al., “Knowing who to listen to in speech recognition: visually guided beamforming,” 1995 International Conference on Acoustics, Speech, and Signal Processing—May 9-12, 1995—Detroit, MI, USA, IEEE—New York, NY, US… [cited by applicant]
Invitation to Pay Additional Fees and, Where Applicable, Protest Fee in International Application No. PCT/US2014/036831, mailed Jul. 28, 2014, 8 pages. [cited by applicant]
Ho, T., et al., “A word shape analysis approach to lexicon based word recognition”, Pattern Recognition Letters 13 (1992) 821-826, Nov. 1992. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2014/036831, mailed Nov. 3, 2014, 16 pages. [cited by applicant]
International Preliminary Report on Patentability in International Application No. PCT/US2014/036831, mailed Dec. 3, 2015, 11 pages. [cited by applicant]