IP Library › Granted Patent US 12,592,237
Granted Patent B2
US 12,592,237 · App. 17/547,917 · Granted Mar 31, 2026

Driver interface with voice and image control

Inventors: Zili Li (San Jose, CA); Cristina Vasconcelos (Toronto, CA)
Assignee: SoundHound AI IP, LLC
G10L15/24G06F18/217G06V10/764G06V10/768G06V10/82G06V20/46G10L15/02G10L15/063G10L15/16G10L15/1822G10L15/187G10L15/22G10L15/30G10L2015/025G10L2015/223G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,237
App. No.
17/547,917
Granted
Mar 31, 2026
Kind
B2
Abstract

A driver interface for use within an automobile provides responses to voice commands issued for example by a driver of the automobile. The interface includes a camera and microphone for capturing image data such as gestures and audio data from the automobile driver. The image data and audio data are processed to extract image and linguistic features from the image and audio data, which image and linguistic features are processed to interpret and infer a meaning of the voice command.

Claims (39)

1 . A driver interface system comprising:

a camera interface enabled to capture images of a driver and an environment surrounding the driver;

an image processor enabled to extract, from captured images of the driver and an environment surrounding the driver, visual features, the visual features being an abstraction of the captured images, the abstraction being a compressed representation of the captured images;

a microphone interface enabled to capture speech;

a speech recognizer enabled to extract, from the captured speech, linguistic features;

a natural language processor enabled to infer a voice command from the linguistic features and the visual features; and

an output interface for rendering a response to the voice command; and

wherein the speech recognizer uses a machine-learned linguistic model based on a combination of audio feature tensors and visual feature tensors.

2 . The driver interface system of claim 1 wherein inferring the voice command from the linguistic features and the visual features is configured by a set of acoustic model configurations parameters.

3 . The driver interface system of claim 1 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the speech recognizer extracts linguistic features by making an application programming interface (API) call to the server through the data interface.

4 . The driver interface system of claim 3 wherein the data interface is a wireless network interface.

5 . The driver interface system of claim 1 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the natural language processor infers the voice command by making an application programming interface (API) call to the server through the data interface.

6 . The driver interface system of claim 5 wherein the data interface is a wireless network interface.

7 . The driver interface system of claim 1 wherein the visual features indicate a gesture.

8 . The driver interface system of claim 1 wherein the response is rendered by projection onto a display screen.

9 . A driver interface system comprising:

a camera interface enabled to capture images of a driver and an environment surrounding the driver;

a neural network image processor enabled to extract, from captured images of the driver and an environment surrounding the driver, visual features, the visual features being an abstraction of the captured images determined by the neural network image processor, the abstraction being a compressed representation of the captured images with reduced information content as compared to the captured images;

a microphone interface enabled to capture speech;

a speech recognizer enabled to extract, from the captured speech, linguistic features; and

a natural language processor enabled to infer a voice command from the linguistic features and the visual features, including visual features absracted from an environment surrounding the driver;

wherein the image processor, the speech recognizer, and the natural language processor are jointly configured by joint training, such that the extraction of the visual feature tensors and audio feature tensors is optimized in combination with the natural language processor to minimize errors in inferring the voice command.

10 . The driver interface system of claim 9 wherein inferring the voice command from the linguistic features and the visual features is configured by a set of acoustic model configurations parameters.

11 . The driver interface system of claim 9 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the speech recognizer extracts linguistic features by making an application programming interface (API) call to the server through the data interface.

12 . The driver interface system of claim 11 wherein the data interface is a wireless network interface.

13 . The driver interface system of claim 9 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the natural language processor infers the voice command by making an application programming interface (API) call to the server through the data interface.

14 . The driver interface system of claim 13 wherein the data interface is a wireless network interface.

15 . The driver interface system of claim 9 wherein the speech recognizer uses a machine-learned linguistic model based on a combination of audio feature tensors and visual feature tensors.

16 . The driver interface system of claim 9 wherein the visual features indicate a gesture.

17 . The driver interface system of claim 9 further comprising an output interface for rendering a response to the voice command.

18 . The driver interface system of claim 17 wherein the response is rendered by projection onto a display screen.

19 . A driver interface system comprising:

a camera interface enabled to capture images of a driver and an environment surrounding the driver;

an image processor enabled to extract, from captured images of the driver and an environment surrounding the driver, visual feature tensors, the visual feature tensors being an abstraction of the captured images, the abstraction being a compressed representation of the captured images with reduced information content as compared to the captured images;

a microphone interface enabled to capture speech;

a speech recognizer enabled to extract, from the captured speech, audio feature tensors;

a natural language processor enabled to concatonate or merge the visual feature vectors and audio feature vectors to infer a voice command;

wherein the image processor, the speech recognizer, and the natural language processor are jointly configured by joint training, such that the extraction of the visual feature tensors and audio feature tensors is optimized in combination with the natural language processor to minimize errors in inferring the voice command; and

an output interface for rendering a response to the voice command.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2021
From: LI, ZILI; VASCONCELOS, CRISTINA
To: SOUNDHOUND, INC.
Reel/Frame 058362/0353 →
Continuity (2)
Continuation 16509029 · Jul 11, 2019
Related Publication 20220139393A1 · May 5, 2022
References Cited (187)
US 6366882B1 · Bijl et al. · 2002 [cited by applicant]
US 6442820B1 · Mason · 2002 [cited by applicant]
US 6633844B1 · Verma et al. · 2003 [cited by applicant]
US 6964023B2 · Maes et al. · 2005 [cited by applicant]
US 7254538B1 · Hermansky et al. · 2007 [cited by applicant]
US 8442820B2 · Kim et al. · 2013 [cited by applicant]
US 8768693B2 · Somekh et al. · 2014 [cited by applicant]
US 9190058B2 · Klein · 2015 [cited by applicant]
US 9418674B2 · Tzirkel-Hancock · 2016 [cited by examiner]
US 9514751B2 · Kim · 2016 [cited by applicant]
US 10595072B2 · Wexler · 2020 [cited by examiner]
US 10614803B2 · Xie · 2020 [cited by examiner]
US 11257493B2 · Vasconcelos · 2022 [cited by examiner]
US 11290518B2 · Kim · 2022 [cited by examiner]
US 20020135618A1 · Maes et al. · 2002 [cited by applicant]
US 20030018475A1 · Basu et al. · 2003 [cited by applicant]
US 20030130840A1 · Forand · 2003 [cited by applicant]
US 20090060351A1 · Li et al. · 2009 [cited by applicant]
US 20100106486A1 · Hua et al. · 2010 [cited by applicant]
US 20110071830A1 · Kim et al. · 2011 [cited by applicant]
US 20110119053A1 · Kuo et al. · 2011 [cited by applicant]
US 20110161076A1 · Davis et al. · 2011 [cited by applicant]
US 20120265528A1 · Gruber et al. · 2012 [cited by applicant]
US 20130311184A1 · Badavne et al. · 2013 [cited by applicant]
US 20140112556A1 · Kalinli-Akbacak · 2014 [cited by examiner]
US 20140214424A1 · Wang et al. · 2014 [cited by applicant]
US 20150161995A1 · Sainath et al. · 2015 [cited by applicant]
US 20150199960A1 · Huo et al. · 2015 [cited by applicant]
US 20150205568A1 · Matsuoka · 2015 [cited by applicant]
US 20150235641A1 · Vanblon et al. · 2015 [cited by applicant]
US 20150340034A1 · Schalkwyk et al. · 2015 [cited by applicant]
US 20160034811A1 · Paulik et al. · 2016 [cited by applicant]
US 20170061966A1 · Marcheret et al. · 2017 [cited by applicant]
US 20170186044A1 · Tal-Israel · 2017 [cited by examiner]
US 20170186432A1 · Aleksic et al. · 2017 [cited by applicant]
US 20170236516A1 · Lane · 2017 [cited by applicant]
US 20180075849A1 · Khoury et al. · 2018 [cited by applicant]
US 20190130200A1 · Stent · 2019 [cited by examiner]
US 20210065712A1 · Holm · 2021 [cited by applicant]
CN 101187990 · 2008 [cited by applicant]
CN 101187990A · 2008 [cited by applicant]
CN 104463250 · 2015 [cited by applicant]
CN 104463250A · 2015 [cited by applicant]
CN 105760852 · 2016 [cited by applicant]
CN 106782545 · 2017 [cited by applicant]
CN 107945789 · 2018 [cited by applicant]
CN 109147763 · 2019 [cited by applicant]
CN 109697976 · 2019 [cited by applicant]
CN 110111783 · 2019 [cited by applicant]
CN 110136698 · 2019 [cited by applicant]
CN 107507612 · 2020 [cited by applicant]
CN 107507612B · 2020 [cited by applicant]
JP 2002268683 · 2002 [cited by applicant]
JP 2002268683A · 2002 [cited by applicant]
JP 2003241788 · 2003 [cited by applicant]
JP 2003271182 · 2003 [cited by applicant]
JP 2003271182A · 2003 [cited by applicant]
JP 2004260641 · 2004 [cited by applicant]
JP 2004260641A · 2004 [cited by applicant]
JP 2004333738 · 2004 [cited by applicant]
JP 2005128307 · 2005 [cited by applicant]
JP 2005128307A · 2005 [cited by applicant]
JP 2007027990 · 2007 [cited by applicant]
JP 2007027990A · 2007 [cited by applicant]
JP 2012022053 · 2012 [cited by applicant]
JP 2012242609 · 2012 [cited by applicant]
JP 2015088099 · 2015 [cited by applicant]
JP 2015088099A · 2015 [cited by applicant]
JP 2015175859 · 2015 [cited by applicant]
JP 2016143050 · 2016 [cited by applicant]
JP 2016170701 · 2016 [cited by applicant]
JP 2017090612 · 2017 [cited by applicant]
JP 2017090612A · 2017 [cited by applicant]
JP 2019097016 · 2019 [cited by applicant]
JP 2019097016A · 2019 [cited by applicant]
JP 2019102081 · 2019 [cited by applicant]
JP 2019102081A · 2019 [cited by applicant]
JP 2019128732 · 2019 [cited by applicant]
JP 2019128732A · 2019 [cited by applicant]
KR 20030077012A · 2003 [cited by applicant]
KR 20090055426 · 2009 [cited by applicant]
KR 20090055426A · 2009 [cited by applicant]
KR 20180111197 · 2018 [cited by applicant]
KR 20180111197A · 2018 [cited by applicant]
KR 20180126353 · 2018 [cited by applicant]
TW 201140558 · 2011 [cited by applicant]
WO 2011111221 · 2011 [cited by applicant]
WO 2011111221A1 · 2011 [cited by applicant]
WO 2014199596 · 2017 [cited by applicant]
WO 2017135148 · 2018 [cited by applicant]
WO 2019159364 · 2020 [cited by applicant]
Tatsuya Shigetomi, et al., “Location Estimation Method Robust Against Environmental Changes Using Deep Learning,” Proceedings of the 22nd Annual Conference of the Virtual Reality Society of Japan, Sep. 2017 (with Englis… [cited by applicant]
English language Abstract of JP2004333738 published Nov. 25, 2004. [cited by applicant]
English language Abstract of JP2012022053 published Feb. 2, 2012. [cited by applicant]
English language Abstract of JP2003241788 published Aug. 29, 2003. [cited by applicant]
English language Abstract for JP2019097016 published Jun. 20, 2019. [cited by applicant]
English language Abstract for JP2003271182 published Sep. 25, 2003. [cited by applicant]
English language Abstract for WO2011111221 published Sep. 15, 2011. [cited by applicant]
English language Abstract for JP2002268683 published Sep. 20, 2002. [cited by applicant]
Somada, “For face recognition that is robust against aging Research on collation methods”, master's thesis for the Department of Information Engineering, Graduate School of Engineering, Mie University, Mar. 2018. [cited by applicant]
Dave Sullivan, GM's Super Cruise Is a Marriage of Culling-Edge Hardware and Software, Forbes, Oct. 23, 2017. [cited by applicant]
Mtank, Multi-Modal Methods: Visual Speech Recognition (Lip Reading), Medium.com, May 3, 2018. [cited by applicant]
Sri Garimella, et al., Robust i-vector based adaptation of DNN acoustic model for speech recognition. Sixteenth Annual Conference of the International Speech Communication Association, Sep. 6-10, 2015. [cited by applicant]
David Snyder, et al., Spoken Language Recognition using X-vectors. InOdyssey Jun. 2018 (pp. 105-111). [cited by applicant]
Najim Dehak, et al., Front-end factor analysis for speaker verification. IEEE Transactions on Audio, Speech, and Language Processing. Aug. 9, 2010; 19(4):788-98. [cited by applicant]
Yannis M. Assael, et al., Lipnet: End-to-end sentence-level lipreading. arXiv preprint arXiv: 1611.01599. Nov. 6, 2016. [cited by applicant]
Nvidia, Nvidia Drive IX, developer.nvidia.com, Jun. 5, 2019. [cited by applicant]
Jesus F. Guitartr Perez, et al., Lip reading for robust speech recognition on embedded devices. InProceedings. ICASSP'05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005. Mar. 23, 2005 (… [cited by applicant]
Triantafyllos Afouras, et al., Deep audio-visual speech recognition. IEEE transactions on pattern analysis and machine intelligence. Dec. 21, 2018. [cited by applicant]
Brendan Shillingford et al., Large-Scale Visual Speech Recognition, 1807.05162v3 [cs.CV] Oct. 1, 2018, DeepMind and Google. [cited by applicant]
Shinichi Hara, et al., Speaker-Adaptive Speech Recognition using Facial Image Identification, IEICE Technical Report, SP2012-55 (Jul. 2012), School of Engineering, Soka University, Tokyo, Japan. [cited by applicant]
Moriya et al., “Multimodal Speaker Adaptation of Acoustic Model and Language Model for ASR Using Speaker Face Embedding.” ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP… [cited by applicant]
English language Abstract for CN107507612 published Aug. 28, 2020. [cited by applicant]
English language Abstract for JP2004260641 published Sep. 16, 2004. [cited by applicant]
English language Abstract for JP2017090612 published May 25, 2017. [cited by applicant]
English language Abstract for JP2019102081 published Jun. 24, 2019. [cited by applicant]
English language Abstract for JP2019128732 published Aug. 1, 2019. [cited by applicant]
English language Abstract for JP2005128307 published May 19, 2005. [cited by applicant]
English language Abstract for JP2007027990 published Feb. 1, 2007. [cited by applicant]
English language Abstract for JP2015088099 published May 7, 2015. [cited by applicant]
Ciprian Chelba, et al., Sparse non-negative matrix language modeling for geo-annotated query session data. In2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) Dec. 13, 2015 (pp. 8-14). IEEE. [cited by applicant]
Oriol Vinyals, et al., Show and tell: A neural image caption generator. InProceedings of the IEEE conference on computer vision and pattern recognition 2015 (pp. 3156-3164). [cited by applicant]
Yusuf Aytar, et al., Soundnet: Learning sound representations from unlabeled video. InAdvances in neural information processing systems 2016 (pp. 892-900). [cited by applicant]
Joon Son Chung, et al., Lip reading sentences in the wild. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Jul. 21, 2017 (pp. 3444-3453). IEEE. [cited by applicant]
Hans-Gunter Hirsch, et al., The Aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions. InASR2000-Automatic Speech Recognition: Challenges for the new Millenium… [cited by applicant]
Wikipedia, tf-idf, Jul. 9, 2019. [cited by applicant]
Michael Johnston, et al., MATCH: An architecture for multimodal dialogue systems. InProceedings of the 40th Annual Meeting on Association for Computational Linguistics Jul. 6, 2002 (pp. 376-383). Association for Computa… [cited by applicant]
X. Chen, et al., Recurrent neural network language model adaptation for multi-genre broadcast speech recognition. InSixteenth Annual Conference of the International Speech Communication Association 2015. [cited by applicant]
Tomas Mikolov, et al., Context dependent recurrent neural network language model. In2012 IEEE Spoken Language Technology Workshop (SLT) Dec. 2, 2012 (pp. 234-239). IEEE. [cited by applicant]
Cong Duy Vu Hoang, et al., Incorporating side information into recurrent neural network language models. InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistic… [cited by applicant]
Yi Luan, et al., LSTM based conversation models. arXiv preprint arXiv:1603.09457. Mar. 31, 2016. [cited by applicant]
Shalini Ghosh, et al., Contextual Istm (clstm) models for large scale nlp tasks. arXiv preprint arXiv:1602.06291. Feb. 19, 2016. [cited by applicant]
Aaron Jaech, et al., Low-rank RNN adaptation for context-aware language modeling. Transactions of the Association for Computational Linguistics. Jul. 2018;6:497-510. [cited by applicant]
English language Abstract for CN101187990 published May 28, 2018. [cited by applicant]
English language Abstract for CN104463250 published Mar. 25, 2015. [cited by applicant]
English language Abstract for KR20180111197 published Oct. 11, 2018. [cited by applicant]
English language Abstract for KR20090055426 published Jun. 2, 2009. [cited by applicant]
Office Action dated Sep. 2, 2021 in U.S. Appl. No. 16/509,029. [cited by applicant]
Response to Office Action filed Sep. 2, 2021 in U.S. Appl. No. 16/509,029. [cited by applicant]
Notice of Allowance dated Dec. 17, 2021 in U.S. Appl. No. 16/509,029. [cited by applicant]
Amendment Under 37 C.F.R. 1.312 dated Oct. 20, 2021 in U.S. Appl. No. 16/509,029. [cited by applicant]
English language Abstract of KR20180126363 published Nov. 27, 2018. [cited by applicant]
Aaron Jaech, et al., “Low-rank RNN_ adaptation for context-aware language modeling”, Transactions of the Association for Computational Linguistics, Jul. 1, 2018; 6:497-510. [cited by applicant]
Vahid Asadpour et al., “Audio-visual speaker identification uing dynamic facial movements and utterance phonetic content.” Applied Soft Computing 11, No. 2, Jul. 13, 2010; 2083-2093. [cited by applicant]
Chao, et al., “Deep Speaker Embedding for Speaker-Targeted Automatic Speech Recognition”, Proceedings of the 2019 3rd International Conference on Natural Language Processing and Information Retrieval—NLPIR 2019, Jun. 28… [cited by applicant]
Ciprian Chelba et al., “Multinomial Loss on Held-Out Data for the Sparse Non-negative Matrix Language Model”, Feb. 22, 2016. [cited by applicant]
Cong Duy, et al., “Incorporating Side Information into Recurrent Neural Network Language Models”, In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics Huma… [cited by applicant]
Dave Sullivan, “GM's Super Cruise Is a Marriage of Cutting-Edge Hardware and Software”, Forbes, Oct. 23, 2017. [cited by applicant]
David Snyder, et al., “Spoken Language Recognition using X-vectors”, InOdyssey Jun. 26, 2018 (pp. 105-111). [cited by applicant]
David Pearce, et al., “The Aurora experimental framework for the performance evaluation of the speech recognition systems under noisy conditions”, In ASR2000-Automatic Speech Recognition: Challenges for the new Milleniu… [cited by applicant]
Jesus F. Guitartr Perez, et al., “Lip Reading for Robust Speech Recognition on Embedded Devices”, InProceedings.(ICASSP'05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005. Mar. 23, 2005… [cited by applicant]
Darsh J Shah et al., “Robust Zero-Shot Cross-Domain Slot Filling with Example Values”, arXiv:1906.06870v1, Jun. 17, 2019. [cited by applicant]
Yasufumi Moriya et al., “Multimodal Speaker Adaptation of Acoustic Model and Language Model for ASR Using Speaker Face Embedding.” ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processin… [cited by applicant]
Michael Johnston, et al., MATCH: An architecture for Multimodal Dialogue Systems, In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics Jul. 6, 2002 (pp. 376-383). [cited by applicant]
Nvidia, Nvidia Drive IX, developer.nvidia.com, 2019. [cited by applicant]
Oriol Vinyals, et al., “Show and Tell: A Neural Image Caption Generator”, In Proceedings of the IEEE Conference on Computer vision and pattern recognition, Jun. 7-12, 2015 (pp. 3156-3164). [cited by applicant]
Shalini Ghosh, et al., “Contextual LSTM (CLSTM) models for Large scale NLP tasks”, arXiv: 1602.06291v2. May 31, 2016. [cited by applicant]
Brendan Shillingford, et al. “Large-scale visual speech recognition.” arXiv:1807.05162v3 Oct. 1, 2018. [cited by applicant]
Daniel R. Flynn, “Multi-Modal Methods: Recent Interactions Between Computer Vision and Natural Language Processing”, The M Tank, Jul. 2, 2018. [cited by applicant]
Tomas Mikolov, et al., “Context Dependent Recurrent Neural Network Language Model”, In 2012 IEEE Spoken Language Technology Workshop (SLT) Dec. 2012 (pp. 234-239); Microsoft Research Technical Report, Jul. 27, 2012. [cited by applicant]
Triantafyllos Afouras, et al., “Deep Audio-Visual Speech Recognition”, IEEE transactions on pattern analysis and machine intelligence. Dec. 2018. [cited by applicant]
K. Chen, et al., “Recurrent Neural Network Language Model Adaptation for Multi-Genre Broadcast Speech Recognition”, Sixteenth Annual Conference of the International Speech Communication Association, Sep. 6-10, 2015. [cited by applicant]
Yajie Miao et al., “Improvements to Speaker Adaptive Training of Deep Neural Networks”, Language Technologies Institute, School of Computer Science, Carnegie Mellon University, © 2014 IEEE Spoken Language Technology Wor… [cited by applicant]
Yannis M. Assael, et al., “Lipnet: End-to-End Sentence-Level Lipreading”, arXiv:1611.01599v2, Dec. 16, 2016. [cited by applicant]
Yi Luan, et al., “LSTM based Conversation Models”, arXiv: 1603.09457v1. Mar. 31, 2016. [cited by applicant]
Yusuf Aytar, et al., “Soundnet: Learning Sound Representations from Unlabeled Video”, In Advances in neural information processing systems 2016 (pp. 892-900); 30th Conference on Neural Information Processing Systems (NI… [cited by applicant]
Yuta Soda, “Master's Thesis—Resarch on collation methods for face recognition that is robust against secular variation”, Master's Program in Information Engineering, Graduate School of Engineering, Mie University, 2017. [cited by applicant]
Jesus F. Guitartr Perez, et al., Lip reading for robust speech recognition on embedded devices. InProceedings. ICASSP'05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005. Mar. 2, 20053 (… [cited by applicant]
Joon Son Chung, et al., Lip reading sentences in the wild. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Jul. 2, 20171 (pp. 3444-3453). IEEE. [cited by applicant]
Michael Johnston, et al., MATCH: An architecture for multimodal dialogue systems. InProceedings of the 40th Annual Meeting on Association for Computational Linguistics Jul. 6, 2002 (pp. 376-383). Association for Computa… [cited by applicant]
Yi Luan, et al., LSTM based conversation models. arXiv preprint arXiv:1603.09457. Mar. 3, 20161. [cited by applicant]
Response to Office Action dated Sep. 2, 2021 in U.S. Appl. No. 16/509,029. [cited by applicant]
Notice of Allowance and Fees Due dated Oct. 4, 2021 in U.S. Appl. No. 16/509,029. [cited by applicant]
Amendment Under 37 C.F.R. 1.1312 dated Oct. 20, 2021 in U.S. Appl. No. 16/509,029. [cited by applicant]
Feng et al., “Audio Visual Speech Recognition with Multimodal Recurrent Neural Networks”, IEEE, 2017 International Joint Conference on Neural Networks (IJCNN), May 14-19, 2017. [cited by applicant]
English language Abstract for WO2017135148 published Nov. 29, 2018. [cited by applicant]
Masayoshi Yoshikawa et al., Multimodal Speech Recognition Based on Lightweight Visual Features, IEICE, Aug. 2012, vol. J95-D, No. 3, p. 618-627. [cited by applicant]
English language Abstract for CN106782545 published May 31, 2017. [cited by applicant]
English language Abstract for CN107945789 published Apr. 20, 2018. [cited by applicant]
English language Abstract for CN109147763 published Jan. 4, 2019. [cited by applicant]
English language Abstract for CN109697976 published Apr. 30, 2019. [cited by applicant]
English language Abstract for CN110111783 published Aug. 9, 2019. [cited by applicant]
English language Abstract for CN110136698 published Aug. 16, 2019. [cited by applicant]
English language Abstract for JP2012242609 published Dec. 10, 2012. [cited by applicant]
English language Abstract for TW201140558 published Nov. 16, 2011. [cited by applicant]
English language Abstract for WO2014199596 published Feb. 3, 2017. [cited by applicant]
Campr et al., “Online Speaker Adaptation of an Acoustic Model using Face Recognition”, International Conference on Text, Speech and Dialogue, Sep. 2013. [cited by applicant]