IP Library Granted Patent US 12,243,515
Granted Patent B2
US 12,243,515 · App. 18/177,717 · Granted Mar 4, 2025

Speech recognition using neural networks

Inventors: Andrew W. Senior (New York, NY); Ignacio L. Moreno (New York, NY)
Assignee: Google LLC
G10L15/16G06N3/02G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,515
App. No.
18/177,717
Granted
Mar 4, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech recognition using neural networks. A feature vector that models audio characteristics of a portion of an utterance is received. Data indicative of latent variables of multivariate factor analysis is received. The feature vector and the data indicative of the latent variables is provided as input to a neural network. A candidate transcription for the utterance is determined based on at least an output of the neural network.

Claims (36)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

while a user is speaking a current utterance:

obtaining feature vectors indicative of audio characteristics of corresponding portions of the current utterance that have been spoken by the user;

for each predetermined duration of new audio received that characterizes a corresponding portion of the current utterance, calculating a corresponding i-vector;

providing, as input to a neural network acoustic model, the feature vectors and the corresponding i-vector calculated for each predetermined duration of new audio received; and

based on the feature vectors and the corresponding i-vector calculated for each predetermined duration of new audio received and provided as input to the neural network acoustic model, determining, as output from an output layer of the neural network acoustic model, a posterior probability distribution of possible speech units representing each feature vector.

2. The computer-implemented method of claim 1 , wherein the corresponding i-vector is calculated using a Gaussian mixture model (GMM).

3. The computer-implemented method of claim 2 , wherein the GMM comprises 39-dimensional Gaussians.

4. The computer-implemented method of claim 1 , wherein the speech units in the posterior probability distribution of possible speech units comprise components of phones.

5. The computer-implemented method of claim 1 , wherein the neural network acoustic model comprises an input layer, multiple hidden layers, and the output layer.

6. The computer-implemented method of claim 1 , wherein the operations further comprise determining a candidate transcription for the current utterance based on the posterior probability distribution of possible speech units representing each feature vector.

7. The computer-implemented method of claim 1 , wherein the operations further comprise:

providing, as input to a hidden markov model (HMM), the posterior probability distribution of possible speech units output from the output layer of the neural network acoustic model; and

determining, as output from the HMM, a word lattice associated with a transcription of the utterance.

8. The computer-implemented method of claim 7 , wherein the HMM is approximated by a set of weighted finite state transducers that represent a language model.

9. The computer-implemented method of claim 1 , wherein the corresponding i-vector calculated for each predetermined duration of new audio received indicates time-independent audio characteristics.

10. The computer-implemented method of claim 1 , wherein the operations further comprise projecting the corresponding i-vector calculated for each predetermined duration of new audio received using linear discriminant analysis (LDA).

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

while a user is speaking a current utterance:

obtaining feature vectors indicative of audio characteristics of corresponding portions of the current utterance that have been spoken by the user;

for each predetermined duration of new audio received that characterizes a corresponding portion of the current utterance, calculating a corresponding i-vector;

providing, as input to a neural network acoustic model, the feature vectors and the corresponding i-vector calculated for each predetermined duration of new audio received; and

based on the feature vectors and the corresponding i-vector calculated for each predetermined duration of new audio received and provided as input to the neural network acoustic model, determining, as output from an output layer of the neural network acoustic model, a posterior probability distribution of possible speech units representing each feature vector.

12. The system of claim 11 , wherein the corresponding i-vector is calculated using a Gaussian mixture model (GMM).

13. The system of claim 11 , wherein the GMM comprises 39-dimensional Gaussians.

14. The system of claim 11 , wherein the speech units in the posterior probability distribution of possible speech units comprise components of phones.

15. The system of claim 11 , wherein the neural network acoustic model comprises an input layer, multiple hidden layers, and the output layer.

16. The system of claim 11 , wherein the operations further comprise determining a candidate transcription for the current utterance based on the posterior probability distribution of possible speech units representing each feature vector.

17. The system of claim 11 , wherein the operations further comprise:

providing, as input to a hidden markov model (HMM), the posterior probability distribution of possible speech units output from the output layer of the neural network acoustic model; and

determining, as output from the HMM, a word lattice associated with a transcription of the utterance.

18. The system of claim 17 , wherein the HMM is approximated by a set of weighted finite state transducers that represent a language model.

19. The system of claim 11 , wherein the corresponding i-vector calculated for each predetermined duration of new audio received indicates time-independent audio characteristics.

20. The system of claim 11 , wherein the operations further comprise projecting the corresponding i-vector calculated for each predetermined duration of new audio received using linear discriminant analysis (LDA).

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2023
From: SENIOR, ANDREW W.; MORENO, IGNACIO L.
To: GOOGLE INC.
Reel/Frame 062863/0569 →
CHANGE OF NAME Recorded Mar 2, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 062941/0421 →
Continuity (4)
Continuation 17154376 · Jan 21, 2021
Continuation 16573232 · Sep 17, 2019
Continuation 13955483 · Jul 31, 2013
Related Publication 20230206909A1 · Jun 29, 2023
References Cited (89)
US 5237515A · Herron et al. · 1993 [cited by applicant]
US 5621857A · Cole et al. · 1997 [cited by applicant]
US 5758022A · Trompf et al. · 1998 [cited by applicant]
US 5774831A · Gupta · 1998 [cited by applicant]
US 5903863A · Wang · 1999 [cited by applicant]
US 5946656A · Rahim et al. · 1999 [cited by applicant]
US 6219642B1 · Asghar et al. · 2001 [cited by applicant]
US 6542866B1 · Jiang et al. · 2003 [cited by applicant]
US 6675145B1 · Yehia et al. · 2004 [cited by applicant]
US 7610199B2 · Abrash et al. · 2009 [cited by applicant]
US 7617101B2 · Chang et al. · 2009 [cited by applicant]
US 7720683B1 · Vermeulen et al. · 2010 [cited by applicant]
US 8386251B2 · Strom et al. · 2013 [cited by applicant]
US 8484022B1 · Vanhoucke · 2013 [cited by applicant]
US 8554562B2 · Aronowitz · 2013 [cited by applicant]
US 8566093B2 · Vair et al. · 2013 [cited by applicant]
US 8762142B2 · Jeong et al. · 2014 [cited by applicant]
US 9240184B1 · Lin et al. · 2016 [cited by applicant]
US 9466292B1 · Lei et al. · 2016 [cited by applicant]
US 20020091518A1 · Baruch et al. · 2002 [cited by applicant]
US 20020156626A1 · Hutchison · 2002 [cited by applicant]
US 20030204394A1 · Garudadri et al. · 2003 [cited by applicant]
US 20040024298A1 · Marshik-Geurts et al. · 2004 [cited by applicant]
US 20040102961A1 · Jensen et al. · 2004 [cited by applicant]
US 20040167778A1 · Valsan · 2004 [cited by examiner]
US 20070271086A1 · Peters et al. · 2007 [cited by applicant]
US 20080050357A1 · Gustafsson et al. · 2008 [cited by applicant]
US 20080114595A1 · Vair · 2008 [cited by examiner]
US 20080208577A1 · Jeong et al. · 2008 [cited by applicant]
US 20090073023A1 · Ammar · 2009 [cited by examiner]
US 20100155243A1 · Schneider et al. · 2010 [cited by applicant]
US 20120119080A1 · Hazebroek et al. · 2012 [cited by applicant]
US 20120259632A1 · Willett · 2012 [cited by applicant]
US 20130225128A1 · Gomar · 2013 [cited by applicant]
US 20130240628A1 · van der Merwe · 2013 [cited by examiner]
US 20140122087A1 · Macho · 2014 [cited by applicant]
US 20140214420A1 · Yao et al. · 2014 [cited by applicant]
US 20140358541A1 · Colibro et al. · 2014 [cited by applicant]
US 20150340039A1 · Gomar et al. · 2015 [cited by applicant]
EP 0574951 · 1993 [cited by applicant]
EP 0574951A2 · 1993 [cited by applicant]
EP 2189976 · 2010 [cited by applicant]
EP 2189976A1 · 2010 [cited by applicant]
WO 2007131530 · 2007 [cited by applicant]
WO 2007131530A1 · 2007 [cited by applicant]
WO 2014029099 · 2014 [cited by applicant]
WO 2014029099A1 · 2014 [cited by applicant]
Hinton, Geoffrey, et al. “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups.” IEEE Signal processing magazine 29.6 (2012): 82-97. (Year: 2012). [cited by examiner]
Thomas, Samuel, et al. “Adaptation transforms of auto-associative neural networks as features for speaker verification.” Odyssey 2012-the Speaker and Language Recognition Workshop. 2012. (Year: 2012). [cited by examiner]
Garimella, Sri, and Hynek Hermansky. “Factor analysis of mixture of auto-associative neural networks for speaker verification.” Odyssey 2012-The Speaker and Language Recognition Workshop. 2012. (Year: 2012). [cited by examiner]
Mohamed Ibnkahla, “Applications of neural networks to digital communications—a survey”, Signal Processing, vol. 80, Issue 7, Jul. 2000, pp. 1185-1215. Retrieved from <https://www.sciencedirect.com/science/article/pii/S0… [cited by examiner]
Jia Min Karen Kua, Julien Epps, Eliathamby Ambikairajah, “i-Vector with sparse representation classification for speaker verification ”, Speech Communication, vol. 55, Issue 5, Jun. 2013. Retrieved from <https://www.sci… [cited by examiner]
Man-Wao. Mak and W. Rao, “Likelihood-ratio empirical kernels for i-vector based PLDA-SVM scoring,” 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada, 2013, pp. 7702-770… [cited by examiner]
Mohamed, Abdel-rahman, George E. Dahl, and Geoffrey Hinton. “Acoustic modeling using deep belief networks.” IEEE transactions on audio, speech, and language processing 20.1 (2011): 14-22. (Year: 2011). [cited by examiner]
Cumani, Sandro, Oldich Plchot, and Martin Karafiát. “Independent component analysis and MLLR transforms for speaker identification.” 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)… [cited by examiner]
P. Matejka et al., “Full-covariance UBM and heavy-tailed PLDA in i-vector speaker verification,” 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Prague, 2011, pp. 4828-4831, doi: … [cited by applicant]
M. Karafiat, L. Burget, P. Matejka, 0. Glembek and J. Cernocky, “iVector-based discriminative adaptation for automatic speech recognition,” 2011 IEEE Workshop on Automatic Speech Recognition & Understanding, Waikoloa, H… [cited by applicant]
Stafylakis, Themas, et al. “Preliminary investigation of boltzmann machine classifiers for speaker recognition.” Odyssey 2012—The Speaker and Language Recognition Workshop. 2012. (Year: 2012). [cited by applicant]
Zhang, Wen-Lin, et al. “Rapid speaker adaptation using compressive sensing.” Speech Communication 55.10 (2013): 950-963. (Year: 2013). [cited by applicant]
Abad, Alberto, “The L2F Language Recognition System for NIST LRE 2011”, The 2011 NIST Language Recognition evaluation (LREI 1) Workshop, Atlanta, US, Dec. 2011, 7 pages. [cited by applicant]
Bahar, Mohammad Hasan et al., “Accent Recognition Using I-Vector, Gaussian Mean Supervector and Gaussian Posterior Probability Supervector for Spontaneous Telephone Speech,” 2013 IEEE International Conference on Acousti… [cited by applicant]
Bahari, Mohammad Hasan et al., “Age Estimation from Telephone Speech using i-vectors”, Interspeech 2012, 13th Annual Conference of the International Speech Communication Association, 4 pages. [cited by applicant]
Coccaro and Jurafsky, “Towards Better Integration of Semantic Predictors in Statistical Language Modeling,” Proceedings for ICSLP-98, vol. 6, DD. 2403-2406. [cited by applicant]
Coccaro, “Latent Semantic Analysis as a Tool to Improve Automatic Speech Recognition Performance,” Doctoral Dissertation, University of Colorado at Boulder, 2005, 102 pages. [cited by applicant]
D'Haro, Luis Ferdinand et al., “Phonotactic Language Recognition using i-vectors and Phoneme Posteriogram Counts”, Interspeech 2012, 13th Annual Conference of the International Speech Communication Association , 4 pages. [cited by applicant]
Dehak, Najim et al., “Front-End Factor Analysis for Speaker Verification”, IEEE Transactions on Audio, Speech and Language Processing, vol. 19, issue 4, Mav 2011, 12 pages. [cited by applicant]
Dehak, Najim et al., “Language Recognition via Ivectors and Dimensionality Reduction”, Interspeech 2011, 12th Annual Conference of the International Speech Communication Association, 4 pages. [cited by applicant]
Factor Analysis from Wikipedia, the free encyclopedia, downloaded from the internet on Jul. 15, 2013, <http://en.wikipedia.org/w/index.php?title=Factor_analvsis&oldid=559725962>, 14 pages. [cited by applicant]
Garcia-Romero et al, “Analysis of I-vector Length Normalization in Speaker Recognition Systems,” Twelfth Annual Conference of the International Speech Communication Association, 2011. Web availability: <http://www.ismmd… [cited by applicant]
Generative model from Wikipedia, the free encyclopedia, Apr. 30, 2015, downloaded from the Internet on Aug. 4, 2015, <https://en.wikipedia.org/wiki/Generative_model> , 2 pages. [cited by applicant]
Hinton, Geoffrey et al., “Deep Neural Networks for Acoustic Modeling in Speech Recognition”, IEEE Signal Processing Magazine, Nov. 2012, 16 pages. [cited by applicant]
International Preliminary Report on Patentability in International Application No. PCT/US2014/044528, mailed Feb. 11, 2016, 7 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2014/044528, mailed Oct. 2, 2014, 10 pages. [cited by applicant]
I-Vectors from ALIZE wiki, downloaded from the internet on Jul. 15, 2013, <http://mistral.univ-/avignon.fr/mediawiki/index.php/1-Vectors>, 2 pages. [cited by applicant]
Larcher, Anthony et al., “I-Vectors in the Context of Phonetically-Constrained Short Utterances for Speaker Verification”, 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4 pages. [cited by applicant]
Latent Variable from Wikipedia, the free encyclopedia, downloaded from the internet on Jul. 15, 2013, <http://en.wikipedia.org/w/index.php?title=Latent_variable&oldid=555584475> , 3 pages. [cited by applicant]
Martinez, David et al., “ IVector-Based Prosodic System for Language Identification”, 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4 pages. [cited by applicant]
Martinez, David et al., “Language Recognition in iVectors Space”, Interspeech 2011, 12th Annual Conference of the International Speech Communication Association , 4 pages. [cited by applicant]
Mikolov and Zweig, “Context Dependent Recurrent Neural Network Language Model,” Spoken Language Technologies, Jul. 2012, pp. 234-239. [cited by applicant]
Mohamed et al., “Deep belief networks using discriminative features for phone recognition,” in Proc. ICASSP, May 2011, pp. 5060-5063. [cited by applicant]
Restricted Boltzmann machine from Wikipedia, the free encyclopedia, Mar. 5, 2015, downloaded from the Internet on Aug. 4, 2015, https://en.wikipedia.org/wiki/Restricted_Boltzmann_ machine, 5 pages. [cited by applicant]
Saul et al, “Maximum Likelihood and Minimum Classification Error Factor Analysis for Automatic Speech Recognition,” IEEE Transactions on Speech and Audio Processing, vol. 8, No. 2, Mar. 2, 2000. (Year: 2000). [cited by applicant]
Singer, Elliot et al., “The MITLL NIST LRE 2011 Language Recognition System”, Odyssey 2012, The Speaker and Language Recognition Workshop, Jun. 25-28, 2012, Singapore, 7 pages. [cited by applicant]
Speech Recognition from Wikipedia, the free encyclopedia, downloaded from the internet on Jul. 15, 2013, http://en.wikipedia.org/w/index.php?title=Speech_recognition&oldid=55508I4I5, <http://en.wikipedia.org/w/index.php… [cited by applicant]
Zhang and Rudnicky, “Improve Latent Semantic Analysis Based Language Model by Integrating Multiple Level Knowledge,” Conference Proceeding, Carnegie Mellon University, 2002, 5 pages. [cited by applicant]
M. Karafiat, L. Burget, P. Matejka, 0. Glembek and J. Cernocky, “iVector-based discriminative adaptation for automatic speech recognition,” 2011 IEEE Workshop on Automatic Speech Recognition & Understanding, Waikoloa, H… [cited by applicant]
Factor Analysis from Wikipedia, the free encyclopedia, downloaded from the internet on Jul. 15, 2013, <http://en.wikipedia.org/w/index.php?title=Factor_analysis&oldid=559725962>, 14 pages. [cited by applicant]
Garcia-Romero et al, “Analysis of I-vector Length Normalization in Speaker Recognition Systems,” Twelfth Annual Conference of the International Speech Communication Association, 2011. Web availability: <http://www.ismmd… [cited by applicant]
Speech Recognition from Wikipedia, the free encyclopedia, downloaded from the internet on Jul. 15, 2013, http://en.wikipedia.org/w/index.php?title=Speech_recognition&oldid-555081415, <http://en.wikipedia.org/w/index.php… [cited by applicant]