IP Library Granted Patent US 12,340,805
Granted Patent B2
US 12,340,805 · App. 18/659,198 · Granted Jun 24, 2025

Providing pre-computed hotword models

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/063G10L15/08G10L15/26G10L15/30G06F3/04842G10L2015/0631G10L2015/0638G10L2015/088G10L15/18G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,805
App. No.
18/659,198
Granted
Jun 24, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, for each of multiple words or sub-words, audio data corresponding to multiple users speaking the word or sub-word; training, for each of the multiple words or sub-words, a pre-computed hotword model for the word or sub-word based on the audio data for the word or sub-word; receiving a candidate hotword from a computing device; identifying one or more pre-computed hotword models that correspond to the candidate hotword; and providing the identified, pre-computed hotword models to the computing device.

Claims (32)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving a candidate hotword from a user device, the candidate hotword input by a user of the user device via a graphical user interface executing on the user device;

providing, for display in the graphical user interface executing on the user device, a graphical button, that when selected by the user, indicates that the user accepts the candidate hotword;

receiving a user input indication indicating the user has selected the graphical button displayed in the graphical user interface;

based on receiving the user input indication, dynamically creating a hotword model that corresponds to the candidate hotword; and

transmitting the dynamically created hotword model to the user device.

2. The computer-implemented method of claim 1 , wherein the graphical user interface executing on the user device is configured to display a prompt requesting the user to provide the candidate hotword.

3. The computer-implemented method of claim 1 , wherein the graphical user interface executing on the user device is configured to display text corresponding to the candidate hotword.

4. The computer-implemented method of claim 1 , wherein the candidate hotword comprises three or more words.

5. The computer-implemented method of claim 1 , wherein the candidate hotword comprises one word.

6. The computer-implemented method of claim 1 , wherein the candidate hotword comprises two words.

7. The computer-implemented method of claim 1 , wherein the candidate hotword is input by the user as a text-based input via the graphical user interface.

8. The computer-implemented method of claim 1 , wherein the user inputs the candidate hotword by speaking the candidate hotword.

9. The computer-implemented method of claim 1 , wherein the hotword model comprises a neural network.

10. The computer-implemented method of claim 1 , wherein the dynamically created hotword model is configured to detect the candidate hotword in spoken utterances.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a candidate hotword from a user device, the candidate hotword input by a user of the user device via a graphical user interface executing on the user device;

providing, for display in the graphical user interface executing on the user device, a graphical button, that when selected by the user, indicates that the user accepts the candidate hotword;

receiving a user input indication indicating the user has selected the graphical button displayed in the graphical user interface;

based on receiving the user input indication, dynamically creating a hotword model that corresponds to the candidate hotword; and

transmitting the dynamically created hotword model to the user device.

12. The system of claim 11 , wherein the graphical user interface executing on the user device is configured to display a prompt requesting the user to provide the candidate hotword.

13. The system of claim 11 , wherein the graphical user interface executing on the user device is configured to display text corresponding to the candidate hotword.

14. The system of claim 11 , wherein the candidate hotword comprises three or more words.

15. The system of claim 11 , wherein the candidate hotword comprises one word.

16. The system of claim 11 , wherein the candidate hotword comprises two words.

17. The system of claim 11 , wherein the candidate hotword is input by the user as a text-based input via the graphical user interface.

18. The system of claim 11 , wherein the user inputs the candidate hotword by speaking the candidate hotword.

19. The system of claim 11 , wherein the hotword model comprises a neural network.

20. The system of claim 11 , wherein the dynamically created hotword model is configured to detect the candidate hotword in spoken utterances.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: SHARIFI, MATTHEW
To: GOOGLE, INC.
Reel/Frame 067359/0072 →
CHANGE OF NAME Recorded May 9, 2024
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 067359/0422 →
Continuity (12)
Continuation 18313756 · May 8, 2023
Continuation 17304459 · Jun 21, 2021
Continuation 16806332 · Mar 2, 2020
Continuation 16669503 · Oct 30, 2019
Continuation 16529300 · Aug 1, 2019
Continuation 16216752 · Dec 11, 2018
Continuation 15875996 · Jan 19, 2018
Continuation 15463786 · Mar 20, 2017
Continuation 15288241 · Oct 7, 2016
Continuation 15001894 · Jan 20, 2016
Continuation 14340833 · Jul 25, 2014
Related Publication 20240290333A1 · Aug 29, 2024
References Cited (61)
US 6026491A · Hiles · 2000 [cited by examiner]
US 7668718B2 · Kahn · 2010 [cited by examiner]
US 7712031B2 · Law et al. · 2010 [cited by applicant]
US 8548812B2 · Whynot · 2013 [cited by applicant]
US 8719039B1 · Sharifi · 2014 [cited by applicant]
US 8768712B1 · Sharifi · 2014 [cited by applicant]
US 8843376B2 · Cross, Jr. · 2014 [cited by examiner]
US 8924219B1 · Bringert et al. · 2014 [cited by applicant]
US 9031847B2 · Sarin et al. · 2015 [cited by applicant]
US 9263042B1 · Sharifi · 2016 [cited by applicant]
US 9536528B2 · Rubin et al. · 2017 [cited by applicant]
US 20080059188A1 · Konopka et al. · 2008 [cited by applicant]
US 20080082338A1 · O'Neil · 2008 [cited by examiner]
US 20130289994A1 · Newman et al. · 2013 [cited by applicant]
US 20140012586A1 · Rubin et al. · 2014 [cited by applicant]
US 20150134332A1 · Liu et al. · 2015 [cited by applicant]
US 20170193995A1 · Sharifi · 2017 [cited by applicant]
CN 103680498A · 2014 [cited by applicant]
EP 1884923A1 · 2008 [cited by applicant]
Ahmed et al, “Training Hierarchical Feed-Forward Visual Recognition Models Using Transfer Learning from Pseudo-Tasks . . . ” ECCV 2008, Part III, LNCS 5304 . . . 69-82, 2008. [cited by applicant]
Bengio, “Deep Learning of Representations for Unsupervised and Transfer Learning,” JMRL: Workshop and Conference Proceedings 27:17-37, 2012. [cited by applicant]
Caruana, “Multitask Learning,” CIVIU-CS-97-203 paper , School of Computer Science, Carnegie Mellon University, Sep. 23, 1997, 255 pages. [cited by applicant]
Caruana, “Multitask Learning,” Machine Learning, 28, 41-75 (1997). [cited by applicant]
Ciresan et al., “Transfer learning for Latin and Chinese characters with Deep Neural Networks,” The 2012 International Joint Conference on Neural Networks (IJCNN), Jun. 10-15, 2012, 1-6 (abstract only), 2 pages. [cited by applicant]
Collobert et al., “A Unified Architecture for Natural Language Processing: Deep Neural Networks with Multitask Leaming,” Proceedings of the 25th International Conference on Machine Leaming, Helsinki, Finland, 2008, 8 pa… [cited by applicant]
Dahl et al., “Large Vocabulary Continuous Speech Recognition with Context-Dependent DBN-HMMS,” IEEE, May 2011, 4 pages. [cited by applicant]
Dean et al, “Large Scale Distributed Deep Networks,” Advances in Neural Information Processing Systems 25, 2012, 1-11. [cited by applicant]
Fernandez et al. “An application of recurent neural network to discriminative keyword spotting,” ICANN'07 Proceedings of the 17th international conference on Artificial neural networks, 2007, 220-229. [cited by applicant]
Grangier et al, “Discriminative Keyword Spotting” Speech and Speaker Recognition: Large Margin and Kernel Methods, 2001, 1-23. [cited by applicant]
Heigold et al., “Multilingual Acoustic Models Using Distributed Deep Neural Networks,” 2013 IEEE International Conference on Acoutics, Speech and Signal Processing (ICASSP), May 26-31, 2013, 8619-8623. [cited by applicant]
Huang et al., “Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers”, in Proc. ICASSP, 2013, 7304-7308. [cited by applicant]
Hughes et al., “Recurrent Neural Networks for Voice Activity Detection”, ICASSP 2013, IEEE 2013, 7378-7382. [cited by applicant]
Jaitly et al., “Application of Pretrained Deep Neural Networks to Large Vocabulary Conversational Speech Recognition,” Department of Computer Science, University of Toronto, UTML TR 2012-001, Mar. 12, 2012, 11 pages. [cited by applicant]
Le et al., “Building High-level Features Using Large Scale Unsupervised Learning,” Proceedings of the 29th International Conference on Machine Learning, Jul. 12, 2012, 11 page. [cited by applicant]
Lei et al., “Accurate and Compact Large Vocabulary Speech Recognition on Mobile Devices,” Interspeech 2013, Aug. 25-29, 2013, 662-665. [cited by applicant]
Li et al., “A Whole World Recurent Neural Network for Keyword Spotting,” IEEE 1992, 81-84. [cited by applicant]
Mamou et al., “Vocabulary Independent Spoken Term Detection,” SIGIR'07, Jul. 23-27, 2007, 8 page. [cited by applicant]
Miller et al., “Rapid and Accurate Spoken Term Detection .” Interspeech 2007, Aug. 27-31, 2007, 314-317. [cited by applicant]
Parlak et al., “Spoken Term Detection for Turkish Broadcast News,” ICASSP, IEEE, 2008, 5244-5247. [cited by applicant]
Rohlicek et al., “Continuous Hidden Markov Modeling for Speaker-Independent Word Spotting.” IEEE 1989, 627-630. [cited by applicant]
Rose et al., “A Hidden Markov Model Based Keyword Recognition System,” IEEE 1990, 129-132. [cited by applicant]
Schalkwyk et al., “Google Search by Voice: A case study,” 1-35, 2010. [cited by applicant]
Science Net [online]. “Deep Learning Workshop ICML 2013.” Oct. 29, 2013 [retrieved on Jan. 24, 2014]. Retrieved from the internet URL<http://blog.sciencenet.cn/blog-701243-737140.html.>, 13 pages. [cited by applicant]
Silaghi et al., “Iterative Posterior-Based Keyword Spotting Without Filler Models,” IEEE 1999, 4 pages. [cited by applicant]
Silaghi, “Spotting Subsequences matching a HMM using the Average Observation Probability Criteria With application to Keyword Spotting,” American Association for Artificial Intelligence, 1118-1123, 2005. [cited by applicant]
Srivastava et al., “Dicriminative Tranfer Learning with Tree-based Priors,” Advances in Neural Information Processing Systems 26 (NIPS 2013) . . . 12 pages. [cited by applicant]
Sutton et al., “Composition of Conditional Random Fields for Transfer Learning,” In Proceedings of HLT/EMNLP, 2005, 7 pages. [cited by applicant]
Swietojanski et al., “Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR,” In Proc. IEEE Workhop on Spoken Language Technology, Miami, Florida, USA, Dec. 2012, 6 pages. [cited by applicant]
Tabibian et al., An Evolutionary based discriminative system for keyword spotting. IEEE, 83-88, 2011. [cited by applicant]
Weintraub, “Keyword-Spotting Using SRI's Decipher™ Large-Vocabuarly Speech-Recognition System.”IEEE 1993. 463-466. [cited by applicant]
Wilpon et al., “Improvements and Applications for Key Word Recognition Using Hidden Markov Modeling Techniques,” IEEE 1991, 309-312. [cited by applicant]
Yu et al., “Deep Learning with Kernel Regularization for Visual Recognition,” Advances in Neural Information Processing System 21, Annual Conference on Neural Information Processing Systems, Dec. 8-11, 2008 I-8. [cited by applicant]
Zeiler et al., “On Rectified Linear Units for Speech Processing,” ICASSP, p. 3517-3521, 2013. [cited by applicant]
Zhang et al., “Transfer Learning for Voice Activity Detection: A Denoising Deep Neural Network Perspective,” INTERSPEECH2013, Mar. 8, 2013, 5 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2015/030501, mailed Jul. 30, 2015, 9 pages. [cited by applicant]
European Search Report in European Application No. 16181747.3, dated Sep. 30, 2016, 7 page. [cited by applicant]
Office Action issued in European Application No. I 5725946.6, mailed on Nov. 03., 2017, 4 pages. [cited by applicant]
Office Action issued European Application No. 16181747.3, mailed on January 3., 2017, 5 pages. [cited by applicant]
European Search Report for the related EP application No. 19170707.4 dated Jul. 4, 2019. [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 17/304,459, dated Nov. 4, 2022. [cited by applicant]
China National Intellectual Property Administration—The First Office Action for the related Application No. 201910962105.8, dated Apr. 6, 2023, 23 pages. [cited by applicant]