IP Library › Granted Patent US 12,573,400
Granted Patent B2
US 12,573,400 · App. 18/617,476 · Granted Mar 10, 2026

Hotword suppression

Inventors: Alexander H. Gruenstein (Mountain View, CA); Taral Pradeep Joglekar (Sunnyvale, CA); Vijayaditya Peddinti (San Jose, CA); Michiel A.U. Bacchiani (Summit, NJ)
Assignee: Google LLC
G10L15/22G10L15/063G10L15/08G10L15/30G10L17/00G10L17/22G10L25/51G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,400
App. No.
18/617,476
Granted
Mar 10, 2026
Kind
B2
Abstract

A method includes adding, by a first computing device, a first audio watermark to first speech data corresponding to playback of a first utterance including a hotword used to invoke an attention of a second computing device. The method includes outputting, by the first computing device, the playback of the first utterance corresponding to the watermarked first speech data. The second computing device is configured to receive the watermarked first speech data and determine to cease processing of the watermarked first speech data.

Claims (35)

1 . A computer-implemented method of processing audio data, the method comprising:

receiving, at a computing device, audio data captured by a microphone of the computing device, the audio data comprising an audio watermark that is imperceptible to a human listener;

determining, by an audio watermark identification model executing on the computing device, the audio data includes a predefined audio watermark pattern;

based on determining the audio data includes the predefined audio watermark pattern, determining, by the audio watermark identification model executing on the computing device, the audio data includes the audio watermark; and

based on determining the audio data includes the audio watermark, determining, by the computing device, that processing of the audio data is not to be continued.

2 . The method of claim 1 , wherein the audio watermark identification model comprises a neural network.

3 . The method of claim 2 , wherein the audio watermark pattern includes a replication of a watermark sign matrix, and wherein the neural network is trained to estimate a cross-correlation of the watermark sign matrix.

4 . The method of claim 2 , wherein the neural network is trained using a multi-task loss function.

5 . The method of claim 2 , wherein an output of the neural network is processed by using a filter.

6 . The method of claim 1 , wherein the predefined watermark pattern includes a frequency pattern.

7 . An apparatus, comprising:

processing circuitry configured to

receive audio data captured by a microphone of the computing device, the audio data comprising an audio watermark that is imperceptible to a human listener;

determine, by an audio watermark identification model executing on the processing circuitry, the audio data includes a predefined audio watermark pattern;

based on determining the audio data includes the predefined audio watermark pattern, determine, by the audio watermark identification model, the audio data includes the audio watermark, and

based on determining the audio data includes the audio watermark, determine that processing of the audio data is not to be continued.

8 . The apparatus of claim 7 , wherein the audio watermark identification model comprises a neural network.

9 . The apparatus of claim 8 , wherein the audio watermark pattern includes a replication of a watermark sign matrix, and wherein the neural network is trained to estimate a cross-correlation of the watermark sign matrix.

10 . The apparatus of claim 8 , wherein the neural network is trained using a multi-task loss function.

11 . The apparatus of claim 8 , wherein an output of the neural network is processed by using a filter.

12 . The apparatus of claim 7 , wherein the predefined watermark pattern includes a frequency pattern.

13 . The apparatus of claim 7 , wherein the processing circuitry is configured to:

determine that the audio data includes the audio watermark based on the audio data including the predefined audio watermark pattern; and

determine that the processing of the audio data is to be continued based on the audio data including the audio watermark.

14 . The apparatus of claim 7 , wherein the processing circuitry is configured to:

determine that the audio data does not include the audio watermark based on the audio data not including the predefined audio watermark pattern; and

determine that the processing of the audio data is not to be continued based on the audio data not including the audio watermark.

15 . A non-transitory computer-readable storage medium including computer executable instructions, wherein the instructions, when executed by a computer, cause the computer to perform a method, the method comprising:

receiving, at a computing device, audio data captured by a microphone of the computing device, the audio data comprising an audio watermark that is imperceptible to a human listener;

determining, by an audio watermark identification model executing on the computing device, the audio data includes a predefined audio watermark pattern;

based on determining the audio data includes the predefined audio watermark pattern, determining, by the audio watermark identification model executing on the computing device, the audio data includes the audio watermark; and

based on determining the audio data includes the audio watermark, determining, by the computing device, that processing of the audio data is not to be continued.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the audio watermark identification model comprises a neural network.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the audio watermark pattern includes a replication of a watermark sign matrix, and wherein the neural network is trained to estimate a cross-correlation of the watermark sign matrix.

18 . The non-transitory computer-readable storage medium of claim 16 , wherein the neural network is trained using a multi-task loss function.

Continuity (5)
Continuation 17849253 · Jun 24, 2022
Continuation 16874646 · May 14, 2020
Continuation 16418415 · May 21, 2019
Provisional Application 62674973 · May 22, 2018
Related Publication 20240242719A1 · Jul 18, 2024
References Cited (164)
US 4363102A · Holmgren et al. · 1982 [cited by applicant]
US 5659665A · Whelpley, Jr. · 1997 [cited by applicant]
US 5897616A · Kanevsky et al. · 1999 [cited by applicant]
US 5983186A · Miyazawa et al. · 1999 [cited by applicant]
US 6023676A · Erell · 2000 [cited by applicant]
US 6141644A · Kuhn · 2000 [cited by applicant]
US 6567775B1 · Maali et al. · 2003 [cited by applicant]
US 6671672B1 · Heck · 2003 [cited by applicant]
US 6700990B1 · Rhoads · 2004 [cited by examiner]
US 6744860B1 · Schrage · 2004 [cited by applicant]
US 6826159B1 · Shaffer et al. · 2004 [cited by applicant]
US 6931375B1 · Bossemeyer, Jr. et al. · 2005 [cited by applicant]
US 6973426B1 · Schier et al. · 2005 [cited by applicant]
US 7016833B2 · Gable et al. · 2006 [cited by applicant]
US 7222072B2 · Chang · 2007 [cited by applicant]
US 7571014B1 · Lamboume et al. · 2009 [cited by applicant]
US 7720012B1 · Borah et al. · 2010 [cited by applicant]
US 7904297B2 · Mirkovic et al. · 2011 [cited by applicant]
US 8099288B2 · Zhang et al. · 2012 [cited by applicant]
US 8194624B2 · Park et al. · 2012 [cited by applicant]
US 8200488B2 · Kemp et al. · 2012 [cited by applicant]
US 8209174B2 · Al-Telmissani · 2012 [cited by applicant]
US 8214447B2 · Deslippe et al. · 2012 [cited by applicant]
US 8340975B1 · Rosenberger · 2012 [cited by applicant]
US 8588949B2 · Lamboume et al. · 2013 [cited by applicant]
US 8670985B2 · Lindahl et al. · 2014 [cited by applicant]
US 8709018B2 · Williams et al. · 2014 [cited by applicant]
US 8713119B2 · Lindahl · 2014 [cited by applicant]
US 8717949B2 · Crinon et al. · 2014 [cited by applicant]
US 8719009B2 · Baldwin et al. · 2014 [cited by applicant]
US 8719018B2 · Dinerstein · 2014 [cited by applicant]
US 8768687B1 · Quasthoff et al. · 2014 [cited by applicant]
US 8775191B1 · Sharifi et al. · 2014 [cited by applicant]
US 8805890B2 · Zhang et al. · 2014 [cited by applicant]
US 8838457B2 · Cerra et al. · 2014 [cited by applicant]
US 8938394B1 · Faaborg et al. · 2015 [cited by applicant]
US 8996372B1 · Secker-Walker et al. · 2015 [cited by applicant]
US 9142218B2 · Schroeter · 2015 [cited by applicant]
US 9354778B2 · Cornaby · 2016 [cited by examiner]
US 9462433B2 · Rodriguez · 2016 [cited by examiner]
US 9484046B2 · Knudson · 2016 [cited by examiner]
US 9548053B1 · Basye · 2017 [cited by examiner]
US 10074371B1 · Wang et al. · 2018 [cited by applicant]
US 10147433B1 · Bradley · 2018 [cited by examiner]
US 10395650B2 · Garcia · 2019 [cited by examiner]
US 10692496B2 · Gruenstein · 2020 [cited by examiner]
US 11170793B2 · Jin · 2021 [cited by examiner]
US 11373652B2 · Gruenstein · 2022 [cited by examiner]
US 11967323B2 · Gruenstein · 2024 [cited by examiner]
US 12094474B1 · Gowal · 2024 [cited by examiner]
US 20020049596A1 · Burchard et al. · 2002 [cited by applicant]
US 20020072905A1 · White et al. · 2002 [cited by applicant]
US 20020123890A1 · Kopp et al. · 2002 [cited by applicant]
US 20020193991A1 · Bennett et al. · 2002 [cited by applicant]
US 20030018479A1 · Oh et al. · 2003 [cited by applicant]
US 20030200090A1 · Kawazoe · 2003 [cited by applicant]
US 20030231746A1 · Hunter et al. · 2003 [cited by applicant]
US 20040101112A1 · Kuo · 2004 [cited by applicant]
US 20050165607A1 · Di Fabbrizio et al. · 2005 [cited by applicant]
US 20060074656A1 · Mathias et al. · 2006 [cited by applicant]
US 20060085188A1 · Goodwin et al. · 2006 [cited by applicant]
US 20060085199A1 · Jain · 2006 [cited by examiner]
US 20060184370A1 · Kwak et al. · 2006 [cited by applicant]
US 20070100620A1 · Tavares · 2007 [cited by applicant]
US 20070198262A1 · Mindlin et al. · 2007 [cited by applicant]
US 20080252595A1 · Boillot · 2008 [cited by applicant]
US 20090106796A1 · McCarthy et al. · 2009 [cited by applicant]
US 20090256972A1 · Ramaswamy et al. · 2009 [cited by applicant]
US 20090258333A1 · Yu · 2009 [cited by applicant]
US 20090292541A1 · Daya et al. · 2009 [cited by applicant]
US 20100057231A1 · Slater et al. · 2010 [cited by applicant]
US 20100070276A1 · Wasserblat et al. · 2010 [cited by applicant]
US 20100110834A1 · Kim et al. · 2010 [cited by applicant]
US 20110026722A1 · Jing et al. · 2011 [cited by applicant]
US 20110054892A1 · Jung et al. · 2011 [cited by applicant]
US 20110060587A1 · Phillips et al. · 2011 [cited by applicant]
US 20110066429A1 · Shperling et al. · 2011 [cited by applicant]
US 20110066437A1 · Luff · 2011 [cited by applicant]
US 20110161076A1 · Davis · 2011 [cited by examiner]
US 20110184730A1 · LeBeau et al. · 2011 [cited by applicant]
US 20110304648A1 · Kim et al. · 2011 [cited by applicant]
US 20120084087A1 · Yang et al. · 2012 [cited by applicant]
US 20120232896A1 · Taleb et al. · 2012 [cited by applicant]
US 20120265528A1 · Gruber et al. · 2012 [cited by applicant]
US 20130024882A1 · Lee et al. · 2013 [cited by applicant]
US 20130060571A1 · Soemo et al. · 2013 [cited by applicant]
US 20130124207A1 · Sarin et al. · 2013 [cited by applicant]
US 20130132086A1 · Xu et al. · 2013 [cited by applicant]
US 20130150177A1 · Rodriguez et al. · 2013 [cited by applicant]
US 20130183944A1 · Mozer et al. · 2013 [cited by applicant]
US 20140012573A1 · Hung et al. · 2014 [cited by applicant]
US 20140012578A1 · Morioka · 2014 [cited by applicant]
US 20140088961A1 · Woodward et al. · 2014 [cited by applicant]
US 20140142958A1 · Sharma et al. · 2014 [cited by applicant]
US 20140222430A1 · Rao · 2014 [cited by applicant]
US 20140257821A1 · Adams et al. · 2014 [cited by applicant]
US 20140278383A1 · Fan · 2014 [cited by applicant]
US 20140278435A1 · Ganong, III et al. · 2014 [cited by applicant]
US 20150154953A1 · Bapat et al. · 2015 [cited by applicant]
US 20150262577A1 · Nomura · 2015 [cited by applicant]
US 20150293743A1 · Yang · 2015 [cited by applicant]
US 20150294666A1 · Miyasaka · 2015 [cited by examiner]
US 20160049153A1 · Kakkirala et al. · 2016 [cited by applicant]
US 20160104483A1 · Foerster et al. · 2016 [cited by applicant]
US 20160104498A1 · Beack et al. · 2016 [cited by applicant]
US 20160260431A1 · Newendorp et al. · 2016 [cited by applicant]
US 20170084277A1 · Sharifi · 2017 [cited by applicant]
US 20170110130A1 · Sharifi et al. · 2017 [cited by applicant]
US 20170110144A1 · Sharifi et al. · 2017 [cited by applicant]
US 20170117108A1 · Richardson et al. · 2017 [cited by applicant]
US 20170251269A1 · Yoshizawa · 2017 [cited by applicant]
US 20180096690A1 · Mixter et al. · 2018 [cited by applicant]
US 20180130469A1 · Gruenstein · 2018 [cited by examiner]
US 20180182397A1 · Carbune et al. · 2018 [cited by applicant]
US 20180350356A1 · Garcia · 2018 [cited by examiner]
US 20190362719A1 · Gruenstein · 2019 [cited by examiner]
US 20200098379A1 · Tai · 2020 [cited by examiner]
US 20200279562A1 · Gruenstein · 2020 [cited by examiner]
US 20220319519A1 · Gruenstein · 2022 [cited by examiner]
US 20240242719A1 · Gruenstein · 2024 [cited by examiner]
US 20250095662A1 · Looney · 2025 [cited by examiner]
US 20250118319A1 · Joglekar · 2025 [cited by examiner]
US 20250149048A1 · Gowal · 2025 [cited by examiner]
JP S59180599 · 1984 [cited by applicant]
JP H1152976 · 1999 [cited by applicant]
JP H11231896 · 1999 [cited by applicant]
JP 2000310999 · 2000 [cited by applicant]
JP 2003263182A · 2003 [cited by applicant]
JP 2006227634A · 2006 [cited by applicant]
JP 2010164992A · 2010 [cited by applicant]
JP 2020526781A · 2020 [cited by applicant]
KR 1020140031391 · 2014 [cited by applicant]
WO 19980404875 · 1998 [cited by applicant]
WO 2014008194 · 2014 [cited by applicant]
WO WO2014112110A1 · 2014 [cited by applicant]
WO 2015025330 · 2015 [cited by applicant]
WO WO2020068401A1 · 2020 [cited by applicant]
Japanese Office Action issued Feb. 25, 2025, in corresponding Japanese Patent Application No. 2023-200953, (with English Translation), 8 pages. [cited by applicant]
Office Action issued Mar. 31, 2022 in Korean Patent Application No. 10-2022-7036730, with English translation. [cited by applicant]
Office Action issued Dec. 9, 2021 in Indian Patent Application No. 202047050167, with concise English translation. [cited by applicant]
Extended European Search Report issued Mar. 31, 2023 in European Patent Application No. 22195235.1, 8 pages. [cited by applicant]
Office Action issued Feb. 27, 2023, in corresponding Korean Patent Application No. 10-2023-7002831 (with English Translation), 6 pages. [cited by applicant]
Notice of Reasons for Rejection issued Apr. 24, 2023 in Japanese Patent Application No. 2020-565375 (with English language translation), 6 pages. [cited by applicant]
PCT International Search Report and Written Opinion in International Appin.No. PCT/US2019/033571 dated Aug. 1, 2019, 12 pages. [cited by applicant]
Calixto et al, “Effectiveness analysis of audio watermark tags for iptv second screen applications and synchronization,” IEEE Xplore, Aug. 2014, 5 pages. [cited by applicant]
Thiagarajan et al, “Analysis of the mpeg-1 layer iii (mp3) algorithm using matlab,” Synthesis Lectures on Algorithms and Software in Engineering, 2011, 131 pages. [cited by applicant]
Lin et al, “Audio Watermarking Techniques,” in Audio Watermark. Springer, 2015, 44 pages. [cited by applicant]
Allen et al, “Image method for efficiently simulating small-room acoustics,” The Journal of the Acoustical Society of America, Jun. 1978, 8 pages. [cited by applicant]
Auckenthaler et al. “Score Normalization for Text-independent Speaker Verification System,” Digital Signal Processing, vol. 10, 2000, 13 pages. [cited by applicant]
Garcia. Digital Watermarking of Audio Signals using a Psychoacoustic Auditory Model and Spread Spectrum Theory, 107th Convention, Audio Engineering Society, New York, NY, Sep. 24-27, 1999, 42 pages. [cited by applicant]
International Search Report and Written Opinion, issued in International Application No. PCT/US2018/022101, mailed on May 25, 2018, 13 pages. [cited by applicant]
Jae-Seung, Choi, “Text-dependent Speaker Recognition using Characteristic Vectors in Speech Signal and Normalized Recognition Method,” Journal of Korean Institute of Information Technology, 10(5), May 2012 (English Abst… [cited by applicant]
Kim et al, “Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home,” Interspeech, Aug. 2017, 5 pages. [cited by applicant]
Kirovski et al, “Robust spread-spectrum audio water-marking,” IEEE International Conference on, 2001, 3: 1345-1348, 4 pages. [cited by applicant]
Maheshwari et al [online], “Burger King ‘O.K.Google’ Ad doesn't seem O.K. with Google,” Apr. 2017, [retrieved May 12, 21, 2019], retrieved from: URL«https://www.nytimes.com/2017/04/12/business/burger-king-tv-ad-google-h… [cited by applicant]
nationalpublicmedia.com <http://nationalpublicmedia.com> [online], “The smart audio report,” NPR and E. Research, 2017, [retrieved on May 21, 2019], retrieved from: URL«https://www.nationalpublicmedia.com/smart-audio-re… [cited by applicant]
nielsen.com <http://nielsen.com> [online], “Super bowl LIi draws 103.4 million TV viewers, 170.7 million social media AAN interactions.” Feb. 2018, [retrieved on May 21, 2019], retrieved from: URL«https://www.nielsen.co… [cited by applicant]
P. Kabal, “An examination and interpretation of itu-r bs. 1387: Perceptual evaluation of audio quality,” TSP Lab Technical Report, Dept. Electrical & Computer Engineering, McGill University, 93 pages. [cited by applicant]
Pan et al, “A tutorial on mpeg/audio compression,” IEEE Multimedia, 1995, 15 pages. [cited by applicant]
Perry et al [online], “A sleeping Alexa can listen for more than just her name,” Feb. 2018, [retrieved May 21, 2019], retrieved from: URL«https://spectrum.ieee.org/view-from-the-valley/consumer-electronics/gadgets/beyon… [cited by applicant]
R. J. Anderson et al, “Information Hiding: An an-notated bibliography,” University of Cambridge, 1999, 65 pages. [cited by applicant]
Sainath et al, “Convolutional neural networks for small footprint keyword spotting,” in Sixteenth Annual Conference of the International Speech Communication Association, 2015, 5 pages. [cited by applicant]
twolame.org <http://twolame.org> [online], “Two LAME audio encoder,” available on or before Feb. 22, 2016, via internet archive: Wayback Machine URL< <https://web.archive.org/web/20060222172741/http://www.twolame.org/>,… [cited by applicant]
wikipedia.org <http://wikipedia.org> [online] “Super bowl,” May 20, 2019 [retrieved May 21, 2019], retrieved from: URL«https://en.wikipedia.org/wiki/Super_Bowl», 21 pages. [cited by applicant]