IP Library Granted Patent US 12,469,498
Granted Patent B2
US 12,469,498 · App. 18/594,833 · Granted Nov 11, 2025

Targeted voice separation by speaker conditioned on spectrogram masking

Inventors: Quan Wang (Hoboken, NJ); Prashant Sridhar (New York, NY); Ignacio Lopez Moreno (New York, NY); Hannah Muckenhim (Martigny, CH)
Assignee: GOOGLE LLC
G10L17/04G10L17/00G10L17/02G10L17/18G10L17/22G10L25/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,498
App. No.
18/594,833
Granted
Nov 11, 2025
Kind
B2
Abstract

Techniques are disclosed that enable processing of audio data to generate one or more refined versions of audio data, where each of the refined versions of audio data isolate one or more utterances of a single respective human speaker. Various implementations generate a refined version of audio data that isolates utterance(s) of a single human speaker by processing a spectrogram representation of the audio data (generated by processing the audio data with a frequency transformation) using a mask generated by processing the spectrogram of the audio data and a speaker embedding for the single human speaker using a trained voice filter model. Output generated over the trained voice filter model is processed using an inverse of the frequency transformation to generate the refined audio data.

Claims (87)

1 . A client device comprising:

one or more processors, and

memory configured to store instructions that, when executed by the one or more processors, cause the one or more processors to perform a method that includes:

invoking an automated assistant client at the client device, wherein invoking the automated assistant client is in response to detecting one or more instances of invocation user interface input;

in response to invoking the automated assistant client:

capturing one or more instances of spoken input via one or more microphones of the client device;

performing certain processing of the one or more instances of spoken input;

generating a responsive action based on the certain processing of the one or more instances of spoken input;

causing performance of the responsive action;

determining that a continued listening mode is activated for the automated assistant client device;

in response to the continued listening mode being activated:

automatically monitoring for additional spoken input including after causing performance of at least part of the responsive action;

receiving audio data during the automatically monitoring;

determining whether the audio data includes any additional spoken input that is from a same human speaker who provided the spoken input, wherein determining whether the audio data includes the additional spoken input that is from the same human speaker comprises:

 identifying a speaker embedding for the human speaker that provided the one or more instances of spoken input, wherein the speaker embedding was generated based on one or more respective instances of speaker audio data corresponding to the human speaker;

 generating a refined version of the audio data, based on the speaker embedding and the one or more instances of spoken input, that isolates any of the audio data that is from the human speaker; and

 determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human.

2 . The client device of claim 1 , wherein the one or more instances of invocation of the user interface input is a selection of a graphical user interface element of the client device.

3 . The client device of claim 1 , wherein determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human comprises:

processing the refined version of the audio data to determine whether the refined version of the audio data includes one or more non-null segments; and

in response to determining the refined version of the audio data includes one or more non-null segments, determining the refined audio data includes any additional spoken input that is from the same human.

4 . The client device of claim 3 , in response to determining the audio data includes any additional spoken input that is from the same human speaker that provided the spoken input and further comprising:

performing one or more additional actions that are based on the refined version of the audio data.

5 . The client device of claim 4 , wherein performing the one or more additional actions that are based on the refined version of the audio data comprises:

generating additional responsive content that is customized for the human speaker and that is based on the refined version of the audio data; and

causing the client device to render output based on the additional responsive content.

6 . The client device of claim 1 , wherein determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human comprises:

processing the refined version of the audio data to determine whether the refined version of the audio data includes one or more non-null segments; and

in response to determining the refined version of the audio data does not include one or more non-null segments, determining the audio data does not include any additional spoken input that is from the same human.

7 . The client device of claim 6 , in response to determining the audio data does not include any additional spoken input that is from the same human that provided the spoken input and further comprising:

performing one or more additional actions that are based on the audio data capturing additional spoken input.

8 . The client device of claim 7 , wherein performing one or more additional actions that are based on the audio data capturing the additional spoken input comprises:

generating additional responsive content that is not customized for the human speaker; and

causing the client device to render output based on the additional responsive content.

9 . A method implemented by one or more processors, the method comprising:

invoking an automated assistant client at a client device, wherein invoking the automated assistant client is in response to detecting one or more instances of invocation user interface input;

in response to invoking the automated assistant client:

capturing one or more instances of spoken input via one or more microphones of the client device;

performing certain processing of the one or more instances of spoken input;

generating a responsive action based on the certain processing of the one or more instances of spoken input;

causing performance of the responsive action;

determining that a continued listening mode is activated for the automated assistant client device;

in response to the continued listening mode being activated:

automatically monitoring for additional spoken input including after causing performance of at least part of the responsive action;

receiving audio data during the automatically monitoring;

determining whether the audio data includes any additional spoken input that is from a same human speaker who provided the spoken input, wherein determining whether the audio data includes the additional spoken input that is from the same human speaker comprises:

 identifying a speaker embedding for the human speaker that provided the one or more instances of spoken input, wherein the speaker embedding was generated based on one or more respective instances of speaker audio data corresponding to the human speaker;

 generating a refined version of the audio data, based on the speaker embedding and the one or more instances of spoken input, that isolates any of the audio data that is from the human speaker; and

 determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human.

10 . The method of claim 9 , wherein the one or more instances of invocation of the user interface input is a selection of a graphical user interface element of the client device.

11 . The method of claim 9 , wherein determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human comprises:

processing the refined version of the audio data to determine whether the refined version of the audio data includes one or more non-null segments; and

in response to determining the refined version of the audio data includes one or more non-null segments, determining the refined audio data includes any additional spoken input that is from the same human.

12 . The method of claim 11 , in response to determining the audio data includes any additional spoken input that is from the same human speaker that provided the spoken input and further comprising:

performing one or more additional actions that are based on the refined version of the audio data.

13 . The method of claim 12 , wherein performing the one or more additional actions that are based on the refined version of the audio data comprises:

generating additional responsive content that is customized for the human speaker and that is based on the refined version of the audio data; and

causing the client device to render output based on the additional responsive content.

14 . The method of claim 9 , wherein determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human comprises:

processing the refined version of the audio data to determine whether the refined version of the audio data includes one or more non-null segments; and

in response to determining the refined version of the audio data does not include one or more non-null segments, determining the audio data does not include any additional spoken input that is from the same human.

15 . The method of claim 14 , in response to determining the audio data does not include any additional spoken input that is from the same human that provided the spoken input and further comprising:

performing one or more additional actions that are based on the audio data capturing additional spoken input.

16 . The method of claim 15 , wherein performing one or more additional actions that are based on the audio data capturing the additional spoken input comprises:

generating additional responsive content that is not customized for the human speaker; and

causing the client device to render output based on the additional responsive content.

17 . A non-transitory computer-readable storage medium storing instructions executable by one or more processors of a client device to perform a method of:

invoking an automated assistant client at the client device, wherein invoking the automated assistant client is in response to detecting one or more instances of invocation user interface input;

in response to invoking the automated assistant client:

capturing one or more instances of spoken input via one or more microphones of the client device;

performing certain processing of the one or more instances of spoken input;

generating a responsive action based on the certain processing of the one or more instances of spoken input;

causing performance of the responsive action;

determining that a continued listening mode is activated for the automated assistant client device;

in response to the continued listening mode being activated:

automatically monitoring for additional spoken input including after causing performance of at least part of the responsive action;

receiving audio data during the automatically monitoring;

determining whether the audio data includes any additional spoken input that is from a same human speaker who provided the spoken input, wherein determining whether the audio data includes the additional spoken input that is from the same human speaker comprises:

identifying a speaker embedding for the human speaker that provided the one or more instances of spoken input, wherein the speaker embedding was generated based on one or more respective instances of speaker audio data corresponding to the human speaker;

generating a refined version of the audio data, based on the speaker embedding and the one or more instances of spoken input, that isolates any of the audio data that is from the human speaker; and

determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the one or more instances of invocation of the user interface input is a selection of a graphical user interface element of the client device.

19 . The non-transitory computer-readable storage medium of claim 17 , wherein determining, based on the refined audio data, whether the audio data includes any additional spoken input that is from the same human comprises:

processing the refined version of the audio data to determine whether the refined version of the audio data includes one or more non-null segments; and

in response to determining the refined version of the audio data includes one or more non-null segments, determining the refined audio data includes any additional spoken input that is from the same human.

20 . The non-transitory computer-readable storage medium of claim 19 , in response to determining the audio data includes any additional spoken input that is from the same human speaker that provided the spoken input and further comprising:

performing one or more additional actions that are based on the refined version of the audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2024
From: WANG, QUAN; SRIDHAR, PRASHANT; MORENO, IGNACIO LOPEZ; MUCKENHIRN, HANNAH
To: GOOGLE LLC
Reel/Frame 066776/0497 →
Continuity (4)
Continuation 17567590 · Jan 3, 2022
Continuation 16598172 · Oct 10, 2019
Provisional Application 62784669 · Dec 24, 2018
Related Publication 20240203426A1 · Jun 20, 2024
References Cited (57)
US 6502067B1 · Hegger · 2002 [cited by examiner]
US 7243060B2 · Atlas et al. · 2007 [cited by applicant]
US 8682667B2 · Haughay · 2014 [cited by examiner]
US 10522167B1 · Ayrapetian · 2019 [cited by applicant]
US 11217254B2 · Wang · 2022 [cited by examiner]
US 20110188671A1 · Anderson · 2011 [cited by applicant]
US 20140350943A1 · Goldstein · 2014 [cited by applicant]
US 20150025887A1 · Sidi · 2015 [cited by examiner]
US 20150046157A1 · Wolff · 2015 [cited by examiner]
US 20160217792A1 · Gorodetski · 2016 [cited by examiner]
US 20160240210A1 · Lou · 2016 [cited by applicant]
US 20170270919A1 · Parthasarathi et al. · 2017 [cited by applicant]
US 20170337924A1 · Yu · 2017 [cited by examiner]
US 20180082692A1 · Khoury · 2018 [cited by applicant]
US 20180088899A1 · Gillespie · 2018 [cited by applicant]
US 20190019505A1 · Nicholson · 2019 [cited by examiner]
US 20190066713A1 · Mesgarani · 2019 [cited by examiner]
US 20190304069A1 · Vogels · 2019 [cited by applicant]
US 20190318754A1 · Le Roux · 2019 [cited by examiner]
US 20200013425A1 · Bergmann · 2020 [cited by applicant]
US 20200135209A1 · Delfarah · 2020 [cited by examiner]
US 20200227064A1 · Xu · 2020 [cited by examiner]
US 20200342857A1 · Moreno · 2020 [cited by examiner]
US 20220122611A1 · Wang et al. · 2022 [cited by applicant]
US 20230116052A1 · Emre · 2023 [cited by applicant]
Huang, P. et al., “Deep learning for monaural speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference; pp. 1562-1566; 2014. [cited by applicant]
Wang, Y. et al., “On training targets for supervised speech separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 22, No. 12, pp. 1849-1858, 2014. [cited by applicant]
Du, J. et al., “A regression approach to single-channel speech separation via high-resolution deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, No. 8, pp. 1424-1437, 2016. [cited by applicant]
Chen, Z. et al., “Deep attractor network for single-microphone speaker separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference; pp. 246-250; 2017. [cited by applicant]
Fu, S. et al., “Raw waveform-based speech enhancement by fully convolutional networks,” arXiv preprint arXiv:1703.02205, 7 pages; 2017. [cited by applicant]
Pascual, S. et al., “Segan: Speech enhancement generative adversarial network,” arXiv preprint arXiv:1703.09452, 5 pages; 2017. [cited by applicant]
Wilson, K. et al., “Exploring tradeoffs in models for low-latency speech enhancement,” in International Workshop on Acoustic Signal Enhancement (IWAENC); 5 pages; 2018. [cited by applicant]
Heigold, G. et al., “End-to-end text-dependent speaker verification,” in Acoustics, Speech and Signal Processing (ICASSP); IEEE; pp. 5115-5119; 2016. [cited by applicant]
Wan, L. et al., “Generalized end-to-end loss for speaker verification,” in Acoustics, Speech and Signal Processing (ICASSP); IEEE; 5 pages; 2018. [cited by applicant]
Zmolikova, K. et al., “Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,” in Interspeech; 5 pages; 2017. [cited by applicant]
Wang, Q. et al., “Speaker diarization with Istm,” in Acoustics, Speech and Signal Processing (ICASSP); IEEE; 5 pages; 2018. [cited by applicant]
Jia, Y. et al., “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Conference on Neural Information Processing Systems (NIPS); 11 pages; 2018. [cited by applicant]
Panayotov, V. et al., “Librispeech: An ASR Corpus Based On Public Domain Audio Books,” in Acoustics, Speech and Signal Processing (ICASSP); IEEE International Conference; pp. 5206-5210; dated 2015. [cited by applicant]
Nagrani, A. et al., “Voxceleb: a large-scale speaker identification dataset,” arXiv preprint arXiv:1706.08612; 5 pages; 2017. [cited by applicant]
Chung, J. et al., “Voxceleb2: Deep speaker recognition,” arXiv preprint arXiv:1806.05622; 6 pages; 2018. [cited by applicant]
Michaely, A. et al., “Unsupervised context learning for speech recognition,” in Spoken Language Technology Workshop (SLT); IEEE; pp. 447-453; 2016. [cited by applicant]
Vincent, E. et al., “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing (TASLP), vol. 14, No. 4, pp. 1462-1469, 2006. [cited by applicant]
Cherry, E., “Some experiments on the recognition of speech, with one and with two ears,” The Journal of the Acoustical Society of America, vol. 25, No. 5, pp. 975-979, 1953. [cited by applicant]
Yu, D. et al., “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE… [cited by applicant]
Dehak, N. et al., “Front-end factor analysis for speaker verification;” IEEE Transactions on Audio, Speech, and Language Processing; vol. 19, No. 4; pp. 788-798; 2010. [cited by applicant]
Wang, J. et al., “Deep extractor network for target speaker recovery from single channel speech mixtures,” arXiv preprint arXiv:1807.08974; 5 pages; 2018. [cited by applicant]
Delcroix, M. et al., “Single channel target speaker extraction and recognition with speaker beam,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE; pp. 5554-5558; 2018. [cited by applicant]
He, Y. et al., “Streaming end-to-end speech recognition for mobile devices,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE; pp. 6381-6385; 2019. [cited by applicant]
Griffin, D. et al., “Signal estimation from modified short-time fourier transform;” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 32, No. 2, pp. 236-243, 1984. [cited by applicant]
Ko, T. et al. “A study on data augmentation of reverberant speech for robust speech recognition,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE; pp. 5220-5224; 2017. [cited by applicant]
Alvarez, R. et al., “On the efficient representation and execution of deep acoustic models,” arXiv preprint arXiv:1607.04683; 5 pages; 2016. [cited by applicant]
Hershey, J. et al., “Deep clustering: Discriminative embeddings for segmentation and separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference; pp. 31-35; 2016. [cited by applicant]
Kolbæk, M. et al., “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 25, … [cited by applicant]
Wang, Q. et al., “Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,” arXiv preprint arXiv:1810.04826; 5 pages; 2018. [cited by applicant]
Xiao, X. et al., “Single-Channel Speech Extraction Using Speaker Inventory and Attention Network;” International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE; pp. 86-90; May 12, 2019 May 12, 2019. [cited by applicant]
Hetherly, J. et al., “Deep Speech Denoising with Vector Space Projections;” Cornell University; arXiv.org; arXiv:1804.10869v1; 5 pages; Apr. 27, 2018 Apr. 27, 2018. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT application Ser. No. PCT/US2019/055539; 12 pages; dated May 8, 2020. [cited by applicant]