IP Library › Granted Patent US 12,417,760
Granted Patent B2
US 12,417,760 · App. 17/962,248 · Granted Sep 16, 2025

Speaker identification, verification, and diarization using neural networks for conversational AI systems and applications

Inventors: Nithin Rao Koluguri (San Jose, CA); Taejin Park (San Jose, CA); Boris Ginsburg (Sunnyvale, CA)
Assignee: NVIDIA Corporation
G10L15/16G06N3/08G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,760
App. No.
17/962,248
Granted
Sep 16, 2025
Kind
B2
Abstract

Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing speaker recognition, verification, and/or diarization. The techniques include applying a neural network (NN) to a speech data to obtain a speaker embedding representative of an association between the speech data and a speaker that produced the speech. The speech data includes a plurality of frames and a plurality of channels representative of spectral content of the speech data. The NN has one or more blocks of neurons that include a first branch performing convolutions of the speech data across the plurality of channels and across the plurality of frames and a second branch performing convolutions of the speech data across the plurality of channels. Obtained speaker embeddings may be used for various tasks of speaker identification, verification, and/or diarization.

Claims (74)

1. A method comprising:

applying speech data representative of a plurality of channels of spectral content to a neural network (NN), the NN including a first branch and a second branch parallel to the first branch; and

computing, using the NN and based at least on the speech data, an output including a speaker embedding representative of an association between the speech data and a speaker that produced speech corresponding to the speech data, the output being computed based at least on:

a first output of the first branch computed using at least a first set of convolutions with respect to the speech data, the first set of convolutions comprising a first subset of convolutions across a plurality of frames of speech data for one or more fixed channels and a second subset of convolutions across the plurality of channels for one or more fixed frames; and

a second output of the second branch computed using at least a second set of convolutions with respect to the speech data and across the plurality of channels.

2. The method of claim 1 , wherein the speaker embedding is indicative of at least one of: an identification of the speaker within a database of speakers, at least one individual speaker in the database of speakers represented using a set of one or more stored speaker embeddings;

a confirmation that the speaker has produced an additional speech characterized using an additional embedding generated using the NN; or

a distinction of the speaker from one or more additional speakers in a common speech episode that includes the speech and one or more additional speech instances produced by the one or more additional speakers.

3. The method of claim 2 , wherein the first branch includes a squeeze-and-excitation (SE) group of neurons, the SE group of neurons to perform operations comprising:

reducing an intermediate data from a first channel dimension to a second channel dimension;

performing one or more operations using the intermediate data;

expanding the intermediate data from the second channel dimension to the first channel dimension; and

combining the intermediate data with the expanded intermediate data.

4. The method of claim 3 , wherein the first output of the first branch includes a third output of the SE group of neurons, and wherein the third output of the SE group of neurons is combined with the second output of the second branch using an average pooling operation.

5. The method of claim 1 , wherein the NN includes two or more blocks of neurons, wherein individual blocks of the two or more blocks of neurons include the first branch and the second branch.

6. The method of claim 1 , wherein the first set of convolutions is performed multiple times.

7. The method of claim 1 , wherein

the second subset of convolutions is sequential to the first subset of convolutions.

8. The method of claim 1 , wherein at least individual channel of the plurality of channels is associated with a corresponding mel-band of a plurality of mel-bands of a respective time interval corresponding to the speech data.

9. The method of claim 1 , wherein the NN is trained using a plurality of training speech segments, and further wherein at least one of the plurality of training speech segments is segmented among a plurality of predetermined segment durations from a common speech episode.

10. The method of claim 9 , wherein the NN is trained, at least, by:

obtaining a plurality of training speaker embeddings, wherein at least one individual training speaker embedding is representative of a predicted speaker of a respective training speech segment of the plurality of training speech segments;

computing an angular softmax margin loss function comprising a plurality of contributions, at least one individual contribution of the plurality of contributions characterizing, for the respective training speech segment, a mismatch between the predicted speaker and a ground truth speaker of the respective training speech segment; and

modifying one or more parameters of the NN using the computed angular softmax margin loss function.

11. The method of claim 9 , wherein the NN is trained, at least, by:

obtaining a first training speaker embedding representative of a first association between a first training speech segment of the plurality of training speech segments and a first speaker;

obtaining a second training speaker embedding representative of a second association between a second training speech segment of the plurality of training speech segments and a second speaker; and

modifying one or more parameters of the NN to:

cause a cosine similarity between the first training speaker embedding and the second training speaker embedding to increase when the first speaker is the same as the second speaker, or

cause the cosine similarity between the first training speaker embedding and the second training speaker embedding to decrease when the first speaker is different from the second speaker.

12. A method comprising:

determining a plurality of training speech segments based at least on randomly segmenting a speech episode, according to at least two predetermined durations;

computing, using a neural network (NN) and based at least on a first training speech segment of the plurality of training speech segments, a first training speaker embedding corresponding to the first training speech segment associated with a first speaker;

computing a loss function characterizing a similarity between the first training speaker embedding and a second training speaker embedding corresponding to a second training speech segment of the plurality of training speech segments, the second training speech segment associated with a second speaker; and

updating one or more parameters of the NN using to cause the computed loss function to change:

in a first direction responsive to the first speaker being the same as the second speaker, or

in a second direction responsive to the first speaker being different from the second speaker.

13. The method of claim 12 , wherein the loss function includes a plurality of contributions, individual contributions of the plurality of contributions characterizing, for a respective training speech segment of the plurality of training speech segments, a mismatch between a predicted speaker and a ground truth speaker of the respective training speech segment.

14. The method of claim 12 , wherein the loss function comprises an angular softmax marginal loss function.

15. The method of claim 12 , wherein individual training speech segments of the plurality of training speech segments are represented using a plurality of frames, wherein individual frames of the plurality of frames (i) are associated with a respective time segment of a plurality of time segments of a corresponding training speech segment and (ii) include a plurality of channels representative of spectral content of the respective time segment of the corresponding training speech segment, and wherein the NN comprises a plurality of blocks, individual blocks of the plurality of blocks including:

a first branch of neurons performing a first set of convolutions across the plurality of channels and across the plurality of frames, and

a second branch of neurons performing a second set of convolutions across the plurality of channels, and wherein the second branch of neurons is parallel to the first branch of neurons.

16. A system comprising:

one or more processing units to:

apply a neural network (NN) to a speech data pertaining to a speech, wherein the NN comprises a first branch and a second branch parallel to the first branch, and wherein the speech data is representative of a plurality of channels and comprises a plurality of frames; and

obtain an output of the NN comprising a speaker embedding representative of an association between the speech and a speaker that produced the speech, wherein the output is based at least on:

an output of the first branch performing at least a first set of convolutions comprising a first subset of convolutions across the plurality of frames of the speech data for one or more fixed channels and a second subset of convolutions across the plurality of channels for one or more fixed frames, and

an output of the second branch performing at least a second set of convolutions of the speech data across the plurality of channels.

17. The system of claim 16 , wherein the speaker embedding is indicative of at least one of:

an identification of the speaker within a database of speakers, each speaker in the database of speakers represented by a set of one or more stored speaker embeddings;

a confirmation that the speaker has produced an additional speech characterized by an additional embedding generated by applying the NN to an additional speech; or

a distinction of the speaker from one or more additional speakers in a common speech episode that comprises:

the speech, and

one or more additional speech instances produced by the one or more additional speakers.

18. The system of claim 16 , wherein the first branch comprises a squeeze-and-excitation (SE) group of neurons, and wherein to perform computations for the SE group of neurons, the one or more processing units are to:

reduce an intermediate data from a first channel dimension to a second channel dimension;

perform one or more operations with the intermediate data;

expand the intermediate data from the second channel dimension to the first channel dimension; and

combine the intermediate data with the expanded intermediate data.

19. The system of claim 16 , wherein the system is comprised in at least one of:

an in-vehicle infotainment system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content;

a system implemented using a robot;

a system for performing conversational AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2023
From: KOLUGURI, NITHIN RAO; PARK, TAEJIN; GINSBURG, BORIS
To: NVIDIA CORPORATION
Reel/Frame 062601/0090 →
Continuity (1)
Related Publication 20240119927A1 · Apr 11, 2024
References Cited (62)
US 11158307B1 · Ghias et al. · 2021 [cited by applicant]
US 12100383B1 · Ezzerg et al. · 2024 [cited by applicant]
US 20200058290A1 · Chae et al. · 2020 [cited by applicant]
US 20200302223A1 · Dutta et al. · 2020 [cited by applicant]
US 20200394997A1 · Trueba et al. · 2020 [cited by applicant]
US 20200402497A1 · Semenov et al. · 2020 [cited by applicant]
US 20210224319A1 · Ingel et al. · 2021 [cited by applicant]
US 20210304769A1 · Ye et al. · 2021 [cited by applicant]
US 20220028367A1 · Shekhar · 2022 [cited by examiner]
US 20220051654A1 · Finkelstein et al. · 2022 [cited by applicant]
US 20220093106A1 · Mosayyebpour Kaskari · 2022 [cited by examiner]
US 20220223144A1 · Sun · 2022 [cited by examiner]
US 20220319018A1 · Gervais et al. · 2022 [cited by applicant]
US 20230150498A1 · Seong et al. · 2023 [cited by applicant]
US 20240037316A1 · Mohanty · 2024 [cited by examiner]
US 20240104055A1 · McAnallen · 2024 [cited by applicant]
US 20240105289A1 · Khan · 2024 [cited by examiner]
JP 2015102667A · 2015 [cited by applicant]
WO 2022063758A1 · 2022 [cited by applicant]
Garcia-Romero, D. et al., “Speaker Diarization Using Deep Neural Network Embeddings”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4930-4934, Mar. 5, 2017. [cited by applicant]
Han, W. et al., “ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context”, arXiv preprint arXiv:2005.03191 (2020). [cited by applicant]
Kingma, D. et al., “Improved Variational Inference with Inverse Autoregressive Flow”, 30th Conference on Neural Information Processing Systems, 9 pages, 2016. [cited by applicant]
Koluguri, N. et al., “Titanet: Neural Model for Speaker Representation With 1D Depth-Wise Separable Convolutions and Global Context”, International Conference on Acoustics, Speech and Signal Processing, pp. 8102-8106, M… [cited by applicant]
Lancucki, Adrian, “Fastpitch: Parallel Text-to-Speech with Pitch Prediction”, In ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6588-6592. IEEE, 2021. [cited by applicant]
Muller, T. et al., “Neural Importance Sampling”, ACM Transactions on Graphics (ToG), vol. 38, No. 5, 2019, pp. 1-19. [cited by applicant]
Park, T. et al., “Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap”, IEEE Signal Processing Letters, vol. 27, pp. 381-385, 2019. [cited by applicant]
Park, T. et al., “Multi-Scale Speaker Diarization with Neural Affinity Score Fusion”, In ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7173-7177. IEEE, 2021. [cited by applicant]
Peng, K. et al., “Non-Autoagressive Neural Text-to-Speech”, International Conference on Machine Learning, pp. 7586-7598, 2020. [cited by applicant]
Ren, Y. et al., “FastSpeech: Fast, Robust and Controllable Text to Speech”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 13 pages. [cited by applicant]
Sasirekha, D. et al., “Text-to-Speech: A Simple Tutorial”, International Journal of Soft Computing and Engineering, vol. 2, Issue 1, 4 pages, Mar. 2012. [cited by applicant]
Shih, K. et al., “Generative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows”, arXiv preprint arXiv:2203.01786 (2022). [cited by applicant]
Valle, R., “Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Sythesis”, arXiv preprint arXiv:2005.05957. May 12, 2020. [cited by applicant]
Chen, J., et al., “VFlow: More Expressive Generative Flows with Variational Data Augmentation,” Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, vol. 119, pp. 1660-1669. Retrived from th… [cited by applicant]
Dumoulin, V., et al., “A Learned Representation For Artistic Style,” Computer Vision and Pattern Recognition, ICLR, 2017, 26 Pages. Retrived from the Internet: [https://arxiv.org/abs/1610.07629]. [cited by applicant]
Dupont, E., et al., “Augmented Neural ODEs,” 33rd Conference on Neural Information Processing Systems, NeurIPS, 2019, vol. 32, pp. 3140-3150. [cited by applicant]
Durkan, C., et al., “Neural Spline Flows,” Advances in Neural Information Processing Systems, 2019, vol. 32, pp. 7511-7522. [cited by applicant]
Huang, C.W., et al., “Augmented Normalizing Flows: Bridging the Gap Between Generative Flows and Latent Variable Models,” Machine Learning, 2022, 27 Pages. Retrived from the Internet: [https://arxiv.org/abs/2002.07101]. [cited by applicant]
Jeong, M., et al., “Diff-TTS: A Denoising Diffusion Model for Text-to-Speech,” Department of Electrical and Computer Engineering and INMC, 2021, 5 Pages. Retrived from the Internet: [https://arxiv.org/pdf/2104.01409.pdf… [cited by applicant]
Kawahara, H., et al., “Nearly Defect-free Fo Trajectory Extraction for Expressive Speech Modifications Based on Straight,” In Ninth European Conference on Speech Communication and Technology, 2005, 5 Pages. [cited by applicant]
Kim, H., et al., “Softflow: Probabilistic Framework for Normalizing Flow on Manifolds,” Advances in Neural Information Processing Systems, 2020, 10 Pages. [cited by applicant]
Kim, J., et al., “Glow-TTS: A Generative Flow for Text-to-speech via Monotonic Alignment Search,” Advances in Neural Information Processing Systems, 2020b, 11 Pages. [cited by applicant]
Kingma, D.P., et al., “Glow: Generative flow with invertible 1×1 Convolutions,” Advances in Neural Information Processing Systems, 2018, vol. 31, 10 Pages. [cited by applicant]
Kong, J., et al., “Hifi-gan: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” Advances in Neural Information Processing Systems, 2020, pp. 1-14. [cited by applicant]
Kong, J., et al., “Hifi-gan: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” 2022, Retrived from the Internet: [https://github. com/jik876/hifi-gan,]. [cited by applicant]
Mauch, M., et al., “Pyin: A Fundamental Frequency Estimator Using Probabilistic Threshold Distributions,” In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (Icassp), 2014, pp. 659-663. [cited by applicant]
McCree, A.V., et al., “A Mixed Excitation Ipc Vocoder Model for Low Bit Rate Speech Coding,” IEEE Transactions on Speech and Audio Processing, 1995, vol. 3(4), pp. 242-250. [cited by applicant]
Miao, C., et al., “Flow-TTS: A Non-autoregressive Network for Text to Speech Based on Flow,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, 5 Pages. [cited by applicant]
Moog, B., “MIDI: Musical Instrument Digital Interface,” Journal of the Audio Engineering Society, 1986, vol. 34(5), pp. 394-404. [cited by applicant]
Nakatani, T., et al., “A Method for Fundamental Frequency Estimation and Voicing Decision: Application to Infant Utterances Recorded in Real Acoustical Environments,” Speech Communication, Mar. 2008, vol. 50(3), pp. 203… [cited by applicant]
Ping, W., et al., “Waveflow: A Compact Flow-based Model for Raw Audio,” In International Conference on Machine Learning, pp. 7706-7716. [cited by applicant]
Popov, V., et al., “Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,” Arxiv Preprint, 2021, 10 Pages. [cited by applicant]
Ren, Y., et al., “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,” Aug. 8, 2022, 15 Pages. Retrived from the Internet: [https://speechresearch.github.io/fastspeech2/]. [cited by applicant]
Ren, Y., et al., “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,” arXiv:2006.04558v1 [eess.AS] Jun. 8, 2020. [cited by applicant]
Shih, K.J., et al., “RAD-TTS: Parallel Flowbased Tts With Robust Alignment Learning and Diverse Synthesis,” In ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models, 2021. [cited by applicant]
Suni, et al., Wavelets for Intonation Modeling in Hmm Speech Synthesis, ISCA, 2013, 7 Pages. [cited by applicant]
Valle, R., et al., “Flowtron: An Autoregressive Flow-based Generative Network for Text-to-speech Synthesis,” In International Conference on Learning Representations, 2020b. [cited by applicant]
Valle, R., et al., “Mellotron: Multispeaker Expressive Voice Synthesis by Conditioning on Rhythm, Pitch and Global Style Tokens,” In ICASSP 2020-2020 Ieee International Conference on Acoustics, Speech and Signal Process… [cited by applicant]
Doty C., et al., “What Is Speaker Diarization,” Deepgram, Aug. 2022, 9 Pages Retrieved from internet URL: https://deepgram.com/learn/what-is-speaker-diarization. [cited by applicant]
Ito, K., et al., “The LJ Speech Dataset,” BibSonomy, 2017. Retrived from the Internet: [https://keithito.com/LJ-Speech-Dataset/]. [cited by applicant]
Park T.J., et al., “Multi-Scale Speaker Diarization With Neural Affinity Score Fusion.” IEEE, May 2021, pp. 7173-7177 Retrieved from Internet URL: https://ieeexplore.ieee.org/stamp/stamp.jsptp=&arnumber=9414578. [cited by applicant]
Wang, Q., et al., “Speaker Diarization With LSTM,” ArXiv, Jan. 2022, 5 Pages Retrieved from internet URL: https://arxiv.org/pdf/1710.10468. [cited by applicant]
Zhang Y., et al., “Audio segmentation based on multi-scale audio classification.” IEEE, Aug. 2004, vol. 4, pp. 349-352 Retrieved from Internet URL: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1326835&am… [cited by applicant]