IP Library Granted Patent US 12,354,608
Granted Patent B2
US 12,354,608 · App. 18/321,353 · Granted Jul 8, 2025

Channel-compensated low-level features for speaker recognition

Inventors: Elie Khoury (Atlanta, GA); Matthew Garland (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L17/20G10L17/02G10L17/04G10L17/18G10L19/028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,608
App. No.
18/321,353
Granted
Jul 8, 2025
Kind
B2
Abstract

A system for generating channel-compensated features of a speech signal includes a channel noise simulator that degrades the speech signal, a feed forward convolutional neural network (CNN) that generates channel-compensated features of the degraded speech signal, and a loss function that computes a difference between the channel-compensated features and handcrafted features for the same raw speech signal. Each loss result may be used to update connection weights of the CNN until a predetermined threshold loss is satisfied, and the CNN may be used as a front-end for a deep neural network (DNN) for speaker recognition/verification. The DNN may include convolutional layers, a bottleneck features layer, multiple fully-connected layers, and an output layer. The bottleneck features may be used to update connection weights of the convolutional layers, and dropout may be applied to the convolutional layers.

Claims (38)

1. A computer-implemented method comprising:

obtaining, by a computer, one or more enrolling speech signals for a registered speaker during speaker registration;

generating, by the computer, one or more degraded enrolling signals by applying an acoustic channel simulator on the one or more enrolling speech signals;

extracting, by the computer, an enrolled voiceprint for the registered speaker by applying a trained neural network on the one or more degraded enrolling signals and the one or more enrolling speech signals;

obtaining, by the computer, a test speech signal for a test speaker during speaker recognition;

extracting, by the computer, channel-compensated features from the test speech signal for the test speaker by applying the trained neural network on the test speech signal; and

identifying, by the computer, the test speaker as the registered speaker in response to determining that a loss between the test voiceprint and the enrolled voiceprint satisfies a distance threshold.

2. The method according to claim 1 , wherein the acoustic channel simulator generates a degraded enrolling speech signal having a type of degradation added to an enrolling speech signal from the registered speaker, and wherein the type of degradation includes one of environmental noise, reverberation, an audio acquisition device characteristic, or an audio channel transcoding artifact.

3. The method according to claim 1 , wherein generating a degraded enrolling speech signal includes:

selecting, by the computer, an environmental noise type as a type of degradation from a plurality of environmental noise types stored in a noise profile database; and

simulating, by the computer, the environmental noise type in the enrolling speech signal applied by the acoustic channel simulator on an enrolling speech signal.

4. The method according to claim 1 , wherein generating the degraded enrolling signal includes simulating, by the computer, a set of one or more audio acquisition device characteristics according to an audio acquisition device profile applied by the acoustic channel simulator to an enrolling speech signal.

5. The method according to claim 4 , wherein each audio acquisition device profile includes at least one of a frequency characteristic, an amplitude characteristic, a filtering characteristic, an electrical noise characteristic, or a physical noise characteristic.

6. The method according to claim 1 , wherein generating the degraded enrolling speech signal includes:

selecting, by the computer, a transcoding profile from a plurality of transcoding profiles stored in a transcoding profile database; and

simulating, by the computer, a set of audio channel transcoding characteristics according to the transcoding profile applied by the acoustic channel simulator to an enrolling speech signal.

7. The method according to claim 1 , wherein generating the degraded enrolling speech signal includes simulating, by the computer, a reverberation according to a direct-to-reverberation ratio (DRR) applied by the acoustic channel simulator to the enrolling speech signal.

8. The method according to claim 1 , further comprising training, by the computer, a neural network to minimize a loss between a set of channel-compensated training features extracted from a training signal by applying the neural network and a set of handcrafted training features extracted from corresponding clean training speech signals, thereby training the trained neural network.

9. The method according to claim 8 , further comprising applying, by the computer, the acoustic channel simulator on a clean training signal to generate the training signal having a type of degradation corresponding to a channel-compensated training feature for training the trained neural network.

10. The method according to claim 8 , wherein a handcrafted training features includes at least one of: Mel-frequency cepstrum coefficients (MFCCs), low-frequency cepstrum coefficients (LFCCs), perceptual linear prediction (PLP) coefficients, linear or Mel filter banks, or glottal features.

11. A computer-implemented method comprising:

obtaining, by a computer, one or more training speech signals containing corresponding training utterances;

for a training speech signal of the one or more training speech signals,

generating, by the computer, a degraded training speech signal having a type of degradation by applying a noise simulator on the training speech signal;

extracting, by the computer, a set of training low-level features for the degraded speech signal by applying a neural network on the degraded speech signal;

extracting, by the computer, a set of training handcrafted features for the speech signal by applying a signal analyzer on the training speech signal; and

updating, by the computer, one or more connection values of the neural network in response to determining that a loss result fails to satisfy a training loss threshold, the loss result based upon the set of training low-level features and the set of training handcrafted features.

12. The method according to claim 11 , further comprising generating, by the computer, the loss result according to a loss function using the set of training low-level features and the set of training handcrafted features.

13. The method according to claim 11 , wherein, for each type of degradation of a plurality of types of degradation, the computer updates the one or more connection values of the neural network according to a loss function to minimize the loss result between the set of training low-level features and the set of training handcrafted features.

14. The method according to claim 11 , wherein, for each training speech of the one or more training speech signals, the computer updates the one or more connection values of the neural network according to a loss function to minimize the loss result between the set of training low-level features and the set of training handcrafted features.

15. The method according to claim 11 , further comprising:

determining, by the computer, that a next loss result for a next training signal satisfies the training loss threshold; and

storing, by the computer, a trained neural network from the neural network in response to determining that the next loss result satisfies the training loss threshold.

16. The method according to claim 11 , wherein generating the degraded training speech signal having the type of degradation includes simulating, by the computer, an environmental noise being applied to the training speech signal according to a type of environmental noise.

17. The method according to claim 11 , wherein generating the degraded training speech signal having the type of degradation includes simulating, by the computer, a reverberation being applied to the training speech signal according to a direct-to-reverberation ratio (DRR).

18. The method according to claim 11 , wherein generating the degraded training speech signal having the type of degradation includes simulating, by the computer, one or more audio device characteristics being applied to the training speech signal according to an audio acquisition device profile.

19. The method according to claim 11 , wherein generating the degraded training speech signal having the type of degradation includes simulating, by the computer, a set of audio channel transcoding characteristics being applied to the training speech signal according to a transcoding profile.

20. The method according to claim 11 , wherein the low-level features include at least one of: Mel-frequency cepstrum coefficients (MFCCs), low-frequency cepstrum coefficients (LFCCs), perceptual linear prediction (PLP) coefficients, linear or Mel filter banks, or glottal features.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2023
From: KHOURY, ELIE; GARLAND, MATTHEW
To: PINDROP SECURITY, INC.
Reel/Frame 063717/0896 →
Continuity (6)
Continuation 17107496 · Nov 30, 2020
Continuation 16505452 · Jul 8, 2019
Continuation 15709024 · Sep 19, 2017
Provisional Application 62396670 · Sep 19, 2016
Provisional Application 62396617 · Sep 19, 2016
Related Publication 20230290357A1 · Sep 14, 2023
References Cited (339)
US 4817156A · Bahl et al. · 1989 [cited by applicant]
US 4829577A · Kuroda et al. · 1989 [cited by applicant]
US 4972485A · Dautrich et al. · 1990 [cited by applicant]
US 5072452A · Brown et al. · 1991 [cited by applicant]
US 5461697A · Nishimura et al. · 1995 [cited by applicant]
US 5475792A · Stanford et al. · 1995 [cited by applicant]
US 5598507A · Kimber et al. · 1997 [cited by applicant]
US 5659662A · Wilcox et al. · 1997 [cited by applicant]
US 5835890A · Matsui et al. · 1998 [cited by applicant]
US 5867562A · Scherer · 1999 [cited by applicant]
US 5949874A · Mark · 1999 [cited by applicant]
US 5995927A · Li · 1999 [cited by applicant]
US 6009392A · Kanevsky et al. · 1999 [cited by applicant]
US 6021119A · Derks et al. · 2000 [cited by applicant]
US 6055498A · Neumeyer et al. · 2000 [cited by applicant]
US 6094632A · Hattori · 2000 [cited by applicant]
US 6141644A · Kuhn et al. · 2000 [cited by applicant]
US 6411930B1 · Burges · 2002 [cited by applicant]
US 6463413B1 · Applebaum et al. · 2002 [cited by applicant]
US 6519561B1 · Farrell et al. · 2003 [cited by applicant]
US 6760701B2 · Sharma et al. · 2004 [cited by applicant]
US 6882972B2 · Kompe et al. · 2005 [cited by applicant]
US 6922668B1 · Downey · 2005 [cited by applicant]
US 6975708B1 · Scherer · 2005 [cited by applicant]
US 7003460B1 · Bub et al. · 2006 [cited by applicant]
US 7209881B2 · Yoshizawa et al. · 2007 [cited by applicant]
US 7295970B1 · Gorin et al. · 2007 [cited by applicant]
US 7318032B1 · Chaudhari et al. · 2008 [cited by applicant]
US 7324941B2 · Choi et al. · 2008 [cited by applicant]
US 7739114B1 · Chen et al. · 2010 [cited by applicant]
US 7813927B2 · Navratil et al. · 2010 [cited by applicant]
US 8046230B1 · McIntosh · 2011 [cited by applicant]
US 8112160B2 · Foster · 2012 [cited by applicant]
US 8160811B2 · Prokhorov · 2012 [cited by applicant]
US 8160877B1 · Nucci et al. · 2012 [cited by applicant]
US 8380503B2 · Gross · 2013 [cited by applicant]
US 8385536B2 · Whitehead · 2013 [cited by applicant]
US 8484023B2 · Kanevsky et al. · 2013 [cited by applicant]
US 8484024B2 · Kanevsky et al. · 2013 [cited by applicant]
US 8554563B2 · Aronowitz · 2013 [cited by applicant]
US 8712760B2 · Hsia et al. · 2014 [cited by applicant]
US 8738442B1 · Liu et al. · 2014 [cited by applicant]
US 8856895B2 · Perrot · 2014 [cited by applicant]
US 8886663B2 · Gainsboro et al. · 2014 [cited by applicant]
US 8903859B2 · Zeppenfeld et al. · 2014 [cited by applicant]
US 9042867B2 · Gomar · 2015 [cited by applicant]
US 9064491B2 · Rachevsky et al. · 2015 [cited by applicant]
US 9277049B1 · Danis · 2016 [cited by applicant]
US 9336781B2 · Scheffer et al. · 2016 [cited by applicant]
US 9338619B2 · Kang · 2016 [cited by applicant]
US 9343067B2 · Ariyaeeinia et al. · 2016 [cited by applicant]
US 9344892B1 · Rodrigues et al. · 2016 [cited by applicant]
US 9355646B2 · Oh et al. · 2016 [cited by applicant]
US 9373330B2 · Cumani et al. · 2016 [cited by applicant]
US 9401143B2 · Senior et al. · 2016 [cited by applicant]
US 9401148B2 · Lei et al. · 2016 [cited by applicant]
US 9406298B2 · Cumani et al. · 2016 [cited by applicant]
US 9431016B2 · Aviles-Casco et al. · 2016 [cited by applicant]
US 9444839B1 · Faulkner et al. · 2016 [cited by applicant]
US 9454958B2 · Li et al. · 2016 [cited by applicant]
US 9460722B2 · Sidi et al. · 2016 [cited by applicant]
US 9466292B1 · Lei et al. · 2016 [cited by applicant]
US 9502038B2 · Wang et al. · 2016 [cited by applicant]
US 9514753B2 · Sharifi et al. · 2016 [cited by applicant]
US 9558755B1 · Laroche et al. · 2017 [cited by applicant]
US 9584946B1 · Lyren et al. · 2017 [cited by applicant]
US 9620145B2 · Bacchiani et al. · 2017 [cited by applicant]
US 9626971B2 · Rodriguez et al. · 2017 [cited by applicant]
US 9633652B2 · Kurniawati et al. · 2017 [cited by applicant]
US 9641954B1 · Typrin et al. · 2017 [cited by applicant]
US 9665823B2 · Saon et al. · 2017 [cited by applicant]
US 9685174B2 · Karam et al. · 2017 [cited by applicant]
US 9729727B1 · Zhang · 2017 [cited by applicant]
US 9818431B2 · Yu · 2017 [cited by applicant]
US 9824692B1 · Khoury et al. · 2017 [cited by applicant]
US 9860367B1 · Jiang et al. · 2018 [cited by applicant]
US 9875739B2 · Ziv et al. · 2018 [cited by applicant]
US 9875742B2 · Gorodetski et al. · 2018 [cited by applicant]
US 9875743B2 · Gorodetski et al. · 2018 [cited by applicant]
US 9881617B2 · Sidi et al. · 2018 [cited by applicant]
US 9930088B1 · Hodge · 2018 [cited by applicant]
US 9984706B2 · Wein · 2018 [cited by applicant]
US 10277628B1 · Jakobsson · 2019 [cited by applicant]
US 10325601B2 · Khoury et al. · 2019 [cited by applicant]
US 10347256B2 · Khoury · 2019 [cited by examiner]
US 10397398B2 · Gupta · 2019 [cited by applicant]
US 10404847B1 · Unger · 2019 [cited by applicant]
US 10462292B1 · Stephens · 2019 [cited by applicant]
US 10506088B1 · Singh · 2019 [cited by applicant]
US 10554821B1 · Koster · 2020 [cited by applicant]
US 10638214B1 · Delhoume et al. · 2020 [cited by applicant]
US 10659605B1 · Braundmeier et al. · 2020 [cited by applicant]
US 10679630B2 · Khoury et al. · 2020 [cited by applicant]
US 10854205B2 · Khoury · 2020 [cited by examiner]
US 11069352B1 · Tang et al. · 2021 [cited by applicant]
US 11670304B2 · Khoury et al. · 2023 [cited by applicant]
US 11854205B2 · Cao · 2023 [cited by examiner]
US 20020095287A1 · Botterweck · 2002 [cited by applicant]
US 20020143539A1 · Botterweck · 2002 [cited by applicant]
US 20030231775A1 · Wark · 2003 [cited by applicant]
US 20030236663A1 · Dimitrova et al. · 2003 [cited by applicant]
US 20040218751A1 · Colson et al. · 2004 [cited by applicant]
US 20040230420A1 · Kadambe et al. · 2004 [cited by applicant]
US 20050038655A1 · Mutel et al. · 2005 [cited by applicant]
US 20050039056A1 · Bagga et al. · 2005 [cited by applicant]
US 20050286688A1 · Scherer · 2005 [cited by applicant]
US 20060058998A1 · Yamamoto et al. · 2006 [cited by applicant]
US 20060111905A1 · Navratil et al. · 2006 [cited by applicant]
US 20060293771A1 · Tazine et al. · 2006 [cited by applicant]
US 20070189479A1 · Scherer · 2007 [cited by applicant]
US 20070198257A1 · Zhang et al. · 2007 [cited by applicant]
US 20070294083A1 · Bellegarda et al. · 2007 [cited by applicant]
US 20080195389A1 · Zhang et al. · 2008 [cited by applicant]
US 20080312926A1 · Vair et al. · 2008 [cited by applicant]
US 20090138712A1 · Driscoll · 2009 [cited by applicant]
US 20090265328A1 · Parekh et al. · 2009 [cited by applicant]
US 20100131273A1 · Aley-Raz et al. · 2010 [cited by applicant]
US 20100217589A1 · Gruhn et al. · 2010 [cited by applicant]
US 20100232619A1 · Uhle · 2010 [cited by examiner]
US 20100262423A1 · Huo et al. · 2010 [cited by applicant]
US 20110010173A1 · Scott et al. · 2011 [cited by applicant]
US 20120173239A1 · Asenjo et al. · 2012 [cited by applicant]
US 20120185418A1 · Capman et al. · 2012 [cited by applicant]
US 20130041660A1 · Waite · 2013 [cited by applicant]
US 20130080165A1 · Wang et al. · 2013 [cited by applicant]
US 20130109358A1 · Balasubramaniyan et al. · 2013 [cited by applicant]
US 20130300939A1 · Chou et al. · 2013 [cited by applicant]
US 20140046878A1 · Lecomte et al. · 2014 [cited by applicant]
US 20140053247A1 · Fadel · 2014 [cited by applicant]
US 20140081640A1 · Farrell et al. · 2014 [cited by applicant]
US 20140195236A1 · Hosom et al. · 2014 [cited by applicant]
US 20140214417A1 · Wang et al. · 2014 [cited by applicant]
US 20140214676A1 · Bukai · 2014 [cited by applicant]
US 20140241513A1 · Springer · 2014 [cited by applicant]
US 20140250512A1 · Goldstone et al. · 2014 [cited by applicant]
US 20140278412A1 · Scheffer et al. · 2014 [cited by applicant]
US 20140288928A1 · Penn et al. · 2014 [cited by applicant]
US 20140337017A1 · Watanabe et al. · 2014 [cited by applicant]
US 20150036813A1 · Ananthakrishnan et al. · 2015 [cited by applicant]
US 20150127336A1 · Lei et al. · 2015 [cited by applicant]
US 20150149165A1 · Saon · 2015 [cited by applicant]
US 20150161522A1 · Saon et al. · 2015 [cited by applicant]
US 20150189086A1 · Romano et al. · 2015 [cited by applicant]
US 20150199960A1 · Huo et al. · 2015 [cited by applicant]
US 20150269931A1 · Senior et al. · 2015 [cited by applicant]
US 20150269941A1 · Jones · 2015 [cited by applicant]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20150310008A1 · Thudor et al. · 2015 [cited by applicant]
US 20150334231A1 · Rybak et al. · 2015 [cited by applicant]
US 20150348571A1 · Koshinaka et al. · 2015 [cited by applicant]
US 20150356630A1 · Hussain · 2015 [cited by applicant]
US 20150365530A1 · Kolbegger et al. · 2015 [cited by applicant]
US 20160019458A1 · Kaufhold · 2016 [cited by applicant]
US 20160019883A1 · Aronowitz · 2016 [cited by applicant]
US 20160028434A1 · Kerpez et al. · 2016 [cited by applicant]
US 20160078863A1 · Chung et al. · 2016 [cited by applicant]
US 20160104480A1 · Sharifi · 2016 [cited by applicant]
US 20160125877A1 · Foerster et al. · 2016 [cited by applicant]
US 20160180214A1 · Kanevsky et al. · 2016 [cited by applicant]
US 20160189707A1 · Donjon · 2016 [cited by examiner]
US 20160240190A1 · Lee et al. · 2016 [cited by applicant]
US 20160275953A1 · Sharifi et al. · 2016 [cited by applicant]
US 20160284346A1 · Visser et al. · 2016 [cited by applicant]
US 20160293167A1 · Chen et al. · 2016 [cited by applicant]
US 20160314790A1 · Tsujikawa et al. · 2016 [cited by applicant]
US 20160343373A1 · Ziv et al. · 2016 [cited by applicant]
US 20170060779A1 · Falk · 2017 [cited by applicant]
US 20170069313A1 · Aronowitz · 2017 [cited by applicant]
US 20170069327A1 · Heigold et al. · 2017 [cited by applicant]
US 20170098444A1 · Song · 2017 [cited by applicant]
US 20170111515A1 · Bandyopadhyay et al. · 2017 [cited by applicant]
US 20170126884A1 · Balasubramaniyan et al. · 2017 [cited by applicant]
US 20170142150A1 · Sandke et al. · 2017 [cited by applicant]
US 20170169816A1 · Blandin et al. · 2017 [cited by applicant]
US 20170230390A1 · Faulkner et al. · 2017 [cited by applicant]
US 20170262837A1 · Gosalia · 2017 [cited by applicant]
US 20180082691A1 · Khoury et al. · 2018 [cited by applicant]
US 20180152558A1 · Chan et al. · 2018 [cited by applicant]
US 20180249006A1 · Dowlatkhah et al. · 2018 [cited by applicant]
US 20180295235A1 · Tatourian et al. · 2018 [cited by applicant]
US 20180337962A1 · Ly et al. · 2018 [cited by applicant]
US 20190037081A1 · Rao et al. · 2019 [cited by applicant]
US 20190172476A1 · Wung et al. · 2019 [cited by applicant]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190180758A1 · Washio · 2019 [cited by applicant]
US 20190238956A1 · Gaubitch et al. · 2019 [cited by applicant]
US 20190297503A1 · Traynor et al. · 2019 [cited by applicant]
US 20200137221A1 · Dellostritto et al. · 2020 [cited by applicant]
US 20200195779A1 · Weisman et al. · 2020 [cited by applicant]
US 20200252510A1 · Ghuge et al. · 2020 [cited by applicant]
US 20200312313A1 · Maddali et al. · 2020 [cited by applicant]
US 20200396332A1 · Gayaldo · 2020 [cited by applicant]
US 20210084147A1 · Kent et al. · 2021 [cited by applicant]
US 20210150010A1 · Strong et al. · 2021 [cited by applicant]
US 20230326462A1 · Khoury et al. · 2023 [cited by applicant]
WO WO2015079885 · 2015 [cited by applicant]
WO WO2016195261A1 · 2016 [cited by applicant]
WO WO2017167900A1 · 2017 [cited by applicant]
Australian Examination Report on AU Appl. Ser. No. 2021286422 dated Oct. 21, 2022 (2 pages). [cited by applicant]
First Examiner's Requisition for CA App. 3,135,210 dated Oct. 5, 2023 (5 pages). [cited by applicant]
First Examiners Requisition on CA Appl. 3,096,378 dated Jun. 12, 2023 (3 pages). [cited by applicant]
Office Action on Japanese Application 2022-104204 dated Jul. 26, 2023 (4 pages). [cited by applicant]
J. H. L. Hansen and T. Hasan, “Speaker Recognition by Machines and Humans: A tutorial review,” in IEEE Signal Processing Magazine, vol. 32, No. 6, pp. 74-99, Nov. 2015, doi: 10.1109/MSP.2015.2462851. (Year: 2015). [cited by applicant]
M. Witkowski, M. Igras, J. Grzybowska, P. Jaci6w, J. Galka and M. Zi6/ko, “Caller identification by voice,” XXII Annual Pacific Voice Conference (PVC), Krakow, Poland, 2014, pp. 1-7, doi: 10.1109/PVC.2014.6845420. (Year… [cited by applicant]
T. F. Zheng, Q. Jin, L. Li, J. Wang and F. Bie, “An overview of robustness related issues in speaker recognition,” Signal and Information Processing Association Annual Summit and Conference (APSIPA), 2014 Asia-Pacific, … [cited by applicant]
First Examiner's Requisition on CA App. 3,179,080 dated Apr. 9, 2024 (3 pages). [cited by applicant]
Third-party submission for U.S. Appl. No. 18/329,138, dated Feb. 9, 2024, pp. 1-6. [cited by applicant]
Ahmad et al., “A Unique Approach in Text Independent Speaker Recognition using MFCC Feature Sets and Probabilistic Neural Network”, Eighth International Conference on Advances in Pattern Recognition (ICAPR}, IEEE, 2015 … [cited by applicant]
Almaadeed, et al., “Speaker identification using multimodal neural networks and wavelet analysis,” IET Biometrics 4.1 (2015), 18-28. [cited by applicant]
Atrey, et al., “Audio based event detection for multimedia surveillance”, Acoustics, Speech and Signal Processing, 2006, ICASSP 2006 Proceedings, 2006 IEEE International Conference on vol. 5, IEEE, 2006. pp. 813-816. [cited by applicant]
Baraniuk, “Compressive Sensing [Lecture Notes]”, IEEE Signal Processing Magazine, vol. 24, Jul. 2007, 4 pages. [cited by applicant]
Bredin, “TristouNet: Triplet Loss for Speaker Turn Embedding”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 14, 2016, XP080726602. [cited by applicant]
Buera et al., “Unsupervised Data-Driven Feature Vector Normalization With Acoustic Model Adaptation for Robust Speech Recognition”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 18, No. 2, Feb. 2010,… [cited by applicant]
Campbell, “Using Deep Belief Networks for Vector-Based Speaker Recognition”, Proceedings of Interspeech 2014, Sep. 14, 2014, pp. 676-680, XP055433784. [cited by applicant]
Castaldo et al., “Compensation of Nuisance Factors for Speaker and Language Recognition,” IEEE Transactions on Audio, Speech and Language Processing, ieeexplore.ieee.org, vol. 15, No. 7, Sep. 2007. pp. 1969-1978. [cited by applicant]
Communication pursuant to Article 94(3) EPC issued in EP Application No. 17 772 184.2-1207 dated Jul. 19, 2019. [cited by applicant]
Communication pursuant to Article 94(3) EPC on EP 17772184.2 dated Jun. 18, 2020. [cited by applicant]
Cumani, et al., “Factorized Sub-space Estimation for Fast and Memory Effective I-Vector Extraction”, IEEE/ACM TASLP, vol. 22, Issue 1, Jan. 2014, 28 pages. [cited by applicant]
D.Etter and C.Domniconi, “Multi2Rank: Multimedia Multiview Ranking,” 2015 IEEE International Conference on Multimedia Big Data, 2015, pp. 80-87. (Year 2015). [cited by applicant]
Dehak, et al., “Front-End Factor Analysis for Speaker Verification”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, No. 4, May 2011, pp. 788-798. [cited by applicant]
Douglas A. Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Examination Report for EP 17778046.7 dated Jun. 16, 2020 (4 pages). [cited by applicant]
Examination Report for IN 201947014575 dated Nov. 16, 2021 (6 pages). [cited by applicant]
Examination Report No. 1 for AU 2017322591 dated Jul. 16, 2021 (2 pages). [cited by applicant]
Final Office Action for U.S. Appl. No. 16/200,283 dated Jun. 11, 2020 (15 pages). [cited by applicant]
Final Office Action for U.S. Appl. No. 16/784,071 dated Oct. 27, 2020 (12 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 15/872,639 dated Jan. 29, 2019 (11 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 16/829,705 dated Mar. 24, 2022 (23 pages). [cited by applicant]
First Office Action issued in KR 10-2019-7010-7010208 dated Jun. 29, 2019. (5 pages). [cited by applicant]
First Office Action issued on CA Application No. 3,036,533 dated Apr. 12, 2019. (4 pages). [cited by applicant]
First Office Action on CA Application No. 3,075,049 dated May 7, 2020. (3 pages). [cited by applicant]
Florian et al., “FaceNet: A unified embedding for face recognition and clustering”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 7, 2015, pp. 815-823, XP032793492, DOI: 10.1109/CVPR… [cited by applicant]
Foreign Action on JP 2019-535198 dated Mar. 1, 2022 (6 pages). [cited by applicant]
Fu et al., “SNR-Aware Convolutional Neural Network Modeling for Speech Enhancement”, Interspeech 2016, Sep. 8-12, 2016, pp. 3768-3772, XP055427533, ISSN: 1990-9772, DOI: 10.21437/Interspeech.2016-211 (5 pages). [cited by applicant]
Gao, et al., “Dimensionality reduction via compressive sensing”, Pattern Recognition Letters 33, Elsevier Science BV 0167-8655, 2012. 8 pages. [cited by applicant]
Garcia-Romero et al., “Unsupervised Domain Adaptation for I-Vector Speaker Recgonition” Odyssey 2014, pp. 260-264. [cited by applicant]
Ghahabi Omid et al., “Restricted Boltzmann Machine Supervectors for Speaker Recognition”, 2015 IEEE International Conference on acoustics, Speech and Signal Processing (ICASSP), IEEE, Apr. 19, 2015, XP033187673, pp. 480… [cited by applicant]
Gish, et al., “Segregation of Speakers for Speech Recognition and Speaker Identification”, Acoustics, Speech, and Signal Processing, 1991, ICASSP-91, 1991 International Conference on IEEE, 1991. pp. 873-876. [cited by applicant]
Hoffer et al., “Deep Metric Learning Using Triplet Network”, 2015, arXiv: 1412.6622v3, retrieved Oct. 4, 2021 from URL: https://deepsense.ai/wp-content/uploads/2017/08/1412.6622-3.pdf (8 pages). [cited by applicant]
Hoffer et al., “Deep Metric Learning Using Triplet Network,” ICLR 2015 (workshop contribution), Mar. 23, 2015, pp. 1-8. [cited by applicant]
Huang, et al., “A Blind Segmentation Approach to Acoustic Event Detection Based on I-Vector”, Interspeech, 2013. pp. 2282-2286. [cited by applicant]
Information Disclosure Statement filed Aug. 8, 2019. [cited by applicant]
Information Disclosure Statement filed Feb. 12, 2020 (4 pages). [cited by applicant]
International Preliminary Report on Patentability for PCT/US2020/017051 dated Aug. 19, 2021 (11 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/039697 dated Jan. 1, 2019 (10 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052293 dated Mar. 19, 2019 (8 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052316 dated Mar. 19, 2019 (7 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052335 dated Mar. 19, 2019 (8 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2020/026992 dated Oct. 21, 2021 (10 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US20/17051 dated Apr. 23, 2020 (12 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US2017/052335 dated Dec. 8, 2017 (10 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US2020/24709 dated Jun. 19, 2020 (10 pages). [cited by applicant]
International Search Report and Written Opinion issued in corresponding International Application No. PCT/US2018/013965 with Date of mailing May 14, 2018. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US20/26992 with Date of mailing Jun. 26, 2020. [cited by applicant]
International Search Report and Written Opinion issued in the corresponding International Application No. PCT/US2017/039697, mailed on Sep. 20, 2017. 17 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority issued in corresponding International Application No. PCT/US2017/050927 with Date of mailing Dec. 11, 2017. [cited by applicant]
International Search Report and Written Opinion on PCT Appl. Ser. No. PCT/US2020/026992 dated Jun. 26, 2020 (11 pages). [cited by applicant]
International Search Report issued in corresponding International Application No. PCT/US2017/052293 dated Dec. 21, 2017. [cited by applicant]
International Search Report issued in corresponding International Application No. PCT/US2017/052316 with Date of mailing Dec. 21, 2017. 3 pages. [cited by applicant]
Kenny et al., “Deep Neural Networks for extracting Baum-Welch statistics for Speaker Recognition”, Jun. 29, 2014, XP055361192, Retrieved from the Internet: URL:http://www.crim.ca/perso/patrick.kenny/stafylakis_odyssey20… [cited by applicant]
Kenny, “A Small Footprint i-Vector Extractor” Proc. Odyssey Speaker and Language Recognition Workshop, Singapore, Jun. 25, 2012. 6 pages. [cited by applicant]
Khoury et al., “Combining Transcription-Based and Acoustic-Basedd Speaker Identifications for Broadcast News”, ICASSP, Kyoto, Japan, 2012, pp. 4377-4380. [cited by applicant]
Khoury, et al., “Improvised Speaker Diariztion System for Meetings”, Acoustics, Speech and Signal Processing, 2009, ICASSP 2009, IEEE International Conference on IEEE, 2009. pp. 4097-4100. [cited by applicant]
Kockmann et al., “Syllable based Feature-Contours for Speaker Recognition,” Proc. 14th International Workshop on Advances, 2008. 4 pages. [cited by applicant]
Korean Office Action (with English summary), dated Jun. 29, 2019, issued in Korean application No. 10-2019-7010208, 6 pages. [cited by applicant]
Lei et al., “A Novel Scheme for Speaker Recognition Using a Phonetically-aware Deep Neural Network”, Proceedings on ICASSP, Florence, Italy, IEEE Press, 2014, pp. 1695-1699. [cited by applicant]
Luque, et al., “Clustering Initialization Based on Spatial Information for Speaker Diarization of Meetings”, Ninth Annual Conference of the International Speech Communication Association, 2008. pp. 383-386. [cited by applicant]
McLaren, et al., “Advances in deep neural network approaches to speaker recognition,” In Proc. 40th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015. [cited by applicant]
McLaren, et al., “Exploring the Role of Phonetic Bottleneck Features for Speaker and Language Recognition”, 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 5575-557… [cited by applicant]
Meignier, et al., “Lium Spkdiarization: an Open Source Toolkit for Diarization” CMU SPUD Workshop, 2010. 7 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,232 dated Jun. 27, 2019 (11 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,232 dated Oct. 5, 2018 (21 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,290 dated Oct. 5, 2018 (12 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/200,283 dated Jan. 7, 2020 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/442,368 dated Oct. 4, 2019 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/784,071 dated May 12, 2020 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/829,705 dated Nov. 10, 2021 (20 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 15/610,378 dated Mar. 1, 2018 (11 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 16/551,327 dated Dec. 9, 2019 (7 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 15/872,639 dated Aug. 23, 2018 (10 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 16/829,705 dated Jul. 21, 2022 (27 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/610,378 dated Aug. 7, 2018 (5 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,024 dated Mar. 18, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,232 dated Feb. 6, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,232 dated Oct. 8, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,290 dated Feb. 6, 2019 (6 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/200,283 dated Aug. 24, 2020 (7 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/442,368 dated Feb. 5, 2020 (7 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/505,452 dated Jul. 23, 2020 (8 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/505,452 dated May 13, 2020 (9 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/784,071 dated Jan. 27, 2021 (14 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/262,748 dated Sep. 13, 2017 (9 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/818,231 dated Mar. 27, 2019 (9 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/872,639 dated Apr. 25, 2019 (5 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/551,327 DTD Mar. 26, 2020. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 17/107,496 dated Jan. 26, 2023 (8 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/536,293 dated Jun. 3, 2022 (9 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/895,750 dated Jan. 25, 2023 (6 pages). [cited by applicant]
Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, mailed May 14, 2018, in corresponding International Application No. PCT/US2018/013965, 13 … [cited by applicant]
Novoselov, et al., “SIC Speaker Recognition System for the NIST i-Vector Challenge.” Odyssey: The Speaker and Language Recognition Workshop. Jun. 16-19, 2014. pp. 231-240. [cited by applicant]
Office Action for CA 3036561 dated Jan. 23, 2020 (5 pages). [cited by applicant]
Oguzhan et al., “Recognition of Acoustic Events Using Deep Neural Networks”, 2014 22nd European Signal Processing Conference (EUSIPCO), Sep. 1, 2014, pp. 506-510 (5 pages). [cited by applicant]
Piegeon, et al., “Applying Logistic Regression to the Fusion of the NIST'99 1-Speaker Submissions”, Digital Signal Processing Oct. 1-3, 2000. pp. 237-248. [cited by applicant]
Prazak et al., “Speaker Diarization Using PLDA-based Speaker Clustering”, The 6th IEEE International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications, Sep. 2011, pp.… [cited by applicant]
Prince, et al., “Probabilistic Linear Discriminant Analysis for Inferences About Identity,” Proceedings of the International Conference on Computer Vision, Oct. 14-21, 2007. 8 pages. [cited by applicant]
Reasons for Refusal for JP 2019-535198 dated Sep. 10, 2021 (7 pages). [cited by applicant]
Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Richardson, et al., “Channel Compensation for Speaker Recognition using MAP Adapted PLDA and Denoising DNNs”, Proc. Speaker Lang. Recognit. Workshop, Jun. 22, 2016, pp. 225-230. [cited by applicant]
Richardson, et al., “Deep Neural Network Approaches to Speaker and Language Recognition”, IEEE Signal Processing Letters, vol. 22, No. 10, Oct. 2015, pp. 1671-1675. [cited by applicant]
Richardson, et al., “Speaker Recognition Using Real vs Synthetic Parallel Data for DNN Channel Compensation”, Interspeech, 2016. 6 pages. [cited by applicant]
Richardson, et al., Speaker Recognition Using Real vs Synthetic Parallel Data for DNN Channel Compensation, Interspeech, 2016, retrieved Sep. 14, 2021 from URL: https://www.ll.mit.edu/sites/default/files/publication/doc… [cited by applicant]
Rouvier et al., “An Open-source State-of-the-art Toolbox for Broadcast News Diarization”, Interspeech, Aug. 2013, pp. 1477-1481 (5 pages). [cited by applicant]
Scheffer et al., “Content matching for short duration speaker recognition”, Interspeech, Sep. 14-18, 2014, pp. 1317-1321. [cited by applicant]
Schmidt et al., “Large-Scale Speaker Identification” ICASSP, 2014, pp. 1650-1654. [cited by applicant]
Seddik, et al., “Text independent speaker recognition using the Mel frequency cepstral coefficients and a neural network classifier.” First International Symposium on Control, Communications and Signal Processing, 2004.… [cited by applicant]
Shajeesh, et al., “Speech Enhancement based on Savitzky-Golay Smoothing Filter”, International Journal of Computer Applications, vol. 57, No. 21, Nov. 2012, pp. 39-44 (6 pages). [cited by applicant]
Shum et al., “Exploiting Intra-Conversation Variability for Speaker Diarization”, Interspeech, Aug. 2011, pp. 945-948 (4 pages). [cited by applicant]
Snyder et al., “Time Delay Deep Nueral Network-Based Universal Background Models for Speaker Recognition”, IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2015, pp. 92-97 (6 pages). [cited by applicant]
Solomonoff, et al., “Nuisance Attribute Projection”, Speech Communication, Elsevier Science BV, Amsterdam, The Netherlands, May 1, 2007. 74 pages. [cited by applicant]
Sturim et al., “Speaker Linking and Applications Using Non-parametric Hashing Methods,” Interspeech, Sep. 2016, 5 pages. [cited by applicant]
Summons to attend oral proceedings pursuant to Rule 115(1) EPC issued in EP Application No. 17 772 184.2-1207 dated Dec. 16, 2019. [cited by applicant]
T.C.Nagavi, S.B. Anusha, P.Monisha and S.P.Poomima, “Content based audio retrieval with MFCC feature extraction, clustering and sort-merge techniquest,” 2013 Fourth International Conference on Computing, Communications … [cited by applicant]
Temko, et al., “Acoustic event detection in meeting-room environments”, Pattern Recognition Letters, vol. 30, No. 14, 2009, pp. 1281-1288. [cited by applicant]
Temko, et al., “Classification of acoustic events using SVM-based clustering schemes”, Pattern Recognition, vol. 39, No. 4, 2006, pp. 682-694. [cited by applicant]
US Non Final Office Action on U.S. Appl. No. 17/107,496 dated Jul. 21, 2022 (7 pages). [cited by applicant]
US Non-Final Office Action on U.S. Appl. No. 16/895,750 dated Oct. 6, 2022 (10 pages). [cited by applicant]
US Notice of Allowance on U.S. Appl. No. 17/107,496 dated Sep. 28, 2022 (8 pages). [cited by applicant]
Uzan et al., “I Know That Voice: Identifying the Voice Actor Behind the Voice”, 2015 International Conference on Biometrics (ICB), 2015, retrieved Oct. 4, 2021 from URL: https://citeseerx.ist.psu.edu/viewdoc/download?do… [cited by applicant]
Variani et al., “Deep neural networks for small footprint text-dependent speaker verification”, 2014 IEEE International Conference On Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2014, pp. 4052-4056, X… [cited by applicant]
Vella et al., “Artificial neural network features for speaker diarization”, 2014 IEEE Spoken Language Technology Workshop (SLT), IEEE Dec. 7, 2014, pp. 402-406, XP032756972, DOI: 10.1109/SLT.2014.7078608. [cited by applicant]
Wang et al., “Learning Fine-Grained Image Similarity with Deep Ranking”, Computer Vision and Pattern Recognition, Jan. 17, 2014, arXiv: 1404.4661v1, retrieved Oct. 4, 2021 from URL: https://arxiv.org/pdf/1404.4661.pdf (… [cited by applicant]
Xiang, et al., “Efficient text-independent speaker verification with structural Gaussian mixture models and neural network.” IEEE Transactions on Speech and Audio Processing 11.5 (2003): 447-456. [cited by applicant]
Xu et al., “Rapid Computation of I-Vector” Odyssey, Bilbao, Spain, Jun. 21-24, 2016. 6 pages. [cited by applicant]
Xue et al., “Fast Query by Example of Enviornmental Sounds via Robust and Efficient Cluster-Based Indexing”, Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2008, pp. 5-8 (4 pages). [cited by applicant]
Yaman et al., “Bottleneck Features for Speaker Recognition”, Proceedings of the Speaker and Language Recognition Workshop 2012, Jun. 28, 2012, pp. 105-108, XP055409424, Retrieved from the Internet: URL:https://pdfs.sema… [cited by applicant]
Yella Sree Harsha et al., “Artificial neural network features for speaker diarization”, 2014 IEEE Spoken Language technology Workshop (SLT), IEEE, Dec. 7, 2014, pp. 402-406, XP032756972, DOI: 10.1109/SLT.2014.7078608. [cited by applicant]
Zhang et al., “Extracting Deep Neural Network Bottleneck Features Using Low-Rank Matrix Factorization”, IEEE, ICASSP, 2014. 5 pages. [cited by applicant]
Zheng, et al., “An Experimental Study of Speech Emotion Recognition Based on Deep Convolutional Neural Networks”, 2015 International Conference on Affective Computing and Intelligent Interaction (ACII), 2015, pp. 827-83… [cited by applicant]
Navratil et al., “An Instantiable Speech Biometrics Module with Natural Language Interface: Implementation in the Telephony Environment,” 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing, a… [cited by applicant]