IP Library Granted Patent US 12,512,101
Granted Patent B2
US 12,512,101 · App. 17/963,091 · Granted Dec 30, 2025

End-to-end speaker recognition using deep neural network

Inventors: Elie Khoury (Atlanta, GA); Matthew Garland (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L17/08G06N3/04G06N3/08G10L15/16G10L17/02G10L17/04G10L17/18G10L17/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,512,101
App. No.
17/963,091
Granted
Dec 30, 2025
Kind
B2
Abstract

The present invention is directed to a deep neural network (DNN) having a triplet network architecture, which is suitable to perform speaker recognition. In particular, the DNN includes three feed-forward neural networks, which are trained according to a batch process utilizing a cohort set of negative training samples. After each batch of training samples is processed, the DNN may be trained according to a loss function, e.g., utilizing a cosine measure of similarity between respective samples, along with positive and negative margins, to provide a robust representation of voiceprints.

Claims (44)

1 . A computer-implemented method comprising:

obtaining, by the computer, one or more enrollment speech samples associated with an enrolled speaker;

executing, by the computer, a neural network on the one or more enrollment speech samples to generate an enrolled voiceprint for the enrolled speaker, wherein the neural network is trained according to positive speech samples and negative speech samples;

obtaining, by the computer, a recognition speech sample associated with a real-time speaker;

executing, by the computer, the neural network on the recognition speech sample to generate a real-time voiceprint for the real-time speaker;

generating, by the computer, a cosine distance between the enrolled voiceprint and the real-time voiceprint indicating an amount of similarity between the enrolled speaker and the real-time speaker; and

authenticating, by the computer, the real-time speaker as the enrolled speaker of a closed set of blacklisted speakers according to the cosine distance.

2 . The method according to claim 1 , further comprising identifying, by the computer, the real-time speaker is the enrolled speaker, in response to determining that the cosine distance satisfies a recognition threshold.

3 . The method according to claim 2 , wherein the enrolled voiceprint of the enrolled speaker is associated with the closed set of blacklisted speakers, and wherein the computer blocks a call of the real-time speaker.

4 . The method according to claim 1 , wherein the neural network is a portion of a triplet neural architecture.

5 . The method according to claim 1 , wherein the neural network is trained by reducing a loss function including a training cosine distance for dual sets of the positive speech samples and a cohort set of the negative speech samples.

6 . The method according to claim 1 , further comprising:

feeding, by the computer, a set of positive training speech samples into one or more feed-forward neural networks including the neural network to generate one or more corresponding embedding vectors;

calculating, by the computer, a loss function based upon the one or more embedding vectors; and

back-propagating, by the computer, the loss function to modify one or more connection weights in each of the feed-forward neural networks.

7 . The method according to claim 6 , wherein the loss function is based upon a positive cosine distance corresponding to a degree of similarity between a first embedding vector and a second embedding vector.

8 . The method according to claim 6 , further comprising initializing, by the computer, at least one feed-forward neural network with random connection weights.

9 . The method according to claim 1 , further comprising pre-processing, by the computer, the speech sample prior to executing the neural network, wherein the speech sample is at least one of an enrollment speech sample or real-time speech sample.

10 . The method according to claim 9 , further comprising:

segmenting, by the computer, the speech sample into windows of a predetermined duration with a predetermined window shift; and

extracting, by the computer, a set of features to be fed into the neural network from each window.

11 . A system comprising:

a non-transitory storage medium storing a plurality of computer program instructions; and

a computer comprising a processor electrically coupled to the non-transitory storage medium and configured to execute the plurality of computer program instructions to:

obtain one or more enrollment speech samples associated with an enrolled speaker;

execute a neural network on the one or more enrollment speech samples to generate an enrolled voiceprint for the enrolled speaker, wherein the neural network is trained according to positive speech samples and negative speech samples;

obtain a recognition speech sample associated with a real-time speaker;

execute the neural network on the recognition speech sample to generate a real-time voiceprint for the real-time speaker;

generate a cosine distance between the enrolled voiceprint and the real-time voiceprint indicating an amount of similarity between the enrolled speaker and the real-time speaker; and

authenticate the real-time speaker as the enrolled speaker of a closed set of blacklisted speakers according to the cosine distance.

12 . The system according to claim 11 , wherein the computer is further configured to identify the real-time speaker is the enrolled speaker, in response to determining that the cosine distance satisfies a recognition threshold.

13 . The system according to claim 12 , wherein the enrolled voiceprint of the enrolled speaker is associated with a closed set of blacklisted speakers, and wherein the computer blocks a call of the real-time speaker.

14 . The system according to claim 11 , wherein the neural network is a portion of a triplet neural architecture.

15 . The system according to claim 11 , wherein the neural network is trained by reducing a loss function including a training cosine distance for dual sets of the positive speech samples and a cohort set of the negative speech samples.

16 . The system according to claim 11 , wherein the computer is further configured to:

feed a set of positive training speech samples into one or more feed-forward neural networks including the neural network to generate one or more corresponding embedding vectors;

calculate a loss function based upon the one or more embedding vectors; and

back-propagate the loss function to modify one or more connection weights in each of the feed-forward neural networks.

17 . The system according to claim 16 , wherein the loss function is based upon a positive cosine distance corresponding to a degree of similarity between a first embedding vector and a second embedding vector.

18 . The system according to claim 16 , wherein the computer is further configured to initialize at least one feed-forward neural network with random connection weights.

19 . The system according to claim 11 , wherein the computer is further configured to pre-process the speech sample prior to executing the neural network, wherein the speech sample is at least one of an enrollment speech sample or real-time speech sample.

20 . The system according to claim 19 , wherein the computer is further configured to:

segment the speech sample into windows of a predetermined duration with a predetermined window shift; and

extract a set of features to be fed into the neural network from each window.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2022
From: KHOURY, ELIE; GARLAND, MATTHEW
To: PINDROP SECURITY, INC.
Reel/Frame 061368/0562 →
Continuity (4)
Continuation 16536293 · Aug 8, 2019
Continuation 15818231 · Nov 20, 2017
Continuation 15262748 · Sep 12, 2016
Related Publication 20230037232A1 · Feb 2, 2023
References Cited (346)
US 4817156A · Bahl et al. · 1989 [cited by applicant]
US 4829577A · Kuroda et al. · 1989 [cited by applicant]
US 4972485A · Dautrich et al. · 1990 [cited by applicant]
US 5072452A · Brown et al. · 1991 [cited by applicant]
US 5461697A · Nishimura et al. · 1995 [cited by applicant]
US 5475792A · Stanford et al. · 1995 [cited by applicant]
US 5598507A · Kimber et al. · 1997 [cited by applicant]
US 5659662A · Wilcox et al. · 1997 [cited by applicant]
US 5835890A · Matsui et al. · 1998 [cited by applicant]
US 5867562A · Scherer · 1999 [cited by applicant]
US 5949874A · Mark · 1999 [cited by applicant]
US 5995927A · Li · 1999 [cited by applicant]
US 6009392A · Kanevsky et al. · 1999 [cited by applicant]
US 6021119A · Derks et al. · 2000 [cited by applicant]
US 6055498A · Neumeyer et al. · 2000 [cited by applicant]
US 6094632A · Hattori · 2000 [cited by applicant]
US 6141644A · Kuhn et al. · 2000 [cited by applicant]
US 6411930B1 · Burges · 2002 [cited by applicant]
US 6463413B1 · Applebaum et al. · 2002 [cited by applicant]
US 6510415B1 · Talmor et al. · 2003 [cited by applicant]
US 6519561B1 · Farrell et al. · 2003 [cited by applicant]
US 6760701B2 · Sharma et al. · 2004 [cited by applicant]
US 6882972B2 · Kompe et al. · 2005 [cited by applicant]
US 6922668B1 · Downey · 2005 [cited by applicant]
US 6975708B1 · Scherer · 2005 [cited by applicant]
US 7003460B1 · Bub et al. · 2006 [cited by applicant]
US 7209881B2 · Yoshizawa et al. · 2007 [cited by applicant]
US 7295970B1 · Gorin et al. · 2007 [cited by applicant]
US 7318032B1 · Chaudhari et al. · 2008 [cited by applicant]
US 7324941B2 · Choi et al. · 2008 [cited by applicant]
US 7739114B1 · Chen et al. · 2010 [cited by applicant]
US 7813927B2 · Navratil et al. · 2010 [cited by applicant]
US 8046230B1 · Mcintosh · 2011 [cited by applicant]
US 8112160B2 · Foster · 2012 [cited by applicant]
US 8160811B2 · Prokhorov · 2012 [cited by applicant]
US 8160877B1 · Nucci et al. · 2012 [cited by applicant]
US 8385536B2 · Whitehead · 2013 [cited by applicant]
US 8484023B2 · Kanevsky et al. · 2013 [cited by applicant]
US 8484024B2 · Kanevsky et al. · 2013 [cited by applicant]
US 8554563B2 · Aronowitz · 2013 [cited by applicant]
US 8712760B2 · Hsia et al. · 2014 [cited by applicant]
US 8738442B1 · Liu et al. · 2014 [cited by applicant]
US 8856895B2 · Perrot · 2014 [cited by applicant]
US 8886663B2 · Gainsboro et al. · 2014 [cited by applicant]
US 8903859B2 · Zeppenfeld et al. · 2014 [cited by applicant]
US 9042867B2 · Gomar · 2015 [cited by applicant]
US 9064491B2 · Rachevsky et al. · 2015 [cited by applicant]
US 9237232B1 · Williams · 2016 [cited by examiner]
US 9277049B1 · Danis · 2016 [cited by applicant]
US 9336781B2 · Scheffer et al. · 2016 [cited by applicant]
US 9338619B2 · Kang · 2016 [cited by applicant]
US 9343067B2 · Ariyaeeinia et al. · 2016 [cited by applicant]
US 9344892B1 · Rodrigues et al. · 2016 [cited by applicant]
US 9355646B2 · Oh et al. · 2016 [cited by applicant]
US 9373330B2 · Cumani et al. · 2016 [cited by applicant]
US 9401143B2 · Senior et al. · 2016 [cited by applicant]
US 9401148B2 · Lei et al. · 2016 [cited by applicant]
US 9406298B2 · Cumani et al. · 2016 [cited by applicant]
US 9431016B2 · Aviles-Casco et al. · 2016 [cited by applicant]
US 9444839B1 · Faulkner et al. · 2016 [cited by applicant]
US 9454958B2 · Li et al. · 2016 [cited by applicant]
US 9460722B2 · Sidi et al. · 2016 [cited by applicant]
US 9466292B1 · Lei et al. · 2016 [cited by applicant]
US 9502038B2 · Wang et al. · 2016 [cited by applicant]
US 9514753B2 · Sharifi et al. · 2016 [cited by applicant]
US 9558755B1 · Laroche et al. · 2017 [cited by applicant]
US 9584946B1 · Lyren et al. · 2017 [cited by applicant]
US 9620145B2 · Bacchiani et al. · 2017 [cited by applicant]
US 9626971B2 · Rodriguez et al. · 2017 [cited by applicant]
US 9633652B2 · Kurniawati et al. · 2017 [cited by applicant]
US 9641954B1 · Typrin et al. · 2017 [cited by applicant]
US 9665823B2 · Saon et al. · 2017 [cited by applicant]
US 9685174B2 · Karam et al. · 2017 [cited by applicant]
US 9729727B1 · Zhang · 2017 [cited by applicant]
US 9818431B2 · Yu · 2017 [cited by applicant]
US 9824692B1 · Khoury et al. · 2017 [cited by applicant]
US 9860367B1 · Jiang et al. · 2018 [cited by applicant]
US 9875739B2 · Ziv et al. · 2018 [cited by applicant]
US 9875742B2 · Gorodetski et al. · 2018 [cited by applicant]
US 9875743B2 · Gorodetski et al. · 2018 [cited by applicant]
US 9881617B2 · Sidi et al. · 2018 [cited by applicant]
US 9930088B1 · Hodge · 2018 [cited by applicant]
US 9984706B2 · Wein · 2018 [cited by applicant]
US 10277628B1 · Jakobsson · 2019 [cited by applicant]
US 10325601B2 · Khoury et al. · 2019 [cited by applicant]
US 10347256B2 · Khoury et al. · 2019 [cited by applicant]
US 10397398B2 · Gupta · 2019 [cited by applicant]
US 10404847B1 · Unger · 2019 [cited by applicant]
US 10462292B1 · Stephens · 2019 [cited by applicant]
US 10506088B1 · Singh · 2019 [cited by applicant]
US 10554821B1 · Koster · 2020 [cited by applicant]
US 10630682B1 · Bhattacharyya et al. · 2020 [cited by applicant]
US 10638214B1 · Delhoume et al. · 2020 [cited by applicant]
US 10659605B1 · Braundmeier et al. · 2020 [cited by applicant]
US 10854205B2 · Khoury et al. · 2020 [cited by applicant]
US 11069352B1 · Tang et al. · 2021 [cited by applicant]
US 11659082B2 · Gupta · 2023 [cited by applicant]
US 20020095287A1 · Botterweck · 2002 [cited by applicant]
US 20020143539A1 · Botterweck · 2002 [cited by applicant]
US 20030231775A1 · Wark · 2003 [cited by applicant]
US 20030236663A1 · Dimitrova et al. · 2003 [cited by applicant]
US 20040218751A1 · Colson et al. · 2004 [cited by applicant]
US 20040230420A1 · Kadambe et al. · 2004 [cited by applicant]
US 20050038655A1 · Mutel et al. · 2005 [cited by applicant]
US 20050039056A1 · Bagga et al. · 2005 [cited by applicant]
US 20050286688A1 · Scherer · 2005 [cited by applicant]
US 20060058998A1 · Yamamoto et al. · 2006 [cited by applicant]
US 20060111905A1 · Navratil et al. · 2006 [cited by applicant]
US 20070189479A1 · Scherer · 2007 [cited by applicant]
US 20070198257A1 · Zhang et al. · 2007 [cited by applicant]
US 20070294083A1 · Bellegarda et al. · 2007 [cited by applicant]
US 20080195389A1 · Zhang et al. · 2008 [cited by applicant]
US 20080312926A1 · Vair et al. · 2008 [cited by applicant]
US 20090138712A1 · Driscoll · 2009 [cited by applicant]
US 20090265328A1 · Parekh et al. · 2009 [cited by applicant]
US 20100131273A1 · Aley-Raz et al. · 2010 [cited by applicant]
US 20100217589A1 · Gruhn et al. · 2010 [cited by applicant]
US 20100232619A1 · Uhle et al. · 2010 [cited by applicant]
US 20100262423A1 · Huo et al. · 2010 [cited by applicant]
US 20110010173A1 · Scott et al. · 2011 [cited by applicant]
US 20120173239A1 · Sanchez Asenjo et al. · 2012 [cited by applicant]
US 20120185418A1 · Capman et al. · 2012 [cited by applicant]
US 20130041660A1 · Waite · 2013 [cited by applicant]
US 20130080165A1 · Wang et al. · 2013 [cited by applicant]
US 20130109358A1 · Balasubramaniyan et al. · 2013 [cited by applicant]
US 20130300939A1 · Chou et al. · 2013 [cited by applicant]
US 20140046878A1 · Lecomte et al. · 2014 [cited by applicant]
US 20140053247A1 · Fadel · 2014 [cited by applicant]
US 20140081640A1 · Farrell et al. · 2014 [cited by applicant]
US 20140136194A1 · Warford · 2014 [cited by examiner]
US 20140195236A1 · Hosom et al. · 2014 [cited by applicant]
US 20140214417A1 · Wang · 2014 [cited by examiner]
US 20140214676A1 · Bukai · 2014 [cited by applicant]
US 20140241513A1 · Springer · 2014 [cited by applicant]
US 20140250512A1 · Goldstone et al. · 2014 [cited by applicant]
US 20140278412A1 · Scheffer et al. · 2014 [cited by applicant]
US 20140288928A1 · Penn et al. · 2014 [cited by applicant]
US 20140337017A1 · Watanabe et al. · 2014 [cited by applicant]
US 20150036813A1 · Ananthakrishnan et al. · 2015 [cited by applicant]
US 20150127336A1 · Lei et al. · 2015 [cited by applicant]
US 20150149165A1 · Saon · 2015 [cited by applicant]
US 20150161522A1 · Saon et al. · 2015 [cited by applicant]
US 20150189086A1 · Romano et al. · 2015 [cited by applicant]
US 20150199960A1 · Huo et al. · 2015 [cited by applicant]
US 20150269931A1 · Senior et al. · 2015 [cited by applicant]
US 20150269941A1 · Jones · 2015 [cited by applicant]
US 20150310008A1 · Thudor et al. · 2015 [cited by applicant]
US 20150334231A1 · Rybak et al. · 2015 [cited by applicant]
US 20150348571A1 · Koshinaka et al. · 2015 [cited by applicant]
US 20150356630A1 · Hussain · 2015 [cited by applicant]
US 20150365530A1 · Kolbegger et al. · 2015 [cited by applicant]
US 20160019458A1 · Kaufhold · 2016 [cited by applicant]
US 20160019883A1 · Aronowitz · 2016 [cited by applicant]
US 20160028434A1 · Kerpez et al. · 2016 [cited by applicant]
US 20160048846A1 · Douglas et al. · 2016 [cited by applicant]
US 20160078863A1 · Chung et al. · 2016 [cited by applicant]
US 20160104480A1 · Sharifi · 2016 [cited by applicant]
US 20160125877A1 · Foerster et al. · 2016 [cited by applicant]
US 20160180214A1 · Kanevsky et al. · 2016 [cited by applicant]
US 20160189707A1 · Donjon et al. · 2016 [cited by applicant]
US 20160240190A1 · Lee et al. · 2016 [cited by applicant]
US 20160275953A1 · Sharifi et al. · 2016 [cited by applicant]
US 20160284346A1 · Visser et al. · 2016 [cited by applicant]
US 20160293167A1 · Chen · 2016 [cited by examiner]
US 20160314790A1 · Tsujikawa et al. · 2016 [cited by applicant]
US 20160343373A1 · Ziv et al. · 2016 [cited by applicant]
US 20170060779A1 · Falk · 2017 [cited by applicant]
US 20170069313A1 · Aronowitz · 2017 [cited by applicant]
US 20170069327A1 · Heigold · 2017 [cited by examiner]
US 20170098444A1 · Song · 2017 [cited by applicant]
US 20170111515A1 · Bandyopadhyay et al. · 2017 [cited by applicant]
US 20170126884A1 · Balasubramaniyan et al. · 2017 [cited by applicant]
US 20170142150A1 · Sandke et al. · 2017 [cited by applicant]
US 20170169816A1 · Blandin et al. · 2017 [cited by applicant]
US 20170230390A1 · Faulkner et al. · 2017 [cited by applicant]
US 20170262837A1 · Gosalia · 2017 [cited by applicant]
US 20180082691A1 · Khoury et al. · 2018 [cited by applicant]
US 20180152558A1 · Chan et al. · 2018 [cited by applicant]
US 20180249006A1 · Dowlatkhah et al. · 2018 [cited by applicant]
US 20180268023A1 · Korpusik et al. · 2018 [cited by applicant]
US 20180295235A1 · Tatourian et al. · 2018 [cited by applicant]
US 20180337962A1 · Ly et al. · 2018 [cited by applicant]
US 20190037081A1 · Rao et al. · 2019 [cited by applicant]
US 20190172476A1 · Wung et al. · 2019 [cited by applicant]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190180758A1 · Washio · 2019 [cited by applicant]
US 20190238956A1 · Gaubitch et al. · 2019 [cited by applicant]
US 20190297503A1 · Traynor et al. · 2019 [cited by applicant]
US 20200035224A1 · Ward et al. · 2020 [cited by applicant]
US 20200137221A1 · Dellostritto et al. · 2020 [cited by applicant]
US 20200195779A1 · Weisman et al. · 2020 [cited by applicant]
US 20200252510A1 · Ghuge et al. · 2020 [cited by applicant]
US 20200312313A1 · Maddali et al. · 2020 [cited by applicant]
US 20200396332A1 · Gayaldo · 2020 [cited by applicant]
US 20210084147A1 · Kent et al. · 2021 [cited by applicant]
US 20210150010A1 · Strong et al. · 2021 [cited by applicant]
JP 2013005205A · 2013 [cited by applicant]
WO WO2015079885 · 2015 [cited by applicant]
WO WO2016195261A1 · 2016 [cited by applicant]
WO WO2017167900A1 · 2017 [cited by applicant]
Deep Metric Learning Using Triplet Network) by Hoffer et al. (Year: 2015). [cited by examiner]
Deep Neural Network Approaches to Speaker and Language Recognition by Richardson et al. (Year: 2014). [cited by examiner]
Anonymous: “Filter bank—Wikipedia”, Apr. 3, 2024 (Apr. 3, 2024), XP093147551, Retrieved from the Internet: URL:https://en.wikipedia.org/wiki/Filter_bank [retrieved on Apr. 3, 2024]. [cited by applicant]
Database Inspec [Online] The Institution of Electrical Engineers, Stevenage, GB; Sep. 10, 2015 (Sep. 10, 2015), Majeed SA et al: “Mel Frequency Cepstral Coefficients (MFCC) Feature Extraction Enhancement In The Applicat… [cited by applicant]
Examiner's Report on EPO App. 20786776.3 dated Apr. 12, 2024 (6 pages). [cited by applicant]
Sivakuntaran P et al: “The use of sub-band cepstrum in speaker verification”, Acoustics, Speech, and Signal Processing, 2000. ICASSP '00. Proceedings. 2000 IEEE International Conference on Jun. 5-9, 2000, Piscataway, NJ… [cited by applicant]
Ahmad et al., “A Unique Approach in Text Independent Speaker Recognition using MFCC Feature Sets and Probabilistic Neural Network”, Eighth International Conference on Advances in Pattern Recognition (ICAPR}, IEEE, 2015 … [cited by applicant]
Almaadeed, et al., “Speaker identification using multimodal neural networks and wavelet analysis,” IET Biometrics 4.1 (2015), 18-28. [cited by applicant]
Atrey, et al., “Audio based event detection for multimedia surveillance”, Acoustics, Speech and Signal Processing, 2006, ICASSP 2006 Proceedings, 2006 IEEE International Conference on vol. 5, IEEE, 2006. pp. 813-816. [cited by applicant]
Australian Examination Report on AU Appl. Ser. No. 2020271801 dated Jul. 25, 2022 (4 pages). [cited by applicant]
Baraniuk, “Compressive Sensing [Lecture Notes]”, IEEE Signal Processing Magazine, vol. 24, Jul. 2007, 4 pages. [cited by applicant]
Bredin, “TristouNet: Triplet Loss for Speaker Turn Embedding”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 14, 2016, XP080726602. [cited by applicant]
Castaldo et al., “Compensation of Nuisance Factors for Speaker and Language Recognition,” IEEE Transactions on Audio, Speech and Language Processing, ieeexplore.ieee.org, vol. 15, No. 7, Sep. 2007. pp. 1969-1978. [cited by applicant]
Communication pursuant to Article 94(3) EPC issued in EP Application No. 17 772 184.2-1207 dated Jul. 19, 2019. [cited by applicant]
Communication pursuant to Article 94(3) EPC on EP 17772184.2 dated Jun. 18, 2020. [cited by applicant]
Cumani, et al., “Factorized Sub-space Estimation for Fast and Memory Effective I-Vector Extraction”, IEEE/ACM TASLP, vol. 22, Issue 1, Jan. 2014, 28 pages. [cited by applicant]
Dehak, et al., “Front-End Factor Analysis for Speaker Verification”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, No. 4, May 2011, pp. 788-798. [cited by applicant]
Douglas A. Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Examination Report for IN 201947014575 dated Nov. 16, 2021 (6 pages). [cited by applicant]
Examination Report No. 1 for AU 2017322591 dated Jul. 16, 2021 (2 pages). [cited by applicant]
Final Office Action for U.S. Appl. No. 16/200,283 dated Jun. 11, 2020 (15 pages). [cited by applicant]
Final Office Action for U.S. Appl. No. 16/784,071 dated Oct. 27, 2020 (12 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 15/872,639 dated Jan. 29, 2019 (11 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 16/829,705 dated Mar. 24, 2022 (23 pages). [cited by applicant]
First Office Action issued in KR 10-2019-7010-7010208 dated Jun. 29, 2019. (5 pages). [cited by applicant]
First Office Action issued on CA Application No. 3,036,533 dated Apr. 12, 2019. (4 pages). [cited by applicant]
First Office Action on CA Application No. 3,075,049 dated May 7, 2020. (3 pages). [cited by applicant]
Florian et al., “FaceNet: A unified embedding for face recognition and clustering”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 7, 2015, pp. 815-823, XP032793492, DOI: 10.1109/CVPR… [cited by applicant]
Foreign Action on JP 2019-535198 dated Mar. 1, 2022 (6 pages). [cited by applicant]
Fu et al., “SNR-Aware Convolutional Neural Network Modeling for Speech Enhancement”, INTERSPEECH 2016, Sep. 8-12, 2016, pp. 3768-3772, XP055427533, ISSN: 1990-9772, DOI: 10.21437/Interspeech.2016-211 (5 pages). [cited by applicant]
Gao, et al., “Dimensionality reduction via compressive sensing”, Pattern Recognition Letters 33, Elsevier Science BV 0167-8655, 2012. 8 pages. [cited by applicant]
Garcia-Romero et al., “Unsupervised Domain Adaptation for I-Vector Speaker Recgonition” Odyssey 2014, pp. 260-264. [cited by applicant]
Ghahabi Omid et al., “Restricted Boltzmann Machine Supervectors for Speaker Recognition”, 2015 IEEE International Conference on acoustics, Speech and Signal Processing (ICASSP), IEEE, Apr. 19, 2015, XP033187673, pp. 480… [cited by applicant]
Gish, et al., “Segregation of Speakers for Speech Recognition and Speaker Identification”, Acoustics, Speech, and Signal Processing, 1991, ICASSP-91, 1991 International Conference on IEEE, 1991. pp. 873-876. [cited by applicant]
Hoffer et al., “Deep Metric Learning Using Triplet Network”, 2015, arXiv: 1412.6622v3, retrieved Oct. 4, 2021 from URL: https://deepsense.ai/wp-content/uploads/2017/08/1412.6622-3.pdf (8 pages). [cited by applicant]
Hoffer et al., “Deep Metric Learning Using Triplet Network,” ICLR 2015 (workshop contribution), Mar. 23, 2015, pp. 1-8. [cited by applicant]
Huang, et al., “A Blind Segmentation Approach to Acoustic Event Detection Based on I-Vector”, INTERSPEECH, 2013. pp. 2282-2286. [cited by applicant]
Information Disclosure Statement filed Aug. 8, 2019. [cited by applicant]
International Preliminary Report on Patentability for PCT/US2020/017051 dated Aug. 19, 2021 (11 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/039697 dated Jan. 1, 2019 (10 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052293 dated Mar. 19, 2019 (8 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052335 dated Mar. 19, 2019 (8 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2020/026992 dated Oct. 21, 2021 (10 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US20/17051 dated Apr. 23, 2020 (12 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US2017/052335 dated Dec. 8, 2017 (10 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US2020/24709 dated Jun. 19, 2020 (10 pages). [cited by applicant]
International Search Report and Written Opinion issued in corresponding International Application No. PCT/US2018/013965 with Date of mailing May 14, 2018. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US20/26992 with Date of mailing Jun. 26, 2020. [cited by applicant]
International Search Report and Written Opinion issued in the corresponding International Application No. PCT/US2017/039697, mailed on Sep. 20, 2017. 17 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority issued in corresponding International Application No. PCT/US2017/050927 with Date of mailing Dec. 11, 2017. [cited by applicant]
International Search Report and Written Opinion on PCT Appl. Ser. No. PCT/US2020/026992 dated Jun. 26, 2020 (11 pages). [cited by applicant]
International Search Report issued in corresponding International Application No. PCT/US2017/052293 dated Dec. 21, 2017. [cited by applicant]
International Search Report issued in corresponding International Application No. PCT/US2017/052316 with Date of mailing Dec. 21, 2017. 3 pages. [cited by applicant]
Kenny et al., “Deep Neural Networks for extracting Baum-Welch statistics for Speaker Recognition”, Jun. 29, 2014, XP055361192, Retrieved from the Internet: URL:http://www.crim.ca/perso/patrick.kenny/stafylakis_odyssey20… [cited by applicant]
Kenny, “A Small Footprint i-Vector Extractor” Proc. Odyssey Speaker and Language Recognition Workshop, Singapore, Jun. 25, 2012. 6 pages. [cited by applicant]
Khoury et al., “Combining Transcription-Based and Acoustic-Basedd Speaker Identifications for Broadcast News”, Icassp, Kyoto, Japan, 2012, pp. 4377-4380. [cited by applicant]
Khoury, et al., “Improvised Speaker Diariztion System for Meetings”, Acoustics, Speech and Signal Processing, 2009, ICASSP 2009, IEEE International Conference on IEEE, 2009. pp. 4097-4100. [cited by applicant]
Kockmann et al., “Syllable based Feature-Contours for Speaker Recognition,” Proc. 14th International Workshop on Advances, 2008. 4 pages. [cited by applicant]
Korean Office Action (with English summary), dated Jun. 29, 2019, issued in Korean application No. 10-2019-7010208, 6 pages. [cited by applicant]
Lei et al., “A Novel Scheme for Speaker Recognition Using a Phonetically-aware Deep Neural Network”, Proceedings on ICASSP, Florence, Italy, IEEE Press, 2014, pp. 1695-1699. [cited by applicant]
Luque, et al., “Clustering Initialization Based on Spatial Information for Speaker Diarization of Meetings”, Ninth Annual Conference of the International Speech Communication Association, 2008. pp. 383-386. [cited by applicant]
McLaren, et al., “Advances in deep neural network approaches to speaker recognition,” In Proc. 40th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015. [cited by applicant]
McLaren, et al., “Exploring the Role of Phonetic Bottleneck Features for Speaker and Language Recognition”, 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 5575-557… [cited by applicant]
Meignier, et al., “Lium Spkdiarization: An Open Source Toolkit for Diarization” CMU SPUD Workshop, 2010. 7 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,290 dated Oct. 5, 2018 (12 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/200,283 dated Jan. 7, 2020 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/442,368 dated Oct. 4, 2019 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/784,071 dated May 12, 2020 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/829,705 dated Nov. 10, 2021 (20 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 15/610,378 dated Mar. 1, 2018. [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 16/551,327 dated Dec. 9, 2019. [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 15/872,639 dated Aug. 23, 2018 (10 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 16/983,967 dated Jun. 30, 2022 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/610,378 dated Aug. 7, 2018 (5 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,290 dated Feb. 6, 2019 (6 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/200,283 dated Aug. 24, 2020 (7 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/442,368 dated Feb. 5, 2020 (7 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/784,071 dated Jan. 27, 2021 (14 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/317,575 dated Jan. 28, 2022 (11 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/262,748 dated Sep. 13, 2017. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/818,231 dated Mar. 27, 2019. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/872,639 dated Apr. 25, 2019. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/551,327 DTD Mar. 26, 2020. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/536,293 dated Jun. 3, 2022 (9 pages). [cited by applicant]
Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, mailed May 14, 2018, in corresponding International Application No. PCT/US2018/013965, 13 … [cited by applicant]
Novoselov, et al., “SIC Speaker Recognition System for the NIST i-Vector Challenge.” Odyssey: The Speaker and Language Recognition Workshop. Jun. 16-19, 2014. pp. 231-240. [cited by applicant]
Oguzhan et al., “Recognition of Acoustic Events Using Deep Neural Networks”, 2014 22nd European Signal Processing Conference (EUSiPCO), Sep. 1, 2014, pp. 506-510 (5 pages). [cited by applicant]
Piegeon, et al., “Applying Logistic Regression to the Fusion of the NIST'99 1-Speaker Submissions”, Digital Signal Processing Oct. 1-3, 2000. pp. 237-248. [cited by applicant]
Prazak et al., “Speaker Diarization Using PLDA-based Speaker Clustering”, The 6th IEEE International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications, Sep. 2011, pp.… [cited by applicant]
Prince, et al., “Probabilistic Linear Discriminant Analysis for Inferences About Identity,” Proceedings of the International Conference on Computer Vision, Oct. 14-21, 2007. 8 pages. [cited by applicant]
Reasons for Refusal for JP 2019-535198 dated Sep. 10, 2021 (7 pages). [cited by applicant]
Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Richardson, et al., “Channel Compensation for Speaker Recognition using MAP Adapted PLDA and Denoising DNNs”, Proc. Speaker Lang. Recognit. Workshop, Jun. 22, 2016, pp. 225-230. [cited by applicant]
Richardson, et al., “Deep Neural Network Approaches to Speaker and Language Recognition”, IEEE Signal Processing Letters, vol. 22, No. 10, Oct. 2015, pp. 1671-1675. [cited by applicant]
Richardson, et al., “Speaker Recognition Using Real vs Synthetic Parallel Data for DNN Channel Compensation”, INTERSPEECH, 2016. 6 pages. [cited by applicant]
Richardson, et al., Speaker Recognition Using Real vs Synthetic Parallel Data for DNN Channel Compensation, INTERSPEECH, 2016, retrieved Sep. 14, 2021 from URL: https://www.ll.mit.edu/sites/default/files/publication/doc… [cited by applicant]
Rouvier et al., “An Open-source State-of-the-art Toolbox for Broadcast News Diarization”, INTERSPEECH, Aug. 2013, pp. 1477-1481 (5 pages). [cited by applicant]
Scheffer et al., “Content matching for short duration speaker recognition”, INTERSPEECH, Sep. 14-18, 2014, pp. 1317-1321. [cited by applicant]
Schmidt et al., “Large-Scale Speaker Identification” ICASSP, 2014, pp. 1650-1654. [cited by applicant]
Seddik, et al., “Text independent speaker recognition using the Mel frequency cepstral coefficients and a neural network classifier.” First International Symposium on Control, Communications and Signal Processing, 2004.… [cited by applicant]
Shajeesh, et al., “Speech Enhancement based on Savitzky-Golay Smoothing Filter”, International Journal of Computer Applications, vol. 57, No. 21, Nov. 2012, pp. 39-44 (6 pages). [cited by applicant]
Shum et al., “Exploiting Intra-Conversation Variability for Speaker Diarization”, INTERSPEECH, Aug. 2011, pp. 945-948 (4 pages). [cited by applicant]
Snyder et al., “Time Delay Deep Nueral Network-Based Universal Background Models for Speaker Recognition”, IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2015, pp. 92-97 (6 pages). [cited by applicant]
Solomonoff, et al., “Nuisance Attribute Projection”, Speech Communication, Elsevier Science BV, Amsterdam, The Netherlands, May 1, 2007. 74 pages. [cited by applicant]
Sturim et al., “Speaker Linking and Applications Using Non-parametric Hashing Methods,” Interspeech, Sep. 2016, 5 pages. [cited by applicant]
Summons to attend oral proceedings pursuant to Rule 115(1) EPC issued in EP Application No. 17 772 184.2-1207 dated Dec. 16, 2019. [cited by applicant]
Temko, et al., “Acoustic event detection in meeting-room environments”, Pattern Recognition Letters, vol. 30, No. 14, 2009, pp. 1281-1288. [cited by applicant]
Temko, et al., “Classification of acoustic events using SVM-based clustering schemes”, Pattern Recognition, vol. 39, No. 4, 2006, pp. 682-694. [cited by applicant]
Uzan et al., “I Know That Voice: Identifying the Voice Actor Behind the Voice”, 2015 International Conference on Biometrics (ICB), 2015, retrieved Oct. 4, 2021 from URL: https://citeseerx.ist.psu.edu/viewdoc/download?do… [cited by applicant]
Vella et al., “Artificial neural network features for speaker diarization”, 2014 IEEE Spoken Language Technology Workshop (SLT), IEEE Dec. 7, 2014, pp. 402-406, XP032756972, DOI: 10.1109/SLT.2014.7078608. [cited by applicant]
Wang et al., “Learning Fine-Grained Image Similarity with Deep Ranking”, Computer Vision and Pattern Recognition, Jan. 17, 2014, arXiv: 1404.4661v1, retrieved Oct. 4, 2021 from URL: https://arxiv.org/pdf/1404.4661.pdf (… [cited by applicant]
Xiang, et al., “Efficient text-independent speaker verification with structural Gaussian mixture models and neural network.” IEEE Transactions on Speech and Audio Processing 11.5 (2003): 447-456. [cited by applicant]
Xu et al., “Rapid Computation of I-Vector” Odyssey, Bilbao, Spain, Jun. 21-24, 2016. 6 pages. [cited by applicant]
Xue et al., “Fast Query By Example of Enviornmental Sounds Via Robust and Efficient Cluster-Based Indexing”, Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2008, pp. 5-8 (4 pages). [cited by applicant]
Yella Sree Harsha et al., “Artificial neural network features for speaker diarization”, 2014 IEEE Spoken Language technology Workshop (SLT), IEEE, Dec. 7, 2014, pp. 402-406, XP032756972, DOI: 10.1109/SLT.2014.7078608. [cited by applicant]
Zhang et al., “Extracting Deep Neural Network Bottleneck Features Using Low-Rank Matrix Factorization”, IEEE, ICASSP, 2014. 5 pages. [cited by applicant]
Zheng, et al., “An Experimental Study of Speech Emotion Recognition Based on Deep Convolutional Neural Networks”, 2015 International Conference on Affective Computing and Intelligent Interaction (ACII), 2015, pp. 827-83… [cited by applicant]
Shun-ichi Amari, (“Backpropagation and stochastic gradient descent method”) (Year: 1992). [cited by applicant]
Bengio et al., “Word Embeddings for Speech Recognition”, Published in Interspeech in 2014, pp. 1053-1057. [cited by applicant]
Notice of Refusal Office Action issued in Japanese Application No. 2022-104204 dated Sep. 19, 2024 (4 Pages). [cited by applicant]
Buera et al., “Unsupervised Data-Driven Feature Vector Normalization With Acoustic Model Adaptation for Robust Speech Recognition”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 18, No. 2, Feb. 2010,… [cited by applicant]
Campbell, “Using Deep Belief Networks for Vector-Based Speaker Recognition”, Proceedings of Interspeech 2014, Sep. 14, 2014, pp. 676-680, XP055433784. [cited by applicant]
Examination Report for EP 17778046.7 dated Jun. 16, 2020 (4 pages). [cited by applicant]
First Examiners Requisition on CA Appl. 3,096,378 dated Jun. 12, 2023 (3 pages). [cited by applicant]
Information Disclosure Statement filed Feb. 12, 2020 (4 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052316 dated Mar. 19, 2019 (7 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,232 dated Jun. 27, 2019 (11 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,232 dated Oct. 5, 2018 (21 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,024 dated Mar. 18, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,232 dated Feb. 6, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,232 dated Oct. 8, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/505,452 dated Jul. 23, 2020 (8 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/505,452 dated May 13, 2020 (9 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 17/107,496 dated Jan. 26, 2023 (8 pages). [cited by applicant]
Office Action for CA 3036561 dated Jan. 23, 2020 (5 pages). [cited by applicant]
Office Action on Japanese Application 2022-104204 dated Jul. 26, 2023 (4 pages). [cited by applicant]
US Non Final Office Action on U.S. Appl. No. 17/107,496 dated Jul. 21, 2022 (7 pages). [cited by applicant]
US Notice of Allowance on U.S. Appl. No. 17/107,496 dated Sep. 28, 2022 (8 pages). [cited by applicant]
Variani et al., “Deep neural networks for small footprint text-dependent speaker verification”, 2014 IEEE International Conference On Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2014, pp. 4052-4056, X… [cited by applicant]
Yaman et al., “Bottleneck Features for Speaker Recognition”, Proceedings of the Speaker and Language Recognition Workshop 2012, Jun. 28, 2012, pp. 105-108, XP055409424, Retrieved from the Internet: URL:https://pdfs.sema… [cited by applicant]
Chen et al., “Res Net and Model Fusion for Automatic Spoofing Detection.” Interspeech. (Year: 2017). [cited by applicant]
Dinkel et al., “End-to-end spoofing detection with raw waveform CLDNNS.” 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE (Year: 2017). [cited by applicant]
First Examiner's Requisition for CA App. 3,135,210 dated Oct. 5, 2023 (5 pages). [cited by applicant]
Ravanelli et al., “Speaker recognition from raw waveform with sincnet.” 2018 IEEE spoken language technology workshop (SLT). IEEE (Year: 2018). [cited by applicant]
Office Action from the Japanese Office Action on App. 2022-104204 dated Jan. 6, 2024 (6 pages). [cited by applicant]
Office Action issued in Japanese Patent Application No. 2025-003657 dated Oct. 22, 2025. [cited by applicant]