IP Library Granted Patent US 12,711,960
Granted Patent B2
US 12,711,960 · App. 18/329,138 · Granted Aug 18, 2026

Speaker recognition in the call center

Inventors: Elie Khoury (Atlanta, GA); Matthew Garland (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L17/00G06N7/01G10L15/07G10L15/19G10L15/26G10L17/04G10L17/08G10L17/24H04M1/271H04M2203/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,960
App. No.
18/329,138
Filed
Jun 5, 2023
Granted
Aug 18, 2026
Kind
B2
Art Unit
2656
USPC
704/246
Abstract

Utterances of at least two speakers in a speech signal may be distinguished and the associated speaker identified by use of diarization together with automatic speech recognition of identifying words and phrases commonly in the speech signal. The diarization process clusters turns of the conversation while recognized special form phrases and entity names identify the speakers. A trained probabilistic model deduces which entity name(s) correspond to the clusters.

Claims (40)

1 . A computer-implemented method comprising:

extracting, by the computer, an enrollment voiceprint for an enrolled speaker using a plurality of enrollment audio features of one or more enrollment speech samples, including a set of enrollment audio features for a keyword sequence occurring in each enrollment speech sample;

extracting, by the computer, a test voiceprint for a test speaker using a plurality test audio features of a test speech sample;

determining, by the computer, an audio similarity score based upon a difference between the enrollment voiceprint and the test voiceprint; and

identifying, by the computer, a content match between the set of enrollment audio features for the keyword sequence and a set of test audio features of the plurality of test audio features of the test speech sample.

2 . The method according to claim 1 , further comprising identifying, by the computer, the test speaker as the enrolled speaker, in response to determining that the audio similarity score satisfied a threshold.

3 . The method according to claim 1 , further comprising identifying, by the computer, the test speaker as the enrolled speaker, in response to identifying the content match.

4 . The method according to claim 1 , wherein the computer obtains the one or more enrollment speech samples and the test speech sample via a telephone channel associated with an agent device.

5 . The method according to claim 1 , wherein the keyword sequence includes repeated content, and wherein each enrollment speech sample includes an instance of the repeated content.

6 . The method according to claim 1 , wherein identifying the content match includes applying, by the computer, a modified dynamic time warping process on the set of enrollment audio features containing the keyword sequence and the set of test audio features.

7 . The method according to claim 1 , further comprising:

for each enrollment speech sample, generating, by the computer, a set of enrollment text features by applying a speaker diarization algorithm on the enrollment speech sample; and

generating, by the computer, a set of test text features by applying the speaker diarization algorithm on the test speech sample,

wherein the computer identifies the content match having the keyword sequence using each set of enrollment text features of the one or more enrollment speech samples and the set of test text features of the test speech sample.

8 . The method according to claim 7 , further comprising:

for each set of enrollment text features, extracting, by the computer, an enrollment entity name associated with the enrolled speaker for the keyword sequence using the enrollment text features; and

extracting, by the computer, a test entity name associated with the test speaker for the keyword sequence using the test text features.

9 . The method according to claim 1 , wherein the computer performs a passive recognition by extracting the plurality of test features from a plurality of test speech samples at a given interval.

10 . The method according to claim 1 , wherein the enrollment audio features and the test audio features comprise at least one of mel-frequency cepstral coefficients (MFCCs), linear predictive cepstral coefficients (LPCCs), or perceptual linear prediction (PLP).

11 . A system comprising:

a non-transitory storage medium storing a plurality of computer program instructions; and

a computer having at least one processor electrically coupled to the non-transitory storage medium and configured to execute the computer program instructions to:

extract an enrollment voiceprint for an enrolled speaker using a plurality of enrollment audio features of one or more enrollment speech samples, including a set of enrollment audio features for a keyword sequence occurring in each enrollment speech sample;

extract a test voiceprint for a test speaker using a plurality test audio features of a test speech sample;

determine an audio similarity score based upon a difference between the enrollment voiceprint and the test voiceprint; and

identify a content match between the set of enrollment audio features for the keyword sequence and a set of test audio features of the plurality of test audio features of the test speech sample.

12 . The system according to claim 11 , wherein the computer is further configured to identify the test speaker as the enrolled speaker, in response to determining that the audio similarity score satisfied a threshold.

13 . The system according to claim 11 , wherein the computer is further configured to identify the test speaker as the enrolled speaker, in response to identifying the content match.

14 . The system according to claim 11 , wherein the computer obtains the one or more enrollment speech samples and the test speech sample via a telephone channel associated with an agent device.

15 . The system according to claim 11 , wherein the keyword sequence includes repeated content, and wherein each enrollment speech sample includes an instance of the repeated content.

16 . The system according to claim 11 , wherein when identifying the content match the computer is further configured to apply a modified dynamic time warping process on the set of enrollment audio features containing the keyword sequence and the set of test audio features.

17 . The system according to claim 11 , wherein the computer is further configured to:

for each enrollment speech sample, generate a set of enrollment text features by applying a speaker diarization algorithm on the enrollment speech sample; and

generate a set of test text features by applying the speaker diarization algorithm on the test speech sample,

wherein the computer identifies the content match having the keyword sequence using each set of enrollment text features of the one or more enrollment speech samples and the set of test text features of the test speech sample.

18 . The system according to claim 17 , wherein the computer is further configured to:

for each set of enrollment text features, extract an enrollment entity name associated with the enrolled speaker for the keyword sequence using the enrollment text features; and

extract a test entity name associated with the test speaker for the keyword sequence using the test text features.

19 . The system according to claim 11 , wherein the computer performs a passive recognition by extracting the plurality of test features from a plurality of test speech samples at a given interval.

20 . The system according to claim 11 , wherein the enrollment audio features and the test audio features comprise at least one of mel-frequency cepstral coefficients (MFCCs), linear predictive cepstral coefficients (LPCCs), or perceptual linear prediction (PLP).

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2023
From: KHOURY, ELIE; GARLAND, MATTHEW
To: PINDROP SECURITY, INC.
Reel/Frame 063855/0176 →
Continuity (5)
Continuation 16895750 · Jun 8, 2020
Continuation 16442368 · Jun 14, 2019
Continuation 15709290 · Sep 19, 2017
Provisional Application 62396670 · Sep 19, 2016
Related Publication 20230326462A1 · Oct 12, 2023
References Cited (344)
US 4817156A · Bahl et al. · 1989 [cited by applicant]
US 4829577A · Kuroda et al. · 1989 [cited by applicant]
US 4972485A · Dautrich et al. · 1990 [cited by applicant]
US 5072452A · Brown et al. · 1991 [cited by applicant]
US 5461697A · Nishimura et al. · 1995 [cited by applicant]
US 5475792A · Stanford et al. · 1995 [cited by applicant]
US 5598507A · Kimber et al. · 1997 [cited by applicant]
US 5659662A · Wilcox et al. · 1997 [cited by applicant]
US 5835890A · Matsui et al. · 1998 [cited by applicant]
US 5867562A · Scherer · 1999 [cited by applicant]
US 5949874A · Mark · 1999 [cited by applicant]
US 5995927A · Li · 1999 [cited by applicant]
US 6009392A · Kanevsky et al. · 1999 [cited by applicant]
US 6021119A · Derks et al. · 2000 [cited by applicant]
US 6055498A · Neumeyer et al. · 2000 [cited by applicant]
US 6094632A · Hattori · 2000 [cited by applicant]
US 6141644A · Kuhn et al. · 2000 [cited by applicant]
US 6411930B1 · Burges · 2002 [cited by applicant]
US 6463413B1 · Applebaum et al. · 2002 [cited by applicant]
US 6519561B1 · Farrell et al. · 2003 [cited by applicant]
US 6760701B2 · Sharma et al. · 2004 [cited by applicant]
US 6882972B2 · Kompe et al. · 2005 [cited by applicant]
US 6922668B1 · Downey · 2005 [cited by applicant]
US 6975708B1 · Scherer · 2005 [cited by applicant]
US 7003460B1 · Bub et al. · 2006 [cited by applicant]
US 7209881B2 · Yoshizawa et al. · 2007 [cited by applicant]
US 7295970B1 · Gorin et al. · 2007 [cited by applicant]
US 7318032B1 · Chaudhari et al. · 2008 [cited by applicant]
US 7324941B2 · Choi et al. · 2008 [cited by applicant]
US 7739114B1 · Chen et al. · 2010 [cited by applicant]
US 7813927B2 · Navratil et al. · 2010 [cited by applicant]
US 8046230B1 · Mcintosh · 2011 [cited by applicant]
US 8112160B2 · Foster · 2012 [cited by applicant]
US 8160811B2 · Prokhorov · 2012 [cited by applicant]
US 8160877B1 · Nucci et al. · 2012 [cited by applicant]
US 8380503B2 · Gross · 2013 [cited by third party]
US 8385536B2 · Whitehead · 2013 [cited by applicant]
US 8484023B2 · Kanevsky et al. · 2013 [cited by applicant]
US 8484024B2 · Kanevsky et al. · 2013 [cited by applicant]
US 8554563B2 · Aronowitz · 2013 [cited by applicant]
US 8712760B2 · Hsia et al. · 2014 [cited by applicant]
US 8738442B1 · Liu et al. · 2014 [cited by applicant]
US 8856895B2 · Perrot · 2014 [cited by applicant]
US 8886663B2 · Gainsboro et al. · 2014 [cited by applicant]
US 8903859B2 · Zeppenfeld et al. · 2014 [cited by applicant]
US 9042867B2 · Gomar · 2015 [cited by applicant]
US 9064491B2 · Rachevsky et al. · 2015 [cited by applicant]
US 9277049B1 · Danis · 2016 [cited by applicant]
US 9336781B2 · Scheffer et al. · 2016 [cited by applicant]
US 9338619B2 · Kang · 2016 [cited by applicant]
US 9343067B2 · Ariyaeeinia et al. · 2016 [cited by applicant]
US 9344892B1 · Rodrigues et al. · 2016 [cited by applicant]
US 9355646B2 · Oh et al. · 2016 [cited by applicant]
US 9373330B2 · Cumani et al. · 2016 [cited by applicant]
US 9401143B2 · Senior et al. · 2016 [cited by applicant]
US 9401148B2 · Lei et al. · 2016 [cited by applicant]
US 9406298B2 · Cumani et al. · 2016 [cited by applicant]
US 9431016B2 · Aviles-Casco et al. · 2016 [cited by applicant]
US 9444839B1 · Faulkner et al. · 2016 [cited by applicant]
US 9454958B2 · Li et al. · 2016 [cited by applicant]
US 9460722B2 · Sidi et al. · 2016 [cited by applicant]
US 9466292B1 · Lei et al. · 2016 [cited by applicant]
US 9502038B2 · Wang et al. · 2016 [cited by applicant]
US 9514753B2 · Sharifi et al. · 2016 [cited by applicant]
US 9558755B1 · Laroche et al. · 2017 [cited by applicant]
US 9584946B1 · Lyren et al. · 2017 [cited by applicant]
US 9620145B2 · Bacchiani et al. · 2017 [cited by applicant]
US 9626971B2 · Rodriguez et al. · 2017 [cited by applicant]
US 9633652B2 · Kurniawati et al. · 2017 [cited by applicant]
US 9641954B1 · Typrin et al. · 2017 [cited by applicant]
US 9665823B2 · Saon et al. · 2017 [cited by applicant]
US 9685174B2 · Karam et al. · 2017 [cited by applicant]
US 9729727B1 · Zhang · 2017 [cited by applicant]
US 9818431B2 · Yu · 2017 [cited by applicant]
US 9824692B1 · Khoury et al. · 2017 [cited by applicant]
US 9860367B1 · Jiang et al. · 2018 [cited by applicant]
US 9875739B2 · Ziv et al. · 2018 [cited by applicant]
US 9875742B2 · Gorodetski et al. · 2018 [cited by applicant]
US 9875743B2 · Gorodetski et al. · 2018 [cited by applicant]
US 9881617B2 · Sidi et al. · 2018 [cited by applicant]
US 9930088B1 · Hodge · 2018 [cited by applicant]
US 9984706B2 · Wein · 2018 [cited by applicant]
US 10277628B1 · Jakobsson · 2019 [cited by applicant]
US 10325601B2 · Khoury et al. · 2019 [cited by applicant]
US 10347256B2 · Khoury et al. · 2019 [cited by applicant]
US 10397398B2 · Gupta · 2019 [cited by applicant]
US 10404847B1 · Unger · 2019 [cited by applicant]
US 10462292B1 · Stephens · 2019 [cited by applicant]
US 10506088B1 · Singh · 2019 [cited by applicant]
US 10554821B1 · Koster · 2020 [cited by applicant]
US 10638214B1 · Delhoume et al. · 2020 [cited by applicant]
US 10659605B1 · Braundmeier et al. · 2020 [cited by applicant]
US 10679630B2 · Khoury · 2020 [cited by examiner]
US 10854205B2 · Khoury et al. · 2020 [cited by applicant]
US 11069352B1 · Tang et al. · 2021 [cited by applicant]
US 11670304B2 · Khoury · 2023 [cited by examiner]
US 11854205B2 · Cao et al. · 2023 [cited by applicant]
US 12175983B2 · Khoury · 2024 [cited by examiner]
US 20020095287A1 · Botterweck · 2002 [cited by applicant]
US 20020143539A1 · Botterweck · 2002 [cited by applicant]
US 20030231775A1 · Wark · 2003 [cited by applicant]
US 20030236663A1 · Dimitrova et al. · 2003 [cited by applicant]
US 20040218751A1 · Colson et al. · 2004 [cited by applicant]
US 20040230420A1 · Kadambe et al. · 2004 [cited by applicant]
US 20050038655A1 · Mutel et al. · 2005 [cited by applicant]
US 20050039056A1 · Bagga et al. · 2005 [cited by applicant]
US 20050286688A1 · Scherer · 2005 [cited by applicant]
US 20060058998A1 · Yamamoto et al. · 2006 [cited by applicant]
US 20060111905A1 · Navratil et al. · 2006 [cited by applicant]
US 20060293771A1 · Tazine et al. · 2006 [cited by applicant]
US 20070189479A1 · Scherer · 2007 [cited by applicant]
US 20070198257A1 · Zhang et al. · 2007 [cited by applicant]
US 20070294083A1 · Bellegarda et al. · 2007 [cited by applicant]
US 20080195389A1 · Zhang et al. · 2008 [cited by applicant]
US 20080312926A1 · Vair et al. · 2008 [cited by applicant]
US 20090138712A1 · Driscoll · 2009 [cited by applicant]
US 20090265328A1 · Parekh et al. · 2009 [cited by applicant]
US 20100131273A1 · Aley-Raz et al. · 2010 [cited by applicant]
US 20100217589A1 · Gruhn et al. · 2010 [cited by applicant]
US 20100232619A1 · Uhle et al. · 2010 [cited by applicant]
US 20100262423A1 · Huo et al. · 2010 [cited by applicant]
US 20110010173A1 · Scott et al. · 2011 [cited by applicant]
US 20120173239A1 · Asenjo et al. · 2012 [cited by applicant]
US 20120185418A1 · Capman et al. · 2012 [cited by applicant]
US 20130041660A1 · Waite · 2013 [cited by applicant]
US 20130080165A1 · Wang et al. · 2013 [cited by applicant]
US 20130109358A1 · Balasubramaniyan et al. · 2013 [cited by applicant]
US 20130300939A1 · Chou et al. · 2013 [cited by applicant]
US 20140046878A1 · Lecomte et al. · 2014 [cited by applicant]
US 20140053247A1 · Fadel · 2014 [cited by applicant]
US 20140081640A1 · Farrell et al. · 2014 [cited by applicant]
US 20140195236A1 · Hosom et al. · 2014 [cited by applicant]
US 20140214417A1 · Wang et al. · 2014 [cited by applicant]
US 20140214676A1 · Bukai · 2014 [cited by applicant]
US 20140241513A1 · Springer · 2014 [cited by applicant]
US 20140250512A1 · Goldstone et al. · 2014 [cited by applicant]
US 20140278412A1 · Scheffer et al. · 2014 [cited by applicant]
US 20140288928A1 · Penn et al. · 2014 [cited by applicant]
US 20140337017A1 · Watanabe et al. · 2014 [cited by applicant]
US 20150036813A1 · Ananthakrishnan et al. · 2015 [cited by applicant]
US 20150127336A1 · Lei et al. · 2015 [cited by applicant]
US 20150149165A1 · Saon · 2015 [cited by applicant]
US 20150161522A1 · Saon et al. · 2015 [cited by applicant]
US 20150189086A1 · Romano et al. · 2015 [cited by applicant]
US 20150199960A1 · Huo et al. · 2015 [cited by applicant]
US 20150269931A1 · Senior et al. · 2015 [cited by applicant]
US 20150269941A1 · Jones · 2015 [cited by applicant]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20150310008A1 · Thudor et al. · 2015 [cited by applicant]
US 20150334231A1 · Rybak et al. · 2015 [cited by applicant]
US 20150348571A1 · Koshinaka et al. · 2015 [cited by applicant]
US 20150356630A1 · Hussain · 2015 [cited by applicant]
US 20150365530A1 · Kolbegger et al. · 2015 [cited by applicant]
US 20160019458A1 · Kaufhold · 2016 [cited by applicant]
US 20160019883A1 · Aronowitz · 2016 [cited by applicant]
US 20160028434A1 · Kerpez et al. · 2016 [cited by applicant]
US 20160078863A1 · Chung et al. · 2016 [cited by applicant]
US 20160104480A1 · Sharifi · 2016 [cited by applicant]
US 20160125877A1 · Foerster et al. · 2016 [cited by applicant]
US 20160180214A1 · Kanevsky et al. · 2016 [cited by applicant]
US 20160189707A1 · Donjon et al. · 2016 [cited by applicant]
US 20160240190A1 · Lee et al. · 2016 [cited by applicant]
US 20160275953A1 · Sharifi et al. · 2016 [cited by applicant]
US 20160284346A1 · Visser et al. · 2016 [cited by applicant]
US 20160293167A1 · Chen et al. · 2016 [cited by applicant]
US 20160314790A1 · Tsujikawa et al. · 2016 [cited by applicant]
US 20160343373A1 · Ziv et al. · 2016 [cited by applicant]
US 20170060779A1 · Falk · 2017 [cited by applicant]
US 20170069313A1 · Aronowitz · 2017 [cited by applicant]
US 20170069327A1 · Heigold et al. · 2017 [cited by applicant]
US 20170098444A1 · Song · 2017 [cited by applicant]
US 20170111515A1 · Bandyopadhyay et al. · 2017 [cited by applicant]
US 20170126884A1 · Balasubramaniyan et al. · 2017 [cited by applicant]
US 20170142150A1 · Sandke et al. · 2017 [cited by applicant]
US 20170169816A1 · Blandin et al. · 2017 [cited by applicant]
US 20170230390A1 · Faulkner et al. · 2017 [cited by applicant]
US 20170262837A1 · Gosalia · 2017 [cited by applicant]
US 20180082691A1 · Khoury et al. · 2018 [cited by applicant]
US 20180152558A1 · Chan et al. · 2018 [cited by applicant]
US 20180249006A1 · Dowlatkhah et al. · 2018 [cited by applicant]
US 20180295235A1 · Tatourian et al. · 2018 [cited by applicant]
US 20180337962A1 · Ly et al. · 2018 [cited by applicant]
US 20190037081A1 · Rao et al. · 2019 [cited by applicant]
US 20190172476A1 · Wung et al. · 2019 [cited by applicant]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190180758A1 · Washio · 2019 [cited by applicant]
US 20190238956A1 · Gaubitch et al. · 2019 [cited by applicant]
US 20190297503A1 · Traynor et al. · 2019 [cited by applicant]
US 20200137221A1 · Dellostritto et al. · 2020 [cited by applicant]
US 20200195779A1 · Weisman et al. · 2020 [cited by applicant]
US 20200252510A1 · Ghuge et al. · 2020 [cited by applicant]
US 20200312313A1 · Maddali et al. · 2020 [cited by applicant]
US 20200396332A1 · Gayaldo · 2020 [cited by applicant]
US 20210084147A1 · Kent et al. · 2021 [cited by applicant]
US 20210150010A1 · Strong et al. · 2021 [cited by applicant]
US 20230326462A1 · Khoury et al. · 2023 [cited by applicant]
WO WO2015079885 · 2015 [cited by applicant]
WO WO2016195261A1 · 2016 [cited by applicant]
WO WO2017167900A1 · 2017 [cited by applicant]
Navratil et al., “An Instantiable Speech Biometrics Module with Natural Language Interface: Implementation in the Telephony Environment,” 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing, a… [cited by applicant]
First Examiners Requisition on CA Appl. 3,096,378 dated Jun. 12, 2023 (3 pages). [cited by applicant]
Office Action on Japanese Application 2022-104204 dated Jul. 26, 2023 (4 pages). [cited by applicant]
J. H. L. Hansen and T. Hasan, “Speaker Recognition by Machines and Humans: A tutorial review,” in IEEE Signal Processing Magazine, vol. 32, No. 6, pp. 74-99, Nov. 2015, doi: 10.1109/MSP.2015.2462851. (Year: 2015). [cited by applicant]
M. Witkowski, M. Igras, J. Grzybowska, P. Jaci6w, J. Galka and M. Zi6/ko, “Caller identification by voice,” XXII Annual Pacific Voice Conference (PVC), Krakow, Poland, 2014, pp. 1-7, doi: 10.1109/PVC.2014.6845420. (Year… [cited by applicant]
T. F. Zheng, Q. Jin, L. Li, J. Wang and F. Bie, “An overview of robustness related issues in speaker recognition,” Signal and Information Processing Association Annual Summit and Conference (APSIPA), 2014 Asia-Pacific, … [cited by applicant]
First Examiner's Requisition for CA App. 3,135,210 dated Oct. 5, 2023 (5 pages). [cited by applicant]
First Examiner's Requisition on CA App. 3,179,080 dated Apr. 9, 2024 (3 pages). [cited by applicant]
Ahmad et al., “A Unique Approach in Text Independent Speaker Recognition using MFCC Feature Sets and Probabilistic Neural Network”, Eighth International Conference on Advances in Pattern Recognition (ICAPR}, IEEE, 2015 … [cited by applicant]
Almaadeed, et al., “Speaker identification using multimodal neural networks and wavelet analysis,” IET Biometrics 4.1 (2015), 18-28. [cited by applicant]
Atrey, et al., “Audio based event detection for multimedia surveillance”, Acoustics, Speech and Signal Processing, 2006, ICASSP 2006 Proceedings, 2006 IEEE International Conference on vol. 5, IEEE, 2006. pp. 813-816. [cited by applicant]
Australian Examination Report on AU Appl. Ser. No. 2021286422 dated Oct. 21, 2022 (2 pages). [cited by applicant]
Baraniuk, “Compressive Sensing [Lecture Notes]”, IEEE Signal Processing Magazine, vol. 24, Jul. 2007, 4 pages. [cited by applicant]
Bredin, “TristouNet: Triplet Loss for Speaker Turn Embedding”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 14, 2016, XP080726602. [cited by applicant]
Buera et al., “Unsupervised Data-Driven Feature Vector Normalization With Acoustic Model Adaptation for Robust Speech Recognition”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 18, No. 2, Feb. 2010,… [cited by applicant]
Campbell, “Using Deep Belief Networks for Vector-Based Speaker Recognition”, Proceedings of Interspeech 2014, Sep. 14, 2014, pp. 676-680, XP055433784. [cited by applicant]
Castaldo et al., “Compensation of Nuisance Factors for Speaker and Language Recognition,” IEEE Transactions on Audio, Speech and Language Processing, ieeexplore.ieee.org, vol. 15, No. 7, Sep. 2007. pp. 1969-1978. [cited by applicant]
Communication pursuant to Article 94(3) EPC issued in EP Application No. 17 772 184.2-1207 dated Jul. 19, 2019. [cited by applicant]
Communication pursuant to Article 94(3) EPC on EP 17772184.2 dated Jun. 18, 2020. [cited by applicant]
Cumani, et al., “Factorized Sub-space Estimation for Fast and Memory Effective I-Vector Extraction”, IEEE/ACM TASLP, vol. 22, Issue 1, Jan. 2014, 28 pages. [cited by applicant]
D.Etter and C.Domniconi, “Multi2Rank: Multimedia Multiview Ranking,” 2015 IEEE International Conference on Multimedia Big Data, 2015, pp. 80-87. (Year 2015). [cited by applicant]
Dehak, et al., “Front-End Factor Analysis for Speaker Verification”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, No. 4, May 2011, pp. 788-798. [cited by applicant]
Douglas A. Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Examination Report for EP 17778046.7 dated Jun. 16, 2020 (4 pages). [cited by applicant]
Examination Report for IN 201947014575 dated Nov. 16, 2021 (6 pages). [cited by applicant]
Examination Report No. 1 for AU 2017322591 dated Jul. 16, 2021 (2 pages). [cited by applicant]
Final Office Action for U.S. Appl. No. 16/200,283 dated Jun. 11, 2020 (15 pages). [cited by applicant]
Final Office Action for U.S. Appl. No. 16/784,071 dated Oct. 27, 2020 (12 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 15/872,639 dated Jan. 29, 2019 (11 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 17/121,291 dated Mar. 27, 2023 (7 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 16/829,705 dated Mar. 24, 2022 (23 pages). [cited by applicant]
First Office Action issued in KR 10-2019-7010-7010208 dated Jun. 29, 2019. (5 pages). [cited by applicant]
First Office Action issued on CA Application No. 3,036,533 dated Apr. 12, 2019. (4 pages). [cited by applicant]
First Office Action on CA Application No. 3,075,049 dated May 7, 2020. (3 pages). [cited by applicant]
Florian et al., “FaceNet: A unified embedding for face recognition and clustering”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 7, 2015, pp. 815-823, XP032793492, DOI: 10.1109/CVPR… [cited by applicant]
Foreign Action on JP 2019-535198 dated Mar. 1, 2022 (6 pages). [cited by applicant]
Fu et al., “SNR-Aware Convolutional Neural Network Modeling for Speech Enhancement”, INTERSPEECH 2016, Sep. 8-12, 2016, pp. 3768-3772, XP055427533, ISSN: 1990-9772, DOI: 10.21437/Interspeech.2016-211 (5 pages). [cited by applicant]
Gao, et al., “Dimensionality reduction via compressive sensing”, Pattern Recognition Letters 33, Elsevier Science BV 0167-8655, 2012. 8 pages. [cited by applicant]
Garcia-Romero et al., “Unsupervised Domain Adaptation for I-Vector Speaker Recgonition” Odyssey 2014, pp. 260-264. [cited by applicant]
Ghahabi Omid et al., “Restricted Boltzmann Machine Supervectors for Speaker Recognition”, 2015 IEEE International Conference on acoustics, Speech and Signal Processing (ICASSP), IEEE, Apr. 19, 2015, XP033187673, pp. 480… [cited by applicant]
Gish, et al., “Segregation of Speakers for Speech Recognition and Speaker Identification”, Acoustics, Speech, and Signal Processing, 1991, ICASSP-91, 1991 International Conference on IEEE, 1991. pp. 873-876. [cited by applicant]
Hoffer et al., “Deep Metric Learning Using Triplet Network”, 2015, arXiv: 1412.6622v3, retrieved Oct. 4, 2021 from URL: https://deepsense.ai/wp-content/uploads/2017/08/1412.6622-3.pdf (8 pages). [cited by applicant]
Hoffer et al., “Deep Metric Learning Using Triplet Network,” ICLR 2015 (workshop contribution), Mar. 23, 2015, pp. 1-8. [cited by applicant]
Huang, et al., “A Blind Segmentation Approach to Acoustic Event Detection Based on I-Vector”, INTERSPEECH, 2013. pp. 2282-2286. [cited by applicant]
Information Disclosure Statement filed Aug. 8, 2019. [cited by applicant]
Information Disclosure Statement filed Feb. 12, 2020 (4 pages). [cited by applicant]
International Preliminary Report on Patentability for PCT/US2020/017051 dated Aug. 19, 2021 (11 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/039697 dated Jan. 1, 2019 (10 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052293 dated Mar. 19, 2019 (8 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052316 dated Mar. 19, 2019 (7 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2017/052335 dated Mar. 19, 2019 (8 pages). [cited by applicant]
International Preliminary Report on Patentability, Ch. I, for PCT/US2020/026992 dated Oct. 21, 2021 (10 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US20/17051 dated Apr. 23, 2020 (12 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US2017/052335 dated Dec. 8, 2017 (10 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US2020/24709 dated Jun. 19, 2020 (10 pages). [cited by applicant]
International Search Report and Written Opinion issued in corresponding International Application No. PCT/US2018/013965 with Date of mailing May 14, 2018. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US20/26992 with Date of mailing Jun. 26, 2020. [cited by applicant]
International Search Report and Written Opinion issued in the corresponding International Application No. PCT/US2017/039697, mailed on Sep. 20, 2017. 17 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority issued in corresponding International Application No. PCT/US2017/050927 with Date of mailing Dec. 11, 2017. [cited by applicant]
International Search Report and Written Opinion on PCT Appl. Ser. No. PCT/US2020/026992 dated Jun. 26, 2020 (11 pages). [cited by applicant]
International Search Report issued in corresponding International Application No. PCT/US2017/052293 dated Dec. 21, 2017. [cited by applicant]
International Search Report issued in corresponding International Application No. PCT/US2017/052316 with Date of mailing Dec. 21, 2017. 3 pages. [cited by applicant]
Kenny et al., “Deep Neural Networks for extracting Baum-Welch statistics for Speaker Recognition”, Jun. 29, 2014, XP055361192, Retrieved from the Internet: URL:http://www.crim.ca/perso/patrick.kenny/stafylakis_odyssey20… [cited by applicant]
Kenny, “A Small Footprint i-Vector Extractor” Proc. Odyssey Speaker and Language Recognition Workshop, Singapore, Jun. 25, 2012. 6 pages. [cited by applicant]
Khoury et al., “Combining Transcription-Based and Acoustic-Basedd Speaker Identifications for Broadcast News”, ICASSP, Kyoto, Japan, 2012, pp. 4377-4380. [cited by applicant]
Khoury, et al., “Improvised Speaker Diariztion System for Meetings”, Acoustics, Speech and Signal Processing, 2009, ICASSP 2009, IEEE International Conference on IEEE, 2009. pp. 4097-4100. [cited by applicant]
Kockmann et al., “Syllable based Feature-Contours for Speaker Recognition,” Proc. 14th International Workshop on Advances, 2008. 4 pages. [cited by applicant]
Korean Office Action (with English summary), dated Jun. 29, 2019, issued in Korean application No. 10-2019-7010208, 6 pages. [cited by applicant]
Lei et al., “A Novel Scheme for Speaker Recognition Using a Phonetically-aware Deep Neural Network”, Proceedings on ICASSP, Florence, Italy, IEEE Press, 2014, pp. 1695-1699. [cited by applicant]
Luque, et al., “Clustering Initialization Based on Spatial Information for Speaker Diarization of Meetings”, Ninth Annual Conference of the International Speech Communication Association, 2008. pp. 383-386. [cited by applicant]
McLaren, et al., “Advances in deep neural network approaches to speaker recognition,” In Proc. 40th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015. [cited by applicant]
McLaren, et al., “Exploring the Role of Phonetic Bottleneck Features for Speaker and Language Recognition”, 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 5575-557… [cited by applicant]
Meignier, et al., “Lium Spkdiarization: An Open Source Toolkit for Diarization” CMU SPUD Workshop, 2010. 7 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,232 dated Jun. 27, 2019 (11 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,232 dated Oct. 5, 2018 (21 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/709,290 dated Oct. 5, 2018 (12 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/200,283 dated Jan. 7, 2020 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/442,368 dated Oct. 4, 2019 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/784,071 dated May 12, 2020 (10 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/829,705 dated Nov. 10, 2021 (20 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 15/610,378 dated Mar. 1, 2018 (11 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 16/551,327 dated Dec. 9, 2019 (7 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 15/872,639 dated Aug. 23, 2018 (10 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 16/829,705 dated Jul. 21, 2022 (27 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/610,378 dated Aug. 7, 2018 (5 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,024 dated Mar. 18, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,232 dated Feb. 6, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,232 dated Oct. 8, 2019 (11 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/709,290 dated Feb. 6, 2019 (6 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/200,283 dated Aug. 24, 2020 (7 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/442,368 dated Feb. 5, 2020 (7 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/505,452 dated Jul. 23, 2020 (8 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/505,452 dated May 13, 2020 (9 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/784,071 dated Jan. 27, 2021 (14 pages). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/317,575 dated Jan. 28, 2022 (11 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/262,748 dated Sep. 13, 2017 (9 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/818,231 dated Mar. 27, 2019 (9 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 15/872,639 dated Apr. 25, 2019 (5 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/551,327 DTD Mar. 26, 2020. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 17/107,496 dated Jan. 26, 2023 (8 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/536,293 dated Jun. 3, 2022 (9 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/829,705 dated Dec. 21, 2022 (14 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/895,750 dated Jan. 25, 2023 (6 pages). [cited by applicant]
Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, mailed May 14, 2018, in corresponding International Application No. PCT/US2018/013965, 13 … [cited by applicant]
Novoselov, et al., “SIC Speaker Recognition System for the NIST i-Vector Challenge.” Odyssey: The Speaker and Language Recognition Workshop. Jun. 16-19, 2014. pp. 231-240. [cited by applicant]
Office Action for CA 3036561 dated Jan. 23, 2020 (5 pages). [cited by applicant]
Oguzhan et al., “Recognition of Acoustic Events Using Deep Neural Networks”, 2014 22nd European Signal Processing Conference (EUSIPCO), Sep. 1, 2014, pp. 506-510 (5 pages). [cited by applicant]
Piegeon, et al., “Applying Logistic Regression to the Fusion of the NIST'99 1-Speaker Submissions”, Digital Signal Processing Oct. 1-3, 2000. pp. 237-248. [cited by applicant]
Prazak et al., “Speaker Diarization Using PLDA-based Speaker Clustering”, The 6th IEEE International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications, Sep. 2011, pp.… [cited by applicant]
Prince, et al., “Probabilistic Linear Discriminant Analysis for Inferences About Identity,” Proceedings of the International Conference on Computer Vision, Oct. 14-21, 2007. 8 pages. [cited by applicant]
Reasons for Refusal for JP 2019-535198 dated Sep. 10, 2021 (7 pages). [cited by applicant]
Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Richardson, et al., “Channel Compensation for Speaker Recognition using MAP Adapted PLDA and Denoising DNNs”, Proc. Speaker Lang. Recognit. Workshop, Jun. 22, 2016, pp. 225-230. [cited by applicant]
Richardson, et al., “Deep Neural Network Approaches to Speaker and Language Recognition”, IEEE Signal Processing Letters, vol. 22, No. 10, Oct. 2015, pp. 1671-1675. [cited by applicant]
Richardson, et al., “Speaker Recognition Using Real vs Synthetic Parallel Data for DNN Channel Compensation”, INTERSPEECH, 2016. 6 pages. [cited by applicant]
Richardson, et al., Speaker Recognition Using Real vs Synthetic Parallel Data for DNN Channel Compensation, INTERSPEECH, 2016, retrieved Sep. 14, 2021 from URL: https://www.ll.mit.edu/sites/default/files/publication/doc… [cited by applicant]
Rouvier et al., “An Open-source State-of-the-art Toolbox for Broadcast News Diarization”, INTERSPEECH, Aug. 2013, pp. 1477-1481 (5 pages). [cited by applicant]
Scheffer et al., “Content matching for short duration speaker recognition”, INTERSPEECH, Sep. 14-18, 2014, pp. 1317-1321. [cited by applicant]
Schmidt et al., “Large-Scale Speaker Identification” ICASSP, 2014, pp. 1650-1654. [cited by applicant]
Seddik, et al., “Text independent speaker recognition using the Mel frequency cepstral coefficients and a neural network classifier.” First International Symposium on Control, Communications and Signal Processing, 2004.… [cited by applicant]
Shajeesh, et al., “Speech Enhancement based on Savitzky-Golay Smoothing Filter”, International Journal of Computer Applications, vol. 57, No. 21, Nov. 2012, pp. 39-44 (6 pages). [cited by applicant]
Shum et al., “Exploiting Intra-Conversation Variability for Speaker Diarization”, INTERSPEECH, Aug. 2011, pp. 945-948 (4 pages). [cited by applicant]
Snyder et al., “Time Delay Deep Nueral Network-Based Universal Background Models for Speaker Recognition”, IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2015, pp. 92-97 (6 pages). [cited by applicant]
Solomonoff, et al., “Nuisance Attribute Projection”, Speech Communication, Elsevier Science BV, Amsterdam, The Netherlands, May 1, 2007. 74 pages. [cited by applicant]
Sturim et al., “Speaker Linking and Applications Using Non-parametric Hashing Methods,” Interspeech, Sep. 2016, 5 pages. [cited by applicant]
Summons to attend oral proceedings pursuant to Rule 115(1) EPC issued in EP Application No. 17 772 184.2-1207 dated Dec. 16, 2019. [cited by applicant]
T.C.Nagavi, S.B. Anusha, P.Monisha and S.P.Poomima, “Content based audio retrieval with MFCC feature extraction, clustering and sort-merge techniquest,” 2013 Fourth International Conference on Computing, Communications … [cited by applicant]
Temko, et al., “Acoustic event detection in meeting-room environments”, Pattern Recognition Letters, vol. 30, No. 14, 2009, pp. 1281-1288. [cited by applicant]
Temko, et al., “Classification of acoustic events using SVM-based clustering schemes”, Pattern Recognition, vol. 39, No. 4, 2006, pp. 682-694. [cited by applicant]
US Non Final Office Action on U.S. Appl. No. 17/107,496 dated Jul. 21, 2022 (7 pages). [cited by applicant]
US Non-Final Office Action on U.S. Appl. No. 17/121,291 dated Oct. 28, 2022 (11 pages). [cited by applicant]
US Non-Final Office Action on U.S. Appl. No. 16/895,750 dated Oct. 6, 2022 (10 pages). [cited by applicant]
US Non-Final Office Action on U.S. Appl. No. 17/706,398 dated Dec. 15, 2022 (5 pages). [cited by applicant]
US Notice of Allowance on U.S. Appl. No. 17/107,496 dated Sep. 28, 2022 (8 pages). [cited by applicant]
Uzan et al., “I Know That Voice: Identifying the Voice Actor Behind the Voice”, 2015 International Conference on Biometrics (ICB), 2015, retrieved Oct. 4, 2021 from URL: https://citeseerx.ist.psu.edu/viewdoc/download?do… [cited by applicant]
Variani et al., “Deep neural networks for small footprint text-dependent speaker verification”, 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2014, pp. 4052-4056, X… [cited by applicant]
Vella et al., “Artificial neural network features for speaker diarization”, 2014 IEEE Spoken Language Technology Workshop (SLT), IEEE Dec. 7, 2014, pp. 402-406, XP032756972, DOI: 10.1109/SLT.2014.7078608. [cited by applicant]
Wang et al., “Learning Fine-Grained Image Similarity with Deep Ranking”, Computer Vision and Pattern Recognition, Jan. 17, 2014, arXiv: 1404.4661v1, retrieved Oct. 4, 2021 from URL: https://arxiv.org/pdf/1404.4661.pdf (… [cited by applicant]
Xiang, et al., “Efficient text-independent speaker verification with structural Gaussian mixture models and neural network.” IEEE Transactions on Speech and Audio Processing 11.5 (2003): 447-456. [cited by applicant]
Xu et al., “Rapid Computation of I-Vector” Odyssey, Bilbao, Spain, Jun. 21-24, 2016. 6 pages. [cited by applicant]
Xue et al., “Fast Query by Example of Enviornmental Sounds Via Robust and Efficient Cluster-Based Indexing”, Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2008, pp. 5-8 (4 pages). [cited by applicant]
Yaman et al., “Bottleneck Features for Speaker Recognition”, Proceedings of the Speaker and Language Recognition Workshop 2012, Jun. 28, 2012, pp. 105-108, XP055409424, Retrieved from the Internet: URL:https://pdfs.sema… [cited by applicant]
Yella Sree Harsha et al., “Artificial neural network features for speaker diarization”, 2014 IEEE Spoken Language technology Workshop (SLT), IEEE, Dec. 7, 2014, pp. 402-406, XP032756972, DOI: 10.1109/SLT.2014.7078608. [cited by applicant]
Zhang et al., “Extracting Deep Neural Network Bottleneck Features Using Low-Rank Matrix Factorization”, IEEE, ICASSP, 2014. 5 pages. [cited by applicant]
Zheng, et al., “An Experimental Study of Speech Emotion Recognition Based on Deep Convolutional Neural Networks”, 2015 International Conference on Affective Computing and Intelligent Interaction (ACII), 2015, pp. 827-83… [cited by applicant]