IP Library Granted Patent US 12,488,072
Granted Patent B2
US 12,488,072 · App. 17/231,672 · Granted Dec 2, 2025

Passive and continuous multi-speaker voice biometrics

Inventors: Elie Khoury (Atlanta, GA); Ganesh Sivaraman (Atlanta, GA); Avrosh Kumar (Atlanta, GA); Ivan Antolic-Soban (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G06F21/32G06N20/00G10L17/04G10L17/18G10L17/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,072
App. No.
17/231,672
Granted
Dec 2, 2025
Kind
B2
Abstract

Embodiments described herein provide for a voice biometrics system execute machine-learning architectures capable of passive, active, continuous, or static operations, or a combination thereof. Systems passively and/or continuously, in some cases in addition to actively and/or statically, enrolling speakers as the speakers speak into or around an edge device (e.g., car, television, radio, phone). The system identifies users on the fly without requiring a new speaker to mirror prompted utterances for reconfiguring operations. The system manages speaker profiles as speakers provide utterances to the system. Machine-learning architectures implement a passive and continuous voice biometrics system, possibly without knowledge of speaker identities. The system creates identities in an unsupervised manner, sometimes passively enrolling and recognizing known or unknown speakers. The system offers personalization and security across a wide range of applications, including media content for over-the-top services and IoT devices (e.g., personal assistants, vehicles), and call centers.

Claims (61)

1 . A computer-implemented method comprising:

extracting, by a computer, an inbound embedding for an inbound speaker by applying a machine-learning model on an inbound audio signal;

generating, by the computer, a similarity score based upon a distance between the inbound embedding for the inbound speaker and a voiceprint stored in a speaker profile in a speaker profile database; and

responsive to the computer determining that the similarity score for the inbound embedding of the inbound speaker fails to satisfy a similarity threshold:

generating, by the computer, in the speaker profile database a new speaker profile for the inbound speaker containing the inbound embedding for the inbound speaker of the inbound audio signal, the new speaker profile is a database record storing the inbound embedding as a new voiceprint for the inbound speaker, wherein the new voiceprint satisfies a maturity threshold and is based upon at least one inbound embedding for the inbound speaker;

responsive to the computer determining that a second similarity score for a second inbound embedding extracted for a second inbound signal satisfies the similarity threshold:

updating, by the computer, the new voiceprint for the inbound speaker based upon the second inbound signal.

2 . The method according to claim 1 , further comprising:

receiving, by the computer, the inbound audio signal from an end-user device via an intermediate server; and

transmitting, by the computer, a new speaker identifier associated with the new speaker profile to the intermediate server.

3 . The method according to claim 2 , wherein the end-user device is at least one of a smart television, a media device coupled to a television, and an edge device.

4 . The method according to claim 1 , further comprising extracting, by the computer, one or more features from the inbound audio signal, wherein the computer generates the inbound embedding by applying the machine-learning model on the one or more features the computer extracted from the inbound audio signal.

5 . The method according to claim 1 , further comprising:

extracting, by the computer, the second inbound embedding from the second inbound signal by applying the machine-learning model on the second inbound audio signal; and

generating, by the computer, the second similarity score based upon the distance between the second inbound embedding and the new voiceprint stored in the new speaker profile.

6 . The method according to claim 1 , further comprising:

receiving, by the computer, a subscriber identifier associated with the inbound audio signal; and

identifying, by the computer, one or more speaker profiles associated with the subscriber identifier stored in the speaker profile database, wherein the computer generates one or more similarity scores for the inbound embedding based upon one or more voiceprints stored in the one or more speaker profiles associated with the subscriber identifier.

7 . The method according to claim 6 , wherein the subscriber identifier is associated with one or more speaker identifiers, and wherein each speaker profile is associated with a corresponding speaker identifier.

8 . The method according to claim 7 , wherein at least one of the subscriber identifier and each speaker identifier is an anonymized identifier.

9 . The method according to claim 1 , wherein the computer generates the new voiceprint based upon one or more inbound embeddings, the method further comprising:

identifying, by the computer, one or more maturity factors for the new voiceprint based upon the one or more inbound embeddings; and

determining, by the computer, a level of maturity for the new voiceprint based upon the one or more maturity factors.

10 . The method according to claim 9 , further comprising updating, by the computer, a new similarity threshold of the new speaker profile in response to the computer determining that the level of maturity for the new voiceprint satisfies a second maturity threshold.

11 . The method according to claim 9 , further comprising:

responsive to the computer determining that the level of maturity fails the maturity threshold:

generating, by the computer, an active enrollment prompts the active enrollment prompt comprising a user interface configured to display a request for an additional inbound audio signal;

extracting, by the computer, an additional embedding from the additional inbound signal; and

updating, by the computer, the new voiceprint according to an additional embedding extracted from the additional inbound signal.

12 . The method according to claim 9 , further comprising updating, by the computer, the new speaker profile from a temporary profile to a permanent profile in response to the computer determining that the level of maturity satisfies the maturity threshold.

13 . A system comprising:

a speaker profile database comprising non-transitory machine-readable storage media configured to store data records containing speaker profiles; and

a computer comprising a processor configured to:

extract an inbound embedding for an inbound speaker by applying a machine-learning model on an inbound audio signal;

generate a similarity score based upon a distance between the inbound embedding for the inbound speaker and a voiceprint stored in a speaker profile in the speaker profile database; and

responsive to the computer determining that the similarity score for the inbound embedding for the inbound speaker fails to satisfy a similarity threshold:

generate in the speaker profile database a new speaker profile for the inbound speaker containing the inbound embedding for the inbound speaker of the inbound audio signal, the new speaker profile is a database record storing the inbound embedding as a new voiceprint for the inbound speaker, wherein the new voiceprint satisfies a maturity threshold and is based upon at least one inbound embedding for the inbound speaker;

responsive to the computer determining that a second similarity score for a second inbound embedding extracted for a second inbound signal satisfies the similarity threshold:

update the new voiceprint for the inbound speaker based upon the second inbound signal.

14 . The system according to claim 13 , wherein the computer is further configured to:

receive the inbound audio signal from an end-user device via an intermediate server; and

transmit a new speaker identifier associated with the new speaker profile to the intermediate server.

15 . The system according to claim 14 , wherein the end-user device is at least one of a smart television, a media device coupled to a television, and an edge device.

16 . The system according to claim 13 , wherein the computer is further configured to extract one or more features from the inbound audio signal, wherein the computer generates the inbound embedding by applying the machine-learning model on the one or more features the computer extracted from the inbound audio signal.

17 . The system according to claim 13 , wherein the computer is further configured to:

extract the second inbound embedding from the second inbound signal by applying the machine-learning model on the second inbound audio signal; and

generate the second similarity score based upon the distance between the second inbound embedding and the new voiceprint stored in the new speaker profile.

18 . The system according to claim 13 , wherein the computer is further configured to:

receive a subscriber identifier associated with the inbound audio signal; and

identify one or more speaker profiles associated with the subscriber identifier stored in the speaker profile database, wherein the computer generates one or more similarity scores for the inbound embedding based upon one or more voiceprints stored in the one or more speaker profiles associated with the subscriber identifier.

19 . The system according to claim 18 , wherein the subscriber identifier is associated with one or more speaker identifiers, and wherein each speaker profile is associated with a corresponding speaker identifier.

20 . The system according to claim 19 , wherein at least one of the subscriber identifier and each speaker identifier is an anonymized identifier.

21 . The system according to claim 13 , wherein the computer generates the new voiceprint based upon one or more inbound embeddings, and wherein the computer is further configured to:

identify one or more maturity factors for the new voiceprint based upon the one or more inbound embeddings; and

determine a level of maturity for the new voiceprint based upon the one or more maturity factors.

22 . The system according to claim 21 , wherein the computer is further configured to update a new similarity threshold of the new speaker profile in response to the computer determining that the level of maturity for the new voiceprint satisfies a second maturity threshold.

23 . The system according to claim 21 , wherein the computer is further configured to:

responsive to determining that the level of maturity fails the maturity threshold:

generate an active enrollment prompts the active enrollment prompt comprising a user interface configured to display a request for an additional inbound audio signal; and

extract an additional embedding from the additional inbound signal; and update the new voiceprint according to an additional embedding extracted from the additional inbound signal.

24 . The system according to claim 21 , wherein the computer is further configured to update the new speaker profile from a temporary profile to a permanent profile in response to the computer determining that the level of maturity satisfies the maturity threshold.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2021
From: KHOURY, ELIE; SIVARAMAN, GANESH; KUMAR, AVROSH; ANTOLIC-SOBAN, IVAN
To: PINDROP SECURITY, INC.
Reel/Frame 055933/0400 →
Continuity (2)
Provisional Application 63010504 · Apr 15, 2020
Related Publication 20210326421A1 · Oct 21, 2021
References Cited (137)
US 5442696A · Lindberg et al. · 1995 [cited by applicant]
US 5570412A · Leblanc · 1996 [cited by applicant]
US 5724404A · Garcia et al. · 1998 [cited by applicant]
US 5825871A · Mark · 1998 [cited by applicant]
US 6041116A · Meyers · 2000 [cited by applicant]
US 6134448A · Shoji et al. · 2000 [cited by applicant]
US 6654459B1 · Bala et al. · 2003 [cited by applicant]
US 6735457B1 · Link et al. · 2004 [cited by applicant]
US 6765531B2 · Anderson · 2004 [cited by applicant]
US 7133792B2 · Murakami et al. · 2006 [cited by applicant]
US 7545961B2 · Ahern et al. · 2009 [cited by applicant]
US 7787598B2 · Agapi et al. · 2010 [cited by applicant]
US 8050393B2 · Apple · 2011 [cited by applicant]
US 8085907B2 · Jaiswal · 2011 [cited by applicant]
US 8145562B2 · Wasserblat et al. · 2012 [cited by applicant]
US 8223755B2 · Jennings et al. · 2012 [cited by applicant]
US 8260350B2 · Jaiswal et al. · 2012 [cited by applicant]
US 8311218B2 · Mehmood et al. · 2012 [cited by applicant]
US 8385888B2 · Labrador et al. · 2013 [cited by applicant]
US 8417289B2 · Jaiswal et al. · 2013 [cited by applicant]
US 8476600B2 · Lee et al. · 2013 [cited by applicant]
US 8768648B2 · Panther et al. · 2014 [cited by applicant]
US 8925058B1 · Dotan et al. · 2014 [cited by applicant]
US 9060057B1 · Danis · 2015 [cited by applicant]
US 9078143B2 · Rodriguez et al. · 2015 [cited by applicant]
US 9372976B2 · Bukai · 2016 [cited by applicant]
US 9405967B2 · Samet · 2016 [cited by applicant]
US 9412365B2 · Biadsy et al. · 2016 [cited by applicant]
US 9502038B2 · Wang et al. · 2016 [cited by applicant]
US 10032451B1 · Mamkina et al. · 2018 [cited by applicant]
US 10157272B2 · Kim et al. · 2018 [cited by applicant]
US 10257591B2 · Gaubitch et al. · 2019 [cited by applicant]
US 10311872B2 · Howard et al. · 2019 [cited by applicant]
US 10388272B1 · Thomson et al. · 2019 [cited by applicant]
US 10418957B1 · Wang et al. · 2019 [cited by applicant]
US 20020181448A1 · Uskela et al. · 2002 [cited by applicant]
US 20030012358A1 · Kurtz et al. · 2003 [cited by applicant]
US 20080212846A1 · Yamamoto et al. · 2008 [cited by applicant]
US 20110051905A1 · Maria Poels · 2011 [cited by applicant]
US 20110082877A1 · Gupta et al. · 2011 [cited by applicant]
US 20110123008A1 · Sarnowski · 2011 [cited by applicant]
US 20140214417A1 · Wang et al. · 2014 [cited by applicant]
US 20140244257A1 · Colibro et al. · 2014 [cited by applicant]
US 20140254778A1 · Zeppenfeld et al. · 2014 [cited by applicant]
US 20150067822A1 · Randall · 2015 [cited by applicant]
US 20150088509A1 · Gimenez et al. · 2015 [cited by applicant]
US 20150112682A1 · Rodriguez et al. · 2015 [cited by applicant]
US 20150120027A1 · Cote et al. · 2015 [cited by applicant]
US 20150127336A1 · Lei et al. · 2015 [cited by applicant]
US 20160063235A1 · Tussy · 2016 [cited by applicant]
US 20160248768A1 · Mclaren et al. · 2016 [cited by applicant]
US 20160269411A1 · Malachi · 2016 [cited by applicant]
US 20160292149A1 · Mote · 2016 [cited by examiner]
US 20160293167A1 · Chen et al. · 2016 [cited by applicant]
US 20160293185A1 · Van Dorn et al. · 2016 [cited by applicant]
US 20160314616A1 · Su · 2016 [cited by applicant]
US 20170069327A1 · Heigold et al. · 2017 [cited by applicant]
US 20170076727A1 · Ding · 2017 [cited by examiner]
US 20170084292A1 · Yoo · 2017 [cited by applicant]
US 20170169113A1 · Bhatnagar et al. · 2017 [cited by applicant]
US 20170169815A1 · Zhan et al. · 2017 [cited by applicant]
US 20170222960A1 · Agarwal et al. · 2017 [cited by applicant]
US 20170256254A1 · Huang et al. · 2017 [cited by applicant]
US 20170302794A1 · Spievak et al. · 2017 [cited by applicant]
US 20170330586A1 · Roblek et al. · 2017 [cited by applicant]
US 20170351907A1 · Bataller et al. · 2017 [cited by applicant]
US 20170359362A1 · Kashi et al. · 2017 [cited by applicant]
US 20170365118A1 · Nurbegovic et al. · 2017 [cited by applicant]
US 20170372128A1 · Owen · 2017 [cited by applicant]
US 20180082689A1 · Khoury et al. · 2018 [cited by applicant]
US 20180082692A1 · Khoury et al. · 2018 [cited by applicant]
US 20180150740A1 · Wang et al. · 2018 [cited by applicant]
US 20180174575A1 · Bengio et al. · 2018 [cited by applicant]
US 20180197548A1 · Palakodety et al. · 2018 [cited by applicant]
US 20180254046A1 · Khoury et al. · 2018 [cited by applicant]
US 20180293221A1 · Finkelstein et al. · 2018 [cited by applicant]
US 20180366125A1 · Liu et al. · 2018 [cited by applicant]
US 20190013013A1 · Mclaren et al. · 2019 [cited by applicant]
US 20190050716A1 · Barkan et al. · 2019 [cited by applicant]
US 20190156210A1 · He et al. · 2019 [cited by applicant]
US 20190180759A1 · Fontaine et al. · 2019 [cited by applicant]
US 20190236102A1 · Wade · 2019 [cited by examiner]
US 20200211561A1 · Degraye · 2020 [cited by examiner]
US 20200365160A1 · Nassar · 2020 [cited by examiner]
US 20200380980A1 · Shum · 2020 [cited by examiner]
US 20220269850A1 · Lebas · 2022 [cited by examiner]
JP 2014170488A · 2014 [cited by applicant]
JP 2018508799A · 2018 [cited by applicant]
JP 2019109503A · 2019 [cited by applicant]
WO WO2018166316A1 · 2018 [cited by applicant]
International Preliminary Report on Patentability for PCT/US2021/027474 dated Oct. 27, 2022 (11 pages). [cited by applicant]
Invitation to Pay Additional Fees and, Where Applicable, Protest Fees for PCT/US2021/027474 dated Jun. 2, 2021 (2 pages). [cited by applicant]
International Search Report and Written Opinion for PCT/US2021/027474 dated Sep. 2, 2021 (14 pages). [cited by applicant]
Malik, Hafiz, “Securing Voice-driven Interfaces against Fake (Cloned) Audio Attacks”, IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), 2019, retrieved Jul. 29, 2021 from URL: https://ieeexplore… [cited by applicant]
Alam et al., “Spoofing Detection on the ASVspoof 2015 Challenge Corpus employing Deep Neural Networks”, Odyssey 2016, vol. 2016, Jun. 21, 2016, pp. 270-276. [cited by applicant]
Alegre et al., “Re-assessing the Threat of Replay Spoofing Attacks Against Automatic Speaker Verification”, 2014 International Conference of BIOSIG, Sep. 2014, pp. 157-168. [cited by applicant]
Buyukyilmaz et al., “Voice Gender Recognition Using Deep Learning”, Advances in Compuer Science Research, vol. 58, pp. 409-411, 2016. [cited by applicant]
Canadian Examination Report dated Oct. 16, 2019, issued in corresponding Canadian Application No. 3,032,807, 3 pages. [cited by applicant]
Douglas A. Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Erbilik, et al. “Improved age prediction from biometric data using multimodal configurations” BIOSIG, Sep. 10-12, 2014, pp. 179-186. [cited by applicant]
Ergunay et al., “On the Vulnerability of Speaker Verification to Realistic Voice Spoofing”, IEEE ICBTAS IEEE, Sep. 2015, pp. 1-6. [cited by applicant]
Finnian Kelly et al., “Score-Aging Calibration for Speaker Verification”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, IEEE, USA, vol. 24, No. 12, Dec. 1, 2016, pp. 2414-2424, XP058309805, ISSN: 2329… [cited by applicant]
Han et al., “Age Estimation from Face Images: Human vs. Machine Performance,” ICB, Jun. 4-7, 2013. [cited by applicant]
Hoffer, et al. “Deep Metric Learning Using Triplet Network”, ICLR 2015 (workshop contribution), Mar. 23, 2015, retrieved from <https://arxiv.org/pdf/1412.6622v3.pdf> on Dec. 8, 2016. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2018/017249, mailed May 3, 2018. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2020/054825 dated Jan. 28, 2021. [cited by applicant]
Written Opinion issued in International Application No. PCT/US2017/044849 dated Jan. 11, 2018. [cited by applicant]
Ioffe. et al, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, Proceedings of the 32nd ICML—vol. 37, pp. 448-456, Lille, France, Jul. 6-11, 2015. [cited by applicant]
Jain et al., “Guidelines for Best Practices in Biometrics Research”, ICB, Phuket, Thailand, May 19-22, 2015. [cited by applicant]
Kelly et al., “Score-Aging Calibration for Speaker Verification,” IEEE/ACM TASLP vol. 24, Iss. 12, Aug. 24, 2016, pp. 2414-2424. [cited by applicant]
Ling et al. “A Study of Face Recognition as People Age”, ICCV, 2007. [cited by applicant]
Nagarsheth et al., “Replay Attack Detection Using DNN for Channel Discrimination”, Interspeech 2017, Aug. 20, 2017, pp. 97-101. [cited by applicant]
International Search Report and the Written Opinion Issued in International Application No. PCT/US2018/020624 dated May 29, 2018. [cited by applicant]
Reynolds et al., “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing 10, 2000, pp. 19-41. [cited by applicant]
Richardson, et al., “Deep Neural Network Approaches to Speaker and Language Recognition”, IEEE Signal Processing Letters, vol. 22, No. 10, Oct. 2015, pp. 1671-1675. [cited by applicant]
Schulzrinne et al., “RTP Payload for DTMF Digits, Telephone Tones, and Telephony Signals” Columbia University, Dec. 2006, <https://tools.ielf.org/html/rfc4733.>. [cited by applicant]
Sdjadi et al., “Speaker Age Estimation on Conversational Telephone Speech Using Senone Posterior Based I-vectors”, IEEE ICASSP, Mar. 20-25, 2016. [cited by applicant]
Snyder et al., “X-Vectors: Robust DNN Embeddings for Speaker Recognition,” Center for Language and Speech Processing & Human Language Technology Center of Excellence, the Johns Hopkins University, 2018. [cited by applicant]
Todisco, et al., “A New Feature for Automatic Speaker Verification Anti-Spoofing: Constant Q Cepstral Coefficients”, Odyssey 2016, Jun. 21-24, 2016, Bilbao, Spain, pp. 283-290. [cited by applicant]
Wu et al., “ASVspoof: The Automatic Speaker Verification Spoofing and Countermeasures Challenge”, IEEE Journal of Selected Topics in Signal Processing, IEEE, US, vol. 11, No. 4, Feb. 17, 2017, pp. 588-604. [cited by applicant]
Yu, et al., “Spoofing Detection in Automatic Speaker Verification Systems Using DNN Classifiers and Dynamic Acoustic Features”, IEEE Transactions on Neural Networks and Learning Systems, vol. PP, Issue 99, Dec. 4, 2017. [cited by applicant]
Zhang et al., “An Investigation of Deep-Learning Frameworks for Speaker Verification Antispoofing”, IEEE Journal of Selected Topics in Signal Processing, IEEE, US, vol. 11, No. 4, Jan. 16, 2017, pp. 684-694. [cited by applicant]
Extended European Search Report dated Mar. 19, 2024 on EPO App. 21789203.3 (7 pages). [cited by applicant]
EPO Examination Report for European Application No. 21789203.3 mailing date Dec. 20, 2024, 12 pages. [cited by applicant]
Garcia-Romero et al., “Speaker Diarization Using Deep Neural Network Embeddings,” 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Mar. 5, 2017 (Mar. 5, 2017), pp. 4930-4934,… [cited by applicant]
Judith A Markowitz: “Voice Biometrics,” ARXIV:2003.08934V1, United States, vol. 43, No. 9, Sep. 1, 2000 (Sep. 1, 2000), pp. 66-73, XP058232480, DOI: 10.1145/348941.348995. [cited by applicant]
Snyder et al: “Deep Neural Network-Based Speaker Embeddings for End-to-End Speaker Verification,” 2016 IEEE Spoken Language Technology Workshop (SL T), IEEE, Dec. 13, 2016 (Dec. 13, 2016), pp. 165-170, XP033061735, DOI:… [cited by applicant]
Yanick et al: “Speaker Identification and Clustering Using Convolutional Neural Networks”, 2016 IEEE 26th International Workshop on Machine Learning for Signal Processing (MLSP), IEEE, Sep. 13, 2016 (Sep. 13, 2016), pp.… [cited by applicant]
Zhang et al., “End-to-End Text-Independent Speaker Verification with Triplet Loss on Short Utterances”, Interspeech 2017, Jan. 1, 2017 (Jan. 1, 2017), pp. 1487-1491, XP093213121, DOI: 10.21437/Interspeech.2017-1608. [cited by applicant]
Garcia-Romero et al.: “Speaker Diarization Using Deep Neural Network Embeddings”, Published in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4930-4934. [cited by applicant]
Markowitz (2000): “Voice Biometrics”, Published in Communications of the ACM on Sep. 2000, vol. 43, Issue 9, pp. 66-73. [cited by applicant]
Lukic et al.: “Speaker Identification and Clustering Using Convolutional Neural Networks”, Published in Conference: 2016 IEEE 26th International Workshop on Machine Learning for Signal Processing (MLSP) on Sep. 2016, Re… [cited by applicant]
Observations by third parties pursuant to Article 115 EPC issued in European Patent Application 21789203.3 dated Sep. 27, 2024 (12 Pages). [cited by applicant]
Snyder et al.: “Deep Neural Network-Based Speaker Embeddings for End-To-End Speaker Verification”, Published in Conference: 2016 IEEE Spoken Language Technology Workshop (SLT) on Dec. 2016, pp. 165-170, Retrieved on Aug… [cited by applicant]
Zhang et al.: “End-to-End Text-Independent Speaker Verification with Triplet Loss on Short Utterances”, Published in Conference: Interspeech 2017 on Aug. 2017, pp. 1487-1491, Retrieved from: http://dx.doi.org/10.21437/I… [cited by applicant]
JP Office Action for Application No. 2022-561448 mailing date May 19, 2025, 6 pages with English translation. [cited by applicant]
Office Action issued in corresponding Singaporean Patent Application No. 11202253750C dated Sep. 4, 2025 (8 pages). [cited by applicant]
Cited By (1)
US 12,567,419