IP Library › Granted Patent US 12,495,107
Granted Patent B2
US 12,495,107 · App. 18/581,804 · Granted Dec 9, 2025

Methods and systems for automatic discovery of fraudulent calls using speaker recognition

Inventors: Zhiyuan Guan (McLean, VA); Carl S. Ashby (Montpelier, VA); Isabelle Alice Yvonne Moulinier (Richfield, MN); Mark E. Dickison (Falls Church, VA)
Assignee: Capital One Services, LLC
H04M1/67G10L17/00H04M3/385H04M2201/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,495,107
App. No.
18/581,804
Granted
Dec 9, 2025
Kind
B2
Abstract

A method for determining potentially undesirable voices, in embodiments, includes: receiving audio recordings comprising voices associated with undesirable activity, and determining audio components of each of the audio recordings. The method may further comprise generating a multi-dimensional vector of the audio components for each of the plurality of audio recordings, and comparing audio components between the multi-dimensional vectors to determine clusters of multi-dimensional vectors, each cluster comprising two or more of the multi-dimensional vectors of audio components, wherein each cluster corresponds to a blacklisted voice. The method may further comprise receiving an audio recording or audio stream, and determining whether the audio recording or audio stream is associated with a voice associated with undesirable activity based on a comparison to the clusters.

Claims (51)

1 . A computer-implemented method for identifying a voice associated with undesirable activity, comprising:

obtaining voiceprint data that includes vector representations of voices from a plurality of audio recordings;

identifying one or more vector representations that are unlikely to cluster with others of the vector representations;

generating a sub-set of vector representations by discarding the one or more identified vector representations that are unlikely to cluster with others of the vector representations;

grouping the sub-set of vector representations into a plurality of clusters by performing a clustering operation:

identifying one or more spurious clusters by performing a filtering operation on the plurality of clusters; and

forming a voiceprint library for identifying a voice associated with undesirable activity by discarding the one or more spurious clusters from the plurality of clusters.

2 . The computer-implemented method of claim 1 , wherein the clustering operation includes iteratively, until no grouping of one or more vector representations has a similarity above a predetermined threshold, performing the operations of:

determining a similarity between each grouping and one or more neighboring groupings; and

merging groupings that have a similarity above the predetermined threshold.

3 . The computer-implemented method of claim 2 , wherein the similarity between each grouping and the one or more neighboring groupings is defined by a complete linkage criterion between all vector representations in the groupings to be merged.

4 . The computer-implemented method of claim 2 , wherein the similarity between each grouping and the one or more neighboring groupings is defined by an average linkage criterion between the groupings to be merged.

5 . The computer-implemented method of claim 1 , wherein the clustering operation includes an agglomerative hierarchical clustering process.

6 . The computer-implemented method of claim 1 , wherein:

the voiceprint data further includes metadata associated with the vector representations; and

one or more of the identifying, grouping, or forming is further based on the metadata.

7 . The computer-implemented method of claim 6 , wherein the metadata includes information regarding text corresponding to at least a portion of speech from a corresponding audio recording.

8 . The computer-implemented method of claim 7 , wherein the information regarding text includes a frequency count for each word in a predetermined set of words.

9 . The computer-implemented method of claim 6 , wherein the metadata includes account information associated with a corresponding audio recording.

10 . The computer-implemented method of claim 1 , wherein the filtering operation is based on one or more of cluster size, cluster coherence, metadata associated with vector representations within each cluster, or a proportion of vector representations within each cluster that have a predetermined association with undesirable activity.

11 . The computer-implemented method of claim 1 , wherein the filtering operation is configured to identify one or more clusters that only contain vector representations associated with a single account, such that forming the voiceprint library includes discarding clusters only associated with a single account.

12 . The computer-implemented method of claim 1 , further comprising:

configuring the voiceprint library as a blacklist, such that a future audio recording is compared against each cluster in the voiceprint library, the voiceprint library configured to blacklist the future audio recording upon a similarity between the audio recording and at least one of the clusters being above a predetermined similarity threshold.

13 . A computer-implemented method for identifying a voice associated with undesirable activity, comprising:

obtaining voiceprint data that includes vector representations of voices from a plurality of audio recordings;

performing a nearest-neighbor analysis on the vector representations to identify one or more vector representations having a similarity below a predetermined threshold;

generating a sub-set of vector representations by discarding the one or more identified vector representations;

grouping the sub-set of vector representations into a plurality of clusters by performing a clustering operation:

identifying one or more spurious clusters by performing a filtering operation on the plurality of clusters, wherein the one or more clusters are identified as spurious based on one or more of cluster size, cluster coherence, number of voices represented by vectors within a same cluster, or number of accounts associated with vector representations within a same cluster; and

forming a voiceprint library for identifying a voice associated with undesirable activity by discarding the one or more spurious clusters from the plurality of clusters.

14 . The computer-implemented method of claim 13 , wherein the nearest-neighbor analysis is an exact nearest-neighbor analysis.

15 . The computer-implemented method of claim 13 , wherein the nearest-neighbor analysis is an approximate nearest-neighbor analysis.

16 . The computer-implemented method of claim 13 , wherein the similarity for each vector representation is based on a similarity score for the vector representation and each of a predetermined number of nearest neighbors being below the predetermined threshold.

17 . The computer-implemented method of claim 13 , wherein the similarity includes one or more of a cosine similarity or a log-likelihood ratio.

18 . The computer-implemented method of claim 13 , wherein the performing of the nearest-neighbor analysis and the generating of the sub-set of vector representations is performed in response to determining that a total quantity of the vector representations is above a predetermined quantity threshold.

19 . The computer-implemented method of claim 13 , wherein the performing of the nearest-neighbor analysis and the generating of the sub-set of vector representations is performed in response to determining that a predicted time to perform the clustering operation on the vector representations is above a predetermined time threshold.

20 . A computer-implemented method for identifying a voice associated with undesirable activity, comprising:

obtaining voiceprint data that includes vector representations of voices from a plurality of audio recordings;

identifying one or more vector representations that are unlikely to cluster with others of the vector representations by performing a nearest-neighbor analysis on the vector representations to identify one or more vector representations having a similarity below a predetermined threshold;

generating a sub-set of vector representations by discarding the one or more identified vector representations that are unlikely to cluster with others of the vector representations;

grouping the sub-set of vector representations into a plurality of clusters by performing a clustering operation, wherein the clustering operation includes iteratively, until no grouping of one or more vector representations has a similarity above a predetermined threshold, performing the operations of:

determining a similarity between each grouping and one or more neighboring groupings; and

merging groupings that have a similarity above the predetermined threshold:

identifying one or more spurious clusters by performing a filtering operation on the plurality of clusters, wherein a spurious cluster is a cluster that one or more of:

(i) includes vector representations of more than one voice;

(ii) includes vector representations that were clustered together due to common background noise or interference in corresponding audio recordings;

(iii) has a cluster coherence below a predetermined threshold;

(iv) includes less than a threshold number of vector representations; or

(v) includes vector representations associated with only one account;

forming a voiceprint library for identifying a voice associated with undesirable activity by discarding the one or more spurious clusters from the plurality of clusters; and

configuring the voiceprint library as a blacklist, such that a future audio recording is compared against each cluster in the voiceprint library, the voiceprint library configured to blacklist the future audio recording upon a similarity between the audio recording and at least one of the clusters being above a predetermined similarity threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2024
From: GUAN, ZHIYUAN; ASHBY, CARL S.; MOULINIER, ISABELLE ALICE YVONNE; DICKISON, MARK E.
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 066552/0130 →
Continuity (4)
Continuation 17645179 · Dec 20, 2021
Continuation 16845171 · Apr 10, 2020
Continuation 16360783 · Mar 21, 2019
Related Publication 20240195905A1 · Jun 13, 2024
References Cited (48)
US 5623539A · Bassenyemukasa et al. · 1997 [cited by applicant]
US 6078807A · Dunn et al. · 2000 [cited by applicant]
US 6442519B1 · Kanevsky et al. · 2002 [cited by applicant]
US 7299177B2 · Broman et al. · 2007 [cited by applicant]
US 8423349B1 · Huynh et al. · 2013 [cited by applicant]
US 8510215B2 · Gutierrez et al. · 2013 [cited by applicant]
US 8812318B2 · Broman et al. · 2014 [cited by applicant]
US 8924285B2 · Rajakumar et al. · 2014 [cited by applicant]
US 9349373B1 · Williams · 2016 [cited by examiner]
US 9620123B2 · Faians · 2017 [cited by examiner]
US 9716791B1 · Moran · 2017 [cited by applicant]
US 9824692B1 · Khoury et al. · 2017 [cited by applicant]
US 10043189B1 · Jones · 2018 [cited by applicant]
US 10043190B1 · Jones · 2018 [cited by applicant]
US 10659588B1 · Guan et al. · 2020 [cited by applicant]
US 11632459B2 · Chawla · 2023 [cited by examiner]
US 20020010587A1 · Pertrushin · 2002 [cited by applicant]
US 20030009333A1 · Sharma · 2003 [cited by examiner]
US 20040240631A1 · Broman et al. · 2004 [cited by applicant]
US 20050060157A1 · Daugherty et al. · 2005 [cited by applicant]
US 20070071200A1 · Brouwer · 2007 [cited by applicant]
US 20070280436A1 · Rajakumar · 2007 [cited by applicant]
US 20070282605A1 · Rajakumar · 2007 [cited by applicant]
US 20080300877A1 · Gilbert et al. · 2008 [cited by applicant]
US 20120084078A1 · Moganti et al. · 2012 [cited by applicant]
US 20120130713A1 · Shin et al. · 2012 [cited by applicant]
US 20140136194A1 · Warford et al. · 2014 [cited by applicant]
US 20140162598A1 · Villa-Real · 2014 [cited by examiner]
US 20140214676A1 · Bukai · 2014 [cited by applicant]
US 20140330563A1 · Faians et al. · 2014 [cited by applicant]
US 20150026580A1 · Kang et al. · 2015 [cited by applicant]
US 20150067822A1 · Randall · 2015 [cited by applicant]
US 20150112682A1 · Rodriguez et al. · 2015 [cited by applicant]
US 20150207934A1 · Tolksdorf et al. · 2015 [cited by applicant]
US 20150269946A1 · Jones · 2015 [cited by examiner]
US 20170118335A1 · Brackett et al. · 2017 [cited by applicant]
US 20180082689A1 · Khoury · 2018 [cited by examiner]
US 20180190296A1 · Williams et al. · 2018 [cited by applicant]
US 20180226079A1 · Khoury · 2018 [cited by examiner]
US 20180330728A1 · Gruenstein et al. · 2018 [cited by applicant]
US 20190037081A1 · Rao · 2019 [cited by examiner]
US 20200099781A1 · Chawla · 2020 [cited by examiner]
US 20200211571A1 · Shoa · 2020 [cited by examiner]
US 20210105358A1 · Jolly et al. · 2021 [cited by applicant]
US 20230061617A1 · Fainstein · 2023 [cited by examiner]
CN 106506524A · 2017 [cited by applicant]
EP 2437477A1 · 2012 [cited by applicant]
WO 2017048360A1 · 2017 [cited by applicant]