IP Library › Granted Patent US 11,240,360
Granted Patent B2
US 11,240,360 · App. 16/845,171 · Granted Feb 1, 2022

Methods and systems for automatic discovery of fraudulent calls using speaker recognition

Inventors: Zhiyuan Guan (McLean, VA); Carl S. Ashby (Montpelier, VA); Isabelle Alice Yvonne Moulinier (Richfield, MN); Mark E. Dickison (Falls Church, VA)
Assignee: Capital One Services, LLC
H04M1/67G10L17/00H04M3/385H04M2201/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,240,360
App. No.
16/845,171
Granted
Feb 1, 2022
Kind
B2
Abstract

A computer-implemented method for determining potentially undesirable voices, according to some embodiments, includes: receiving a plurality of audio recordings, the plurality of audio recordings comprising voices associated with undesirable activity, and determining a plurality of audio components of each of the plurality of audio recordings. The method may further comprise generating a multi-dimensional vector of audio components, from the plurality of audio components, for each of the plurality of audio recordings to generate a plurality of multi-dimensional vectors of audio components, and comparing audio components between the plurality of multi-dimensional vectors of audio components to determine a plurality of clusters of multi-dimensional vectors, each cluster of the plurality of clusters comprising two or more of the plurality of multi-dimensional vectors of audio components, wherein each cluster of the plurality of clusters corresponds to a blacklisted voice. The method may further comprise receiving an audio recording or audio stream, and determining whether the audio recording or audio stream is associated with a voice associated with undesirable activity based on a comparison to the plurality of clusters.

Claims (89)

1. A computer-implemented method for identifying a voice associated with undesirable activity, comprising:

receiving a plurality of audio recordings associated with one or more voices;

determining a respective plurality of audio components of each of the plurality of audio recordings;

clustering the plurality of audio recordings, based on the determined respective pluralities of audio components, such that each cluster of the plurality of audio recordings is determined to be associated with a different voice;

applying at least one predetermined criterion to each cluster, wherein:

the at least one predetermined criterion includes a determination that the cluster of the plurality of audio recordings is associated with a plurality of accounts associated with different people; and

the determination that the cluster is associated with the plurality of accounts associated with different people is based on metadata associating each audio recording in the cluster with a respective account; and

in response to a respective cluster satisfying the at least one predetermined criterion, identifying the voice associated with the respective cluster as associated with undesirable activity.

2. The computer-implemented method of claim 1 , further comprising:

adding the voice identified as associated with undesirable activity to a voice blacklist.

3. The computer-implemented method of claim 2 , further comprising:

receiving a further audio recording or audio stream associated with a further voice;

determining a further plurality of audio components of the further audio recording or audio stream;

determining, based on the further plurality of audio components, whether the further voice is associated with a cluster of the plurality of audio recordings; and

in response to determining that the further voice is associated with the cluster, adding the further audio recording or audio stream to the associated cluster.

4. The computer-implemented method of claim 3 , further comprising:

determining whether the further voice is associated with the voice in the voice blacklist; and

in response to determining that the further voice is associated with the voice in the voice blacklist, determining a genuine account based on the further audio recording or audio stream, and one or more of:

determining a verified owner of the genuine account, and transmitting an alert to the verified owner;

adding one or more of a restriction or a verification requirement to the genuine account; or

determining an entity associated with the genuine account, and transmitting an alert to the entity.

5. The computer-implemented method of claim 1 , further comprising:

determining a genuine account associated with the voice identified as associated with undesirable activity; and

one or more of:

determining, based on the genuine account, a genuine identity of a fraudulent caller associated with the voice identified as associated with undesirable activity;

flagging the genuine account as associated with a fraudulent caller;

adding a restriction to the genuine account; or

transmitting information associated with the genuine account to a law enforcement entity.

6. The computer-implemented method of claim 1 , wherein the at least one predetermined criterion further includes one or more of:

a number of audio recordings included in the cluster;

a coherence of the cluster; or

a number or proportion of audio recordings in the cluster flagged as associated with suspicious or fraudulent activity.

7. The computer-implemented method of claim 1 , wherein clustering the plurality of audio recordings includes iteratively performing a clustering process on the plurality of audio recordings until a stopping criterion is reached.

8. The computer-implemented method of claim 1 , wherein clustering the plurality of audio recordings includes:

generating a respective multi-dimensional vector based on the respective plurality of audio components; and

comparing the respective multi-dimensional vectors.

9. The computer-implemented method of claim 8 , wherein comparing the respective multi-dimensional vectors includes one or more of:

determining a minimum similarity between each of the respective multi-dimensional vectors; or

determining an average similarity between each of the respective multi-dimensional vectors.

10. The computer-implemented method of claim 1 , further comprising:

receiving a further audio recording or audio stream associated with a further voice;

determining a further plurality of audio components of the further audio recording or audio stream;

determining a genuine account based on the further audio recording or audio stream;

identifying a cluster of the plurality of audio recordings associated with the genuine account;

determining, based on the further plurality of audio components, whether the further voice corresponds with the voice associated with the identified cluster;

in response to determining that the further voice does not correspond to the voice associated with the identified cluster, determining whether the further voice is associated with a voice included in a voice whitelist; and

in response to determining that the further voice is not associated with a voice included in the voice whitelist, identifying the further voice as associated with undesirable activity.

11. A computer system for identifying a voice associated with undesirable activity, comprising:

a memory storing instructions; and

a processor operatively connected to the memory, and configured to execute the instructions so as to perform acts, including:

receiving a plurality of audio recordings associated with one or more voices;

determining a respective plurality of audio components of each of the plurality of audio recordings;

clustering the plurality of audio recordings, based on the determined respective pluralities of audio components, such that each cluster of the plurality of audio recordings is determined to be associated with a different voice;

applying at least one predetermined criterion to each cluster, wherein the at least one predetermined criterion includes a determination, based on metadata associated with audio recordings in the cluster, that the cluster is associated with a plurality of accounts associated with different people; and

in response to a respective cluster satisfying the at least one predetermined criterion, identifying the voice associated with the respective cluster as associated with undesirable activity.

12. The system of claim 11 , wherein the acts further include adding the voice identified as associated with undesirable activity to a voice blacklist.

13. The system of claim 12 , wherein the acts further include:

receiving a further audio recording or audio stream associated with a further voice;

determining a further plurality of audio components of the further audio recording or audio stream;

determining, based on the further plurality of audio components, whether the further voice is associated with a cluster of the plurality of audio recordings; and

in response to determining that the further voice is associated with the cluster, adding the further audio recording or audio stream to the associated cluster.

14. The system of claim 13 , wherein the acts further include:

determining whether the further voice is associated with the voice in the voice blacklist; and

in response to determining that the further voice is associated with the voice in the voice blacklist, determining a genuine account based on the further audio recording or audio stream, and one or more of:

determining a verified owner of the genuine account, and transmitting an alert to the verified owner;

adding one or more of a restriction or a verification requirement to the genuine account; or

determining an entity associated with the genuine account, and transmitting an alert to the entity.

15. The system of claim 11 , wherein the acts further include:

determining a genuine account associated with the voice identified as associated with undesirable activity; and

one or more of:

determining, based on the genuine account, a genuine identity of a fraudulent caller associated with the voice identified as associated with undesirable activity;

flagging the genuine account as associated with a fraudulent caller;

adding a restriction to the genuine account; or

transmitting information associated with the genuine account to a law enforcement entity.

16. The system of claim 11 , wherein clustering the plurality of audio recordings includes iteratively performing a clustering process on the plurality of audio recordings until a stopping criterion is reached.

17. The system of claim 11 , wherein the acts further include:

receiving a further audio recording or audio stream associated with a further voice;

determining a further plurality of audio components of the further audio recording or audio stream;

determining a genuine account based on the further audio recording or audio stream;

identifying a cluster of the plurality of audio recordings associated with the genuine account;

determining, based on the further plurality of audio components, whether the further voice corresponds with the voice associated with the identified cluster;

in response to determining that the further voice does not correspond to the voice associated with the identified cluster, determining whether the further voice is associated with a voice included in a voice whitelist; and

in response to determining that the further voice is not associated with a voice included in the voice whitelist, identifying the further voice as associated with undesirable activity.

18. A computer implemented method for identifying a voice associated with undesirable activity, comprising:

receiving a plurality of audio recordings associated with one or more voices;

determining a respective plurality of audio components of each of the plurality of audio recordings;

clustering the plurality of audio recordings, based on the determined respective pluralities of audio components, such that each cluster of the plurality of audio recordings is determined to be associated with a different voice, wherein clustering the plurality of audio recordings includes iteratively performing a clustering process on the plurality of audio recordings until a stopping criterion is reached;

applying at least one predetermined criterion to each cluster, wherein the at least one predetermined criterion includes a determination, based on metadata associated with audio recordings in the cluster, that the cluster is associated with a plurality of accounts associated with different people; and

in response to a respective cluster satisfying the at least one predetermined criterion, identifying the voice associated with the respective cluster as associated with undesirable activity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2020
From: GUAN, ZHIYUAN; ASHBY, CARL S.; MOULINIER, ISABELLE ALICE YVONNE; DICKISON, MARK E.
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 052476/0425 →
Continuity (2)
Continuation 16360783 · Mar 21, 2019
Related Publication 20200304622A1 · Sep 24, 2020
Cited By (1)
US 12,615,335