IP Library › Granted Patent US 11,943,383
Granted Patent B2
US 11,943,383 · App. 17/645,179 · Granted Mar 26, 2024

Methods and systems for automatic discovery of fraudulent calls using speaker recognition

Inventors: Zhiyuan Guan (McLean, VA); Carl S. Ashby (Montpelier, VA); Isabelle Alice Yvonne Moulinier (Richfield, MN); Mark E. Dickison (Falls Church, VA)
Assignee: Capital One Services, LLC
H04M1/67G10L17/00H04M3/385H04M2201/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,943,383
App. No.
17/645,179
Granted
Mar 26, 2024
Kind
B2
Abstract

A method for determining potentially undesirable voices, in embodiments, includes: receiving audio recordings comprising voices associated with undesirable activity, and determining audio components of each of the audio recordings. The method may further comprise generating a multi-dimensional vector of the audio components for each of the plurality of audio recordings, and comparing audio components between the multi-dimensional vectors to determine clusters of multi-dimensional vectors, each cluster comprising two or more of the multi-dimensional vectors of audio components, wherein each cluster corresponds to a blacklisted voice. The method may further comprise receiving an audio recording or audio stream, and determining whether the audio recording or audio stream is associated with a voice associated with undesirable activity based on a comparison to the clusters.

Claims (88)

1. A computer-implemented method for identifying a voice associated with undesirable activity, comprising:

receiving an audio recording or audio stream associated with a voice;

determining a plurality of audio components of the audio recording or audio stream;

comparing the voice, using the determined plurality of audio components, with a library of voices, each voice in the library of voices defined by a cluster of audio recordings clustered together based on audio components of the audio recordings such that audio recordings associated with a same voice but different accounts are associated with a same cluster;

based on the comparing of the voice to the library of voices, identifying a cluster likely to correspond to a same voice as the voice associated with the received audio recording or audio stream;

determining, based on account information associated with the received audio recording or audio stream and metadata associating each audio recording in the identified cluster with a respective account, whether the voice associated with the identified cluster is associated with a plurality of accounts associated with different people; and

in response to determining that the identified cluster is associated with a plurality of accounts associated with different people, identifying the received audio recording or audio stream as associated with undesirable activity.

2. The computer-implemented method of claim 1 , further comprising:

adding the identified cluster to a voice blacklist.

3. The computer-implemented method of claim 1 , further comprising:

in response to determining that the received audio recording or audio stream is associated with undesirable activity, determining a genuine account associated with the received audio recording or audio stream, and one or more of:

determining a verified owner of the genuine account, and transmitting an alert to the verified owner;

adding one or more of a restriction or a verification requirement to the genuine account; or

determining an entity associated with the genuine account, and transmitting an alert to the entity.

4. The computer-implemented method of claim 1 , further comprising:

in response to determining that the received audio recording or audio stream is associated with undesirable activity, determining a genuine account associated with the voice associated with the received audio recording or audio stream; and

one or more of:

determining, based on the genuine account, a genuine identity of a fraudulent caller associated with the received audio recording or audio stream;

flagging the genuine account as associated with a fraudulent caller;

adding a restriction to the genuine account; or

transmitting information associated with the genuine account to a law enforcement entity.

5. The computer-implemented method of claim 1 , further comprising:

adding a further audio recording based on the received audio recording or audio stream to the library of voices; and

updating a clustering of the audio recordings in the library of voices.

6. The computer-implemented method of claim 5 , wherein updating the clustering of the audio recordings includes:

generating a respective multi-dimensional vector based on the respective plurality of audio components of each audio recording; and

comparing the respective multi-dimensional vectors.

7. The computer-implemented method of claim 6 , wherein comparing the respective multi-dimensional vectors includes one or more of:

determining a minimum similarity between each of the respective multi-dimensional vectors; or

determining an average similarity between each of the respective multi-dimensional vectors.

8. The computer-implemented method of claim 1 , wherein identifying the cluster likely to correspond to a same voice as the voice associated with the received audio recording or audio stream includes:

determining that the voice associated with the received audio recording or audio stream does not correspond to any of the clusters in the library of voices; and

generating a new cluster based on the received audio recording or audio stream.

9. The computer-implemented method of claim 1 , wherein identifying the received audio recording or audio stream as associated with undesirable activity is further performed in response to determining that the identified cluster is not included in a voice whitelist.

10. The computer-implemented method of claim 1 , wherein:

the library of voices represents each audio recording as an n-dimensional vector formed by the plurality of audio components of the audio recording; and

comparing the voice, using the determined plurality of audio components, includes:

representing the determined plurality of audio components as a further n-dimensional vector; and

performing a clustering operation of the further n-dimensional vector with the n-dimensional vectors of the library of voices.

11. A system for identifying a voice associated with undesirable activity, comprising:

at least one memory storing instructions and a library of voices, each voice in the library of voices defined by a cluster of audio recordings clustered together based on audio components of the audio recordings such that audio recordings associated with a same voice but different accounts are associated with a same cluster; and

at least one processor operatively connected to the at least one memory, and configured to execute the instructions to perform operations, including:

receiving an audio recording or audio stream associated with a voice;

determining a plurality of audio components of the audio recording or audio stream;

comparing the voice, using the determined plurality of audio components, with the library of voices;

based on the comparing of the voice to the library of voices, identifying a cluster likely to correspond to a same voice as the voice associated with the received audio recording or audio stream;

determining, based on account information associated with the received audio recording or audio stream and metadata associating each audio recording in the identified cluster with a respective account, whether the voice associated with the identified cluster is associated with a plurality of accounts associated with different people; and

in response to determining that the identified cluster is associated with a plurality of accounts associated with different people, identifying the received audio recording or audio stream as associated with undesirable activity.

12. The system of claim 11 , wherein the operations further include:

adding the identified cluster to a voice blacklist.

13. The system of claim 11 , wherein the operations further include:

in response to determining that the received audio recording or audio stream is associated with undesirable activity, determining a genuine account associated with the received audio recording or audio stream, and one or more of:

determining a verified owner of the genuine account, and transmitting an alert to the verified owner;

adding one or more of a restriction or a verification requirement to the genuine account; or

determining an entity associated with the genuine account, and transmitting an alert to the entity.

14. The system of claim 11 , wherein the operations further include:

in response to determining that the received audio recording or audio stream is associated with undesirable activity, determining a genuine account associated with the voice associated with the received audio recording or audio stream; and

one or more of:

determining, based on the genuine account, a genuine identity of a fraudulent caller associated with the received audio recording or audio stream;

flagging the genuine account as associated with a fraudulent caller;

adding a restriction to the genuine account; or

transmitting information associated with the genuine account to a law enforcement entity.

15. The system of claim 11 , wherein the operations further include:

adding a further audio recording based on the received audio recording or audio stream to the library of voices; and

updating a clustering of the audio recordings in the library of voices.

16. The system of claim 15 , wherein updating the clustering of the audio recordings includes:

generating a respective multi-dimensional vector based on the respective plurality of audio components of each audio recording; and

comparing the respective multi-dimensional vectors by one or more of:

determining a minimum similarity between each of the respective multi-dimensional vectors; or

determining an average similarity between each of the respective multi-dimensional vectors.

17. The system of claim 11 , wherein identifying the cluster likely to correspond to a same voice as the voice associated with the received audio recording or audio stream includes:

determining that the voice associated with the received audio recording or audio stream does not correspond to any of the clusters in the library of voices; and

generating a new cluster based on the received audio recording or audio stream.

18. The system of claim 11 , wherein identifying the received audio recording or audio stream as associated with undesirable activity is further performed in response to determining that the identified cluster is not included in a voice whitelist.

19. The system of claim 11 , wherein:

the library of voices represents each audio recording as an n-dimensional vector formed by the plurality of audio components of the audio recording; and

comparing the voice, using the determined plurality of audio components, includes:

representing the determined plurality of audio components as a further n-dimensional vector; and

performing a clustering operation of the further n-dimensional vector with the n-dimensional vectors of the library of voices.

20. A computer-implemented method for identifying a voice associated with undesirable activity, comprising:

receiving an audio recording or audio stream associated with a voice;

determining a plurality of audio components of the audio recording or audio stream;

comparing the voice, using the determined plurality of audio components, with a library of voices, each voice in the library of voices defined by a cluster of audio recordings clustered together based on audio components of the audio recordings such that audio recordings associated with a same voice but different accounts are associated with a same cluster;

based on the comparing of the voice to the library of voices, identifying a cluster likely to correspond to a same voice as the voice associated with the received audio recording or audio stream;

determining, based on account information associated with the received audio recording or audio stream and metadata associating each audio recording in the identified cluster with a respective account, whether the voice associated with the identified cluster is associated with a plurality of accounts associated with different people;

in response to determining that the identified cluster is associated with a plurality of accounts associated with different people, identifying the received audio recording or audio stream as associated with undesirable activity;

adding a further audio recording based on the received audio recording or audio stream to the identified cluster; and

adding the identified cluster to a voice blacklist.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2022
From: GUAN, ZHIYUAN; ASHBY, CARL S.; MOULINIER, ISABELLE ALICE YVONNE; DICKISON, MARK E.
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 058584/0270 →
Continuity (3)
Continuation 16845171 · Apr 10, 2020
Continuation 16360783 · Mar 21, 2019
Related Publication 20220116493A1 · Apr 14, 2022
Cited By (1)
US 12,615,335