IP Library Granted Patent US 11,245,791
Granted Patent B2
US 11,245,791 · App. 17/086,284 · Granted Feb 8, 2022

Detecting robocalls using biometric voice fingerprints

Inventors: William Li (Seattle, WA); Nam Kim (Seattle, WA); Michael Pruitt (Seattle, WA); Mark Corley (Seattle, WA)
Assignee: Marchex, Inc.
H04M3/4365G10L17/26H04M3/2281H04M3/42042H04M3/42059
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,245,791
App. No.
17/086,284
Granted
Feb 8, 2022
Kind
B2
Abstract

The disclosed system and method detect robocalls using biometric voice fingerprints. The system receives audio input representing a plurality of telephone calls. For at least a portion of the telephone calls, the system analyzes the received audio based on a voice biometrics detection model to identify one or more biometric indicators characterizing a speaker in the analyzed telephone call. The system generates and stores a voice fingerprint characterizing the speaker based on the biometric indicators, and a time of the analyzed telephone call. The system analyzes stored voice fingerprints and times corresponding to speakers in the analyzed telephone calls to determine a frequency of occurrence of each voice fingerprint within an analyzed timeframe. If the frequency of occurrence of a voice fingerprint exceeds a threshold call quantity within the analyzed timeframe, the voice fingerprint is characterized as being associated with a robocaller.

Claims (48)

1. A computer-implemented method for identifying a robocaller in a telephony network, the method comprising:

receiving audio input representing a plurality of telephone calls made in a telephony network during a time period;

for at least some of the plurality of telephone calls:

analyzing, based on a voice biometrics detection model, a portion of the received audio input corresponding to a telephone call to identify one or more biometric indicators characterizing a speaker in the analyzed telephone call;

generating, based on the identified one or more biometric indicators, a voice fingerprint characterizing the speaker in the analyzed telephone call; and

storing in a dataset an indication of the voice fingerprint and a time of the analyzed telephone call within the time period;

determining, within an analyzed timeframe of the dataset, a frequency of occurrence of each voice fingerprint within the dataset; and

characterizing a voice fingerprint within the dataset as being associated with a robocaller when the frequency of occurrence exceeds a threshold call quantity limit.

2. The method of claim 1 , further comprising:

receiving audio input representing a subsequent telephone call;

generating a voice fingerprint associated with the subsequent telephone call;

searching the dataset to identify the generated voice fingerprint; and

taking a corrective action associated with the analyzed telephone call when the generated voice fingerprint is identified as a robocaller in the dataset.

3. The method of claim 2 , wherein the corrective action includes is terminating the subsequent telephone call.

4. The method of claim 1 , wherein the voice biometrics detection model is generated based on one or more artificial intelligence (AI) speech data processing models.

5. The method of claim 1 , wherein the one or more biometric indicators characterizing a speaker in the analyzed telephone call include one or more of volume, speaking rate, pitch, length of pauses, or duration of pauses.

6. The method of claim 1 , wherein the received audio input includes a caller channel and a called channel, and wherein the one or biometric indicators are identified for a speaker only on the caller channel of the analyzed telephone call.

7. The method of claim 1 , wherein determining a frequency of occurrence of each voice fingerprint within the dataset includes calculating a probability that two or more voice fingerprints within the dataset correspond to a common speaker.

8. The method of claim 1 , wherein the analyzed timeframe and the time period are the same.

9. A computer-implemented method for identifying a caller in a telephone call, the method comprising:

receiving an audio input directed to a call recipient, wherein the audio input includes real or simulated human speech;

analyzing, based on a voice biometrics detection model, the received audio input to identify one or more biometric indicators characterizing a speaker in the audio input;

generating a speaker voice fingerprint based on the one or more biometric indicators characterizing the speaker;

querying a dataset comprising a plurality of biometric voice fingerprints corresponding to known callers to determine whether the generated speaker fingerprint matches one or more of the plurality of biometric voice fingerprints corresponding to known callers;

generating and transmitting to the call recipient, in response to determining that the generated speaker voice fingerprint does not match one or more of the plurality of biometric voice fingerprints corresponding to known callers, a caller confirmation request,

wherein the caller confirmation requests that the call recipient provide an indication whether the speaker in the received call is associated with a known caller type;

receiving, in response to the caller confirmation request, an indication from the call recipient that the speaker in the audio input is associated with the known caller type;

adding, in response to the received indication, the generated speaker voice fingerprint to the dataset; and

associating, in the dataset, the generated speaker voice fingerprint with the known caller type.

10. The method of claim 9 , wherein the voice biometrics detection model is generated based on one or more artificial intelligence (AI) speech data processing models, and wherein the one or more biometric indicators characterizing a speaker in the audio input include volume, speaking rate, pitch, length of pauses, and duration of pauses.

11. The method of claim 9 , wherein the known caller type is a type of caller classified as a legitimate caller.

12. The method of claim 9 , wherein the known caller type is a type of caller classified as a spam caller or robocaller.

13. The method of claim 9 , wherein the caller confirmation request is transmitted to the call recipient via a graphical user interface (GUI), a text message, or an email.

14. A non-transitory computer-readable medium comprising instructions configured to cause one or more processors to perform a method for identifying a robocaller in a telephony network, the method comprising:

receiving audio input representing a plurality of telephone calls made in a telephony network during a defined time period;

for at least some of the plurality of telephone calls:

analyzing, based on a voice biometrics detection model, a portion of the received audio input corresponding to a telephone call to identify one or more biometric indicators characterizing a speaker in the analyzed telephone call;

generating, based on the identified one or more biometric indicators, a voice fingerprint characterizing the speaker in the analyzed telephone call; and

storing in a dataset an indication of the voice fingerprint and a time of the analyzed telephone call within the defined time period;

determining, within an analyzed timeframe of the dataset, a frequency of occurrence of each voice fingerprint within the dataset; and

characterizing a voice fingerprint within the dataset as being associated with a robocaller when the frequency of occurrence exceeds a threshold call quantity limit.

15. The computer-readable medium of claim 14 , wherein the method is performed during an analyzed telephone call, and wherein a voice fingerprint characterizing a speaker in the analyzed telephone call is characterized as being associated with a robocaller based on the frequency of occurrence of the voice fingerprint exceeding the threshold call quantity limit, the method further comprising:

taking corrective action to terminate the analyzed telephone call.

16. The computer-readable medium of claim 15 , wherein taking corrective action to terminate the analyzed telephone call includes generating and transmitting to a user a robocaller confirmation request, and receiving from the user an indication in response to the robocaller confirmation request.

17. The computer-readable medium of claim 14 , wherein the voice biometrics detection model is generated based on one or more artificial intelligence (AI) speech data processing models, and wherein the one or more biometric indicators characterizing a speaker in the analyzed telephone call include volume, speaking rate, pitch, length of pauses, and duration of pauses.

18. The computer-readable medium of claim 14 , wherein analyzing a portion of the received audio input corresponding to a telephone call to identify one or more biometric indicators characterizing a speaker in the analyzed telephone call includes identifying a caller channel and a called channel within the analyzed telephone call, and wherein the one or biometric indicators are identified for a speaker only on the caller channel of the analyzed telephone call.

19. The computer-readable medium of claim 14 , wherein determining a frequency of occurrence of each voice fingerprint within the dataset includes calculating a probability that two or more voice fingerprints within the dataset correspond to a common speaker.

20. The computer-readable medium of claim 14 , wherein the analyzed timeframe and the defined time period are the same.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 24, 2020
From: LI, WILLIAM; KIM, NAM; PRUITT, MICHAEL; CORLEY, MARK
To: MARCHEX, INC.
Reel/Frame 054456/0179 →
Continuity (2)
Provisional Application 62928222 · Oct 30, 2019
Related Publication 20210136200A1 · May 6, 2021