IP Library Granted Patent US 10,679,630
Granted Patent B2
US 10,679,630 · App. 16/442,368 · Granted Jun 9, 2020

Speaker recognition in the call center

Inventors: Elie Khoury (Atlanta, GA); Matthew Garland (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L17/005G06N7/005G10L15/07G10L15/19G10L15/26G10L17/04G10L17/08G10L17/24H04M1/271H04M2203/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,679,630
App. No.
16/442,368
Granted
Jun 9, 2020
Kind
B2
Abstract

Utterances of at least two speakers in a speech signal may be distinguished and the associated speaker identified by use of diarization together with automatic speech recognition of identifying words and phrases commonly in the speech signal. The diarization process clusters turns of the conversation while recognized special form phrases and entity names identify the speakers. A trained probabilistic model deduces which entity name(s) correspond to the clusters.

Claims (40)

1. A computer-implemented method comprising:

extracting, by a computer, a first set of audio features from a first audio signal containing a first utterance according to a speaker diarization algorithm, the first audio signal received from an enrollee electronic device;

extracting, by the computer, a second set of audio features from a second audio signal containing a second utterance according to the speaker diarization algorithm, the second audio signal received from a caller electronic device;

performing, by the computer, a phonetic and acoustic comparison of the first utterance and the second utterance based upon the first set of audio features and the second set of audio features; and

determining, by the computer based upon the phonetic and acoustic comparison, at least a partial keyword sequence match and a speaker match between the first utterance and the second utterance.

2. The computer-implemented method of claim 1 , wherein the first set of audio features and the second set of audio features comprise at least one of mel-frequency cepstral coefficients (MFCCs), linear predictive cepstral coefficients (LPCCs), or perceptual linear prediction (PLP).

3. The computer-implemented method of claim 1 , wherein the step of performing the phonetic and acoustic comparison further comprises:

executing, by the computer, a modified dynamic time warping process on at least a portion of the first audio signal; and

executing, by the computer, the modified dynamic time warping process on at least a portion of the second audio signal.

4. The computer-implemented method of claim 3 , further comprising:

time hopping, by the computer, on the portion of the first audio signal during the modified dynamic time warping process based upon a first hop size of the first set of audio features; and

time hopping, by the computer, on the portion of the second audio signal during the modified dynamic time warping process based upon a second hop size of the second set of audio features.

5. The computer-implemented method of claim 4 , wherein the partial keyword sequence match includes a first set of one or more keywords in the first utterance and a second set of one or more keywords in the second utterance, the first set of one or more keywords being the same as the second set of one or more keywords.

6. The computer-implemented method of claim 5 , wherein the first set of one or more keywords in the first utterance begins at a different time than the second set of one or more keywords in the second utterance.

7. The computer-implemented method of claim 1 , wherein the first audio signal is an enrollment sample.

8. The computer-implemented method of claim 1 , wherein the second audio signal is a test sample.

9. The computer-implemented method of claim 1 , wherein the second utterance in the second audio signal is from a speaker to be identified.

10. The computer-implemented method of claim 9 , further comprising:

identifying, by the computer, the speaker based upon the partial keyword sequence match and the speaker match.

11. A system comprising:

a non-transitory storage medium storing a plurality of computer program instructions; and

a processor electrically coupled to the non-transitory storage medium and configured to execute the computer program instructions to:

extract a first set of audio features from a first audio signal containing a first utterance according to a speaker diarization algorithm, the first audio signal received from an enrollee electronic device;

extract a second set of audio features from a second audio signal containing a second utterance according to the speaker diarization algorithm, the second audio signal received from a caller electronic device;

perform a phonetic and acoustic comparison of the first utterance and the second utterance based upon the first set of audio features and the second set of audio features; and

determine based upon the phonetic and acoustic comparison, at least a partial keyword sequence match and a speaker match between the first utterance and the second utterance.

12. The system of claim 11 , wherein the first set of audio features and the second set of audio features comprise at least one of mel-frequency cepstral coefficients (MFCCs), linear predictive cepstral coefficients (LPCCs), or perceptual linear prediction (PLP).

13. The system of claim 11 , wherein the processor is configured to further execute the computer program instructions to:

execute a modified dynamic time warping process on at least a portion of the first audio signal; and

execute the modified dynamic time warping process on at least a portion of the second audio signal.

14. The system of claim 13 , wherein the processor is configured to further execute the computer program instructions to:

time hop on the portion of the first audio signal during the modified dynamic time warping process based upon a first hop size of the first set of audio features; and

time hop on the portion of the second audio signal during the modified dynamic time warping process based upon a second hop size of the second set of audio features.

15. The system of claim 14 , wherein the partial keyword sequence match includes a first set of one or more keywords in the first utterance and a second set of one or more keywords in the second utterance, the first set of one or more keywords being the same as the second set of one or more keywords.

16. The system of claim 15 , wherein the first set of one or more keywords in the first utterance begins at a different time than the second set of one or more keywords in the second utterance.

17. The system of claim 11 , wherein the first audio signal is an enrollment sample.

18. The system of claim 11 , wherein the second audio signal is a test sample.

19. The system of claim 11 , wherein the second utterance in the second audio signal is from a speaker to be identified.

20. The system of claim 19 , wherein the processor is configured to further execute the computer program instructions to:

identify the speaker based upon the partial keyword sequence match and the speaker match.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2019
From: KHOURY, ELIE; GARLAND, MATTHEW
To: PINDROP SECURITY, INC.
Reel/Frame 049478/0937 →
Continuity (3)
Continuation 15709290 · Sep 19, 2017
Provisional Application 62396670 · Sep 19, 2016
Related Publication 20190304468A1 · Oct 3, 2019
Cited By (4)
US 12,256,040 US 12,354,608 US 12,525,244 US 12,711,960