IP Library Granted Patent US 11,670,304
Granted Patent B2
US 11,670,304 · App. 16/895,750 · Granted Jun 6, 2023

Speaker recognition in the call center

Inventors: Elie Khoury (Atlanta, GA); Matthew Garland (Atlanta, GA)
Assignee: PINDROP SECURITY, INC.
G10L17/00G06N7/01G10L15/07G10L15/19G10L15/26G10L17/04G10L17/08G10L17/24H04M1/271H04M2203/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,670,304
App. No.
16/895,750
Granted
Jun 6, 2023
Kind
B2
Abstract

Utterances of at least two speakers in a speech signal may be distinguished and the associated speaker identified by use of diarization together with automatic speech recognition of identifying words and phrases commonly in the speech signal. The diarization process clusters turns of the conversation while recognized special form phrases and entity names identify the speakers. A trained probabilistic model deduces which entity name(s) correspond to the clusters.

Claims (44)

1. A computer-implemented method comprising:

extracting, by a computer, from one or more audio signals, a first set of audio features associated with a first speaker and a second set of audio features associated with a second speaker;

generating, by the computer, a first cluster associated with the first speaker based upon the first set of features;

generating, by the computer, a second cluster associated with the second speaker based upon the second set of features;

extracting, by the computer, from each respective set of audio features an entity name and a phrase; and

generating, by the computer, a trained probabilistic model for each respective speaker based upon the respective cluster, the respective entity name, and the respective phrase.

2. The method according to claim 1 , further comprising executing, by the computer, a speech activity detector on an audio signal, the speech activity detector generating one or more speech portions and removing non-speech portions.

3. The method according to claim 1 , further comprising performing, by the computer, speaker diarization on an audio signal, the speaker diarization generating the first cluster and the second cluster.

4. The method according to claim 1 , wherein the computer extracts the one or more features from the one or more audio signals at a given interval, and wherein the computer partitionally clusters the one or more features into the respective clusters for the first speaker and the second speaker at the given interval.

5. The method according to claim 1 , wherein extracting the entity name and the phrase comprises:

identifying, by the computer, text content in each of the clusters extracted from an audio signal by executing an automatic speech recognition process configured to recognize the text content of speech in each of the clusters,

wherein the computer extracts each entity name and each phrase from the text content identified in the first cluster and in the second cluster.

6. The method according to claim 5 , wherein at least one phrase is associated with a respective prior probability that an adjacent utterance is the entity name.

7. The method according to claim 1 , further comprising:

receiving, by the computer, a second audio signal involving at least one of the first speaker and the second speaker; and

generating, by the computer, a next cluster using the probabilistic model associated with the first speaker or the probabilistic model associated with the second speaker.

8. The method according to claim 1 , further comprising identifying, by the computer, at least one of the first speaker and the second speaker as a caller based upon at least one of the trained probabilistic models generated by the computer.

9. The method according to claim 1 , further comprising identifying, by the computer, at least one of the first speaker and the second speaker as an agent based upon at least one of the trained probabilistic models generated by the computer.

10. The method according to claim 1 , wherein the first set of audio features and the second set of audio features each comprise at least one of: mel-frequency cepstral coefficients (MFCCs), linear predictive cepstral coefficients (LPCCs), and perceptual linear prediction (PLP).

11. A system comprising:

a non-transitory storage medium storing a plurality of computer program instructions; and

a processor electrically coupled to the non-transitory storage medium and configured to execute the computer program instructions to:

extract from one or more audio signals, a first set of audio features associated with a first speaker and a second set of audio features associated with a second speaker;

generate a first cluster associated with the first speaker based upon the first set of features;

generate a second cluster associated with the second speaker based upon the second set of features;

extract from each respective set of audio features, an entity name and a phrase; and

generate a trained probabilistic model for each respective speaker based upon the respective cluster, the respective entity name, and the respective phrase.

12. The system according to claim 11 , wherein the processor is further configured to:

execute a speech activity detector an audio signal, the speech activity detector configured to generate one or more speech portions and removing non-speech portions.

13. The system according to claim 11 , wherein the processor is further configured to:

perform speaker diarization on an audio signal, the speaker diarization configured to generate the first cluster and the second cluster.

14. The system according to claim 11 , wherein the processor is configured to extract the one or more features from the one or more audio signals at a given interval, and wherein the processor is configured to partitionally cluster the one or more features into the respective clusters for the first speaker and the second speaker at the given interval.

15. The system according to claim 11 , wherein to extract the entity name and the phrase the processor is configured to:

identify text content in each of the clusters extracted from an audio signal by executing an automatic speech recognition process configured to recognize the text content of speech in each of the clusters,

wherein the processor extracts each entity name and each phrase from the text content identified in the first cluster and in the second cluster.

16. The system according to claim 15 , wherein at least one phrase is associated with a respective prior probability that an adjacent utterance is the entity name.

17. The system according to claim 11 , wherein the processor is further configured to:

receive a second audio signal involving at least one of the first speaker and the second speaker; and

generate a next cluster using the probabilistic model associated with the first speaker or the probabilistic model associated with the second speaker.

18. The system according to claim 11 , wherein the processor is further configured to:

identify at least one of the first speaker and the second speaker as a caller based upon at least one of the trained probabilistic models generated by the computer.

19. The system according to claim 11 , wherein the processor is further configured to:

identify at least one of the first speaker and the second speaker as an agent based upon at least one of the trained probabilistic models generated by the computer.

20. The system according to claim 11 , wherein the first set of audio features and the second set of audio features each comprise at least one of: mel-frequency cepstral coefficients (MFCCs), linear predictive cepstral coefficients (LPCCs), and perceptual linear prediction (PLP).

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2020
From: KHOURY, ELIE; GARLAND, MATTHEW
To: PINDROP SECURITY, INC.
Reel/Frame 052868/0993 →
Continuity (4)
Continuation 16442368 · Jun 14, 2019
Continuation 15709290 · Sep 19, 2017
Provisional Application 62396670 · Sep 19, 2016
Related Publication 20200302939A1 · Sep 24, 2020
Cited By (3)
US 12,223,945 US 12,354,608 US 12,711,960