IP Library Granted Patent US 9,043,207
Granted Patent B2
US 9,043,207 · App. 13/509,606 · Granted May 26, 2015

Speaker recognition from telephone calls

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,043,207
App. No.
13/509,606
Granted
May 26, 2015
Kind
B2
Abstract

The present invention relates to a method for speaker recognition, comprising the steps of obtaining and storing speaker information for at least one target speaker; obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker; classifying the speech samples according to the at least one unknown speaker thereby providing speaker-dependent classes of speech samples; extracting speaker information for the speech samples of each of the speaker-dependent classes of speech samples; combining the extracted speaker information for each of the speaker-dependent classes of speech samples; comparing the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and determining whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.

Claims (26)

1. A method for speaker recognition, comprising the steps of

obtaining and storing, in a database on a computer, speaker information for at least one target speaker;

obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker;

classifying, using software stored and operating on the computer, the speech samples according to the at least one unknown speaker thereby providing one, two or more speaker-dependent classes of speech samples;

extracting, using software stored and operating on the computer, speaker information for the speech samples of each of the speaker-dependent classes of speech samples;

combining, using software stored and operating on the computer, the extracted speaker information for each of the speaker-dependent classes of speech samples;

comparing, using software stored and operating on the computer, the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and

determining, using software stored and operating on the computer, whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.

2. The method of claim 1 , further comprising grouping of the telephone calls according to the telephone numbers of the telephone calls.

3. The method of claim 2 , wherein the speaker information for the at least one target speaker are obtained by obtaining a plurality of speech samples of the at least one target speaker.

4. The method of claim 3 , wherein at least one of the plurality of speech samples of the at least one target speaker is obtained from a telephone call of the at least one target speaker.

5. The method of claim 4 , wherein the speech samples according to the at least one unknown speaker are classified by a speaker clustering technique, in particular, by Agglomerative Hierarchical Clustering.

6. The method of claim 5 , wherein the speaker clustering technique is based on a Gaussian Mixture Model and a Gaussian Mixture Model metric.

7. The method of claim 6 , wherein the speaker clustering technique employs a Joint Factor Analysis.

8. The method of claim 1 , wherein combining the extracted speaker information for each of the speaker-dependent classes of speech samples comprises generating for a particular class a combined Gaussian Mixture Model from the extracted speaker information of the speech samples of that class.

9. The method of claim 6 , wherein the combined Gaussian Mixture Model is generated from Gaussian Mixture Models of the speech samples of that class.

10. The method of claim 1 , wherein combining the extracted speaker information for each of the speaker-dependent classes of speech samples comprises combining feature vectors obtained for one or more speech samples of a speaker-dependent class with feature vectors of one or more other speech samples of the same speaker-dependent class, in particular, by summation of at least some of the feature vectors, more particularly, comprising adding a feature vector of one speech sample of the speaker-dependent class and another feature vector of another speech sample of the speaker-dependent class, if they are close to each other within predetermined limits.

11. A system for performing speaker recognition, comprising:

a database stored and operating on a computer and configured to store speaker information for a target speaker;

software means stored and operating on the computer and configured to classify speech samples of telephone calls according to at least one unknown speaker thereby providing one, two or more speaker-dependent classes of speech samples;

software means stored and operating on the computer and confugured to extract speaker information for the speech samples of each of the speaker-dependent classes of speech samples;

software means stored and operating on the computer and configured to combine the extracted speaker information for each of the speaker-dependent classes of speech samples;

software means stored and operating on the computer and configured to compare the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and

software means stored and operating on the computer and configured to determine whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.

12. The system of claim 11 , further comprising software means stored and operating on the computer and configured to receive telephone calls from at least one unknown speaker.

13. The system of claim 11 , further comprising software means stored and operating on the computer and configured to group the telephone calls according to the telephone numbers of the telephone calls.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2012
From: BRUMMER, JOHAN NIKOLAAS LANGEHOVEN; RODRIGUEZ, LUIS BUERA; GOMAR, MARTA GARCIA
To: AGNITIO SL
Reel/Frame 028332/0056 →