IP Library Patent Application 16007092
Patent Application
App. No. 16/007,092

SPEAKER RECOGNITION BASED ON DISCRIMINANT ANALYSIS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/007,092
Abstract

A method for speaker recognition, an electronic device and a speaker recognition system are disclosed. An example method includes receiving speech data corresponding to one or more utterances from a plurality of speakers that include a plurality of voice features. A plurality of variability factors is extracted from the speech data. The dimensionality of the plurality of variability factors is reduced using a non parametric analysis, thereby generating dimensionality reduced features. A score space is defined based at least on the dimensionality reduced features.

Claims (39)

1 . A method for speaker recognition, comprising:

receiving speech data corresponding to one or more utterances from a plurality of speakers that include a plurality of voice features;

extracting a plurality of variability factors from the speech data;

reducing dimensionality of the plurality of variability factors using a non-parametric analysis, thereby generating dimensionality reduced features; and

defining a score space based at least on the dimensionality reduced features.

2 . The method of claim 1 , wherein the variability factors include speaker-dependent factors and session-dependent factors.

3 . The method of claim 1 , further comprising:

receiving subsequent speech data from a target speaker;

scoring multiple variability factors of the target speaker using the score space; and

identifying the target speaker based at least on a score of the multiple variability factors.

4 . The method of claim 1 , wherein the non-parametric analysis is a Nearest Neighbor Discriminant Analysis (NNDA).

5 . The method of claim 4 , further comprising using a nearest neighbor rule which maintains within-class and between-class variations of the plurality of variability factors to reduce dimensionality.

6 . The method of claim 1 , comprising defining the score space using a probabilistic discriminant analysis of the dimensionality reduced features.

7 . The method of claim 1 , comprising extracting the plurality of variability factors using a total variability matrix trained by a Universal Background Model (UBM) trained by a Gaussian Mixture Model (GMM).

8 . The method of claim 7 , wherein the total variability matrix is further trained using Baum-Welch statistics of the plurality of voice features.

9 . The method of claim 1 , wherein the plurality of voice features are determined using Mel frequency cepstral coefficients (MFCC).

10 . An electronic device comprising:

an extractor configured to extract a plurality of variability factors from speech data; and

an analyzer configured to reduce dimensionality of the plurality of variability factors using a non-parametric analysis, thereby generating dimensionality reduced features, and define a score space using a probabilistic discriminant analysis on the dimensionality reduced features.

11 . The electronic device of claim 10 , further comprising a scorer configured to:

receive, from the extractor, multiple variability factors extracted from subsequently received speech data of a target speaker;

score at the multiple variability factors of the target speaker using the score space; and

identify the target speaker based at least on a score of the multiple variability factors.

12 . The electronic device of claim 10 , wherein the analyzer is configured to reduce dimensionality using a Nearest Neighbor Discriminant Analysis (NNDA).

13 . The electronic device of claim 10 , wherein the analyzer is configured to define the score space using a probabilistic discriminant analysis of the dimensionality reduced features.

14 . A computer-readable medium having computer-executable instructions stored thereon that, when executed by a computer, cause the computer to perform corresponding functions, the functions comprising:

receiving speech data corresponding to one or more utterances from a plurality of speakers that include a plurality of voice features;

extracting a plurality of variability factors from the speech data;

reducing dimensionality of the plurality of variability factors using a non-parametric analysis, thereby generating dimensionality reduced features; and

defining a score space based at least on the dimensionality reduced features.

15 . The computer-readable medium of claim 14 , wherein the instructions further comprise instructions that, when executed by the computer, cause the computer to perform corresponding functions, the functions comprising:

receiving subsequent speech data from a target speaker;

scoring multiple variability factors of the target speaker using the score space; and

identifying the target speaker based at least on a score of the multiple variability factors.

16 . The computer-readable medium of claim 14 , wherein the instructions further comprise instructions that, when executed by the computer, cause the computer to perform corresponding functions, the functions comprising reducing dimensionality by computing local sample averages of a number of samples in a neighborhood of each individual sample of the plurality of variability factors.

17 . The computer-readable medium of claim 14 , wherein the instructions further comprise instructions that, when executed by the computer, cause the computer to perform corresponding functions, the functions comprising defining the score space using a probabilistic discriminant analysis of the dimensionality reduced features.

18 . The computer-readable medium of claim 14 , wherein the instructions further comprise instructions that, when executed by the computer, cause the computer to perform corresponding functions, the functions comprising extracting the plurality of variability factors using a total variability matrix trained by a Universal Background Model (UBM) trained by a Gaussian Mixture Model (GMM).

19 . The computer-readable medium of claim 14 , wherein the variability factors include speaker-dependent factors and session-dependent factors.

20 . The computer-readable medium of claim 14 , wherein the non-parametric analysis is a Nearest Neighbor Discriminant Analysis (NNDA).

Assignments (2)
SECURITY AGREEMENT Recorded Jul 9, 2021
From: MAXLINEAR, INC.; MAXLINEAR COMMUNICATIONS, LLC; EXAR CORPORATION
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 056816/0089 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2020
From: INTEL CORPORATION
To: MAXLINEAR, INC.
Reel/Frame 053626/0636 →