IP Library Granted Patent US 11,348,591
Granted Patent B1
US 11,348,591 · App. 17/482,639 · Granted May 31, 2022

Dialect based speaker identification

Inventors: Muhammad Moinuddin (Jeddah, SA); Ubaid M. Al-Saggaf (Jeddah, SA); Shahid Munir Shah (Karachi, PK); Rizwan Ahmed Khan (Karachi, PK); Zahraa Ubaid Al-Saggaf (Jeddah, SA)
Assignee: King Abdulaziz University
G10L17/18G06N3/084G06N7/005G10L17/04G10L25/21G10L25/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,348,591
App. No.
17/482,639
Granted
May 31, 2022
Kind
B1
Abstract

A speaker identification system and method to identify a speaker based on the speaker's voice is disclosed. In an exemplary embodiment, the speaker identification system comprises a Gaussian Mixture Model (GMM) for speaker accent and dialect identification for a given speech signal input by the speaker and an Artificial Neural Network (ANN) to identify the speaker based on the identified dialect, in which the output of the GMM is input to the ANN.

Claims (48)

1. A speaker identification system to identify an unknown speaker based on a voice of the speaker, comprising:

a sound input device for inputting a speech signal of the voice of the unknown speaker;

Gaussian Mixture Model (GMM) circuitry configured to perform identification of speaker dialect for the speech signal by way of a mixture of a finite number of Gaussian distributions with unknown parameters;

Artificial Neural Network (ANN) circuitry having an input for receiving the identified speaker dialect and configured to identify the unknown speaker based on the identified dialect together with the speech signal; and

an output for indicating the identified speaker,

wherein an input to the Gaussian Mixture Model circuitry includes Mel-frequency cepstral coefficients of the speech signal, and pitch and energy of the speech signal and is also the input to the Artificial Neural Network circuitry, and

wherein the Artificial Neural Network circuitry is trained with a combination of a dialect code obtained as the output from the Gaussian Mixture Model circuitry, the Mel-frequency cepstral coefficients of the speech signal, and the pitch and the energy of the speech signal.

2. The speaker identification system of claim 1 , wherein the Gaussian Mixture Model circuitry is configured to perform dialect identification with spectral and prosodic features of the speech signal as the input speech signal and the output of the GMM is the identified dialect.

3. The speaker identification system of claim 1 , further comprising a speech signal analyzer configured to extract speech spectral features of Mel-frequency cepstral coefficient from the speech signal.

4. The speaker identification system of claim 1 , further comprising a speech signal analyzer configured to extract speech prosodic features of pitch and energy of speech signals from the speech signal.

5. The speaker identification system of claim 1 , wherein the Gaussian Mixture Model circuitry is configured to learn parameters by supervised learning using an expectation maximization algorithm.

6. The speaker identification system of claim 3 , wherein the speech signal analyzer is configured to extract Mel-frequency cepstral coefficients by employing a linear cosine transform of a log power spectrum on a nonlinear Mel scale of frequency.

7. The speaker identification system of claim 2 , wherein the GMM circuitry is configured to perform dialect identification by grouping the features of the speech signal using (1) Mel-frequency cepstral coefficient (MFCC) features only (2) pitch and Energy features and (3) a combination of MFCC, Pitch, and Energy Features.

8. The speaker identification system of claim 1 , wherein

the GMM circuitry is configured to perform a training phase in which training feature vectors of a plurality of speech dialects are input to obtain a trained GMM model,

the trained GMM model generates binary codes for the plurality of speech dialects, and

the ANN circuitry is configured to receive training speech feature vectors along with generated binary codes, an error between desired speaker and generated speaker is then utilized for training an ANN model using a back-propagation algorithm.

9. A speaker identification method to identify an unknown speaker based on a voice of the speaker, the method comprising:

inputting, by a sound input device, a speech signal of the voice of the unknown speaker;

identifying speaker dialect for the speech signal, by Gaussian Mixture Model (GMM) circuitry;

identifying, by Artificial Neural Network (ANN) circuitry, the unknown speaker based on the identified dialect, in which an output of the GMM is input to the ANN;

outputting an indication for the identified speaker; and

inputting to the Gaussian Mixture Model as well as to the Artificial Neural Network Mel-frequency cepstral coefficients of the speech signal, and pitch and energy of the speech signal, and training the Artificial Neural Network circuitry with a combination of a dialect code obtained as the output from the Gaussian Mixture Model circuitry, the Mel-frequency cepstral coefficients of the speech signal, and the pitch and energy of the speech signal.

10. The speaker identification method of claim 9 , wherein the identifying the speaker dialect, by the Gaussian Mixture Model circuitry, includes

using spectral and prosodic features of the speech signal as the input speech signal and the output of the GMM is the identified dialect.

11. The speaker identification method of claim 9 , further comprising:

extracting, by a speech signal analyzer, speech spectral features of Mel-frequency cepstral coefficient from the speech signal.

12. The speaker identification method of claim 9 , further comprising:

extracting, by a speech signal analyzer, speech prosodic features of pitch and energy of speech signals from the speech signal.

13. The speaker identification method of claim 9 , further comprising:

learning parameters of the Gaussian Mixture Model circuitry by supervised learning using an expectation maximization algorithm.

14. The speaker identification method of claim 11 , further comprising:

extracting, by the speech signal analyzer, Mel-frequency cepstral coefficients by employing a linear cosine transform of a log power spectrum on a nonlinear Mel scale of frequency.

15. The speaker identification method of claim 10 , wherein the dialect identification is performed by grouping the features of the speech signal using one of (1) Mel-frequency cepstral coefficient (MFCC) features only (2) pitch and Energy features and (3) a combination of MFCC, Pitch, and Energy Features.

16. The speaker identification method of claim 9 , further comprising:

performing, by the GMM circuitry, a training phase in which training feature vectors of a plurality of speech dialects are input to obtain a trained GMM model;

generating, by the trained GMM model, a binary code for a speech dialect;

receiving, by the ANN circuitry, training speech feature vectors along with the generated binary code; and

training an ANN model using a back-propagation algorithm based on an error between a known speaker and generated speaker.

17. A non-transitory computer-readable storage medium storing a program, which when executed by a computer performs a speaker identification method to identify an unknown speaker based on a voice of the speaker, comprising:

performing, by Gaussian Mixture Model (GMM) circuitry, a supervised training phase in which training feature vectors of a plurality of speech dialects are input to obtain a trained GMM model;

generating, by the trained GMM model, binary codes for the plurality of speech dialects;

receiving, by Artificial Neural Network (ANN) circuitry, training speech feature vectors along with the generated binary codes; and

training an ANN model using a back-propagation algorithm, based on the training speech feature vectors along with the generated binary codes and based on an error between a known speaker and generated speaker;

inputting, by a sound input device, a speech signal of the voice of the unknown speaker;

identifying speaker dialect for the speech signal, by the Gaussian Mixture Model (GMM) circuitry;

identifying, by the Artificial Neural Network (ANN) circuitry, the unknown speaker based on the identified dialect, in which an output of the GMM is input to the ANN; and

outputting an indication for the identified speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2021
From: MOINUDDIN, MUHAMMAD; AL-SAGGAF, UBAID M.; SHAH, SHAHID MUNIR; KHAN, RIZWAN AHMED; AL-SAGGAF, ZAHRAA UBAID
To: KING ABDULAZIZ UNIVERSITY
Reel/Frame 057572/0487 →
Cited By (3)
US 12,373,836 US 12,488,201 US 12,706,086