IP Library Granted Patent US 10,535,000
Granted Patent B2
US 10,535,000 · App. 15/727,498 · Granted Jan 14, 2020

System and method for speaker change detection

Inventors: Zhenhao Ge (Indianapolis, IN); Ananth Nagaraja Iyer (Indianapolis, IN); Srinath Cheluvaraja (Indianapolis, IN); Aravind Ganapathiraju (Hyderabad, IN)
G06N3/084G10L17/005G10L17/04G10L17/18G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,535,000
App. No.
15/727,498
Granted
Jan 14, 2020
Kind
B2
Abstract

A method for training a neural network of a neural network based speaker classifier for use in speaker change detection. The method comprises: a) preprocessing input speech data; b) extracting a plurality of feature frames from the preprocessed input speech data; c) normalizing the extracted feature frames of each speaker within the preprocessed input speech data with each speaker's mean and variance; d) concatenating the normalized feature frames to form overlapped longer frames having a frame length and a hop size; e) inputting the overlapped longer frames to the neural network based speaker classifier; and f) training the neural network through forward-backward propagation.

Claims (73)

1. A method for training a neural network of a neural network based speaker classifier for use in speaker change detection, the method comprising:

a) preprocessing input speech data;

b) extracting a plurality of feature frames from the preprocessed input speech data;

c) normalizing the extracted feature frames of each speaker within the preprocessed input speech data with each speaker's mean and variance;

d) concatenating the normalized feature frames to form overlapped longer frames having a frame length and a hop size;

e) inputting the overlapped longer frames to the neural network based speaker classifier; and

f) training the neural network through forward-backward propagation, wherein step (a) comprises:

a.1) scaling a maximum of absolute amplitude of the input speech data to one;

a.2) performing voice activity detection on the scaled input speech data; and

a.3) eliminating unvoiced portions of the scaled input speech data, and

wherein step (a.2) comprises:

a.2.1) segmenting the scaled input speech data into overlapped frames with a window size and a hop size;

a.2.2) determine a short-term energy E of each frame;

a.2.3) determine a spectral centroid C of each frame;

a.2.4) remove any frame from the scaled input speech data in which the short-term energy E of the frame is below a predetermined threshold TE; and

a.2.5) remove any frame from the scaled input speech data in which the spectral centroid C of the frame is below a predetermined threshold TC.

2. The method of claim 1 , wherein the window size of the overlapped frames is 50 ms and the hop size of the overlapped frames is 25 ms.

3. The method of claim 1 , wherein the short-term energy E comprises:

E

=

1

N

n

=

1

N

s

(

n

)

2

,

where s(n) is the frame data with N samples.

4. The method of claim 1 , wherein the spectral centroid C comprises:

C

=

k

=

1

K

kS

(

k

)

k

=

1

K

S

(

k

)

,

wherein S(k) is a Discrete Fourier Transform (DFT) of frame data s(n) with k frequency components.

5. The method of claim 1 , wherein:

the predetermined threshold TE is a weighted average of the local maxima in a distribution histogram of the short-term energy E; and

the predetermined threshold TC is a weighted average of the local maxima in a distribution histogram of the spectral centroid C.

6. The method of claim 1 , wherein step (b) comprises generating 39-dimensional Mel-Frequency Cepstral Coefficients (MFCCs) from the preprocessed input speech data using overlapping Hamming windows and a hop size.

7. The method of claim 6 , wherein the Hamming window is 25 ms and the hop size of the Hamming windows is 10 ms.

8. The method of claim 1 , wherein the overlapped longer frame length is 10 frames and the hop size of the overlapped longer frames is 3 frames.

Assignments (4)
CHANGE OF NAME Recorded Jun 6, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067646/0448 →
SECURITY AGREEMENT Recorded Feb 12, 2020
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 051902/0850 →
MERGER Recorded Jul 1, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 046463/0839 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2017
From: GE, ZHENHAO; IYER, ANANTH NAGARAJA; CHELUVARAJA, SRINATH; GANAPATHIRAJU, ARAVIND
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 043809/0195 →
Continuity (2)
Provisional Application 62372057 · Aug 8, 2016
Related Publication 20180039888A1 · Feb 8, 2018
Cited By (1)
US 12,518,762