IP Library Granted Patent US 11,580,967
Granted Patent B2
US 11,580,967 · App. 17/253,434 · Granted Feb 14, 2023

Speech feature extraction apparatus, speech feature extraction method, and computer-readable storage medium

Inventors: Qiongqiong Wang (Tokyo, JP); Koji Okabe (Tokyo, JP); Kong Aik Lee (Tokyo, JP); Takafumi Koshinaka (Tokyo, JP)
Assignee: NEC CORPORATION
G10L15/20G10L15/02G10L25/30G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,967
App. No.
17/253,434
Granted
Feb 14, 2023
Kind
B2
Abstract

A speech feature extraction apparatus 100 includes a voice activity detection unit 103 that drops non-voice frames from frames corresponding to an input speech utterance, and calculates a posterior of being voiced for each frame, a voice activity detection process unit 106 calculates a function value as weights in pooling frames to produce an utterance-level feature, from a given a voice activity detection posterior, and an utterance-level feature extraction unit 112 that extracts an utterance-level feature, from the frame on a basis of multiple frame-level features, using the function values.

Claims (28)

1. A speech feature extraction apparatus comprising:

a processor; and

a memory device storing instructions executable by the processor to:

drop non-voice frames from frames corresponding to an input speech utterance, and calculate posterior of voice for the frames;

calculate function values from the posteriors; and

extract an utterance-level feature from the frames based on multiple frame-level features by using the function values as weights for pooling the frames in a pooling layer of a neural network.

2. The speech feature extraction apparatus according to claim 1 , wherein the instructions are executable by the processor to further train extraction of the utterance-level feature to generate utterance-level feature extraction parameters using the multiple frame-level features and the weights.

3. The speech feature extraction apparatus according to claim 2 ,

wherein from function values of the second posteriors used as weights for dropping the non-voice frames are further used to train the extraction.

4. The speech feature extraction apparatus according to claim 2 ,

wherein voice activity detection used for obtaining weights for dropping the non-voice frames is further used to train the extraction.

5. The speech feature extraction apparatus according to claim 1 , wherein the instructions are executable by the processor to further:

calculate second posteriors of voice for the frames; and

wherein function values of the second posteriors are used as weights for dropping the non-voice frames.

6. The speech feature extraction apparatus according to claim 5 ,

wherein voice activity detection is further used for obtaining the weights for dropping the non-voice frames.

7. The speech feature extraction apparatus according to claim 1 ,

wherein the a monotonically increasing and non-linear function defined as one of normalized Odds, and normalized log Odds, is used to calculate the function values from the posteriors, an i-vector is extracted as the utterance-level feature.

8. The speech feature extraction apparatus according to claim 1 ,

wherein the a monotonically increasing function is used to calculate the function values from the posteriors, and the utterance-level feature is extracted using the neural network.

9. A speech feature extraction method comprising:

dropping non-voice frames from frames corresponding to an input speech utterance, and calculating posteriors of voice for the frames;

calculating function values as weights from the posteriors;

extracting an utterance-level feature from the frames based on multiple frame-level features by using the function values as weights for pooling the frames in a pooling layer of a neural network.

10. A non-transitory computer-readable storage medium storing a program that includes commands for causing a computer to execute:

dropping non-voice frames from frames corresponding to an input speech utterance, and calculating posteriors of voice for the frames;

calculating function values as weights from the posteriors;

extracting an utterance-level feature from the frames based on multiple frame-level features by using the function values as weights for pooling the frames in a pooling layer of a neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2022
From: WANG, QIONGQIONG; OKABE, KOJI; LEE, KONG AIK; KOSHINAKA, TAKAFUMI
To: NEC CORPORATION
Reel/Frame 061510/0469 →
Continuity (1)
Related Publication 20210256970A1 · Aug 19, 2021