IP Library › Granted Patent US 10,991,362
Granted Patent B2
US 10,991,362 · App. 16/849,321 · Granted Apr 27, 2021

Online target-speech extraction method based on auxiliary function for robust automatic speech recognition

Inventors: Hyung Min Park (Seoul, KR); Seoyoung Lee (Seoul, KR); Seung-Yun Kim (Seoul, KR); Byung Joon Cho (Seoul, KR); Uihyeop Shin (Changwon-si, KR)
Assignee: INDUSTRY-UNIVERSITY COOPERATION FOUNDATION SOGANG UNIVERSITY
G10L15/08H04R1/326H04R2430/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,991,362
App. No.
16/849,321
Granted
Apr 27, 2021
Kind
B2
Abstract

Provided is a target speech signal extraction method for robust speech recognition including: receiving information on a direction of arrival of the target speech source with respect to the microphones; generating a nullformer by using the information on the direction of arrival of the target speech source to remove the target speech signal from the input signals and to estimate noise; setting a real output of the target speech source using an adaptive vector as a first channel and setting a dummy output by the nullformer as a remaining channel; setting a cost function for minimizing dependency between the real output of the target speech source and the dummy output using the nullformer by performing independent component analysis (ICA) or independent vector analysis (IVA); setting an auxiliary function to the cost function; and estimating the target speech signal by using the cost function and the auxiliary function.

Claims (63)

1. A target speech signal extraction method of extracting a target speech signal from input signals input to at least two or more microphones for robust speech recognition, by a processor of a speech recognition apparatus, comprising:

(a) receiving information on a direction of arrival of the target speech source with respect to the microphones;

(b) generating a nullformer for removing the target speech signal from the input signals and estimating noise by using the information on the direction of arrival of the target speech source;

(c) setting a real output of the target speech source using an adaptive vector w(k) as a first channel and setting a dummy output by the nullformer as a remaining channel;

(d) setting a cost function for minimizing dependency between the real output of the target speech source and the dummy output using the nullformer by performing independent component analysis (ICA) or independent vector analysis (IVA);

(e) setting an auxiliary function to the cost function; and

(f) estimating the target speech signal by using the cost function and the auxiliary function, thereby extracting the target speech signal from the input signals,

wherein the auxiliary function is set an inequality relation so that the auxiliary function has always values greater than or same as that of the cost function.

2. The target speech signal extraction method according to claim 1 , wherein the direction of arrival of the target speech source is a separation angle θ target formed between a vertical line in the microphone and the target speech source.

3. The target speech signal extraction method according to claim 1 , wherein the nullformer is a “delay-subtract nullformer”, and

wherein the (b) includes of obtaining a relative ratio of target speech signals by using the information on the direction of arrival (DOA) of the target speech source, multiplying the relative ratio and an input signal of a microphone and subtracting the multiplied value from input signals of a pair of microphones to cancel out the target speech source component from the input signal of a microphone.

4. The target speech signal extraction method according to claim 1 ,

wherein a probability density function of the cost function is modeling by a generalized Gaussian distribution.

5. The target speech signal extraction method according to claim 4 ,

wherein the generalized gaussian distribution has a varying variance with regard to time-frequency or one of time and frequency, and

wherein the (e) includes of updating the varying variance λ and the adaptive vector w(k) alternately, and estimating the target speech signal by using the updated varying variance and the adaptive vector.

6. The target speech signal extraction method according to claim 4 ,

wherein the generalized gaussian distribution has a constant variance, and

wherein the (e) includes of learning the cost function to update the adaptive vector w(k), and estimating the target speech signal by using the updated adaptive vector.

7. The target speech signal extraction method according to claim 4 , the target speech signal extraction method further comprises (f) applying a minimal distortion principle (MDP) using a target speech element of a diagonal elements in an inverse matrix of a separating matrix, to the estimated the target speech signal in the (e).

8. The target speech signal extraction method according to claim 1 ,

wherein a time domain waveform y(k) of an estimated target speech signal is expressed by the following Mathematical Formula, and

y

⁡

(

t

)

=

∑

τ

⁢

∑

k

=

1

K

⁢

⁢

Y

⁡

(

τ

,

k

)

⁢

e

j

⁢

⁢

ω

k

⁡

(

t

-

τ

⁢

⁢

H

)

wherein Y(k, τ)=w(k)×(k, τ), w(k) denotes an adaptive vector for generating a real output with respect to the target speech source, and k and τ denote a frequency bin number and a frame number, respectively.

9. A non-transitory computer readable storage media having program instructions that, when executed by a processor of a speech recognition apparatus, cause the processor to perform the target speech signal extraction method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2020
From: PARK, HYUNG MIN; LEE, SEOYOUNG; KIM, SEUNG-YUN; CHO, BYUNG JOON; SHIN, UIHYEOP
To: INDUSTRY-UNIVERSITY COOPERATION FOUNDATION SOGANG UNIVERSITY
Reel/Frame 052405/0317 →
Priority Claims (1)
KR 10-2015-0037314 · Mar 18, 2015 · national
Continuity (3)
Continuation In Part 16181798 · Nov 6, 2018
Continuation In Part 15071594 · Mar 16, 2016
Related Publication 20200243072A1 · Jul 30, 2020