IP Library › Granted Patent US 11,423,906
Granted Patent B2
US 11,423,906 · App. 16/926,138 · Granted Aug 23, 2022

Multi-tap minimum variance distortionless response beamformer with neural networks for target speech separation

Inventors: Yong Xu (Bellevue, WA); Meng Yu (Bellevue, WA); Shi-Xiong Zhang (Redmond, WA); Chao Weng (Fremont, CA); Jianming Liu (Bellevue, WA); Dong Yu (Bothell, WA)
Assignee: TENCENT AMERICA LLC
G10L15/25
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,906
App. No.
16/926,138
Granted
Aug 23, 2022
Kind
B2
Abstract

A method, computer system, and computer readable medium are provided for automatic speech recognition. Video data and audio data corresponding to one or more speakers is received. A minimum variance distortionless response function is applied to the received audio and video data. A predicted target waveform corresponding to a target speaker from among the one or more speakers is generated based on back-propagating the output of the applied minimum variance distortionless response function.

Claims (40)

1. A method of automatic speech recognition, executable by a processor, comprising:

receiving video data and audio data corresponding to one or more speakers;

applying a minimum variance distortionless response function to the received audio and video data; and

generating a predicted target waveform corresponding to a target speaker from among the one or more speakers based on back-propagating an output of the applied minimum variance distortionless response function,

wherein the minimum variance distortionless response function generates a covariance matrix by replacing a real-value mask with a complex-value mask, and

wherein the predicted target waveform is generated by estimating the complex-value mask with a linear activation function and multiplying the complex-value mask with a complex spectrum of the received audio and video data.

2. The method of claim 1 , wherein generating the predicted target waveform comprises:

computing a scale-invariant source-to-noise ratio loss value based on the generated predicted target waveform;

back-propagating the computed scale-invariant source-to-noise ratio loss value; and

generating the predicted target waveform based on the back-propagated scale-invariant source-to-noise ratio loss value.

3. The method of claim 1 , wherein the scale-invariant source-to-noise ratio loss is optimized on the predicted target waveform rather than on the complex-value mask.

4. The method of claim 1 , wherein the video data corresponds to lip movement data captured by one or more cameras and the audio data corresponds to speech captured by one or more microphones.

5. The method of claim 4 , wherein the predicted target waveform is generated based on an inter-microphone correlation factor between the one or more microphones and an inter-frame correlation factor value between one or more frames of the captured lip movement data.

6. A computer system for automatic speech recognition, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including:

receiving code configured to cause the one or more computer processors to receive video data and audio data corresponding to one or more speakers;

applying code configured to cause the one or more computer processors to apply a minimum variance distortionless response function to the received audio and video data; and

first generating code configured to cause the one or more computer processors to generate a predicted target waveform corresponding to a target speaker from among the one or more speakers based on back-propagating an output of the applied minimum variance distortionless response function,

wherein the minimum variance distortionless response function generates a covariance matrix by replacing a real-value mask with a complex-value mask, and

wherein the predicted target waveform is generated by estimating the complex-value mask with a linear activation function and multiplying the complex-value mask with a complex spectrum of the received audio and video data.

7. The computer system of claim 6 , wherein generating the predicted target waveform comprises:

computing code configured to cause the one or more computer processors to compute a scale-invariant source-to-noise ratio loss value based on the generated predicted target waveform;

back-propagating code configured to cause the one or more computer processors to back-propagate the computed scale-invariant source-to-noise ratio loss value; and

second generating code configured to cause the one or more computer processors to generate the predicted target waveform based on the back-propagated scale-invariant source-to-noise ratio loss value.

8. The computer system of claim 6 , wherein the scale-invariant source-to-noise ratio loss is optimized on the predicted target waveform rather than on the complex-value mask.

9. The computer system of claim 6 , wherein the video data corresponds to lip movement data captured by one or more cameras and the audio data corresponds to speech captured by one or more microphones.

10. The computer system of claim 9 , wherein the predicted target waveform is generated based on an inter-microphone correlation factor between the one or more microphones and an inter-frame correlation factor value between one or more frames of the captured lip movement data.

11. A non-transitory computer readable medium having stored thereon a computer program for automatic speech recognition, the computer program configured to cause one or more computer processors to:

receive video data and audio data corresponding to one or more speakers;

apply a minimum variance distortionless response function to the received audio and video data; and

generate a predicted target waveform corresponding to a target speaker from among the one or more speakers based on back-propagating an output of the applied minimum variance distortionless response function,

wherein the minimum variance distortionless response function generates a covariance matrix by replacing a real-value mask with a complex-value mask, and

wherein the predicted target waveform is generated by estimating the complex-value mask with a linear activation function and multiplying the complex-value mask with a complex spectrum of the received audio and video data.

12. The non-transitory computer readable medium of claim 11 , wherein the computer program is further configured to cause the one or more computer processors to:

compute a scale-invariant source-to-noise ratio loss value based on the generated predicted target waveform;

back-propagate the computed scale-invariant source-to-noise ratio loss value; and

generate the predicted target waveform based on the back-propagated scale-invariant source-to-noise ratio loss value.

13. The non-transitory computer readable medium of claim 11 , wherein the video data corresponds to lip movement data captured by one or more cameras and the audio data corresponds to speech captured by one or more microphones.

14. The non-transitory computer readable medium of claim 13 , wherein the predicted target waveform is generated based on an inter-microphone correlation factor between the one or more microphones and an inter-frame correlation factor value between one or more frames of the captured lip movement data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2020
From: XU, YONG; YU, MENG; ZHANG, SHI-XIONG; WENG, CHAO; LIU, JIANMING; YU, DONG
To: TENCENT AMERICA LLC
Reel/Frame 053177/0398 →
Continuity (1)
Related Publication 20220013123A1 · Jan 13, 2022
Cited By (1)
US 12,217,761