IP Library › Granted Patent US 12,347,449
Granted Patent B2
US 12,347,449 · App. 18/160,296 · Granted Jul 1, 2025

Spatio-temporal beamformer

Inventors: Saeed Mosayyebpour Kaskari (Irvine, CA); Alireza Masnadi-Shirazi (Irvine, CA)
Assignee: Synaptics Incorporated
G10L21/0232H04R3/005G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,449
App. No.
18/160,296
Granted
Jul 1, 2025
Kind
B2
Abstract

This disclosure provides methods, devices, and systems for signal processing. The present implementations relate more specifically to a spatio-temporal beamformer. In some aspects, a beamforming system may receive an audio signal via a plurality of microphones, the audio signal including a number (B) of frames for each of the plurality of microphones, each of the B frames for each of the plurality of microphones including a number (N) of time-domain samples. For a first microphone, the beamforming system may transform the B*N time-domain samples into B*N/2 first frequency-domain samples; transform the B*N/2 first frequency-domain samples into B*N/2 second frequency-domain samples; and determine a probability of speech associated with the B*N/2 second frequency-domain samples based on a neural network model. The beamformer system may determine a minimum variance distortionless response (MVDR) beamforming filter based at least in part on the probability of speech for the first microphone.

Claims (60)

1. A method of processing an audio signal, comprising:

receiving a first audio signal via a plurality of microphones, the first audio signal including a number (B) of frames for each of the plurality of microphones, each of the B frames for each of the plurality of microphones including a number (N) of time-domain samples;

for a first microphone included in the plurality of microphones:

transforming the B*N time-domain samples into B*N/2 first frequency-domain samples based on an N-point fast Fourier transform (FFT);

transforming the B*N/2 first frequency-domain samples into B*N/2 second frequency-domain samples based on a B-point FFT; and

determining a probability of speech associated with the B*N/2 second frequency-domain samples based on a neural network model;

determining a minimum variance distortionless response (MVDR) beamforming filter based at least in part on the probability of speech for the first microphone; and

processing the first audio signal based on the MVDR beamforming filter.

2. The method of claim 1 , further comprising:

generating a first speech signal based on the probability of speech for the first microphone and the B*N/2 second frequency-domain samples;

transforming the first speech signal into a second speech signal based on a B-point inverse FFT;

transforming the second speech signal into a third speech signal, wherein the third speech signal includes a first number of frequency-domain samples associated with a first frequency bin and a second number of frequency-domain samples associated with a second frequency bin, wherein the first and second numbers are different; and,

determining a probability of speech associated with the third speech signal.

3. The method of claim 2 , wherein the determining of the MVDR beamforming filter comprises determining the MVDR beamforming filter based on the probability of speech associated with the third speech signal.

4. The method of claim 3 , further comprising generating a second audio signal based on the B*N/2 first frequency-domain samples, wherein the second audio signal includes the first number of frequency-domain samples associated with the first frequency bin and the second number of frequency-domain samples associated with the second frequency bin.

5. The method of claim 4 , further comprising generating a reconstructed probability of speech based on the probability of speech associated with the B*N/2 second frequency-domain samples.

6. The method of claim 5 , wherein the reconstructed probability of speech comprises:

for the first frequency bin in the probability of speech associated with the B*N/2 second frequency-domain samples:

a first plurality of probability values included in the probability of speech associated with the B*N/2 second frequency-domain samples and corresponding to a first plurality of second frequency-domain samples associated with the first frequency bin;

a second plurality of probability values included in the probability of speech associated with the B*N/2 second frequency-domain samples and corresponding to a second plurality of second frequency-domain samples associated with a third frequency bin preceding the first frequency bin; and

a third plurality of probability values included in the probability of speech associated with the B*N/2 second frequency-domain samples and corresponding to a third plurality of second frequency-domain samples associated with a fourth frequency bin succeeding the first frequency bin.

7. The method of claim 6 , wherein each of the second plurality of probability values is weighted by a respective first weight, and each of the third plurality of probability values is weighted by a respective second weight.

8. The method of claim 1 , wherein the transforming of the B*N time-domain samples into the B*N/2 first frequency-domain samples comprises:

buffering the B frames; and

applying the N-point FFT to the buffered frames.

9. The method of claim 1 , wherein the determining of the probability of speech associated with the B*N/2 second frequency-domain samples comprises decimating the B*N/2 second frequency-domain samples by a decimation factor (D), the probability of speech associated with the B*N/2 second frequency-domain samples being determined based on the B*N/2D decimated second frequency-domain samples.

10. The method of claim 9 , wherein D=2.

11. The method of claim 9 , wherein the decimating of the B*N/2 second frequency-domain samples comprises:

retaining B/2D second frequency-domain samples associated with a first frequency bin; and

discarding B/2D second frequency-domain samples associated with the first frequency bin.

12. The method of claim 1 , further comprising:

determining an average probability of speech for each frequency bin associated with the B*N/2 second frequency-domain samples; and

determining a probability of speech associated with the B*N/2 first frequency-domain samples based on the average probabilities of speech.

13. A beamforming system, comprising:

a processing system; and

a memory storing instructions that, when executed by the processing system, causes the speech enhancement system to:

receive a first audio signal via a plurality of microphones, the first audio signal including a number (B) of frames for each of the plurality of microphones, each of the B frames for each of the plurality of microphones including a number (N) of time-domain samples;

for a first microphone included in the plurality of microphones:

transform the B*N time-domain samples into B*N/2 first frequency-domain samples based on an N-point fast Fourier transform (FFT);

transform the B*N/2 first frequency-domain samples into B*N/2 second frequency-domain samples based on a B-point FFT; and

determine a probability of speech associated with the B*N/2 second frequency-domain samples based on a neural network model;

determine a minimum variance distortionless response (MVDR) beamforming filter based at least in part on the probability of speech for the first microphone; and

process the first audio signal based on the MVDR beamforming filter.

14. The beamforming system of claim 13 , wherein execution of the instructions further causes the beamforming system to:

generate a first speech signal based on the probability of speech for the first microphone and the B*N/2 second frequency-domain samples;

transform the first speech signal into a second speech signal based on a B-point inverse FFT;

transform the second speech signal into a third speech signal, wherein the third speech signal includes a first number of frequency-domain samples associated with a first frequency bin and a second number of frequency-domain samples associated with a second frequency bin, wherein the first and second numbers are different; and,

determine a probability of speech associated with the third speech signal.

15. The beamforming system of claim 14 , wherein execution of the instructions further causes the beamforming system to determine the MVDR beamforming filter based on the probability of speech associated with the third speech signal.

16. The beamforming system of claim 15 , wherein execution of the instructions further causes the beamforming system to generate a second audio signal based on the B*N/2 first frequency-domain samples, wherein the second audio signal includes the first number of frequency-domain samples associated with the first frequency bin and the second number of frequency-domain samples associated with the second frequency bin.

17. The beamforming system of claim 16 , wherein execution of the instructions further causes the beamforming system to generate a reconstructed probability of speech based on the probability of speech associated with the B*N/2 second frequency-domain samples.

18. The beamforming system of claim 17 , wherein the reconstructed probability of speech comprises:

for the first frequency bin in the probability of speech associated with the B*N/2 second frequency-domain samples:

a first plurality of probability values included in the probability of speech associated with the B*N/2 second frequency-domain samples and corresponding to a first plurality of second frequency-domain samples associated with the first frequency bin;

a second plurality of probability values included in the probability of speech associated with the B*N/2 second frequency-domain samples and corresponding to a second plurality of second frequency-domain samples associated with a third frequency bin preceding the first frequency bin; and

a third plurality of probability values included in the probability of speech associated with the B*N/2 second frequency-domain samples and corresponding to a third plurality of second frequency-domain samples associated with a fourth frequency bin succeeding the first frequency bin.

19. The beamforming system of claim 13 , wherein execution of the instructions further causes the beamforming system to:

buffer the B frames; and

apply the N-point FFT to the buffered frames.

20. The beamforming system of claim 13 , wherein execution of the instructions further causes the beamforming system to decimate the B*N/2 second frequency-domain samples by a decimation factor (D), the probability of speech associated with the B*N/2 second frequency-domain samples being determined based on the B*N/2D decimated second frequency-domain samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2023
From: MOSAYYEBPOUR KASKARI, SAEED; MASNADI-SHIRAZI, ALIREZA
To: SYNAPTICS INCORPORATED
Reel/Frame 062504/0182 →
Continuity (1)
Related Publication 20240257822A1 · Aug 1, 2024
References Cited (9)
US 20080240463A1 · Florencio · 2008 [cited by examiner]
US 20080247274A1 · Seltzer · 2008 [cited by examiner]
US 20220148611A1 · Slapak · 2022 [cited by examiner]
US 20220240026A1 · Zahedi · 2022 [cited by examiner]
Souden et al., “A study of the LCMV and MVDR Noise Reductionfilters,” IEEE Trans. Signal Process., vol. 58, No. 9, pp. 4925-4935, Sep. 2010. [cited by applicant]
Tashev et al., “Unified Framework for Single Channel Speech Enhancement,” IEEE Pacific Rim Conference on Communications, Computers and Signal Processing, 6 pages, Aug. 2009. [cited by applicant]
Hsu et al. “Robust Voice Activity Detection Algorithm Based on Feature of Frequency Modulation of Harmonics and its DSP Implementation”, IEICE Trans. Inf. & Syst., vol. E98-D, No. 10 Oct. 2015 (Year: 2015). [cited by applicant]
Park et al. “Voice Activity Detection in Noisy Environments Based on Double-Combined Fourier Transform and Line Fitting”, Sci. World Journal, 2014 (Year: 2014). [cited by applicant]
Raychowdhury et al. “A 2.3 nJ/Frame Voice Activity Detector-Based Audio Front-End for Context-Aware System-On-Chip Applications in 32-nm CMOS”, IEEE Journal of Solid-State Circuits vol. 48, No. 8, Aug. 2013 (Year: 2013). [cited by applicant]
Cited By (1)
US 12,482,446