IP Library Granted Patent US 12,626,713
Granted Patent B2
US 12,626,713 · App. 17/819,177 · Granted May 12, 2026

Dynamic voice nullformer

Inventors: Yang Liu (Boston, MA); Abinaya Subramaniam (Westborough, MA); Trevor Caldwell (Menlo Park, CA); Douglas George Morton (Southborough, MA)
Assignee: Bose Corporation
G10L21/0232G10L25/84G10L2021/02166H04R1/406H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,713
App. No.
17/819,177
Granted
May 12, 2026
Kind
B2
Abstract

A voice capture system including a first and second voice beamformer, a voice mixer, a voice rejected noise beamformer, a noise beamformer adjustor, a jammer suppressor, and a speech enhancer is provided. The first and second voice beamformer and the voice mixer generate a voice enhanced reference signal based on a first and second frequency domain microphone signal. The voice rejected noise beamformer includes filter weights and generates a noise reference signal based on the first and second frequency domain microphone signal. The noise beamformer adjustor adjusts the one or more filter weights of the voice rejected noise beamformer to account for fit variation. The jammer suppressor generates a jammer suppressed signal based on the voice enhanced reference signal and the noise reference signal. The speech enhancer dynamically generates an output voice signal by applying a dynamic noise suppression signal to each frequency bin of the jammer suppressed signal.

Claims (68)

1 . A voice capture system, comprising:

a minimum variance distortionless response (MVDR) beamformer configured to generate a first voice beamformer signal based on a first frequency domain microphone signal and a second frequency domain microphone signal;

a delay and sum beamformer configured to generate a second voice beamformer signal based on the first frequency domain microphone signal and the second frequency domain microphone signal;

a voice mixer configured to blend the first voice beamformer signal with the second voice beamformer signal to generate a voice enhanced reference signal, wherein an amount of the first beamformer signal as compared to the second beamformer signal in the voice enhanced reference signal is dynamically adjusted based on amplitudes of the first and the second beamformer signals;

a voice rejected noise beamformer comprising one or more filter weights, the voice rejected noise beamformer configured to generate a noise reference signal based on the first frequency domain microphone signal and the second frequency domain microphone signal;

noise beamformer adjustor configured to adjust the one or more filter weights of the voice rejected noise beamformer to account for fit variation, wherein the filter weights are adjusted based on the first frequency domain microphone signal, the second frequency domain microphone signal, and the noise reference signal;

a jammer suppressor configured to generate a jammer suppressed signal based on the voice enhanced reference signal and the noise reference signal; and

a speech enhancer configured to generate an output voice signal based on the jammer suppressed signal, the noise reference signal, and a voice detection signal.

2 . The voice capture system of claim 1 , further comprising a filter bank configured to:

generate the first frequency domain microphone signal based on a first time domain microphone signal; and

generate the second frequency domain microphone signal based on a second time domain microphone signal.

3 . The voice capture system of claim 2 , further comprising:

a first microphone configured to generate the first time domain microphone signal; and

a second microphone configured to generate the second time domain microphone signal.

4 . The voice capture system of claim 1 , wherein the voice detection signal is generated by a voice activity detector based on the voice enhanced reference signal and the noise reference signal.

5 . The voice capture system of claim 1 , wherein the voice rejected noise beamformer is a Wiener delay and subtract noise beamformer.

6 . The voice capture system of claim 1 , wherein the one or more filter weights of the voice rejected noise beamformer correspond to a stock voice direction or a wearer-specific voice direction.

7 . The voice capture system of claim 1 , wherein the noise beamformer adjustor is configured to:

generate a signal-to-noise ratio (SNR) quality check signal based on the second frequency domain microphone signal;

generate, via a quality check voice activity detector, a voice detection quality check signal;

store, via a first data accumulator, first voice data corresponding to a relationship between the first frequency domain microphone signal and the second frequency domain microphone signal;

store, via a second data accumulator, second voice data corresponding to an energy level of the first frequency domain microphone signal; and

dynamically update, if the SNR quality check signal exceeds an SNR quality threshold, the voice detection quality check signal exceeds a voice detection quality threshold, and the first voice data or the second voice data exceeds a storage threshold, the one or more filter weights of the voice rejected noise beamformer based on the first frequency domain microphone signal and the second frequency domain microphone signal.

8 . The voice capture system of claim 1 , where the speech enhancer is configured to generate the output voice signal by:

determining a series of speech signal-to-noise ratios (SNR) corresponding to a series of frequency bins based on the jammer suppressed signal and the noise reference signal;

comparing the speech SNRs of each frequency bin to a set of speech enhancer thresholds;

applying a noise suppression signal to each frequency bin of the jammer suppressed signal, wherein an amplitude of the noise suppression signal applied to a frequency bin of the jammer suppressed signal is related to the SNR corresponding to the frequency bin.

9 . A wearable audio device comprising:

a first microphone configured to generate a first time domain microphone signal;

a second microphone configured to generate a second time domain microphone signal;

a filter bank configured to generate a first frequency domain microphone signal based on the first time domain microphone signal and a second frequency domain microphone signal based on the second time domain microphone signal;

a minimum variance distortionless response (MVDR) beamformer configured to generate a first voice beamformer signal based on the first frequency domain microphone signal and the second frequency domain microphone signal;

a delay and sum beamformer configured to generate a second voice beamformer signal based on the first frequency domain microphone signal and the second frequency domain microphone signal;

a voice rejected noise beamformer comprising one or more filter weights, the voice rejected noise beamformer configured to generate a noise reference signal based on the first frequency domain microphone signal and the second frequency domain microphone signal;

a noise beamformer adjustor configured to adjust the one or more filter weights of the voice rejected noise beamformer to account for fit variation, wherein the filter weights are adjusted based on the first frequency domain microphone signal, the second frequency domain microphone signal, and the noise reference signal;

a voice mixer configured to blend the first voice beamformer signal with the second voice beamformer signal to generate a voice enhanced reference signal, wherein an amount of the first beamformer signal as compared to the second beamformer signal in the voice enhanced reference signal is dynamically adjusted based on amplitudes of the first and the second beamformer signals;

a voice activity detector configured to generate a voice detection signal based on the voice enhanced reference signal and the noise reference signal;

a jammer suppressor configured to generate a jammer suppressed signal based on the voice enhanced reference signal and the noise reference signal; and

a speech enhancer configured to generate an output voice signal based on the jammer suppressed signal, the noise reference signal, and the voice detection signal.

10 . The wearable audio device of claim 9 , wherein the wearable audio device is a single side wearable device.

11 . The wearable audio device of claim 9 , wherein the noise beamformer adjustor is configured to:

generate a signal-to-noise ratio (SNR) quality check signal based on the second frequency domain microphone signal;

generate, via a quality check voice activity detector, a voice detection quality check signal;

store, via a first data accumulator, first voice data corresponding to a relationship between the first frequency domain microphone signal and the second frequency domain microphone signal;

store, via a second data accumulator, second voice data corresponding to an energy level of the first frequency domain microphone signal; and

dynamically update, if the SNR quality check signal exceeds an SNR quality threshold, the voice detection quality check signal exceeds a voice detection quality threshold, and the first voice data or the second voice data exceeds a storage threshold, the one or more filter weights of the voice rejected noise beamformer based on the first frequency domain microphone signal and the second frequency domain microphone signal.

12 . The wearable audio device of claim 9 , where the speech enhancer is configured to generate the output voice signal by:

determining a series of speech signal-to-noise ratios (SNR) corresponding to a series of frequency bins based on the jammer suppressed signal and the noise reference signal;

comparing the speech SNRs of each frequency bin to a set of speech enhancer thresholds;

applying a noise suppression signal to each frequency bin of the jammer suppressed signal, wherein an amplitude of the noise suppression signal applied to a frequency bin of the jammer suppressed signal is related to the SNR corresponding to the frequency bin.

13 . A method for voice capture, comprising:

generating, via a minimum variance distortionless response (MVDR) beamformer, a first voice beamformer signal based on a first frequency domain microphone signal and a second frequency domain microphone signal;

generating, via a delay and sum beamformer, a second voice beamformer signal based on the first frequency domain microphone signal and the second frequency domain microphone signal:

blending, via a voice mixer, the first voice beamformer signal with the second voice beamformer signal to generate a voice enhanced reference signal, wherein an amount of the first beamformer signal as compared to the second beamformer signal in the voice enhanced reference signal is dynamically adjusted based on amplitudes of the first and the second beamformer signals;

generating, via a voice rejected noise beamformer, a noise reference signal based on the first frequency domain microphone signal and the second frequency domain microphone signal;

adjusting, via a noise beamformer adjustor, one or more filter weights of a voice rejected noise beamformer to account for fit variation, wherein the filter weights are adjusted based on the first frequency domain microphone signal, the second frequency domain microphone signal, and the noise reference signal;

generating, via a jammer suppressor, a jammer suppressed signal based on the voice enhanced reference signal and the noise reference signal; and

generating, via a speech enhancer, an output voice signal based on the jammer suppressed signal, the noise reference signal, and a voice detection signal.

14 . The method of claim 13 , further comprising:

generating a signal-to-noise ratio (SNR) quality check signal based on the second frequency domain microphone signal;

generating, via a quality check voice activity detector, a voice detection quality check signal based on a frequency domain feedback microphone signal or the second frequency domain microphone signal;

storing, via a first data accumulator, first voice data corresponding to a relationship between the first frequency domain microphone signal and the second frequency domain microphone signal;

storing, via a second data accumulator, second voice data corresponding to an energy level of the first frequency domain microphone signal; and

dynamically updating, if the SNR quality check signal exceeds an SNR quality threshold, the voice detection quality check signal exceeds a voice detection quality threshold, and the first voice data or the second voice data exceeds a storage threshold, the one or more filter weights of the voice rejected noise beamformer based on the first frequency domain microphone signal and the second frequency domain microphone signal.

15 . The method of claim 13 , where the speech enhancer is configured to generate the output voice signal by:

determining a series of speech signal-to-noise ratios (SNR) corresponding to a series of frequency bins based on the jammer suppressed signal and the noise reference signal;

comparing the speech SNRs of each frequency bin to a set of speech enhancer thresholds;

applying a noise suppression signal to each frequency bin of the jammer suppressed signal, wherein an amplitude of the noise suppression signal applied to a frequency bin of the jammer suppressed signal is related to the SNR corresponding to the frequency bin.

Assignments (2)
SECURITY INTEREST Recorded Feb 28, 2025
From: BOSE CORPORATION
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 070438/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2022
From: LIU, YANG; SUBRAMANIAM, ABINAYA; CALDWELL, TREVOR; MORTON, DOUGLAS GEORGE
To: BOSE CORPORATION
Reel/Frame 061185/0260 →
Continuity (1)
Related Publication 20240055011A1 · Feb 15, 2024
References Cited (16)
US 10424315B1 · Ganeshkumar · 2019 [cited by examiner]
US 11140469B1 · Miller et al. · 2021 [cited by applicant]
US 11315586B2 · Huang · 2022 [cited by examiner]
US 11749262B2 · Gao et al. · 2023 [cited by applicant]
US 20080317259A1 · Zhang · 2008 [cited by examiner]
US 20140064514A1 · Mikami · 2014 [cited by examiner]
US 20140278394A1 · Bastyr · 2014 [cited by examiner]
US 20180359572A1 · Jensen · 2018 [cited by examiner]
US 20210337306A1 · Pedersen · 2021 [cited by examiner]
CN 117153180A · 2022 [cited by examiner]
DE 2146519A1 · 2010 [cited by examiner]
WO 2014143439A1 · 2014 [cited by applicant]
WO 2018175317A1 · 2018 [cited by applicant]
Siow Yong Low et al., “Robust microphone array using subband adaptive beamformer and spectral subtraction,” 2002 (Year: 2002). [cited by examiner]
Vamsynagh Pedamallu, “Microphone Array Wiener Beamforming with emphasis on Reverberation” (Year: 2012). [cited by examiner]
International Search Report and the Written Opinion of the International Searching Authority, PCT Application No. PCT/US2023/034718, dated Apr. 26, 2024, pp. 1-12. [cited by applicant]