IP Library › Granted Patent US 12,475,915
Granted Patent B2
US 12,475,915 · App. 18/370,387 · Granted Nov 18, 2025

Voice activity detection method and system, and voice enhancement method and system

Inventors: Le Xiao (Shenzhen, CN); Chengqian Zhang (Shenzhen, CN); Fengyun Liao (Shenzhen, CN); Xin Qi (Shenzhen, CN)
Assignee: Shenzhen Shokz Co., Ltd.
G10L25/78H04R1/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,915
App. No.
18/370,387
Granted
Nov 18, 2025
Kind
B2
Abstract

A voice activity detection method and system and a voice enhancement method and system are provided. A voice presence probability of a target voice signal present in microphone signals may be determined by calculating a linear correlation between a signal subspace where the microphone signals are located and a target subspace where the target voice signal is located. The voice enhancement method and system may be used to calculate filter coefficients based on the voice presence probability, so as to perform voice enhancement on the microphone signals. The calculation accuracy of the voice presence probability is improved, and the voice enhancement effect is also improved.

Claims (88)

1 . A voice activity detection system, comprising:

at least one non-transitory storage medium storing a set of instructions for voice activity detection; and

at least one processor in communication with the at least one non-transitory storage medium, wherein during an operation of voice activity detection for M microphones distributed in a preset array shape, wherein M is an integer greater than 1, the at least one processor executes the set of instructions to:

obtain microphone signals output by the M microphones,

determine, based on the microphone signals, a signal subspace formed by the microphone signals,

determine a target subspace formed by a target voice signal,

determine, based on a volume correlation between the signal subspace and the target subspace, a linear correlation coefficient between the signal subspace and the target subspace; and

use the linear correlation coefficient as a voice presence probability of the target voice signal being present in the microphone signals, and output the voice presence probability.

2 . The voice activity detection system according to claim 1 , wherein to determine, based on the microphone signals, the signal subspace formed by the microphone signal, the at least one processor executes the set of instructions to:

determine a sample covariance matrix of the microphone signals based on the microphone signals;

perform eigendecomposition on the sample covariance matrix to determine a plurality of eigenvectors of the sample covariance matrix; and

use a matrix formed by at least some eigenvectors of the plurality of eigenvectors as a basis matrix of the signal subspace.

3 . The voice activity detection system according to claim 1 , wherein to determine, based on the microphone signals, the signal subspace formed by the microphone signals, the at least one processor executes the set of instructions to:

determine a signal steering vector of the microphone signals by determining an azimuth of a signal source of the microphone signals based on the microphone signals by using a spatial estimation method, wherein the spatial estimation method includes at least one of a direction of arrival (DOA) estimation method, or a spatial spectrum estimation method; and

determine that the signal steering vector is a basis matrix of the signal subspace.

4 . The voice activity detection system according to claim 1 , wherein to determine the target subspace formed by the target voice signal, the at least one processor executes the set of instructions to:

determine that a preset target steering vector corresponding to the target voice signal is a basis matrix of the target subspace.

5 . The voice activity detection system according to claim 1 , wherein to determine, based on a volume correlation between the signal subspace and the target subspace, a linear correlation coefficient between the signal subspace and the target subspace, the at least one processor executes the set of instructions to:

determine a volume correlation function between the signal subspace and the target subspace; and

determine the linear correlation coefficient between the signal subspace and the target subspace based on the volume correlation function, wherein the linear correlation coefficient is negatively correlated with the volume correlation function.

6 . The voice activity detection system according to claim 5 , wherein to determine the linear correlation coefficient between the signal subspace and the target subspace based on the volume correlation function, the at least one processor executes the set of instructions to perform of the following steps:

determining that the volume correlation function is greater than a first threshold, and determining that the linear correlation coefficient is 0;

determining that the volume correlation function is less than a second threshold, and determining that the linear correlation coefficient is 1, wherein the second threshold is less than the first threshold; and

determining that the volume correlation function is between the first threshold and the second threshold, and determining that the linear correlation coefficient is between 0 and 1 and that the linear correlation coefficient is a negative correlation function of the volume correlation function.

7 . A voice activity detection method for M microphones distributed in a preset array shape, wherein M is an integer greater than 1, the voice activity detection method comprising:

obtaining microphone signals output by the M microphones;

determining, based on the microphone signals, a signal subspace formed by the microphone signals;

determining a target subspace formed by a target voice signal;

determining, based on a volume correlation between the signal subspace and the target subspace, a linear correlation coefficient between the signal subspace and the target subspace; and

using the linear correlation coefficient as a voice presence probability of the target voice signal being present in the microphone signals, and outputting the voice presence probability.

8 . The voice activity detection method according to claim 7 , wherein the determining, based on a volume correlation between the signal subspace and the target subspace, a linear correlation coefficient between the signal subspace and the target subspace include:

determining a volume correlation function between the signal subspace and the target subspace; and

determining the linear correlation coefficient between the signal subspace and the target subspace based on the volume correlation function, wherein the linear correlation coefficient is negatively correlated with the volume correlation function.

9 . The voice activity detection method according to claim 8 , wherein the determining of the linear correlation coefficient between the signal subspace and the target subspace based on the volume correlation function includes one of the following steps:

determining that the volume correlation function is greater than a first threshold, and determining that the linear correlation coefficient is 0;

determining that the volume correlation function is less than a second threshold, and determining that the linear correlation coefficient is 1, wherein the second threshold is less than the first threshold; and

determining that the volume correlation function is between the first threshold and the second threshold, and determining that the linear correlation coefficient is between 0 and 1 and that the linear correlation coefficient is a negative correlation function of the volume correlation function.

10 . A voice enhancement system, comprising:

at least non-transitory one storage medium storing a set of instructions for voice enhancement; and

at least one processor in communication with the at least one non-transitory storage medium, wherein during an operation of voice enhancement for M microphones distributed in a preset array shape, wherein M is an integer greater than 1, the at least one processor executes the set of instructions to:

obtain microphone signals output by the M microphones,

determine a voice presence probability of a target voice signal being present in the microphone signals,

determine, based on the voice presence probability, filter coefficient vectors corresponding to the microphone signals, and

combine the microphone signals based on the filter coefficient vectors to obtain a target audio signal and output the target audio signal, wherein

to determine the voice presence probability of the target voice signal being present in the microphone signals, the at least one processor executes the set of instructions to:

determine, based on the microphone signals, a signal subspace formed by the microphone signals,

determine a target subspace formed by a target voice signal, and

determine, based on a linear correlation between the signal subspace and the target subspace, the voice presence probability of the target voice signal being present in the microphone signals, and output the voice presence probability.

11 . The voice enhancement system according to claim 10 , wherein to determine, based on the microphone signals, the signal subspace formed by the microphone signal, the at least one processor executes the set of instructions to:

determine a sample covariance matrix of the microphone signals based on the microphone signals;

perform eigendecomposition on the sample covariance matrix to determine a plurality of eigenvectors of the sample covariance matrix; and

use a matrix formed by at least some eigenvectors of the plurality of eigenvectors as a basis matrix of the signal subspace.

12 . The voice enhancement system according to claim 10 , wherein to determine, based on the microphone signals, the signal subspace formed by the microphone signals, the at least one processor executes the set of instructions to:

determine a signal steering vector of the microphone signals by determining an azimuth of a signal source of the microphone signals based on the microphone signals by using a spatial estimation method, wherein the spatial estimation method includes at least one of a direction of arrival (DOA) estimation method, or a spatial spectrum estimation method; and

determine that the signal steering vector is a basis matrix of the signal subspace.

13 . The voice enhancement system according to claim 10 , wherein to determine the target subspace formed by the target voice signal, the at least one processor executes the set of instructions to:

determine that a preset target steering vector corresponding to the target voice signal is a basis matrix of the target subspace.

14 . The voice enhancement system according to claim 10 , wherein to determine, based on the linear correlation between the signal subspace and the target subspace, the voice presence probability of the target voice signal being present in the microphone signal and output the voice presence probability, the at least one processor executes the set of instructions to:

determine a volume correlation function between the signal subspace and the target subspace;

determine a linear correlation coefficient between the signal subspace and the target subspace based on the volume correlation function, wherein the linear correlation coefficient is negatively correlated with the volume correlation function; and

use the linear correlation coefficient as the voice presence probability, and output the voice presence probability.

15 . The voice enhancement system according to claim 14 , wherein to determine the linear correlation coefficient between the signal subspace and the target subspace based on the volume correlation function, the at least one processor executes the set of instructions to perform of the following steps:

determining that the volume correlation function is greater than a first threshold, and determining that the linear correlation coefficient is 0;

determining that the volume correlation function is less than a second threshold, and determining that the linear correlation coefficient is 1, wherein the second threshold is less than the first threshold; and

determining that the volume correlation function is between the first threshold and the second threshold, and determining that the linear correlation coefficient is between 0 and 1 and that the linear correlation coefficient is a negative correlation function of the volume correlation function.

16 . The voice enhancement system according to claim 10 , wherein to determine, based on the voice presence probability, the filter coefficient vectors corresponding to the microphone signals, the at least one processor executes the set of instructions to:

determine noise covariance matrices of the microphone signals based on the voice presence probability; and

determine the filter coefficient vectors based on a minimum variance distortionless response (MVDR) method and the noise covariance matrices.

17 . The voice enhancement system according to claim 10 , wherein to determine, based on the voice presence probability, the filter coefficient vectors corresponding to the microphone signals, the at least one processor executes the set of instructions to:

use the voice presence probability as a filter coefficient corresponding to a target microphone signal of the microphone signals, wherein the target microphone signal includes one of the microphone signals that has a highest signal-to-noise ratio; and

determine that filter coefficients corresponding to other microphone signals than the target microphone signal among the microphone signals are 0, wherein

the filter coefficient vectors include vectors formed by the filter coefficient corresponding to the target microphone signal and the filter coefficients corresponding to the other microphone signals.

18 . A voice enhancement method for M microphones distributed in a preset array shape, wherein M is an integer greater than 1, the voice enhancement method comprising:

obtaining microphone signals output by the M microphones;

determining a voice presence probability of a target voice signal being present in the microphone signals;

determining, based on the voice presence probability, filter coefficient vectors corresponding to the microphone signals; and

combining the microphone signals based on the filter coefficient vectors to obtain a target audio signal and outputting the target audio signal, wherein

the determining of the voice presence probability of the target voice signal being present in the microphone signals includes:

determining, based on the microphone signals, a signal subspace formed by the microphone signals,

determining a target subspace formed by a target voice signal, and

determining, based on a linear correlation between the signal subspace and the target subspace, the voice presence probability of the target voice signal being present in the microphone signals, and outputting the voice presence probability.

19 . The voice enhancement method according to claim 18 , wherein the determining, based on the voice presence probability, the filter coefficient vectors corresponding to the microphone signals includes:

determining noise covariance matrices of the microphone signals based on the voice presence probability; and

determining the filter coefficient vectors based on a minimum variance distortionless response (MVDR) method and the noise covariance matrices.

20 . The voice enhancement method according to claim 18 , wherein the determining, based on the voice presence probability, the filter coefficient vectors corresponding to the microphone signals includes:

using the voice presence probability as a filter coefficient corresponding to a target microphone signal of the microphone signals, wherein the target microphone signal includes one of the microphone signals that has a highest signal-to-noise ratio; and

determining that filter coefficients corresponding to other microphone signals than the target microphone signal among the microphone signals are 0, wherein

the filter coefficient vectors include vectors formed by the filter coefficient corresponding to the target microphone signal and the filter coefficients corresponding to the other microphone signals.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: XIAO, LE; ZHANG, CHENGQIAN; LIAO, FENGYUN; QI, XIN
To: SHENZHEN SHOKZ CO., LTD.
Reel/Frame 065188/0631 →
Continuity (2)
Continuation PCTCN2021139747 · Dec 20, 2021
Related Publication 20240038257A1 · Feb 1, 2024
References Cited (26)
US 20130297305A1 · Turnbull · 2013 [cited by applicant]
US 20190208318A1 · Chowdhary et al. · 2019 [cited by applicant]
US 20220237766A1 · Jiang · 2022 [cited by examiner]
CN 101778322A · 2010 [cited by applicant]
CN 105513605A · 2016 [cited by applicant]
CN 108028977A · 2018 [cited by applicant]
CN 108538306A · 2018 [cited by applicant]
CN 108831499A · 2018 [cited by applicant]
CN 110858488A · 2020 [cited by applicant]
CN 111308436A · 2020 [cited by applicant]
CN 116982112A · 2023 [cited by applicant]
EP 1473964A2 · 2004 [cited by examiner]
JP 2002091467A · 2002 [cited by applicant]
KR 20110120788A · 2011 [cited by applicant]
TW 200843541A · 2008 [cited by applicant]
WO 2017094862A1 · 2017 [cited by applicant]
International Search Report of PCT/CN2021/139747(Sep. 14, 2022). [cited by applicant]
Kim Dong Kook et al: “A subspace approach based on embedded prewhitening for voice activity detection”, The Journal of the Acoustical Society of America, American Institute of Physics, 2 Huntington Quadrangle, Melville,… [cited by applicant]
Hioka Yusuke et al: “Voice activity detection with array signal processing in the wavelet domain”, 2010 18th European Signal Processing Conference, IEEE, Sep. 3, 2002 (Sep. 3, 2002), pp. 1-4, XP032754214, ISSN: 2219-549… [cited by applicant]
Hioka Yusuke: “Voice activity detection with array signal processing in the wavelet domain” IEICE TRA NS. Fundamentals, (Nov. 2003) , vol. E86-A, No. 11 , p. 2802-2811. [cited by applicant]
Hioka Yusuke et al: “Voice activity detection with array signal processing in the wavelet domain” Technical Research Report of Electronic Information and Communication Society, Japan, Electronic Information and Communic… [cited by applicant]
Araki Shoko et al: “Microphone array speech processing techniques for conversation scene analysis” Technical Research Report of Electronic Information and Communication Society, Japan, Electronic Information and Communi… [cited by applicant]
Hassani et al: “LCMV beamforming with subspace projection for multi-speaker speech enhancement”, 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) May 19, 2016 , DOI: 10.1109/ICASSP.… [cited by applicant]
Qian Jin et al., “Study on Signal Subspace Approach for Speech Enhancement” China Master's Theses Full-Text Database, CMFD (Information Technology Section), Sep. 15, 2011. [cited by applicant]
Hailong Shi et al. “The Volume-Correlation Subspace Detector” https://arxiv.org/abs/1406.1 286v2, Dec. 16, 2015. [cited by applicant]
Le Xiao et al., “The Research Voice Wakeup Technologu Based on Transfer Learning” China Master's Theses Full-Text Database, CMFD (Information Technology Section), Jul. 15, 2020. [cited by applicant]
Cited By (1)
US 12,651,608