IP Library › Granted Patent US 12,651,608
Granted Patent B2
US 12,651,608 · App. 18/370,381 · Granted Jun 9, 2026

Voice activity detection method and system, and voice enhancement method and system

Inventors: Le Xiao (Shenzhen, CN); Chengqian Zhang (Shenzhen, CN); Fengyun Liao (Shenzhen, CN); Xin Qi (Shenzhen, CN)
Assignee: Shenzhen Shokz Co., Ltd.
G10L25/78G10L21/0216G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,608
App. No.
18/370,381
Granted
Jun 9, 2026
Kind
B2
Abstract

Microphone signals output by the microphone array satisfy a first model corresponding to a noise signal or a second model corresponding to a target voice signal mixed with a noise signal. A method and system may optimize the first model and the second model respectively by using maximization of a likelihood function and rank minimization of a noise covariance matrix as joint optimization objectives, and determine a first estimate of a noise covariance matrix of the first model and a second estimate of a noise covariance matrix of the second model; and determine, by using a statistical hypothesis testing method, whether the microphone signals satisfy the first model or the second model, so as to determine whether the target voice signal is present in the microphone signals, determine a noise covariance matrix of the microphone signals, and further perform voice enhancement on the microphone signals.

Claims (35)

1 . A voice activity detection system, comprising:

at least one storage medium storing a set of instructions for voice activity detection; and

at least one processor in communication with the at least one storage medium, wherein during a process of voice activity detection for M microphones distributed in a preset array shape, wherein M is an integer greater than 1, the at least one processor executes the set of instructions to:

obtain microphone signals output by the M microphones, wherein the microphone signals satisfy a first model corresponding to absence of a target voice signal or a second model corresponding to presence of a target voice signal,

optimize the first model and the second model respectively by using maximization of a likelihood function and rank minimization of a noise covariance matrix as joint optimization objectives, and determine a first estimate of a noise covariance matrix of the first model and a second estimate of a noise covariance matrix of the second model, and

determine that the microphone signals satisfy a target model from the first model and the second model and a noise covariance matrix corresponding to the microphone signals, wherein the noise covariance matrix of the microphone signals is a noise covariance matrix of the target model; and

applying the noise covariance matrix in an operation to process the microphone signals.

2 . The voice activity detection system according to claim 1 , wherein the microphone signals include K frames of continuous audio signals, K is a positive integer greater than 1, and the microphone signals include an M×K data matrix.

3 . The voice activity detection system according to claim 2 , wherein the microphone signals are complete observation signals or incomplete observation signals, all data in the M×K data matrix in the complete observation signals is complete, and a part of data in the M×K data matrix in the incomplete observation signals is missing, and when the microphone signals are the incomplete observation signals, to obtain the microphone signals output by the M microphones, the at least one processor executes the set of instructions to:

obtain the incomplete observation signals, and

perform row-column permutation on the microphone signals based on a position of missing data in each column in the M×K data matrix, and divide the microphone signals into at least one sub microphone signal, wherein the microphone signals include the at least one sub microphone signal.

4 . The voice activity detection system according to claim 1 , wherein to optimize the first model and the second model respectively by using maximization of the likelihood function and rank minimization of the noise covariance matrix as the joint optimization objectives, the at least one processor executes the set of instructions to:

establish a first likelihood function corresponding to the first model by using the microphone signals as sample data, wherein the likelihood function includes the first likelihood function;

optimize the first model by using maximization of the first likelihood function and rank minimization of the noise covariance matrix of the first model as optimization objectives, and determine the first estimate;

establish a second likelihood function corresponding to the second model by using the microphone signals as sample data, wherein the likelihood function includes the second likelihood function; and

optimize the second model by using maximization of the second likelihood function and rank minimization of the noise covariance matrix of the second model as optimization objectives, and determine the second estimate and an estimate of an amplitude of the target voice signal.

5 . The voice activity detection system according to claim 4 , wherein the microphone signals include a noise signal, the noise signal conforms to a Gaussian distribution, and the noise signal includes at least:

a colored noise signal conforming to a zero-mean Gaussian distribution, wherein a noise covariance matrix corresponding to the colored noise signal is a low-rank semi-positive definite matrix.

6 . The voice activity detection system according to claim 4 , wherein to determine the target model and the noise covariance matrix corresponding to the microphone signals, the at least one processor executes the set of instructions to:

establish a binary hypothesis testing model based on the microphone signals, wherein an original hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the first model, and an alternative hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the second model;

substitute the first estimate, the second estimate, and the estimate of the amplitude into a decision criterion of a detector of the binary hypothesis testing model to obtain a test statistic; and

determine the target model of the microphone signals based on the test statistic.

7 . The voice activity detection system according to claim 6 , wherein to determine the target model of the microphone signals based on the test statistic, the at least one processor executes the set of instructions to:

determine that the test statistic is greater than a preset decision threshold, determine that the target voice signal is present in the microphone signals, and determine that the target model is the second model and that the noise covariance matrix of the microphone signals is the second estimate; or

determine that the test statistic is less than the preset decision threshold, determine that the target voice signal is absent in the microphone signals, and determine that the target model is the first model and that the noise covariance matrix of the microphone signals is the first estimate.

8 . The voice activity detection system according to claim 6 , wherein the detector includes at least one of a generalized likelihood ratio test (GLRT) detector, a Rao detector, or a Wald detector.

9 . A voice activity detection method, wherein the method is for M microphones distributed in a preset array shape, and M is an integer greater than 1, the voice activity detection method comprising:

obtaining microphone signals output by the M microphones, wherein the microphone signals satisfy a first model corresponding to absence of a target voice signal or a second model corresponding to presence of a target voice signal;

optimizing the first model and the second model respectively by using maximization of a likelihood function and rank minimization of a noise covariance matrix as joint optimization objectives, and determining a first estimate of a noise covariance matrix of the first model and a second estimate of a noise covariance matrix of the second model; and

determining that the microphone signals satisfy a target model from the first model and the second model and a noise covariance matrix corresponding to the microphone signals, wherein the noise covariance matrix of the microphone signals is a noise covariance matrix of the target model; and

applying the noise covariance matrix in an operation to process the microphone signals.

10 . The voice activity detection method according to claim 9 , wherein the determining of the target model and the noise covariance matrix corresponding to the microphone signals includes:

establishing a binary hypothesis testing model based on the microphone signals, wherein an original hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the first model, and an alternative hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the second model;

substituting the first estimate, the second estimate, and an amplitude estimate into a decision criterion of a detector of the binary hypothesis testing model to obtain a test statistic; and

determining the target model of the microphone signals based on the test statistic.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ATTACHMENTS PREVIOUSLY RECORDED ON REEL 65188 FRAME 539. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 20, 2025
From: XIAO, LE; ZHANG, CHENGQIAN; LIAO, FENGYUN; QI, XIN
To: SHENZHEN SHOKZ CO., LTD.
Reel/Frame 070567/0496 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: XIAO, LE; ZHANG, CHENGQIAN; LIAO, FENGYUN; QI, XIN
To: SHENZHEN SHOKZ CO., LTD.
Reel/Frame 065188/0539 →
Continuity (2)
Continuation PCTCN2021130035 · Nov 11, 2021
Related Publication 20240046956A1 · Feb 8, 2024
References Cited (58)
US 6256336B1 · Rademacher · 2001 [cited by examiner]
US 6615170B1 · Liu · 2003 [cited by examiner]
US 6717979B2 · Ribeiro Dias · 2004 [cited by examiner]
US 7478043B1 · Preuss · 2009 [cited by examiner]
US 8082286B1 · Picciolo · 2011 [cited by examiner]
US 9173025B2 · Dickins · 2015 [cited by examiner]
US 9549253B2 · Alexandridis · 2017 [cited by examiner]
US 9949040B2 · Bergmann · 2018 [cited by examiner]
US 11341981B2 · Moon · 2022 [cited by examiner]
US 12475915B2 · Xiao · 2025 [cited by examiner]
US 20030012262A1 · Ribeiro Dias · 2003 [cited by examiner]
US 20030200097A1 · Brand · 2003 [cited by examiner]
US 20040008803A1 · Aldrovandi · 2004 [cited by examiner]
US 20040240587A1 · Ozen · 2004 [cited by examiner]
US 20050227663A1 · He · 2005 [cited by examiner]
US 20070036202A1 · Ge · 2007 [cited by examiner]
US 20090196386A1 · Ikram · 2009 [cited by examiner]
US 20090292966A1 · Liva · 2009 [cited by examiner]
US 20130057433A1 · Sadler · 2013 [cited by examiner]
US 20130311137A1 · Schaum · 2013 [cited by examiner]
US 20140056435A1 · Kjems · 2014 [cited by examiner]
US 20140372091A1 · Larimore · 2014 [cited by examiner]
US 20150082377A1 · Chari · 2015 [cited by examiner]
US 20160180852A1 · Huang · 2016 [cited by examiner]
US 20180102135A1 · Ebenezer · 2018 [cited by examiner]
US 20180359572A1 · Jensen · 2018 [cited by examiner]
US 20190158522A1 · Amirmazlaghani · 2019 [cited by examiner]
US 20200219530A1 · Nesta · 2020 [cited by examiner]
US 20210375306A1 · Li · 2021 [cited by examiner]
US 20210390952A1 · Masnadi-Shirazi · 2021 [cited by examiner]
US 20230077396A1 · Ma · 2023 [cited by examiner]
US 20230077621A1 · Ono · 2023 [cited by examiner]
US 20230260529A1 · Xiao · 2023 [cited by examiner]
US 20240038257A1 · Xiao · 2024 [cited by examiner]
US 20240046956A1 · Xiao · 2024 [cited by examiner]
CN 102855880A · 2013 [cited by applicant]
CN 102938254A · 2013 [cited by applicant]
CN 109087664A · 2018 [cited by applicant]
CN 110148420A · 2019 [cited by applicant]
CN 110164452U · 2019 [cited by applicant]
JP 2006154819A · 2006 [cited by applicant]
JP 2015135437A · 2015 [cited by applicant]
JP 2018036332A · 2018 [cited by applicant]
JP 2019508719A · 2019 [cited by applicant]
KR 1020190073852A · 2019 [cited by applicant]
KR 1020190091061A · 2019 [cited by applicant]
KR 1020200078584A · 2020 [cited by applicant]
WO 2006114100A1 · 2006 [cited by applicant]
M. Viberg, P. Stoica, and B. Ottersten. Maximum Like- lihood Array Processing in Spatially Correlated Noise Fields Using Parameterized Signals. IEEE Trans. Sig-nal Processing, 45(4):996-1004,Apr. 1997. (Year: 1997). [cited by examiner]
Philippe Forster, Thierry Aste. Maximum likelihood multichannel estimation under reduced rank constraint. In Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP '98, … [cited by examiner]
International Search Report of PCT/CN2021/130035 (May 27, 2022). [cited by applicant]
Zhou Hong et al. “Application of Differential Evolution Optimization Based Gaussian Mixture Models to Speaker Recognition” The 26th Chinese Control and Decision Conference (2014 CCDC), Jul. 14, 2014 (Jul. 14, 2014). [cited by applicant]
Hoang Poul et al: “Joint Maximum Likelihood Estimation of Power Spectral Densities and Relative Acoustic Transfer Functions for Acoustic Beamforming”, ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech … [cited by applicant]
Mehdi Zohourian et al: “Binaural Speaker Localization Integrated Into an Adaptive Beamformer for Hearing Aids”, ARXIV:1806.04885V2,, vol. 26, No. 3, Mar. 1, 2018 (Mar. 1, 2018), pp. 515-528, XP058381893, DOI: 10.1109/TA… [cited by applicant]
Martin-Donas Juan Manuel et al: “Online Multichannel Speech Enhancement Based on Recursive EM andDNN-Based Speech Presence Estimation”, ARXIV:1806.04885V2,, vol. 28, Nov. 7, 2020 (Nov. 7, 2020), pp. 3080-3094, XP0118234… [cited by applicant]
Liao, F, “Research of Measurement of Radiated Noise of Submarines and Passive Ranging” Harbin Engineering University, Apr. 15, 2007. [cited by applicant]
Zhu, W, “Determining the No. of Sources and Angle Estimation Based on Uniform Circular Acoustic Vector Sensor Array” Harbin Engineering University, Mar. 31, 2018. [cited by applicant]
Xu, Y, “Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust ASR” 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Dec. 31, 2019. [cited by applicant]