Voice activity detection method and system, and voice enhancement method and system
Microphone signals output by the microphone array satisfy a first model corresponding to a noise signal or a second model corresponding to a target voice signal mixed with a noise signal. A method and system may optimize the first model and the second model respectively by using maximization of a likelihood function and rank minimization of a noise covariance matrix as joint optimization objectives, and determine a first estimate of a noise covariance matrix of the first model and a second estimate of a noise covariance matrix of the second model; and determine, by using a statistical hypothesis testing method, whether the microphone signals satisfy the first model or the second model, so as to determine whether the target voice signal is present in the microphone signals, determine a noise covariance matrix of the microphone signals, and further perform voice enhancement on the microphone signals.
1 . A voice activity detection system, comprising:
at least one storage medium storing a set of instructions for voice activity detection; and
at least one processor in communication with the at least one storage medium, wherein during a process of voice activity detection for M microphones distributed in a preset array shape, wherein M is an integer greater than 1, the at least one processor executes the set of instructions to:
obtain microphone signals output by the M microphones, wherein the microphone signals satisfy a first model corresponding to absence of a target voice signal or a second model corresponding to presence of a target voice signal,
optimize the first model and the second model respectively by using maximization of a likelihood function and rank minimization of a noise covariance matrix as joint optimization objectives, and determine a first estimate of a noise covariance matrix of the first model and a second estimate of a noise covariance matrix of the second model, and
determine that the microphone signals satisfy a target model from the first model and the second model and a noise covariance matrix corresponding to the microphone signals, wherein the noise covariance matrix of the microphone signals is a noise covariance matrix of the target model; and
applying the noise covariance matrix in an operation to process the microphone signals.
2 . The voice activity detection system according to claim 1 , wherein the microphone signals include K frames of continuous audio signals, K is a positive integer greater than 1, and the microphone signals include an M×K data matrix.
3 . The voice activity detection system according to claim 2 , wherein the microphone signals are complete observation signals or incomplete observation signals, all data in the M×K data matrix in the complete observation signals is complete, and a part of data in the M×K data matrix in the incomplete observation signals is missing, and when the microphone signals are the incomplete observation signals, to obtain the microphone signals output by the M microphones, the at least one processor executes the set of instructions to:
obtain the incomplete observation signals, and
perform row-column permutation on the microphone signals based on a position of missing data in each column in the M×K data matrix, and divide the microphone signals into at least one sub microphone signal, wherein the microphone signals include the at least one sub microphone signal.
4 . The voice activity detection system according to claim 1 , wherein to optimize the first model and the second model respectively by using maximization of the likelihood function and rank minimization of the noise covariance matrix as the joint optimization objectives, the at least one processor executes the set of instructions to:
establish a first likelihood function corresponding to the first model by using the microphone signals as sample data, wherein the likelihood function includes the first likelihood function;
optimize the first model by using maximization of the first likelihood function and rank minimization of the noise covariance matrix of the first model as optimization objectives, and determine the first estimate;
establish a second likelihood function corresponding to the second model by using the microphone signals as sample data, wherein the likelihood function includes the second likelihood function; and
optimize the second model by using maximization of the second likelihood function and rank minimization of the noise covariance matrix of the second model as optimization objectives, and determine the second estimate and an estimate of an amplitude of the target voice signal.
5 . The voice activity detection system according to claim 4 , wherein the microphone signals include a noise signal, the noise signal conforms to a Gaussian distribution, and the noise signal includes at least:
a colored noise signal conforming to a zero-mean Gaussian distribution, wherein a noise covariance matrix corresponding to the colored noise signal is a low-rank semi-positive definite matrix.
6 . The voice activity detection system according to claim 4 , wherein to determine the target model and the noise covariance matrix corresponding to the microphone signals, the at least one processor executes the set of instructions to:
establish a binary hypothesis testing model based on the microphone signals, wherein an original hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the first model, and an alternative hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the second model;
substitute the first estimate, the second estimate, and the estimate of the amplitude into a decision criterion of a detector of the binary hypothesis testing model to obtain a test statistic; and
determine the target model of the microphone signals based on the test statistic.
7 . The voice activity detection system according to claim 6 , wherein to determine the target model of the microphone signals based on the test statistic, the at least one processor executes the set of instructions to:
determine that the test statistic is greater than a preset decision threshold, determine that the target voice signal is present in the microphone signals, and determine that the target model is the second model and that the noise covariance matrix of the microphone signals is the second estimate; or
determine that the test statistic is less than the preset decision threshold, determine that the target voice signal is absent in the microphone signals, and determine that the target model is the first model and that the noise covariance matrix of the microphone signals is the first estimate.
8 . The voice activity detection system according to claim 6 , wherein the detector includes at least one of a generalized likelihood ratio test (GLRT) detector, a Rao detector, or a Wald detector.
9 . A voice activity detection method, wherein the method is for M microphones distributed in a preset array shape, and M is an integer greater than 1, the voice activity detection method comprising:
obtaining microphone signals output by the M microphones, wherein the microphone signals satisfy a first model corresponding to absence of a target voice signal or a second model corresponding to presence of a target voice signal;
optimizing the first model and the second model respectively by using maximization of a likelihood function and rank minimization of a noise covariance matrix as joint optimization objectives, and determining a first estimate of a noise covariance matrix of the first model and a second estimate of a noise covariance matrix of the second model; and
determining that the microphone signals satisfy a target model from the first model and the second model and a noise covariance matrix corresponding to the microphone signals, wherein the noise covariance matrix of the microphone signals is a noise covariance matrix of the target model; and
applying the noise covariance matrix in an operation to process the microphone signals.
10 . The voice activity detection method according to claim 9 , wherein the determining of the target model and the noise covariance matrix corresponding to the microphone signals includes:
establishing a binary hypothesis testing model based on the microphone signals, wherein an original hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the first model, and an alternative hypothesis of the binary hypothesis testing model includes that the microphone signals satisfy the second model;
substituting the first estimate, the second estimate, and an amplitude estimate into a decision criterion of a detector of the binary hypothesis testing model to obtain a test statistic; and
determining the target model of the microphone signals based on the test statistic.