Spatial region based audio separation
A method includes receiving target audio data captured by a first audio input device, the target audio data comprising a target audio signal and a first version of an interfering audio signal, and receiving reference audio data captured by a second audio input device different from the first audio input device, the reference audio data comprising a second version of the interfering audio signal. The method also includes processing, using a trained neural network, the target audio data and the reference audio data to generate enhanced audio data, the neural network attenuating the interfering audio signal in the enhanced audio data.
1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:
receiving target audio data captured by a first audio input device, the target audio data comprising a target audio signal and a first version of an interfering audio signal; receiving reference audio data captured by a second audio input device different from the first audio input device, the reference audio data comprising a second version of the interfering audio signal;
processing, using a trained neural network, the target audio data and the reference audio data to generate enhanced audio data;
determining at least one of a delay contrast or a magnitude contrast between the first version of the interfering audio signal and the second version of the interfering audio signal; and
attenuating, using the trained neural network, at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio signal based on the at least one of the delay contrast or the magnitude contrast.
2 . The computer-implemented method of claim 1 , wherein:
the target audio signal originates from a first region, the first region defined by a first set of angles; and
the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of angles different from the first set of angles.
3 . The computer-implemented method of claim 2 , wherein the first audio input device and the second audio input device are symmetrically arranged relative to the first region.
4 . The computer-implemented method of claim 1 , wherein the delay contrast represents an angular separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.
5 . The computer-implemented method of claim 1 , wherein:
attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the delay contrast satisfies a threshold; and
the operations further comprise:
receiving an input representing a time shift; and
time shifting the reference audio data by the time shift to effectively adjust a value of the threshold.
6 . The computer-implemented method of claim 1 , wherein:
the target audio signal originates from a first region, the first region defined by a first set of distances; and
the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of distances different from the first set of distances.
7 . The computer-implemented method of claim 1 , wherein the magnitude contrast represents a distance separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.
8 . The computer-implemented method of claim 1 , wherein:
attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the magnitude contrast satisfies a threshold; and
the operations further comprise:
receiving an input representing a scalar; and
multiplying the reference audio data by the scalar to effectively change the threshold.
9 . The computer-implemented method of claim 1 , wherein a training process trains the neural network by:
obtaining training target audio data comprising sampled speech of interest and a sampled first version of an interfering audio signal;
obtaining training reference audio data comprising a sampled second version of the interfering audio signal;
processing, using the neural network, the training target audio data and the training reference audio data to generate predicted enhanced audio data; and
training the neural network based on a loss term computed based on the predicted enhanced audio data and the sampled speech of interest.
10 . The computer-implemented method of claim 1 , wherein the neural network comprises a U-net model architecture.
11 . The computer-implemented method of claim 10 , wherein the U-net model architecture comprises:
a Fourier transform layer;
a contracting path comprising a plurality of two-dimensional (2D) convolution layers trained to successively reduce spatial information while increasing feature information;
an expansion path comprising a plurality of 2D convolution layers trained to combine feature and spatial information through a sequence of up-convolutions and concatenations with high-resolution features from the contracting path; and
and an inverse Fourier transform layer.
12 . A system comprising:
data processing hardware; and
memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:
receiving target audio data captured by a first audio input device, the target audio data comprising a target audio signal and a first version of an interfering audio signal;
receiving reference audio data captured by a second audio input device different from the first audio input device, the reference audio data comprising a second version of the interfering audio signal;
processing, using a trained neural network, the target audio data and the reference audio data to generate enhanced audio data;
determining at least one of a delay contrast or a magnitude contrast between the first version of the interfering audio signal and the second version of the interfering audio signal; and
attenuating, using the trained neural network, at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio signal based on the at least one of the delay contrast or the magnitude contrast.
13 . The system of claim 12 , wherein:
the target audio signal originates from a first region, the first region defined by a first set of angles; and
the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of angles different from the first set of angles.
14 . The system of claim 13 , wherein the first audio input device and the second audio input device are symmetrically arranged relative to the first region.
15 . The system of claim 12 , wherein the delay contrast represents an angular separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.
16 . The system of claim 12 , wherein:
attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the delay contrast satisfies a threshold; and
the operations further comprise:
receiving an input representing a time shift; and time shifting the reference audio data by the time shift to effectively adjust a value of the threshold.
17 . The system of claim 12 , wherein:
the target audio signal originates from a first region, the first region defined by a first set of distances; and
the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of distances different from the first set of distances.
18 . The system of claim 12 , wherein the magnitude contrast represents a distance separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.
19 . The system of claim 12 , wherein:
attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the magnitude contrast satisfies a threshold; and
the operations further comprise:
receiving an input representing a scalar; and
multiplying the reference audio data by the scalar to effectively change the threshold.
20 . The system of claim 12 , wherein a training process trains the neural network by:
obtaining training target audio data comprising sampled speech of interest and a sampled first version of an interfering audio signal;
obtaining training reference audio data comprising a sampled second version of the interfering audio signal;
processing, using the neural network, the training target audio data and the training reference audio data to generate predicted enhanced audio data; and
training the neural network based on a loss term computed based on the predicted enhanced audio data and the sampled speech of interest.
21 . The system of claim 12 , wherein the neural network comprises a U-net model architecture.
22 . The system of claim 21 , wherein the U-net model architecture comprises:
a Fourier transform layer;
a contracting path comprising a plurality of two-dimensional (2D) convolution layers trained to successively reduce spatial information while increasing feature information;
an expansion path comprising a plurality of 2D convolution layers trained to combine feature and spatial information through a sequence of up-convolutions and concatenations with high-resolution features from the contracting path; and
and an inverse Fourier transform layer.