Adaptive beam cancellation
A system that improves beam cancellation by refining reference beam selection and dynamically controlling an adaptation speed of adaptive filters. For example, a device may control the adaptation speed by dynamically determining a variable step-size parameter based on a relative strength of a microphone signal and/or a signal-to-noise ratio (SNR) value associated with an individual beam. The device may compare current energy levels of the microphone signal to a range of energy levels to determine a microphone step-size value, track a minimum noise floor to determine an SNR step-size value, and determine the variable step-size parameter using a sigmoid curve. Additionally or alternatively, the device can use a hybrid approach to select reference beams and can perform neighbor exclusion to further protect speech. For example, the device may exclude reference beams that are adjacent to a target beam to prevent the speech from being represented in the reference beams.
1 . A computer-implemented method, comprising:
receiving microphone audio data associated with a microphone array of a device;
generating directional audio data using the microphone audio data, the directional audio data comprising:
first audio data corresponding to a first direction relative to the device, and
second audio data corresponding to a second direction relative to the device, the second direction different from the first direction;
determining, using the second audio data and a first plurality of filter coefficient values associated with a first adaptive filter, first reference data;
determining first error data using the first reference data and the first audio data;
determining a first range of values associated with the microphone audio data and a first time window;
determining, using the first range of values and the microphone audio data, a first step-size value of the first adaptive filter, wherein the first step-size value represents a rate at which the first adaptive filter updates filter coefficients over time; and
determining, using the first error data, the first step-size value, the second audio data, and the first plurality of filter coefficient values, a second plurality of filter coefficient values associated with the first adaptive filter.
2 . The computer-implemented method of claim 1 , wherein determining the first step-size value further comprises:
determining, using the first range of values and a first value of the microphone audio data, a second step-size value;
determining a signal quality metric value associated with the first audio data;
determining, using the signal quality metric value, a third step-size value; and
determining the first step-size value using the second step-size value and the third step-size value.
3 . The computer-implemented method of claim 1 , wherein determining the first step-size value further comprises:
determining, using the first range of values and a first value of the microphone audio data, a second step-size value;
determining, using the second step-size value, a second value; and
determining the first step-size value using the second value and a sigmoid function.
4 . The computer-implemented method of claim 1 , wherein determining the first step-size value further comprises:
determining a first value representing a lowest value of the first range of values;
determining a second value representing a highest value of the first range of values;
determining a third value representing an energy level of the microphone audio data;
determining a first difference value by subtracting the first value from the third value;
determining a second difference value by subtracting the first value from the second value; and
determining the first step-size value by dividing the first difference value by the second difference value.
5 . The computer-implemented method of claim 1 , further comprising:
determining, using a portion of the first audio data, first power values;
determining, using the first power values, first noise floor data;
determining a lowest value represented in the first noise floor data; and
determining, using the lowest value, a second step-size value, wherein the first step-size value is determined using the second step-size value.
6 . The computer-implemented method of claim 1 , wherein determining the first reference data further comprises:
determining an association between the first direction and the second direction;
determining that a noise source is associated with third audio data corresponding to a third direction relative to the device, the third direction different from the second direction; and
determining, using the second audio data, the third audio data, and the first plurality of filter coefficient values, the first reference data.
7 . The computer-implemented method of claim 1 , wherein determining the first reference data further comprises:
determining that a first noise level associated with the second audio data satisfies a condition;
determining that a second noise level associated with third audio data satisfies the condition, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;
determining an angular value indicating a separation between the first direction and the third direction;
determining that the angular value is below a threshold value; and
determining, using the second audio data and the first plurality of filter coefficient values, the first reference data.
8 . The computer-implemented method of claim 1 , wherein determining the first reference data further comprises:
determining that third audio data is associated with a highest noise level represented in the directional audio data, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;
determining that a first angular value does not satisfy a condition, the first angular value representing a separation between the first direction and the third direction;
determining that the second audio data is associated with a second highest noise level represented in the directional audio data;
determining that a second angular value satisfies the condition, the second angular value representing a separation between the first direction and the second direction; and
determining, using the second audio data and the first plurality of filter coefficient values, the first reference data.
9 . The computer-implemented method of claim 1 , further comprising:
determining that third audio data is associated with a highest noise level represented in the directional audio data, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;
determining that a first angular value satisfies a condition, the first angular value representing a separation between the second direction and the third direction;
determining second reference data using the third audio data and a third plurality of filter coefficient values associated with a second adaptive filter;
determining second error data using the second reference data and the second audio data; and
determining, using the second error data, the first step-size value, the third audio data, and the third plurality of filter coefficient values, a fourth plurality of filter coefficient values associated with the second adaptive filter.
10 . The computer-implemented method of claim 1 , further comprising:
determining, using a portion of the second audio data and the second plurality of filter coefficient values, second reference data;
determining second error data using the second reference data and a portion of the first audio data; and
determining, using the second error data, the first step-size value, the second audio data, and the second plurality of filter coefficient values, a third plurality of filter coefficient values associated with the first adaptive filter.
11 . The computer-implemented method of claim 1 , further comprising:
determining a second range of values associated with the microphone audio data and a second time window;
determining, using the second range of values and the microphone audio data, a second step-size value associated with the first adaptive filter; and
determining, using second error data, the second step-size value, the second audio data, and the second plurality of filter coefficient values, a third plurality of filter coefficient values associated with the first adaptive filter.
12 . The computer-implemented method of claim 1 , wherein the directional audio data is generated using a beamformer component of the device, the first reference data includes a representation of an audible sound associated with the second direction, and the first error data is determined by subtracting the first reference data from the first audio data.
13 . A system comprising:
at least one processor; and
memory including instructions operable to be executed by the at least one processor to cause the system to:
receive microphone audio data associated with a microphone array of a device;
generate directional audio data using the microphone audio data, the directional audio data comprising:
first audio data corresponding to a first direction relative to the device, and
second audio data corresponding to a second direction relative to the device, the second direction different from the first direction;
determine, using the second audio data and a first plurality of filter coefficient values associated with a first adaptive filter, first reference data;
determine first error data using the first reference data and the first audio data;
determine a first range of values associated with the microphone audio data and a first time window;
determine, using the first range of values and the microphone audio data, a first step-size value of the first adaptive filter, wherein the first step-size value represents a rate at which the first adaptive filter updates filter coefficients over time; and
determine, using the first error data, the first step-size value, the second audio data, and the first plurality of filter coefficient values, a second plurality of filter coefficient values associated with the first adaptive filter.
14 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine, using the first range of values and a first value of the microphone audio data, a second step-size value;
determine a signal quality metric value associated with the first audio data;
determine, using the signal quality metric value, a third step-size value; and
determine the first step-size value using the second step-size value and the third step-size value.
15 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine, using the first range of values and a first value of the microphone audio data, a second step-size value;
determine, using the second step-size value, a second value; and
determine the first step-size value using the second value and a sigmoid function.
16 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine a first value representing a lowest value of the first range of values;
determine a second value representing a highest value of the first range of values;
determine a third value representing an energy level of the microphone audio data;
determine a first difference value by subtracting the first value from the third value;
determine a second difference value by subtracting the first value from the second value; and
determine the first step-size value by dividing the first difference value by the second difference value.
17 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine, using a portion of the first audio data, first power values;
determine, using the first power values, first noise floor data;
determine a lowest value represented in the first noise floor data; and
determine, using the lowest value, a second step-size value, wherein the first step-size value is determined using the second step-size value.
18 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine an association between the first direction and the second direction;
determine that a noise source is associated with third audio data corresponding to a third direction relative to the device, the third direction different from the second direction; and
determine, using the second audio data, the third audio data, and the first plurality of filter coefficient values, the first reference data.
19 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine that a first noise level associated with the second audio data satisfies a condition;
determine that a second noise level associated with third audio data satisfies the condition, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;
determine an angular value indicating a separation between the first direction and the third direction;
determine that the angular value is below a threshold value; and
determine, using the second audio data and the first plurality of filter coefficient values, the first reference data.
20 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine, using a portion of the second audio data and the second plurality of filter coefficient values, second reference data;
determine second error data using the second reference data and a portion of the first audio data; and
determine, using the second error data, the first step-size value, the second audio data, and the second plurality of filter coefficient values, a third plurality of filter coefficient values associated with the first adaptive filter.