IP Library › Granted Patent US 12,739,564
Granted Patent B1
US 12,739,564 · App. 18/466,272 · Granted Sep 15, 2026

Adaptive beam cancellation

Inventors: Ian Ernan Liu (San Diego, CA); Robert Ayrapetian (Morgan Hil, CA); Pradeep Kumar Govindaraju (San Jose, CA); Madhuri Saraf (Sunnyvale, CA); Carlo Murgia (Santa Clara, CA)
Assignee: Amazon Technologies, Inc.
H04R5/04H04R1/406
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,564
App. No.
18/466,272
Granted
Sep 15, 2026
Kind
B1
Abstract

A system that improves beam cancellation by refining reference beam selection and dynamically controlling an adaptation speed of adaptive filters. For example, a device may control the adaptation speed by dynamically determining a variable step-size parameter based on a relative strength of a microphone signal and/or a signal-to-noise ratio (SNR) value associated with an individual beam. The device may compare current energy levels of the microphone signal to a range of energy levels to determine a microphone step-size value, track a minimum noise floor to determine an SNR step-size value, and determine the variable step-size parameter using a sigmoid curve. Additionally or alternatively, the device can use a hybrid approach to select reference beams and can perform neighbor exclusion to further protect speech. For example, the device may exclude reference beams that are adjacent to a target beam to prevent the speech from being represented in the reference beams.

Claims (109)

1 . A computer-implemented method, comprising:

receiving microphone audio data associated with a microphone array of a device;

generating directional audio data using the microphone audio data, the directional audio data comprising:

first audio data corresponding to a first direction relative to the device, and

second audio data corresponding to a second direction relative to the device, the second direction different from the first direction;

determining, using the second audio data and a first plurality of filter coefficient values associated with a first adaptive filter, first reference data;

determining first error data using the first reference data and the first audio data;

determining a first range of values associated with the microphone audio data and a first time window;

determining, using the first range of values and the microphone audio data, a first step-size value of the first adaptive filter, wherein the first step-size value represents a rate at which the first adaptive filter updates filter coefficients over time; and

determining, using the first error data, the first step-size value, the second audio data, and the first plurality of filter coefficient values, a second plurality of filter coefficient values associated with the first adaptive filter.

2 . The computer-implemented method of claim 1 , wherein determining the first step-size value further comprises:

determining, using the first range of values and a first value of the microphone audio data, a second step-size value;

determining a signal quality metric value associated with the first audio data;

determining, using the signal quality metric value, a third step-size value; and

determining the first step-size value using the second step-size value and the third step-size value.

3 . The computer-implemented method of claim 1 , wherein determining the first step-size value further comprises:

determining, using the first range of values and a first value of the microphone audio data, a second step-size value;

determining, using the second step-size value, a second value; and

determining the first step-size value using the second value and a sigmoid function.

4 . The computer-implemented method of claim 1 , wherein determining the first step-size value further comprises:

determining a first value representing a lowest value of the first range of values;

determining a second value representing a highest value of the first range of values;

determining a third value representing an energy level of the microphone audio data;

determining a first difference value by subtracting the first value from the third value;

determining a second difference value by subtracting the first value from the second value; and

determining the first step-size value by dividing the first difference value by the second difference value.

5 . The computer-implemented method of claim 1 , further comprising:

determining, using a portion of the first audio data, first power values;

determining, using the first power values, first noise floor data;

determining a lowest value represented in the first noise floor data; and

determining, using the lowest value, a second step-size value, wherein the first step-size value is determined using the second step-size value.

6 . The computer-implemented method of claim 1 , wherein determining the first reference data further comprises:

determining an association between the first direction and the second direction;

determining that a noise source is associated with third audio data corresponding to a third direction relative to the device, the third direction different from the second direction; and

determining, using the second audio data, the third audio data, and the first plurality of filter coefficient values, the first reference data.

7 . The computer-implemented method of claim 1 , wherein determining the first reference data further comprises:

determining that a first noise level associated with the second audio data satisfies a condition;

determining that a second noise level associated with third audio data satisfies the condition, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;

determining an angular value indicating a separation between the first direction and the third direction;

determining that the angular value is below a threshold value; and

determining, using the second audio data and the first plurality of filter coefficient values, the first reference data.

8 . The computer-implemented method of claim 1 , wherein determining the first reference data further comprises:

determining that third audio data is associated with a highest noise level represented in the directional audio data, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;

determining that a first angular value does not satisfy a condition, the first angular value representing a separation between the first direction and the third direction;

determining that the second audio data is associated with a second highest noise level represented in the directional audio data;

determining that a second angular value satisfies the condition, the second angular value representing a separation between the first direction and the second direction; and

determining, using the second audio data and the first plurality of filter coefficient values, the first reference data.

9 . The computer-implemented method of claim 1 , further comprising:

determining that third audio data is associated with a highest noise level represented in the directional audio data, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;

determining that a first angular value satisfies a condition, the first angular value representing a separation between the second direction and the third direction;

determining second reference data using the third audio data and a third plurality of filter coefficient values associated with a second adaptive filter;

determining second error data using the second reference data and the second audio data; and

determining, using the second error data, the first step-size value, the third audio data, and the third plurality of filter coefficient values, a fourth plurality of filter coefficient values associated with the second adaptive filter.

10 . The computer-implemented method of claim 1 , further comprising:

determining, using a portion of the second audio data and the second plurality of filter coefficient values, second reference data;

determining second error data using the second reference data and a portion of the first audio data; and

determining, using the second error data, the first step-size value, the second audio data, and the second plurality of filter coefficient values, a third plurality of filter coefficient values associated with the first adaptive filter.

11 . The computer-implemented method of claim 1 , further comprising:

determining a second range of values associated with the microphone audio data and a second time window;

determining, using the second range of values and the microphone audio data, a second step-size value associated with the first adaptive filter; and

determining, using second error data, the second step-size value, the second audio data, and the second plurality of filter coefficient values, a third plurality of filter coefficient values associated with the first adaptive filter.

12 . The computer-implemented method of claim 1 , wherein the directional audio data is generated using a beamformer component of the device, the first reference data includes a representation of an audible sound associated with the second direction, and the first error data is determined by subtracting the first reference data from the first audio data.

13 . A system comprising:

at least one processor; and

memory including instructions operable to be executed by the at least one processor to cause the system to:

receive microphone audio data associated with a microphone array of a device;

generate directional audio data using the microphone audio data, the directional audio data comprising:

first audio data corresponding to a first direction relative to the device, and

second audio data corresponding to a second direction relative to the device, the second direction different from the first direction;

determine, using the second audio data and a first plurality of filter coefficient values associated with a first adaptive filter, first reference data;

determine first error data using the first reference data and the first audio data;

determine a first range of values associated with the microphone audio data and a first time window;

determine, using the first range of values and the microphone audio data, a first step-size value of the first adaptive filter, wherein the first step-size value represents a rate at which the first adaptive filter updates filter coefficients over time; and

determine, using the first error data, the first step-size value, the second audio data, and the first plurality of filter coefficient values, a second plurality of filter coefficient values associated with the first adaptive filter.

14 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using the first range of values and a first value of the microphone audio data, a second step-size value;

determine a signal quality metric value associated with the first audio data;

determine, using the signal quality metric value, a third step-size value; and

determine the first step-size value using the second step-size value and the third step-size value.

15 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using the first range of values and a first value of the microphone audio data, a second step-size value;

determine, using the second step-size value, a second value; and

determine the first step-size value using the second value and a sigmoid function.

16 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first value representing a lowest value of the first range of values;

determine a second value representing a highest value of the first range of values;

determine a third value representing an energy level of the microphone audio data;

determine a first difference value by subtracting the first value from the third value;

determine a second difference value by subtracting the first value from the second value; and

determine the first step-size value by dividing the first difference value by the second difference value.

17 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using a portion of the first audio data, first power values;

determine, using the first power values, first noise floor data;

determine a lowest value represented in the first noise floor data; and

determine, using the lowest value, a second step-size value, wherein the first step-size value is determined using the second step-size value.

18 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine an association between the first direction and the second direction;

determine that a noise source is associated with third audio data corresponding to a third direction relative to the device, the third direction different from the second direction; and

determine, using the second audio data, the third audio data, and the first plurality of filter coefficient values, the first reference data.

19 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine that a first noise level associated with the second audio data satisfies a condition;

determine that a second noise level associated with third audio data satisfies the condition, the third audio data corresponding to a third direction relative to the device, the third direction different from the second direction;

determine an angular value indicating a separation between the first direction and the third direction;

determine that the angular value is below a threshold value; and

determine, using the second audio data and the first plurality of filter coefficient values, the first reference data.

20 . The system of claim 13 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using a portion of the second audio data and the second plurality of filter coefficient values, second reference data;

determine second error data using the second reference data and a portion of the first audio data; and

determine, using the second error data, the first step-size value, the second audio data, and the second plurality of filter coefficient values, a third plurality of filter coefficient values associated with the first adaptive filter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2023
From: LIU, IAN ERNAN; AYRAPETIAN, ROBERT; GOVINDARAJU, PRADEEP KUMAR; SARAF, MADHURI; MURGIA, CARLO
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064890/0147 →
References Cited (73)
US 5808967A · Yu et al. · 1998 [cited by applicant]
US 6049607A · Marash et al. · 2000 [cited by applicant]
US 6836243B2 · Kajala et al. · 2004 [cited by applicant]
US 7117145B1 · Venkatesh et al. · 2006 [cited by applicant]
US 7174022B1 · Zhang et al. · 2007 [cited by applicant]
US 7190775B2 · Rambo · 2007 [cited by applicant]
US 7359520B2 · Brennan et al. · 2008 [cited by applicant]
US 8139787B2 · Haykin et al. · 2012 [cited by applicant]
US 8160273B2 · Visser et al. · 2012 [cited by applicant]
US 8175291B2 · Chan et al. · 2012 [cited by applicant]
US 8296143B2 · Kudoh · 2012 [cited by applicant]
US 8321214B2 · Chan et al. · 2012 [cited by applicant]
US 8538749B2 · Visser et al. · 2013 [cited by applicant]
US 8620672B2 · Visser et al. · 2013 [cited by applicant]
US 8744849B2 · Liao · 2014 [cited by applicant]
US 8831936B2 · Toman et al. · 2014 [cited by applicant]
US 8849657B2 · Shin · 2014 [cited by applicant]
US 8861756B2 · Zhu et al. · 2014 [cited by applicant]
US 8929564B2 · Kikkeri · 2015 [cited by applicant]
US 8954324B2 · Wang et al. · 2015 [cited by applicant]
US 9048942B2 · Hershey et al. · 2015 [cited by applicant]
US 9173025B2 · Dickins et al. · 2015 [cited by applicant]
US 9224393B2 · Kjems et al. · 2015 [cited by applicant]
US 9275642B2 · Abdossalami et al. · 2016 [cited by applicant]
US 9338551B2 · Thyssen et al. · 2016 [cited by applicant]
US 9432769B1 · Sundaram et al. · 2016 [cited by applicant]
US 9456276B1 · Chhetri · 2016 [cited by applicant]
US 9530406B2 · Oh · 2016 [cited by applicant]
US 9653060B1 · Hilmes et al. · 2017 [cited by applicant]
US 9659555B1 · Hilmes et al. · 2017 [cited by applicant]
US 9689960B1 · Barton et al. · 2017 [cited by applicant]
US 9711131B2 · Christoph · 2017 [cited by applicant]
US 9747920B2 · Ayrapetian et al. · 2017 [cited by applicant]
US 9818425B1 · Ayrapetian et al. · 2017 [cited by applicant]
US 9966059B1 · Ayrapetian et al. · 2018 [cited by applicant]
US 9966086B1 · Piersol et al. · 2018 [cited by applicant]
US 9967661B1 · Hilmes et al. · 2018 [cited by applicant]
US 9973849B1 · Zhang et al. · 2018 [cited by applicant]
US 10187721B1 · Mansour · 2019 [cited by applicant]
US 10306361B2 · Morton et al. · 2019 [cited by applicant]
US 10339954B2 · Kamdar et al. · 2019 [cited by applicant]
US 10366702B2 · Morton et al. · 2019 [cited by applicant]
US 10403299B2 · Wung et al. · 2019 [cited by applicant]
US 10475471B2 · Ebenezer · 2019 [cited by applicant]
US 10499139B2 · Ganeshkumar · 2019 [cited by applicant]
US 10522167B1 · Ayrapetian et al. · 2019 [cited by applicant]
US 10598543B1 · Mansour et al. · 2020 [cited by applicant]
US 10657981B1 · Mansour et al. · 2020 [cited by applicant]
US 10771894B2 · Janse et al. · 2020 [cited by applicant]
US 10887709B1 · Mansour et al. · 2021 [cited by applicant]
US 11094334B2 · Wang et al. · 2021 [cited by applicant]
US 11200908B2 · Liu et al. · 2021 [cited by applicant]
US 11539833B1 · Nakagawa · 2022 [cited by examiner]
US 11657829B2 · Popovic et al. · 2023 [cited by applicant]
US 12531048B1 · Pulugurtha · 2026 [cited by examiner]
US 20040175006A1 · Kim et al. · 2004 [cited by applicant]
US 20080208538A1 · Visser et al. · 2008 [cited by applicant]
US 20080312918A1 · Kim · 2008 [cited by applicant]
US 20090034752A1 · Zhang et al. · 2009 [cited by applicant]
US 20130301840A1 · Yemdji · 2013 [cited by examiner]
US 20130304476A1 · Kim et al. · 2013 [cited by applicant]
US 20140025374A1 · Lou · 2014 [cited by applicant]
US 20160205263A1 · Liu et al. · 2016 [cited by applicant]
US 20190394576A1 · Petersen · 2019 [cited by examiner]
US 20200372891A1 · Bou Daher · 2020 [cited by examiner]
US 20210249005A1 · Bromand · 2021 [cited by examiner]
US 20210312936A1 · Hu · 2021 [cited by applicant]
US 20220406286A1 · Yamanashi et al. · 2022 [cited by applicant]
US 20230055257A1 · Li · 2023 [cited by applicant]
U.S. Appl. No. 17/488,471, filed Sep. 29, 2021. [cited by applicant]
U.S. Appl. No. 17/470,035, filed Sep. 9, 2021. [cited by applicant]
U.S. Appl. No. 18/323,697, filed May 25, 2023. [cited by applicant]
U.S. Appl. No. 18/460,955, filed Sep. 5, 2023. [cited by applicant]