IP Library Granted Patent US 12,401,942
Granted Patent B1
US 12,401,942 · App. 18/323,697 · Granted Aug 26, 2025

Group beam selection and beam merging

Inventors: Robert Ayrapetian (Morgan Hill, CA); Gautam Shreedhar Bhat (Sunnyvale, CA); Pradeep Kumar Govindaraju (San Jose, CA)
Assignee: Amazon Technologies, Inc.
H04R1/32H04B11/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,401,942
App. No.
18/323,697
Granted
Aug 26, 2025
Kind
B1
Abstract

A system that performs beam selection and beam merging using beam-specific signal quality metrics corresponding to a minimum noise floor. For example, a device may track a minimum noise floor for each beam, determine a highest minimum noise floor across the beams, and determine a noise floor ratio between the beam-specific minimum noise floor and the highest minimum noise floor. Using a combination of the noise floor ratio and signal-to-noise ratio (SNR) values, the device may perform beam selection by prioritizing low background noise as well as high SNR to select a pre-defined beam group. In addition, the device may use the noise floor ratio to perform beam merging and generate single-channel output audio data using the selected beam group. For example, the device may scale the beams based on a combination of the SNR value and the noise floor ratio.

Claims (112)

1. A computer-implemented method, the method comprising:

receiving first audio data including a first portion corresponding to a first direction and a second portion corresponding to a second direction, the first audio data including a first representation of speech;

determining, using the first portion of the first audio data, a first value representing a first noise floor of the first portion of the first audio data;

determining, using the second portion of the first audio data, a second value representing a second noise floor of the second portion of the first audio data;

determining, using the first value, a first signal quality metric value corresponding to the first portion of the first audio data;

determining, using the first value and the second value, a first ratio value; and

determining, using the first signal quality metric value and the first ratio value, second audio data corresponding to a first plurality of directions including the first direction, the second audio data including a second representation of the speech.

2. The computer-implemented method of claim 1 , wherein determining the second audio data further comprises:

determining, using a third portion of the first audio data corresponding to a third direction, a third value representing a third noise floor of the third portion;

determining, using the third value, a second signal quality metric value corresponding to the third portion;

determining, using the third value and the second value, a second ratio value;

determining a first portion of the second audio data using the first portion of the first audio data, the first signal quality metric value, the first ratio value, the second signal quality metric value, and the second ratio value; and

determining a second portion of the second audio data using the third portion of the first audio data, the first signal quality metric value, the first ratio value, the second signal quality metric value, and the second ratio value.

3. The computer-implemented method of claim 1 , wherein determining the second audio data further comprises:

determining a first product of the first ratio value and the first signal quality metric value;

determining a second product of a second ratio value and a second signal quality metric value, the second ratio value and the second signal quality metric value associated with a third portion of the first audio data corresponding to a third direction;

determining a first weight value using the first product and a sum of the first product and the second product;

determining a second weight value using the second product and the sum of the first product and the second product;

determining, using the first weight value and the first portion of the first audio data, a first portion of the second audio data; and

determining, using the second weight value and the third portion of the first audio data, a second portion of the second audio data.

4. The computer-implemented method of claim 1 , wherein determining the second audio data further comprises:

determining a first group value using the first signal quality metric value and the first ratio value, the first group value corresponding to the first plurality of directions;

determining that the first group value satisfies a condition; and

determining, using the first plurality of directions and the first audio data, the second audio data.

5. The computer-implemented method of claim 1 , wherein determining the second audio data further comprises:

determining a first group value using the first signal quality metric value and the first ratio value, the first group value corresponding to the first plurality of directions;

determining a second group value corresponding to a second plurality of directions including the second direction;

determining that the first group value is higher than the second group value; and

determining, using the first plurality of directions and the first audio data, the second audio data.

6. The computer-implemented method of claim 1 , further comprising:

determining, using a third portion of the first audio data corresponding to a third direction, a third value representing a third noise floor of the third portion of the first audio data;

determining that the second value is greater than the first value and the third value;

determining, using the second value, a second ratio value corresponding to the second portion of the first audio data; and

determining, using the third value and the second value, a third ratio value corresponding to the third portion of the first audio data.

7. The computer-implemented method of claim 6 , wherein determining the second audio data further comprises:

determining a first product of the first ratio value and the first signal quality metric value;

determining a second product of the third ratio value and a second signal quality metric value associated with the third portion of the first audio data;

determining that a sum of the first product and the second product satisfies a condition; and

determining, using the first plurality of directions and the first audio data, the second audio data, wherein the first plurality of directions includes the first direction and the third direction.

8. The computer-implemented method of claim 6 , wherein determining the second audio data further comprises:

determining a first product of the first ratio value and the first signal quality metric value;

determining a second product of the third ratio value and a second signal quality metric value associated with the third portion of the first audio data;

determining, using the first product and the first portion of the first audio data, a first portion of the second audio data; and

determining, using the second product and the third portion of the first audio data, a second portion of the second audio data.

9. The computer-implemented method of claim 1 , wherein determining the second audio data further comprises:

determining a first group value using the first signal quality metric value and the first ratio value;

determining a second group value corresponding to a second plurality of directions that includes the second direction;

determining a third group value using the first group value and a first weight value associated with the first plurality of directions;

determining a fourth group value using the second group value and a second weight value associated with the second plurality of directions;

determining that the third group value is higher than the fourth group value; and

determining, using the first plurality of directions and the first audio data, the second audio data.

10. The computer-implemented method of claim 1 , wherein determining the first value further comprises:

determining, using the first portion of the first audio data, first power values;

determining, using the first power values, first noise floor data; and

determining the first value using a lowest value represented in the first noise floor data.

11. A system comprising:

at least one processor; and

memory including instructions operable to be executed by the at least one processor to cause the system to:

receive first audio data including a first portion corresponding to a first direction and a second portion corresponding to a second direction, the first audio data including a first representation of speech;

determine, using the first portion of the first audio data, a first value representing a first noise floor of the first portion of the first audio data;

determine, using the second portion of the first audio data, a second value representing a second noise floor of the second portion of the first audio data;

determine, using the first value, a first signal quality metric value corresponding to the first portion of the first audio data;

determine, using the first value and the second value, a first ratio value; and

determine, using the first signal quality metric value and the first ratio value, second audio data corresponding to a first plurality of directions including the first direction, the second audio data including a second representation of the speech.

12. The system of claim 11 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using a third portion of the first audio data corresponding to a third direction, a third value representing a third noise floor of the third portion;

determine, using the third value, a second signal quality metric value corresponding to the third portion;

determine, using the third value and the second value, a second ratio value;

determine a first portion of the second audio data using the first portion of the first audio data, the first signal quality metric value, the first ratio value, the second signal quality metric value, and the second ratio value; and

determine a second portion of the second audio data using the third portion of the first audio data, the first signal quality metric value, the first ratio value, the second signal quality metric value, and the second ratio value.

13. The system of claim 11 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first product of the first ratio value and the first signal quality metric value;

determine a second product of a second ratio value and a second signal quality metric value, the second ratio value and the second signal quality metric value associated with a third portion of the first audio data corresponding to a third direction;

determine a first weight value using the first product and a sum of the first product and the second product;

determine a second weight value using the second product and the sum of the first product and the second product;

determine, using the first weight value and the first portion of the first audio data, a first portion of the second audio data; and

determine, using the second weight value and the third portion of the first audio data, a second portion of the second audio data.

14. The system of claim 11 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first group value using the first signal quality metric value and the first ratio value, the first group value corresponding to the first plurality of directions;

determine that the first group value satisfies a condition; and

determine, using the first plurality of directions and the first audio data, the second audio data.

15. The system of claim 11 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first group value using the first signal quality metric value and the first ratio value, the first group value corresponding to the first plurality of directions;

determine a second group value corresponding to a second plurality of directions including the second direction;

determine that the first group value is higher than the second group value; and

determine, using the first plurality of directions and the first audio data, the second audio data.

16. The system of claim 11 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using a third portion of the first audio data corresponding to a third direction, a third value representing a third noise floor of the third portion of the first audio data;

determine that the second value is greater than the first value and the third value;

determine, using the second value, a second ratio value corresponding to the second portion of the first audio data; and

determine, using the third value and the second value, a third ratio value corresponding to the third portion of the first audio data.

17. The system of claim 16 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first product of the first ratio value and the first signal quality metric value;

determine a second product of the third ratio value and a second signal quality metric value associated with the third portion of the first audio data;

determine that a sum of the first product and the second product satisfies a condition; and

determine, using the first plurality of directions and the first audio data, the second audio data, wherein the first plurality of directions includes the first direction and the third direction.

18. The system of claim 16 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first product of the first ratio value and the first signal quality metric value;

determine a second product of the third ratio value and a second signal quality metric value associated with the third portion of the first audio data;

determine, using the first product and the first portion of the first audio data, a first portion of the second audio data; and

determine, using the second product and the third portion of the first audio data, a second portion of the second audio data.

19. The system of claim 11 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first group value using the first signal quality metric value and the first ratio value;

determine a second group value corresponding to a second plurality of directions that includes the second direction;

determine a third group value using the first group value and a first weight value associated with the first plurality of directions;

determine a fourth group value using the second group value and a second weight value associated with the second plurality of directions;

determine that the third group value is higher than the fourth group value; and

determine, using the first plurality of directions and the first audio data, the second audio data.

20. The system of claim 11 , wherein the memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using the first portion of the first audio data, first power values;

determine, using the first power values, first noise floor data; and

determine the first value using a lowest value represented in the first noise floor data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2023
From: AYRAPETIAN, ROBERT; SHREEDHAR BHAT, GAUTAM; GOVINDARAJU, PRADEEP KUMAR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 063763/0768 →
References Cited (66)
US 5808967A · Yu et al. · 1998 [cited by applicant]
US 6049607A · Marash et al. · 2000 [cited by applicant]
US 6836243B2 · Kajala et al. · 2004 [cited by applicant]
US 7117145B1 · Venkatesh et al. · 2006 [cited by applicant]
US 7174022B1 · Zhang et al. · 2007 [cited by applicant]
US 7190775B2 · Rambo · 2007 [cited by applicant]
US 7359520B2 · Brennan et al. · 2008 [cited by applicant]
US 8139787B2 · Haykin et al. · 2012 [cited by applicant]
US 8160273B2 · Visser et al. · 2012 [cited by applicant]
US 8175291B2 · Chan et al. · 2012 [cited by applicant]
US 8296143B2 · Kudoh · 2012 [cited by applicant]
US 8321214B2 · Chan et al. · 2012 [cited by applicant]
US 8538749B2 · Visser et al. · 2013 [cited by applicant]
US 8620672B2 · Visser et al. · 2013 [cited by applicant]
US 8744849B2 · Liao · 2014 [cited by applicant]
US 8831936B2 · Toman et al. · 2014 [cited by applicant]
US 8849657B2 · Shin · 2014 [cited by applicant]
US 8861756B2 · Zhu et al. · 2014 [cited by applicant]
US 8929564B2 · Kikkeri · 2015 [cited by applicant]
US 8954324B2 · Wang et al. · 2015 [cited by applicant]
US 9048942B2 · Hershey et al. · 2015 [cited by applicant]
US 9173025B2 · Dickins et al. · 2015 [cited by applicant]
US 9224393B2 · Kjems et al. · 2015 [cited by applicant]
US 9275642B2 · Abdossalami et al. · 2016 [cited by applicant]
US 9338551B2 · Thyssen et al. · 2016 [cited by applicant]
US 9432769B1 · Sundaram et al. · 2016 [cited by applicant]
US 9456276B1 · Chhetri · 2016 [cited by applicant]
US 9530406B2 · Oh · 2016 [cited by applicant]
US 9653060B1 · Hilmes et al. · 2017 [cited by applicant]
US 9659555B1 · Hilmes et al. · 2017 [cited by applicant]
US 9689960B1 · Barton et al. · 2017 [cited by applicant]
US 9711131B2 · Christoph · 2017 [cited by applicant]
US 9747920B2 · Ayrapetian et al. · 2017 [cited by applicant]
US 9818425B1 · Ayrapetian et al. · 2017 [cited by applicant]
US 9966059B1 · Ayrapetian et al. · 2018 [cited by applicant]
US 9966086B1 · Piersol et al. · 2018 [cited by applicant]
US 9967661B1 · Hilmes et al. · 2018 [cited by applicant]
US 9973849B1 · Zhang et al. · 2018 [cited by applicant]
US 10187721B1 · Mansour · 2019 [cited by applicant]
US 10306361B2 · Morton et al. · 2019 [cited by applicant]
US 10339954B2 · Kamdar et al. · 2019 [cited by applicant]
US 10366702B2 · Morton et al. · 2019 [cited by applicant]
US 10403299B2 · Wung et al. · 2019 [cited by applicant]
US 10475471B2 · Ebenezer · 2019 [cited by applicant]
US 10499139B2 · Ganeshkumar · 2019 [cited by applicant]
US 10522167B1 · Ayrapetian et al. · 2019 [cited by applicant]
US 10598543B1 · Mansour et al. · 2020 [cited by applicant]
US 10657981B1 · Mansour et al. · 2020 [cited by applicant]
US 10771894B2 · Janse et al. · 2020 [cited by applicant]
US 10887709B1 · Mansour et al. · 2021 [cited by applicant]
US 11094334B2 · Wang et al. · 2021 [cited by applicant]
US 11200908B2 · Liu et al. · 2021 [cited by applicant]
US 11657829B2 · Popovic et al. · 2023 [cited by applicant]
US 20040175006A1 · Kim et al. · 2004 [cited by applicant]
US 20080208538A1 · Visser et al. · 2008 [cited by applicant]
US 20080312918A1 · Kim · 2008 [cited by applicant]
US 20090034752A1 · Zhang et al. · 2009 [cited by applicant]
US 20130304476A1 · Kim et al. · 2013 [cited by applicant]
US 20140025374A1 · Lou · 2014 [cited by applicant]
US 20160205263A1 · Liu et al. · 2016 [cited by applicant]
US 20210312936A1 · Hu · 2021 [cited by applicant]
US 20220109929A1 · Ayrapetian · 2022 [cited by examiner]
US 20220406286A1 · Yamanashi et al. · 2022 [cited by applicant]
US 20230055257A1 · Li · 2023 [cited by applicant]
U.S. Appl. No. 17/488,471, filed Sep. 29, 2021. [cited by applicant]
U.S. Appl. No. 17/470,035, filed Sep. 9, 2021. [cited by applicant]