IP Library › Granted Patent US 12,387,727
Granted Patent B1
US 12,387,727 · App. 18/439,412 · Granted Aug 12, 2025

Speech processing optimizations based on microphone array

Inventors: Shiva Kumar Sundaram (Mountain View, CA); Minhua Wu (San Jose, CA); Anirudh Raju (San Jose, CA); Spyridon Matsoukas (Hopkinton, MA); Arindam Mandal (Redwood City, CA); Kenichi Kumatani (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G10L15/22G06F40/40G10L15/187G10L15/26G10L15/30G10L21/0208H04R3/005G10L2015/088G10L2015/223G10L2021/02166H04W4/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,727
App. No.
18/439,412
Granted
Aug 12, 2025
Kind
B1
Abstract

Systems and methods for utilizing microphone array information for acoustic modeling are disclosed. Audio data may be received from a device having a microphone array configuration. Microphone configuration data may also be received that indicates the configuration of the microphone array. The microphone configuration data may be utilized as an input vector to an acoustic model, along with the audio data, to generate phoneme data. Additionally, the microphone configuration data may be utilized to train and/or generate acoustic models, select an acoustic model to perform speech recognition with, and/or to improve trigger sound detection.

Claims (50)

1. A method comprising:

receiving first input data from a first device having a first device location of a building;

determining first contextual data indicating one or more first physical objects disposed in association with the first device location;

inputting, into a first model, the first input data and the first contextual data;

generating, using the first model and based at least in part on the first input data and the first contextual data, first output data;

determining that the first device is in a second device location of the building, wherein the second device location is associated with second contextual data indicating one or more second physical objects that differ from the one or more first physical objects;

generating, using the first model and based at least in part on second input data and the second contextual data, second output data that differs at least in part from the first output data, the second output data indicating an action to be performed by the first device responsive to a user query, wherein the action includes controlling a playback parameter of an media output of the first device associated with the user query; and

sending, based at least in part on generating the second output data, a command to the first device that causes the first device to perform the action.

2. The method of claim 1 , wherein generating the first output data comprises generating text data based at least in part on the first input data and the first contextual data.

3. The method of claim 1 , wherein the first contextual data indicates a distance between the first device and a second device of the building.

4. The method of claim 3 , wherein generating the first output data using the first model comprises generating the first output data using the first model with the distance as an input to the first model.

5. The method of claim 1 , wherein the first input data indicates the first device location with respect to a source of the first contextual data.

6. The method of claim 1 , further comprising:

identifying additional input data;

generating third contextual data based at least in part on the additional input data; and

wherein generating the second output data comprises generating the second output data with the third contextual data as input to the first model.

7. The method of claim 1 , further comprising:

determining an environmental context associated with at least a portion of the first input data; and

wherein generating the first output data comprises generating the first output data with the environmental context as input to the first model.

8. The method of claim 1 , further comprising:

determining that a distance between the first device and a second device in the building satisfies a threshold distance; and

wherein generating the first output data comprises generating the first output data based at least in part on the distance satisfying the threshold distance.

9. The method of claim 1 , further comprising determining characteristics of the second device location, wherein the second contextual data is based at least in part on the characteristics of the second device location.

10. The method of claim 1 , wherein the playback parameter includes one or more of mute, volume control, or play.

11. A system comprising:

one or more processors; and

non-transitory computer-readable media storing instructions that, when executed by on the one or more processors, cause the one or more processors to perform operations comprising:

receiving first input data from a first device having a first device location of a building;

determining first contextual data indicating one or more first physical objects disposed in association with the first device location;

inputting, into a first model, the first input data and the first contextual data;

generating, using the first model and based at least in part on the first input data and the first contextual data, first output data;

determining that the first device is in a second device location of the building, wherein the second device location is associated with second contextual data indicating one or more second physical objects that differ from the one or more first physical objects;

generating, using the first model and based at least in part on second input data and the second contextual data, second output data that differs at least in part from the first output data, the second output data indicating an action to be performed by the first device responsive to a user query, wherein the action includes instructions controlling a playback parameter of a media output of the first device that is associated with the user query; and

sending, based at least in part on generating the second output data, a command to the first device that causes the first device to perform the action.

12. The system of claim 11 , wherein generating the first output data comprises generating text data based at least in part on the first input data and the first contextual data.

13. The system of claim 11 , wherein the first contextual data indicates a distance between the first device and a second device of the building.

14. The system of claim 13 , wherein generating the first output data using the first model comprises generating the first output data using the first model with the distance as an input to the first model.

15. The system of claim 11 , wherein the first input data indicates the first device location with respect to a source of the first contextual data.

16. The system of claim 11 , the operations further comprising:

identifying additional input data;

generating third contextual data based at least in part on the additional input data; and

wherein generating the second output data comprises generating the second output data with the third contextual data as input to the first model.

17. The system of claim 11 , the operations further comprising:

determining an environmental context associated with at least a portion of the first input data; and

wherein generating the first output data comprises generating the first output data with the environmental context as input to the first model.

18. The system of claim 11 , the operations further comprising:

determining that a distance between the first device and a second device in the building satisfies a threshold distance; and

wherein generating the first output data comprises generating the first output data based at least in part on the distance satisfying the threshold distance.

19. The system of claim 11 , the operations further comprising determining characteristics of the second device location, wherein the second contextual data is based at least in part on the characteristics of the second device location.

20. The system of claim 11 , wherein the playback parameter includes one or more of mute, volume control, or play.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2024
From: SUNDARAM, SHIVA KUMAR; WU, MINHUA; RAJU, ANIRUDH; MATSOUKAS, SPYRIDON; MANDAL, ARINDAM; KUMATANI, KENICHI
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 067254/0375 →
Continuity (2)
Continuation 16895377 · Jun 8, 2020
Continuation 15927764 · Mar 21, 2018
References Cited (28)
US 7515721B2 · Tashev et al. · 2009 [cited by applicant]
US 10152968B1 · Agrusa et al. · 2018 [cited by applicant]
US 20040059575A1 · Brookes et al. · 2004 [cited by applicant]
US 20050080632A1 · Endo et al. · 2005 [cited by applicant]
US 20100049516A1 · Talwar et al. · 2010 [cited by applicant]
US 20120191449A1 · Lloyd et al. · 2012 [cited by applicant]
US 20140278394A1 · Bastyr et al. · 2014 [cited by applicant]
US 20150058003A1 · Mohideen et al. · 2015 [cited by applicant]
US 20150221300A1 · Sukhomlinov · 2015 [cited by applicant]
US 20160171976A1 · Sun et al. · 2016 [cited by applicant]
US 20160196491A1 · Chandrasekaran et al. · 2016 [cited by applicant]
US 20160322055A1 · Sainath et al. · 2016 [cited by applicant]
US 20160350280A1 · Lavallee et al. · 2016 [cited by applicant]
US 20170061100A1 · Sati · 2017 [cited by applicant]
US 20170187566A1 · Fu · 2017 [cited by examiner]
US 20170187711A1 · Joo et al. · 2017 [cited by applicant]
US 20170289766A1 · Scott et al. · 2017 [cited by applicant]
US 20180052824A1 · Ferrydiansyah et al. · 2018 [cited by applicant]
US 20180061276A1 · Baca · 2018 [cited by examiner]
US 20180121389A1 · Jochim et al. · 2018 [cited by applicant]
US 20180174580A1 · Kim et al. · 2018 [cited by applicant]
US 20180253968A1 · Yalla · 2018 [cited by examiner]
Office Action for U.S. Appl. No. 16/895,377, mailed on Jan. 26, 2023, Shiva Kumar Sundaram, “Speech Processing Optimizations Based on Microphone Array”, 13 pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/895,377, mailed on Sep. 1, 2021, Sundaram, “Speech Processing Optimizations Based on Microphone Array”, 14 Pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/895,377, mailed on Sep. 27, 2022, Shiva Kumar Sundaram, “Speech Processing Optimizations Based on Microphone Array”, 12 pages. [cited by applicant]
Non Final Office Action dated Oct. 15, 2019 for U.S. Appl. No. 15/927,764 “Speech Processing Optimizations Based on Microphone Array” Sundaram, 17 pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/895,377, mailed on Dec. 24, 2021, Sundaram, “Speech Processing Optimizations Based on Microphone Array”, 15 Pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/895,377, mailed Apr. 8, 2022, Sundaram, “Speech Processing Optimizations Based on Microphone Array”, 10 pages. [cited by applicant]