IP Library Granted Patent US 12,445,793
Granted Patent B2
US 12,445,793 · App. 18/255,554 · Granted Oct 14, 2025

Automatic localization of audio devices

Inventors: Daniel Arteaga (Barcelona, ES); Davide Scaini (Barcelona, ES); Mark R. P. Thomas (Walnut Creek, CA); Avery Bruni (San Francisco, CA); Olha Michelle Townsend (San Francisco, CA)
Assignees: Dolby Laboratories Licensing Corporation; Dolby International AB
H04S7/301H04R3/005H04R5/02H04S7/303H04R2430/23H04S2400/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,793
App. No.
18/255,554
Granted
Oct 14, 2025
Kind
B2
Abstract

A method may involve: receiving direction of arrival (DOA) data corresponding to sound emitted by at least a first smart audio device of the audio environment that includes a first audio transmitter and a first audio receiver, the DOA data corresponding to sound received by at least a second smart audio device of the audio environment that includes a second audio transmitter and a second audio receiver, the DOA data corresponding to sound emitted by at least the second smart audio device and received by at least the first smart audio device; receiving one or more configuration parameters corresponding to the audio environment, to one or more audio devices, or both; and minimizing a cost function based at least in part on the DOA data and the configuration parameter(s), to estimate a position and an orientation of at least the first smart audio device and the second smart audio device.

Claims (25)

1. A method for localizing audio devices in an audio environment, the method comprising:

obtaining, by a control system, direction of arrival (DOA) data corresponding to sound emitted by at least a first smart audio device of the audio environment, the first smart audio device including a first audio transmitter and a first audio receiver, the DOA data corresponding to sound received by at least a second smart audio device of the audio environment, the second smart audio device including a second audio transmitter and a second audio receiver, the DOA data also corresponding to sound emitted by at least the second smart audio device and received by at least the first smart audio device;

receiving, by the control system, configuration parameters, the configuration parameters corresponding to the audio environment, corresponding to one or more audio devices of the audio environment, or corresponding to both the audio environment and the one or more audio devices of the audio environment; and

minimizing, by the control system, a cost function based at least in part on the DOA data and the configuration parameters, to estimate a position and an orientation of at least the first smart audio device and the second smart audio device.

2. The method of claim 1 , wherein the DOA data also corresponds to sound received by one or more passive audio receivers of the audio environment, each of the one or more passive audio receivers including a microphone array but lacking an audio emitter, and wherein minimizing the cost function also provides an estimated location and orientation of each of the one or more passive audio receivers.

3. The method of claim 1 , wherein the DOA data also corresponds to sound emitted by one or more audio emitters of the audio environment, each of the one or more audio emitters including at least one sound-emitting transducer but lacking a microphone array, and wherein minimizing the cost function also provides an estimated location of each of the one or more audio emitters.

4. The method of claim 1 , wherein the DOA data also corresponds to sound emitted by third through N th smart audio devices of the audio environment, N corresponding to a total number of smart audio devices of the audio environment, wherein the DOA data also corresponds to sound received by each of the first through N th smart audio devices from all other smart audio devices of the audio environment and wherein minimizing the cost function involves estimating a position and an orientation of the third through N th smart audio devices.

5. The method of claim 1 , wherein the configuration parameters include at least one of a number of audio devices in the audio environment, one or more dimensions of the audio environment, one or more constraints on audio device location or orientation, or disambiguation data for at least one of rotation, translation or scaling.

6. The method of claim 1 , further comprising receiving, by the control system, a seed layout for the cost function, the seed layout specifying a correct number of audio transmitters and receivers in the audio environment and an arbitrary location and orientation for each of the audio transmitters and receivers in the audio environment.

7. The method of claim 1 , further comprising receiving, by the control system, a weight factor associated with one or more elements of the DOA data, the weight factor indicating at least one of the availability or reliability of the one or more elements.

8. The method of claim 1 , further comprising obtaining, by the control system, one or more elements of the DOA data using at least one of a beamforming method, a steered powered response method, a time difference of arrival method or a structured signal method.

9. The method of claim 1 , further comprising receiving, by the control system, time of arrival (TOA) data corresponding to sound emitted by at least one audio device of the audio environment and received by at least one other audio device of the audio environment and wherein the cost function is based at least in part on the TOA data.

10. The method of claim 9 , further comprising estimating at least one playback latency, estimating at least one recording latency, or estimating at least one playback latency and at least one recording latency.

11. The method of claim 10 , wherein the cost function operates with at least one of a rescaled position, a rescaled latency or a rescaled time of arrival.

12. The method of claim 9 , wherein the cost function includes a first term depending on the DOA data only and second term depending on the TOA data only.

13. The method of claim 12 , wherein the first term includes a first weight factor and wherein the second term includes a second weight factor.

14. The method of claim 12 , wherein one or more TOA elements of the second term has a TOA element weight factor indicating the availability or reliability of each of the one or more TOA elements.

15. The method of claim 1 , wherein the configuration parameters include at least one of: playback latency data; recording latency data; data for disambiguating latency symmetry; disambiguation data for rotation; disambiguation data for translation; or

disambiguation data for scaling.

16. An apparatus comprising:

a control system configured to:

obtain direction of arrival (DOA) data corresponding to sound emitted by at least a first smart audio device of the audio environment, the first smart audio device including a first audio transmitter and a first audio receiver, the DOA data corresponding to sound received by at least a second smart device of the audio environment, the second smart audio device including a second audio transmitter and a second audio receiver, the DOA data also corresponding to sound emitted by at least the second smart audio device and received by at least the first smart audio device,

receive configuration parameters, the configuration parameters corresponding to the audio environment, corresponding to one or more audio devices of the audio environment, or corresponding to both the audio environment and the one or more audio devices of the audio environment, and

minimize a cost function based at least in part on the DOA data and the configuration parameters, to estimate a position and an orientation of at least the first smart audio device and the second smart audio device.

17. A non-transitory computer readable medium containing instructions that when executed by a processor perform the method of claim 1 .

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2025
From: ARTEAGA, DANIEL; SCAINI, DAVIDE; THOMAS, MARK R. P.; BRUNI, AVERY; TOWNSEND, OLHA MICHELLE
To: DOLBY LABORATORIES LICENSING CORPORATION; DOLBY INTERNATIONAL AB
Reel/Frame 072264/0933 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2025
From: ARTEAGA, DANIEL; SCAINI, DAVIDE; THOMAS, MARK R. P.; BRUNI, AVERY; TOWNSEND, OLHA MICHELLE
To: DOLBY LABORATORIES LICENSING CORPORATION; DOLBY INTERNATIONAL AB
Reel/Frame 071667/0507 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2024
From: ARTEAGA, DANIEL; SCAINI, DAVID; THOMAS, MARK R.P.; BRUNI, AVERY; TOWNSEND, OLHA MICHELLE
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 066211/0940 →
Priority Claims (2)
ES 202031212 · Dec 3, 2020 · national
ES 202130458 · May 20, 2021 · national
Continuity (4)
Provisional Application 63224778 · Jul 22, 2021
Provisional Application 63203403 · Jul 21, 2021
Provisional Application 63155369 · Mar 2, 2021
Related Publication 20240022869A1 · Jan 18, 2024
References Cited (37)
US 8682675B2 · Togami · 2014 [cited by applicant]
US 8743658B2 · Claussen · 2014 [cited by applicant]
US 8861756B2 · Zhu · 2014 [cited by applicant]
US 8879741B2 · Fukuyama · 2014 [cited by applicant]
US 9031268B2 · Fejzo · 2015 [cited by applicant]
US 9197978B2 · Usami · 2015 [cited by applicant]
US 9408011B2 · Kim · 2016 [cited by applicant]
US 9497544B2 · Mohammad · 2016 [cited by applicant]
US 9549253B2 · Alexandridis · 2017 [cited by applicant]
US 9609141B2 · Beaucoup · 2017 [cited by applicant]
US 9788119B2 · Vilermo · 2017 [cited by applicant]
US 9788120B2 · Miyasaka · 2017 [cited by applicant]
US 9971012B2 · Nakamura · 2018 [cited by applicant]
US 10270642B2 · Zhang · 2019 [cited by applicant]
US 10331396B2 · Habets · 2019 [cited by applicant]
US 10748544B2 · Nakadai · 2020 [cited by applicant]
US 20100217590A1 · Nemer · 2010 [cited by applicant]
US 20110091055A1 · Leblanc · 2011 [cited by applicant]
US 20160241955A1 · Thyssen · 2016 [cited by applicant]
US 20180299527A1 · Helwani · 2018 [cited by applicant]
US 20190132685A1 · Skoglund · 2019 [cited by applicant]
US 20190253801A1 · Arteaga · 2019 [cited by applicant]
US 20190355373A1 · Nesta · 2019 [cited by applicant]
US 20200066295A1 · Karimian-Azari · 2020 [cited by applicant]
US 20200288262A1 · Eronen · 2020 [cited by applicant]
US 20230040846A1 · Thomas · 2023 [cited by applicant]
US 20250008262A1 · Bruni · 2025 [cited by examiner]
JP 2020184788A · 2020 [cited by applicant]
RU 2546717C2 · 2015 [cited by applicant]
RU 2734231C1 · 2020 [cited by applicant]
WO 2018029341A1 · 2018 [cited by applicant]
WO 2019078816A1 · 2019 [cited by applicant]
WO 2020210084A1 · 2020 [cited by applicant]
Kozintsev, I. et al., “Position calibration of microphones and loudspeakers in distributed computing platforms”, IEEE Transactions on Speech and Audio Processing, Year: 2005 vol. 13 , Issue: 1 pp. 70-83. [cited by applicant]
Nadiri, O. et al.; “Localization of Multiple Speakers Under High Reverberation Using a Spherical Microphone Array and the Direct-Path Dominance Test”; Oct. 2014; IEEE/ACM Transactions on Audio, Speech, and Language Proc… [cited by applicant]
Tashev, Ivan J. et al; “Cost Function for Sound Source Localization With Arbitrary Microphone Arrays”; 2017; IEEE; Hands-free Speech Communications and Microphone Arrays (HSCMA); pp. 74-80. [cited by applicant]
Tehrani, Ali Kafaei et al.; “Sound Source Localization Using Time Differences of Arrival; Euclidean Distance Matrices Based Approach”; 2018; 9th International Symposium on Telecommunications; p. 91-95. [cited by applicant]