IP Library › Granted Patent US 12,621,572
Granted Patent B2
US 12,621,572 · App. 18/774,697 · Granted May 5, 2026

Systems and methods for talker tracking and camera positioning in the presence of acoustic reflections

Inventors: Zachary Kane (Chicago, IL); Christopher George Rieger (Chicago, IL)
Assignee: Shure Acquisition Holdings, Inc.
H04N23/695H04N7/147H04N23/611H04N23/66H04N23/69G06T7/20G06T7/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,572
App. No.
18/774,697
Granted
May 5, 2026
Kind
B2
Abstract

Systems and methods configured to generate talker coordinates for directing a camera towards an active talker in the presence of acoustic reflections are disclosed. One method comprises receiving sound location information for a detected audio source from a microphone; determining, based on the sound location information, a first set of coordinates representing an estimated talker location; determining, based on the sound location information and a height of the environment, a second set of coordinates representing a corrected talker location; calculating a weighted height coordinate based on a first height coordinate of the first set of coordinates, a second height coordinate of the second set of coordinates, and stored height coordinates from previously detected audio sources; and transmitting, to a camera, a third set of coordinates comprising the weighted height coordinate and representing a final talker location, to cause the camera to point its image capturing component towards the received location.

Claims (54)

1 . A method performed by one or more processors in communication with each of a camera and at least one microphone disposed in an environment, the method comprising:

receiving, from the at least one microphone, sound location information for an audio source detected by the at least one microphone;

determining, based on the sound location information, a first set of coordinates representing an estimated talker location for the audio source;

determining, based on the sound location information and a height measurement of the environment, a second set of coordinates representing a corrected talker location for the audio source;

calculating a weighted height coordinate based on a first height coordinate of the first set of coordinates, a second height coordinate of the second set of coordinates, and stored height coordinates obtained for previously detected audio sources; and

transmitting, to the camera, a third set of coordinates comprising the weighted height coordinate and representing a final talker location for the audio source, wherein receipt of the third set of coordinates causes the camera to point an image capturing component of the camera towards the final talker location.

2 . The method of claim 1 , further comprising:

determining an amount of discrepancy between the first set of coordinates and the second set of coordinates; and

upon determining that the discrepancy exceeds a threshold, generating the third set of coordinates by replacing the height coordinate of the first set of coordinates with the weighted height coordinate.

3 . The method of claim 2 , wherein the threshold is configured to identify whether an acoustic reflection is present in the environment.

4 . The method of claim 1 , wherein calculating the weighted height coordinate comprises:

determining a first weight value for the first height coordinate based on the stored height coordinates;

determining a second weight value for the second height coordinate based on the stored height coordinates; and

calculating, using the first weight value and the second weight value, a weighted average of the first height coordinate and the second height coordinate.

5 . The method of claim 1 , wherein the camera is configured to point the image capturing component towards the final talker location by adjusting one or more of an angle, a tilt, a zoom, and a framing of the camera.

6 . The method of claim 1 , further comprising: determining the sound location information using an audio localization algorithm executed by an audio activity localizer.

7 . The method of claim 1 , wherein the height measurement comprises a height of the at least one microphone relative to a floor of the environment.

8 . A system comprising:

at least one microphone disposed in an environment and configured to determine sound location information for an audio source detected by the at least one microphone;

a camera disposed in the environment and comprising an image capturing component; and

one or more processors communicatively coupled to each of the at least one microphone and the camera, the one or more processors configured to:

receive the sound location information from the at least one microphone;

determine, based on the sound location information, a first set of coordinates representing an estimated talker location for the audio source;

determine, based on the sound location information and a height measurement of the environment, a second set of coordinates representing a corrected talker location for the audio source;

calculate a weighted height coordinate based on a first height coordinate of the first set of coordinates, a second height coordinate of the second set of coordinates, and stored height coordinates obtained for previously detected audio sources; and

transmit, to the camera, a third set of coordinates comprising the weighted height coordinate and representing a final talker location for the audio source,

wherein responsive to receiving the third set of coordinates, the camera is configured to point the image capturing component towards the final talker location.

9 . The system of claim 8 , wherein the one or more processors are further configured to:

determine an amount of discrepancy between the first set of coordinates and the second set of coordinates; and

upon determining that the discrepancy exceeds a threshold, generate the third set of coordinates by replacing the height coordinate of the first set of coordinates with the weighted height coordinate.

10 . The system of claim 9 , wherein the threshold is configured to identify whether an acoustic reflection is present in the environment.

11 . The system of claim 8 , wherein calculating the weighted height coordinate comprises:

determining a first weight value for the first height coordinate based on the stored height coordinates;

determining a second weight value for the second height coordinate based on the stored height coordinates; and

calculating, using the first weight value and the second weight value, a weighted average of the first height coordinate and the second height coordinate.

12 . The system of claim 8 , wherein the camera is configured to point the image capturing component towards the final talker location by adjusting one or more of an angle, a tilt, a zoom, and a framing of the camera.

13 . The system of claim 8 , further comprising an audio activity localizer configured to determine the sound location information using an audio localization algorithm executed by the audio activity localizer.

14 . The system of claim 8 , wherein the height measurement comprises a height of the at least one microphone relative to a floor of the environment.

15 . A non-transitory computer-readable storage medium comprising instructions that, when executed by one or more processors in communication with each of at least one microphone, and a camera, cause the one or more processors to perform the following:

receive sound location information for an audio source detected by the at least one microphone;

determine, based on the sound location information, a first set of coordinates representing an estimated talker location for the audio source;

determine, based on the sound location information and a height measurement of the environment, a second set of coordinates representing a corrected talker location for the audio source;

calculate a weighted height coordinate based on a first height coordinate of the first set of coordinates, a second height coordinate of the second set of coordinates, and stored height coordinates obtained for previously detected audio sources; and

transmit, to the camera, a third set of coordinates comprising the weighted height coordinate and representing a final talker location for the audio source, wherein receipt of the third set of coordinates causes the camera to point an image capturing component of the camera towards the final talker location.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the instructions further cause the one or more processors to:

determine an amount of discrepancy between the first set of coordinates and the second set of coordinates; and

upon determining that the discrepancy exceeds a threshold, generate the third set of coordinates by replacing the height coordinate of the first set of coordinates with the weighted height coordinate.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the threshold is configured to identify whether an acoustic reflection is present in the environment.

18 . The non-transitory computer-readable storage medium of claim 15 , wherein calculating the weighted height coordinate comprises:

determining a first weight value for the first height coordinate based on the stored height coordinates;

determining a second weight value for the second height coordinate based on the stored height coordinates; and

calculating, using the first weight value and the second weight value, a weighted average of the first height coordinate and the second height coordinate.

19 . The non-transitory computer-readable storage medium of claim 15 , wherein the camera is configured to point the image capturing component towards the final talker location by adjusting one or more of an angle, a tilt, a zoom, and a framing of the camera.

20 . The non-transitory computer-readable storage medium of claim 15 , wherein the instructions further cause the one or more processors to determine the sound location information using an audio localization algorithm executed by an audio activity localizer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2024
From: KANE, ZACHARY; RIEGER, CHRISTOPHER GEORGE
To: SHURE ACQUISITION HOLDINGS, INC.
Reel/Frame 068060/0671 →
Continuity (2)
Provisional Application 63514046 · Jul 17, 2023
Related Publication 20250030947A1 · Jan 23, 2025
References Cited (118)
US 5206721A · Ashida · 1993 [cited by applicant]
US 5686957A · Baker · 1997 [cited by applicant]
US 6731334B1 · Maeng · 2004 [cited by applicant]
US 6980485B2 · McCaskill · 2005 [cited by applicant]
US 7305095B2 · Rui · 2007 [cited by applicant]
US 7349008B2 · Rui · 2008 [cited by applicant]
US 7403217B2 · Schulz · 2008 [cited by applicant]
US 7487056B2 · Tashev · 2009 [cited by applicant]
US 7586513B2 · Muren · 2009 [cited by examiner]
US 7852369B2 · Cutler · 2010 [cited by applicant]
US 8189807B2 · Cutler · 2012 [cited by applicant]
US 8238573B2 · Ishibashi · 2012 [cited by applicant]
US 8248448B2 · Feng · 2012 [cited by applicant]
US 9030520B2 · Chu · 2015 [cited by applicant]
US 9179098B2 · Buckler · 2015 [cited by applicant]
US 9338549B2 · Haulick · 2016 [cited by applicant]
US 9350940B1 · Baker · 2016 [cited by applicant]
US 9392221B2 · Feng · 2016 [cited by applicant]
US 9489948B1 · Chu · 2016 [cited by applicant]
US 9621795B1 · Whyte · 2017 [cited by examiner]
US 9633270B1 · Tangeland · 2017 [cited by applicant]
US 9674453B1 · Tangeland · 2017 [cited by applicant]
US 9769424B2 · Michot · 2017 [cited by applicant]
US 9769552B2 · Choisel · 2017 [cited by applicant]
US 9980040B2 · Whyte · 2018 [cited by applicant]
US 10091412B1 · Feng · 2018 [cited by applicant]
US 10491809B2 · Feng · 2019 [cited by applicant]
US 10582117B1 · Tanaka · 2020 [cited by applicant]
US 10966022B1 · Chu · 2021 [cited by applicant]
US 11127401B2 · Sarkar · 2021 [cited by applicant]
US 11381906B2 · Rollow, IV · 2022 [cited by applicant]
US 11438691B2 · Veselinovic · 2022 [cited by applicant]
US 11601731B1 · Slotznick · 2023 [cited by applicant]
US 11695900B2 · Childress, Jr. · 2023 [cited by applicant]
US 11902656B2 · Muthiah · 2024 [cited by applicant]
US 12143806B2 · Mcelveen · 2024 [cited by applicant]
US 20050008169A1 · Muren · 2005 [cited by applicant]
US 20050140779A1 · Schulz · 2005 [cited by applicant]
US 20050283328A1 · Tashev · 2005 [cited by applicant]
US 20060005136A1 · Wallick · 2006 [cited by examiner]
US 20080037802A1 · Posa · 2008 [cited by applicant]
US 20090323981A1 · Cutler · 2009 [cited by applicant]
US 20110135125A1 · Zhan · 2011 [cited by applicant]
US 20110285808A1 · Feng · 2011 [cited by examiner]
US 20110285809A1 · Feng · 2011 [cited by applicant]
US 20110317522A1 · Florencio · 2011 [cited by applicant]
US 20120294118A1 · Haulick · 2012 [cited by applicant]
US 20130271559A1 · Feng · 2013 [cited by applicant]
US 20130300820A1 · Liu · 2013 [cited by applicant]
US 20140362163A1 · Winterstein · 2014 [cited by applicant]
US 20150146078A1 · Aarrestad · 2015 [cited by applicant]
US 20160277712A1 · Michot · 2016 [cited by examiner]
US 20170201825A1 · Whyte · 2017 [cited by applicant]
US 20180167581A1 · Goesnar · 2018 [cited by applicant]
US 20180270451A1 · Dickins · 2018 [cited by examiner]
US 20190158733A1 · Feng · 2019 [cited by applicant]
US 20200084366A1 · Fujiwara · 2020 [cited by applicant]
US 20200351435A1 · Therkelsen · 2020 [cited by applicant]
US 20210051397A1 · Veselinovic · 2021 [cited by applicant]
US 20210097995A1 · Sarkar · 2021 [cited by applicant]
US 20210152930A1 · Bryans · 2021 [cited by applicant]
US 20210266680A1 · Frieding · 2021 [cited by applicant]
US 20210345040A1 · Meyer · 2021 [cited by applicant]
US 20210360193A1 · Childress, Jr. · 2021 [cited by applicant]
US 20220201421A1 · Mcelveen · 2022 [cited by applicant]
US 20220353465A1 · Smith · 2022 [cited by applicant]
US 20230025997A1 · Liu · 2023 [cited by applicant]
US 20230053202A1 · Chu · 2023 [cited by applicant]
US 20230086490A1 · Abraham · 2023 [cited by applicant]
US 20230216988A1 · Yan · 2023 [cited by examiner]
US 20240007744A1 · Muthiah · 2024 [cited by examiner]
CN 102843540 · 2012 [cited by applicant]
CN 107809596 · 2018 [cited by applicant]
CN 108063910 · 2018 [cited by applicant]
CN 108370470 · 2018 [cited by applicant]
CN 112311999 · 2021 [cited by applicant]
CN 112689092 · 2021 [cited by applicant]
CN 113099160 · 2021 [cited by applicant]
EP 0765084 · 1997 [cited by applicant]
EP 1705911 · 2006 [cited by applicant]
JP H01180193 · 1989 [cited by applicant]
JP H0314386 · 1991 [cited by applicant]
JP H1042264 · 1998 [cited by applicant]
JP 2000041228 · 2000 [cited by applicant]
JP 3739673 · 2006 [cited by applicant]
JP 2011193392 · 2011 [cited by applicant]
JP 2020043456 · 2020 [cited by applicant]
KR 101884446 · 2018 [cited by applicant]
KR 102407872 · 2022 [cited by applicant]
WO 2008047804 · 2008 [cited by applicant]
ATND1061DAN, Beamforming Array Microphone, User Manual, Audio-technica, 2022, 107 pp. [cited by applicant]
AVer Press Release, “AVer Partners with Sennheiser to Transform Video and Audio Collaboration,” Nov. 18, 2021, 1 p. [cited by applicant]
AVer PTZ Link, User Manual, AVer Information Inc., 2022, 71 pp. [cited by applicant]
Dong, et al., “Research on TDOA based microphone array acoustic localization,” IEEE 12th International Conference on Electronic Measurement & Instruments, 2015, 5 pp. [cited by applicant]
HDL300 System, Nureva Product Page, Downloaded from webpage <https://www.nureva.com/audio-conferencing/hdl300> on Dec. 4, 2024, 10 pp. [cited by applicant]
International Search Report and Written Opinion for PCT/US2024/036908 dated Nov. 13, 2024, 14 pp. [cited by applicant]
Introducing StreamSet, Audio-technica, Product Page, Downloaded from webpage <https://www.audio-technica.com/en-us/> on May 15, 2023, 7 pp. [cited by applicant]
Kuhn, “Adaptive Audio/Video Zone Creation in Online Meeting Environment,” ip.com, Jun. 22, 2022, 5 pp. [cited by applicant]
Mitchell, “Poly Studio review: Take the pain out of videoconferencing,” ITPro, Sep. 17, 2019, 9 pp. [cited by applicant]
MXA920 Command Strings, Shure, User Manual, 2024, 41 pp. [cited by applicant]
Vissonic Audio Solutions, Array Microphone Conference Solution, Downloaded from <https://www.vissonic.com/solution/220.html> on Dec. 4, 2024, 7 pp. [cited by applicant]
Webex Help Center, “Differences between the presenter and the audience, briefing room, and classroom setups,” Downloaded from website <https://help.webex.com/en-us/article/n4522teb/Set-up-classroom-on-Room-Series> on Fe… [cited by applicant]
Yealink, MeetingEye 800, 4K UHD Video Conferencing Endpoint, Information Sheet, 2022, 4 pp. [cited by applicant]
Aarabi, “The Fusion of Distributed Microphone Arrays for Sound Localization,” EURASIP Journal on Applied Signal Processing, vol. 2003, No. 4, Jan. 2003, 10 pp. [cited by applicant]
Dibiase, et al., “Robust Localization in Reverberant Rooms,” Microphone Arrays, 2001, 24 pp. [cited by applicant]
Gabriel, et al., “Design and assessment of multiple-sound source localization using microphone arrays,” Proceedings of the 2019 IEEE/SICE Intl. Symposium on System Integration, Jan. 2019, 6 pp. [cited by applicant]
International Search Report and Written Opinion for PCT/US2022/076815 dated Jan. 5, 2023, 12 pp. [cited by applicant]
International Search Report and Written Opinion for PCT/US2023/026753 dated Sep. 15, 2023, 13 pp. [cited by applicant]
International Search Report and Written Opinion for PCT/US2023/030647 dated Nov. 9, 2023, 11 pp. [cited by applicant]
Ishi, et al., “Using multiple microphone arrays and reflections for 3D localization of sound sources,” 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, 6 pp. [cited by applicant]
Ishi, et al., “Speech activity detection and face orientation estimation using multiple microphone arrays and human position information,” IEEE/RSJ Intl. Conference on Intelligent Robots and Systems, 2015, 6 pp. [cited by applicant]
Lathoud, et al., “Short-Term Spatio-Temporal Clustering Applied to Multiple Moving Speakers,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 15, No. 5, Jul. 2007, 15 pp. [cited by applicant]
Legg, et al., “A Combined Microphone and Camera Calibration Technique with Application to Acoustic Imaging,” IEEE Transactions on Image Processing, vol. 22, No. 10, pp. 4028-4039, Oct. 1, 2013, 12 pp. [cited by applicant]
Nishiura, et al., “Collaborative Steering of Microphone Array and Video Camera Toward Multi-lingual Tele-conference Through Speech-to-Speech Translation,” IEEE Workshop on Automatic Speech Recognition and Understanding,… [cited by applicant]
Ronzhin, et al., “Audiovisual Speaker Localization in Medium Smart Meeting Room,” IEEE ICICS 8th International Conference, Dec. 2011, 5 pp. [cited by applicant]
Sharma, “The Ultimate Guide to K-Means Clustering: Definition, Methods and Applications,” Analytics Vidhya, Nov. 3, 2023, 19 pp. [cited by applicant]
Thiergart, et al., “Localization of Sound Sources in Reverberant Environments Based on Directional Audio Coding Parameters,” Audio Engineering Society Convention 127, New York, Oct. 2009, 14 pp. [cited by applicant]
Wang, et al., “Voice Source Localization for Automatic Camera Pointing System in Videoconferencing,” 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing, vol. 1, 4 pp. [cited by applicant]