IP Library › Granted Patent US 12,591,350
Granted Patent B2
US 12,591,350 · App. 17/307,407 · Granted Mar 31, 2026

Techniques for positioning speakers within a venue

Inventors: Sambuddha Saha (Burdwan, IN); Vipul Sharma (Bengaluru, IN); George Georgallis (Los Angeles, CA)
Assignee: Harman International Industries, Incorporated
G06F3/04815G06F3/162G06T17/00H04R5/02G06T2200/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,350
App. No.
17/307,407
Granted
Mar 31, 2026
Kind
B2
Abstract

Various embodiments set forth techniques for positioning speakers within a venue. The techniques include generating, via a machine learning model, at least one of a two-dimensional (2D) representation or a three-dimensional (3D) representation of a venue based on one or more images of the venue. The techniques further include determining one or more parameters associated with one or more speakers to be placed within the venue based on the at least one of the 2D representation or the 3D representation.

Claims (34)

1 . A computer-implemented method for positioning one or more physical speakers within a venue, the method comprising:

generating, via a machine learning model, at least one of a two-dimensional (2D) representation or a three-dimensional (3D) representation of a venue based on one or more images of the venue captured from a position on a stage in the venue showing one or more audience seating areas of the venue, the 2D representation or the 3D representation comprising a plurality of lines associated with a plurality of slopes, the plurality of lines representing the one or more audience seating areas of the venue;

determining, via at least one computing device, a head line based on the plurality of lines, the plurality of slopes, and an average height parameter associated with members of an audience; and

determining, via one or more models, one or more parameters for placing one or more physical speakers within the venue based on the head line and at least one of the 2D representation or the 3D representation.

2 . The computer-implemented method of claim 1 , wherein determining the one or more parameters comprises processing the at least one of the 2D representation or the 3D representation and one or more additional inputs via a plurality of models that optimize different parameters.

3 . The computer-implemented method of claim 2 , wherein each model included in the plurality of models comprises one of a regression model or a deep learning model.

4 . The computer-implemented method of claim 1 , wherein determining the one or more parameters comprises optimizing the one or more parameters with respect to sound pressure levels over a range of frequencies and a center of gravity associated with a line array that includes the one or more physical speakers.

5 . The computer-implemented method of claim 2 , wherein the one or more additional inputs comprise at least one of one or more measurements of the venue, a blueprint of the venue, a type of the one or more physical speakers, whether the one or more physical speakers are grounded or suspended, a temperature, a humidity, a cable weight, a top frame type, a suspension mode, suspension points, whether an extension bar or pull back frame is used, or a budget.

6 . The computer-implemented method of claim 1 , wherein the one or more physical speakers are included in at least one of a suspended line array or a line array placed on a platform.

7 . The computer-implemented method of claim 1 , wherein the one or more images are images of the venue from a position on a stage of the venue.

8 . The computer-implemented method of claim 1 , wherein the one or more parameters include at least one of a number of the one or more physical speakers, a position of a line array that includes the one or more physical speakers, a bottom frame angle or a base plate angle of a line array that includes the one or more physical speakers, a curvature of each physical speaker included in the one or more physical speakers in the line array, attachment points, a pull back frame, an extension bar location, an extension bar position, or an estimated cost.

9 . The computer-implemented method of claim 1 , further comprising modifying at least one of the one or more parameters based on user input.

10 . One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processing units, cause the one or more processing units to position one or more physical speakers within a venue, by performing the steps of:

generating, via a machine learning model, at least one of a two-dimensional (2D) representation or a three-dimensional (3D) representation of a venue based on one or more images of the venue captured from a position on a stage in the venue showing one or more audience seating areas of the venue, the 2D representation or the 3D representation comprising a plurality of lines associated with a plurality of slopes, the plurality of lines representing the one or more audience seating areas of the venue;

determining, via at least one computing device, a head line based on the plurality of lines, the plurality of slopes, and an average height parameter associated with members of an audience; and

determining, via one or more models, one or more parameters for placing one or more physical speakers within the venue based on the head line and at least one of the 2D representation or the 3D representation.

11 . The one or more non-transitory computer-readable storage media of claim 10 , wherein determining the one or more parameters comprises processing the at least one of the 2D representation or the 3D representation and one or more additional inputs via a plurality of models that optimize different parameters.

12 . The one or more non-transitory computer-readable storage media of claim 11 , wherein each model included in the plurality of models comprises one of a regression model or a deep learning model.

13 . The one or more non-transitory computer-readable storage media of claim 10 , wherein determining the one or more parameters comprises optimizing the one or more parameters with respect to sound pressure levels over a range of frequencies and a center of gravity associated with a line array that includes the one or more physical speakers.

14 . The one or more non-transitory computer-readable storage media of claim 10 , wherein determining the one or more parameters comprises optimizing at least a number of the one or more physical speakers and a position of the one or more physical speakers prior to optimizing at least one other parameter associated with the one or more physical speakers.

15 . The one or more non-transitory computer-readable storage media of claim 14 , wherein the at least the number of the one or more physical speakers is optimized via a first model, the at least one other parameter is optimized via a second model, and at least one output of the first model is input into the second model.

16 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the one or more images comprises a wide-angle image of the venue or a pair of stereo images.

17 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the one or more parameters include at least one of a number of the one or more physical speakers, a position of a line array that includes the one or more physical speakers, a bottom frame angle or a base plate angle of a line array that includes the one or more physical speakers, a curvature of each physical speaker included in the one or more physical speakers in the line array, attachment points, a pull back frame, an extension bar location, an extension bar position, or an estimated cost.

18 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the at least one of the 2D representation or the 3D representation indicates one or more slopes associated with one or more seating areas in the venue.

19 . A system, comprising:

one or more memories that include instructions; and

one or more processors that are coupled to the one or more memories and, when executing the instructions:

generate, via a machine learning model, at least one of a two-dimensional (2D) representation or a three-dimensional (3D) representation of a venue based on one or more images of the venue captured from a position on a stage in the venue showing one or more audience seating areas of the venue, the 2D representation or the 3D representation comprising a plurality of lines associated with a plurality of slopes, the plurality of lines representing the one or more audience seating areas of the venue;

determine, via at least one computing device, a head line based on the plurality of lines, the plurality of slopes, and an average height parameter associated with members of an audience; and

determine, via one or more models, one or more parameters for placing one or more physical speakers within the venue based on the head line and at least one of the 2D representation or the 3D representation.

20 . The system of claim 19 , wherein determining the one or more parameters comprises processing the at least one of the 2D representation or the 3D representation and one or more additional inputs via a plurality of models that optimize different parameters.

21 . The computer-implemented method of claim 1 , wherein determining the one or more parameters comprises optimizing the one or more parameters with respect the average height parameter associated for sound pressure levels over a range of frequencies.

22 . The computer-implemented method of claim 1 , further comprising adjusting a center of gravity associated with a line array based on the average height parameter.

23 . The computer-implemented method of claim 1 , wherein generating the head line further comprises adding the average height parameter to each slope from the plurality of slopes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 4, 2021
From: SAHA, SAMBUDDHA; SHARMA, VIPUL; GEORGALLIS, GEORGE
To: HARMAN INTERNATIONAL INDUSTRIES, INCORPORATED
Reel/Frame 056135/0911 →
Continuity (1)
Related Publication 20220357834A1 · Nov 10, 2022
References Cited (54)
US 9239992B2 · Valentino · 2016 [cited by examiner]
US 9843881B1 · Oates, III · 2017 [cited by examiner]
US 10223740B1 · Gore · 2019 [cited by examiner]
US 10645275B1 · Chuah · 2020 [cited by examiner]
US 10827322B1 · Lal · 2020 [cited by examiner]
US 11265668B1 · Patil · 2022 [cited by examiner]
US 20030123692A1 · Ueki · 2003 [cited by examiner]
US 20090003613A1 · Christensen · 2009 [cited by examiner]
US 20090316938A1 · Matsumoto · 2009 [cited by examiner]
US 20120027226A1 · Desenberg · 2012 [cited by examiner]
US 20120051568A1 · Kim · 2012 [cited by examiner]
US 20120117502A1 · Nguyen · 2012 [cited by examiner]
US 20130039515A1 · Jin · 2013 [cited by examiner]
US 20130222393A1 · Merrell · 2013 [cited by examiner]
US 20140133683A1 · Robinson · 2014 [cited by examiner]
US 20140168477A1 · David · 2014 [cited by examiner]
US 20140314256A1 · Fincham · 2014 [cited by examiner]
US 20140329567A1 · Chan · 2014 [cited by examiner]
US 20160026737A1 · Feistel · 2016 [cited by examiner]
US 20160247364A1 · Herman · 2016 [cited by examiner]
US 20160301373A1 · Herman · 2016 [cited by examiner]
US 20170073988A1 · Sallent Puigcercos · 2017 [cited by examiner]
US 20170099557A1 · Saunders · 2017 [cited by examiner]
US 20170182406A1 · Castiglia · 2017 [cited by examiner]
US 20180115825A1 · Milne · 2018 [cited by examiner]
US 20180137215A1 · Lee · 2018 [cited by examiner]
US 20180197551A1 · Mcdowell · 2018 [cited by examiner]
US 20180374276A1 · Powers · 2018 [cited by examiner]
US 20190164340A1 · Pejic · 2019 [cited by examiner]
US 20190208314A1 · Sakagushi · 2019 [cited by examiner]
US 20190385373A1 · Mittleman · 2019 [cited by examiner]
US 20200281521A1 · Cail · 2020 [cited by examiner]
US 20200302681A1 · Totty · 2020 [cited by examiner]
US 20210019453A1 · Yang · 2021 [cited by examiner]
US 20210073449A1 · Segev · 2021 [cited by examiner]
US 20210250669A1 · Andrews · 2021 [cited by examiner]
US 20210334535A1 · Perciful · 2021 [cited by examiner]
US 20210398353A1 · Luo · 2021 [cited by examiner]
US 20220108383A1 · Kunikyo · 2022 [cited by examiner]
US 20220129974A1 · Delgado · 2022 [cited by examiner]
US 20220147563A1 · Perumalla · 2022 [cited by examiner]
US 20220224833A1 · Cier · 2022 [cited by examiner]
US 20220269888A1 · Stoeva · 2022 [cited by examiner]
US 20220327608A1 · Assouline · 2022 [cited by examiner]
US 20230154395A1 · Mcnelley · 2023 [cited by examiner]
US 20230308610A1 · Henderson · 2023 [cited by examiner]
“JBL Line Array Calculator 3.2.0 Now Available”, https://www.jpro.co.nz/news-article/jbl-line-array-calculator-320-now-available-1020, 5 pages. [cited by applicant]
Godard, Clément, Oisin Mac Aodha, and Gabriel J. Brostow. “Unsupervised monocular depth estimation with left-right consistency.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 270-… [cited by applicant]
Roy, Anirban, and Sinisa Todorovic. “Monocular depth estimation using neural regression forest.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5506-5514. 2016. [cited by applicant]
Fu, Huan, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. “Deep ordinal regression network for monocular depth estimation.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Godard, Clément, Oisin Mac Aodha, Michael Firman, and Gabriel J. Brostow. “Digging into self-supervised monocular depth estimation.” In Proceedings of the IEEE International Conference on Computer Vision, pp. 3828-3838.… [cited by applicant]
Shah, Shishir, and J. K. Aggarwal. “Depth estimation using stereo fish-eye lenses.” In Proceedings of 1st International Conference on Image Processing, vol. 2, pp. 740-744. IEEE, 1994. [cited by applicant]
Geiger, Andreas, Julius Ziegler, and Christoph Stiller. “Stereoscan: Dense 3d reconstruction in real-time.” In 2011 IEEE Intelligent Vehicles Symposium (IV), pp. 963-968. Ieee, 2011. [cited by applicant]
Liu, Chen, Kihwan Kim, Jinwei Gu, Yasutaka Furukawa, and Jan Kautz. “Planercnn: 3d plane detection and reconstruction from a single image.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognitio… [cited by applicant]