IP Library › Granted Patent US 10,841,723
Granted Patent B2
US 10,841,723 · App. 16/025,986 · Granted Nov 17, 2020

Dynamic sweet spot calibration

Inventors: Pratyush Sahay (Bangalore, IN); Srinivas Kruthiventi Subrahmanyeswara Sai (Bangalore, IN); Arindam Dasgupta (Kolkata, IN); Pranjal Chakraborty (Bangalore, IN); Debojyoti Majumder (Bangalore, IN)
Assignee: Harman International Industries, Incorporated
H04S7/303G06K9/00369G06K9/00624G06K9/6262G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,841,723
App. No.
16/025,986
Granted
Nov 17, 2020
Kind
B2
Abstract

A technique for dynamic sweet spot calibration. The technique includes receiving an image of a listening environment, which may have been captured under poor lighting conditions, and generating a crowd-density map based on the image. The technique further includes setting at least one audio parameter associated with an audio system based on the crowd-density map. At least one audio output signal may be generated based on the at least one audio parameter.

Claims (53)

1. A method comprising:

receiving an image of a listening environment;

enhancing the image of the listening environment based on a plurality of training images of the listening environment that includes a plurality of different lighting conditions to generate an enhanced image;

generating a crowd-density map based on the enhanced image; and

setting at least one audio parameter associated with an audio system based on the crowd-density map, wherein at least one audio output signal is generated based on the at least one audio parameter,

wherein:

enhancing the image of the listening environment comprises enhancing the image of the listening environment via a convolutional neural network,

the convolutional neural network is trained with the plurality of training images, and

the plurality of training images comprises (i) a first training image of the listening environment illuminated by a first level of light, and (ii) a second training image of the listening environment illuminated by a second level of light greater than the first level of light.

2. The method of claim 1 , wherein setting the at least one audio parameter comprises:

determining a target location in the listening environment based on the crowd-density map;

determining values for the at least one audio parameter to configure the audio system to produce a sweet spot at the target location; and

setting the at least one audio parameter based, at least in part, on the values.

3. The method of claim 2 , wherein the target location corresponds to a centroid of the crowd-density map.

4. The method of claim 1 , wherein generating the crowd-density map further comprises:

detecting, via at least one machine learning algorithm, at least one physical feature of individual persons included in the image of the listening environment;

determining a crowd density based, at least in part, on the physical features of the individual persons; and

generating the crowd-density map using the crowd density.

5. The method of claim 1 , wherein the at least one audio parameter comprises at least one of phase or power of the at least one audio output signal.

6. The method of claim 1 , further comprising dynamically modifying the at least one audio parameter in response to a change in the crowd-density map.

7. The method of claim 2 , wherein determining the target location in the listening environment further comprises converting pixel locations in the crowd-density map to a real-world coordinate system based, at least in part, on physical dimensions of the listening environment.

8. A system comprising:

a memory storing an application; and

a processor that is coupled to the memory and, when executing the application, is configured to:

receive an image of a listening environment;

enhance the image of the listening environment based on a plurality of training images of the listening environment that includes a plurality of different lighting conditions to generate an enhanced image;

generate a crowd-density map based, at least in part, on the enhanced image; and

set at least one audio parameter associated with the system based, at least in part, on the crowd-density map,

wherein:

the processor is configured to enhance the image of the listening environment via a convolutional neural network,

the convolutional neural network is trained with the plurality of training images, and

the plurality of training images comprises (i) a first training image of the listening environment illuminated by a first level of light and (ii) a second training image of the listening environment illuminated by a second level of light greater than the first level of light.

9. The system of claim 8 , wherein the processor is further configured to:

determine a sweet spot location in the listening environment based, at least in part, on the crowd-density map;

determine at least one value for the at least one audio parameter to configure the system to produce a sweet spot at the sweet spot location; and

set the at least one audio parameter based, at least in part, on the at least one value.

10. The system of claim 8 , wherein the processor is further configured to dynamically modify in real-time the at least one audio parameter in response to a change in the crowd-density map.

11. The system of claim 8 , wherein the crowd-density map comprises a heat map.

12. A non-transitory computer-readable storage medium including instructions that, when executed by a processor, causes the processor to perform the steps of:

receiving an image of a listening environment;

enhancing the image of the listening environment based on a plurality of training images of the listening environment that includes a plurality of different lighting conditions to generate an enhanced image;

generating a crowd-density map based, at least in part, on the enhanced image;

determining a location of a centroid or other substantially central distribution of the crowd-density map, wherein the location is with respect to the listening environment; and

determining at least one audio parameter based, at least in part, on the location;

wherein:

enhancing the image of the listening environment comprises enhancing the image of the listening environment via a convolutional neural network,

the convolutional neural network is trained with the plurality of training images, and

the plurality of training images comprises (i) a first training image of the listening environment illuminated by a first level of light and (ii) a second training image of the listening environment illuminated by a second level of light greater than the first level of light.

13. The non-transitory computer-readable storage medium of claim 12 , wherein the instructions, when executed by the processor, further cause the processor to perform the step of transmitting the at least one audio parameter to an audio system.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the instructions, when executed by the processor, further cause the processor to perform the step of determining at least one value for the at least one audio parameter to configure the audio system to produce a sweet spot at the location.

15. The non-transitory computer-readable storage medium of claim 12 , wherein the processor is further configured to dynamically modify in real-time the at least one audio parameter in response to a change in the crowd-density map.

16. The non-transitory computer-readable storage medium of claim 12 , wherein the crowd-density map is generated via a neural network.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the neural network is trained based on a plurality of training images that includes a plurality of different listening environments, a plurality of different crowd densities, and a plurality of different lighting conditions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2019
From: SAHAY, PRATYUSH; SAI, SRINIVAS KRUTHIVENTI SUBRAHMANYESWARA; DASGUPTA, ARINDAM; CHAKRABORTY, PRANJAL; MAJUMDER, DEBOJYOTI
To: HARMAN INTERNATIONAL INDUSTRIES, INCORPORATED
Reel/Frame 048350/0188 →
Continuity (1)
Related Publication 20200008002A1 · Jan 2, 2020
Cited By (1)
US 12,401,962