Object Based Audio Rendering Using Visual Tracking of at Least One Listener
In some embodiments, a method and system for generating image data, processing the image data to generate listener data indicative of at least one listener characteristic (e.g., position and/or size of each listener), and rendering at least one audio object (e.g., rendering an object based audio program) in response to the listener data (and optionally also listener identification data). For rendering a program indicative of audio objects, at least one speaker feed may be generated for driving at least one speaker to emit sound indicative of one of the objects and additional sound indicative of another one of the objects, where the sound is intended to be perceived by a listener at a first position with balance and delay appropriate to the first position, and the additional sound is intended to be perceived by a listener at a second position with balance and delay appropriate to the second position.
1 . A method for rendering an audio program comprising one or more audio objects for playback in an environment including a speaker array comprising at least one speaker, said method including the steps of:
(a) generating image data indicative of at least one listener in the environment;
(b) processing the image data to generate listener data indicative of at least one characteristic of at least one said listener; and
(c) rendering at least one of the audio objects in response to the listener data.
2 . The method of claim 1 , wherein the listener data is indicative of position of at least one said listener, and step (c) includes a step of generating at least one speaker feed for driving at least one speaker of the array to emit sound intended to be perceived by one said listener with balance and delay appropriate to the position of said listener.
3 . The method of claim 1 , wherein the listener data is indicative of a first position of a first listener and a second position of a second listener, the audio program comprises at least two audio objects, and step (c) includes a step of generating at least one speaker feed for driving at least one speaker of the array to emit first sound indicative of one of the audio objects and additional sound indicative of another one of the audio objects, wherein the first sound which is intended to be perceived by the first listener at the first position with balance and delay appropriate to a listener at said first position, and the additional sound is intended to be perceived by the second listener at the second position with balance and delay appropriate to a listener at said second position.
4 . The method of claim 3 , wherein the audio program is an object based audio program indicative of the at least two audio objects.
5 . The method of claim 1 , wherein the listener data is indicative of position and size of at least one said listener.
6 . The method of claim 1 , wherein the audio program includes metadata, and step (c) includes a step of rendering the audio program in response to the listener data and the metadata.
7 . The method of claim 1 , wherein step (c) includes a step of rendering the audio program in response to the listener data and in response to listener identification data.
8 . The method of claim 7 , wherein the listener identification data is indicative of hearing capability of at least one said listener.
9 . The method of claim 8 , wherein the listener data is indicative of position and size of at least one said listener, and step (c) includes a step of determining from the listener identification data and the listener data that one said listener whose size is indicated by the listener data has a hearing capability which is indicated by the listener identification data.
10 . A system for rendering an audio program comprising one or more audio objects for playback in an environment including a speaker array comprising at least one speaker, said system including:
a camera subsystem, including at least one camera, wherein the camera subsystem is configured to generate image data indicative of at least one listener in a field of view of at least one camera of the camera subsystem;
a visual tracking subsystem coupled and configured to process the image data to generate listener data indicative of at least one listener characteristic; and
a rendering subsystem coupled and configured to render at least one of the audio objects in response to the listener data.
11 . The system of claim 10 , wherein the listener data is indicative of position of at least one said listener, and the rendering subsystem is configured to generate at least one speaker feed for driving at least one speaker of the array to emit sound intended to be perceived by one said listener with balance and delay appropriate to the position of said listener.
12 . The system of claim 10 , wherein the listener data is indicative of a first position of a first listener and a second position of a second listener, the audio program comprises at least two audio objects, and the rendering subsystem is configured to generate at least one speaker feed for driving at least one speaker of the array to emit first sound indicative of one of the audio objects and additional sound indicative of another one of the audio objects, wherein the first sound which is intended to be perceived by the first listener at the first position with balance and delay appropriate to a listener at said first position, and the additional sound is intended to be perceived by the second listener at the second position with balance and delay appropriate to a listener at said second position.
13 . The system of claim 12 , wherein the audio program is an object based audio program indicative of the at least two audio objects.
14 . The system of claim 10 , wherein the listener data is indicative of position and size of at least one said listener.
15 . The system of claim 10 , wherein the audio program includes metadata, and the rendering subsystem is configured to render the audio program in response to the listener data and the metadata.
16 . The system of claim 10 , wherein the rendering subsystem is configured to render the audio program in response to the listener data and in response to listener identification data.
17 . The system of claim 16 , wherein the listener identification data is indicative of hearing capability of at least one said listener.
18 . The system of claim 17 , wherein the listener data is indicative of position and size of at least one said listener, and the rendering subsystem is configured to determine from the listener identification data and the listener data that one said listener whose size is indicated by the listener data has a hearing capability which is indicated by the listener identification data.
19 . The system of claim 10 , including a processor coupled to the camera subsystem, wherein the processor is configured to implement both the visual tracking subsystem and the rendering subsystem.
20 . A non-transitory computer readable storage medium that is readable by a device and that records a program of instructions executable by the device to perform a method for rendering an audio program comprising one or more audio objects for playback in an environment including a speaker array comprising at least one speaker, said method including the steps of:
(a) generating image data indicative of at least one listener in the environment;
(b) processing the image data to generate listener data indicative of at least one characteristic of at least one said listener; and
(c) rendering at least one of the audio objects in response to the listener data.