IP Library › Granted Patent US 12,335,718
Granted Patent B2
US 12,335,718 · App. 18/634,825 · Granted Jun 17, 2025

System and method for adaptive audio signal generation, coding and rendering

Inventors: Charles Q. Robinson (Piedmont, CA); Nicolas R. Tsingos (San Francisco, CA); Christophe Chabanne (Carpentras, FR)
Assignee: Dolby Laboratories Licensing Corporation
H04S7/308G10L19/008G10L19/20H04R5/02H04R5/04H04R27/00H04S3/008H04S5/005H04S7/30H04S7/305H04S5/00H04S7/302H04S2400/01H04S2400/03H04S2400/11H04S2420/01H04S2420/03H04S2420/11H04S2420/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,335,718
App. No.
18/634,825
Granted
Jun 17, 2025
Kind
B2
Abstract

Embodiments are described for an adaptive audio system that processes audio data comprising a number of independent monophonic audio streams. One or more of the streams has associated with it metadata that specifies whether the stream is a channel-based or object-based stream. Channel-based streams have rendering information encoded by means of channel name; and the object-based streams have location information encoded through location expressions encoded in the associated metadata. A codec packages the independent audio streams into a single serial bitstream that contains all of the audio data. This configuration allows for the sound to be rendered according to an allocentric frame of reference, in which the rendering location of a sound is based on the characteristics of the playback environment (e.g., room size, shape, etc.) to correspond to the mixer's intent. The object position metadata contains the appropriate allocentric frame of reference information required to play the sound correctly using the available speaker positions in a room that is set up to play the adaptive audio content.

Claims (20)

1. A system for processing audio signals, comprising a rendering system configured to:

receive a bitstream comprising encoded audio data representing a plurality of monophonic audio streams, and further comprising metadata associated with each of the monophonic audio streams and indicating a playback location of a respective monophonic audio stream, wherein at least some of the plurality of monophonic audio streams are identified as object-based audio, and wherein the playback location of an object-based monophonic audio stream comprises a location in a three-dimensional space, wherein at least some other of the plurality of monophonic audio streams are identified as channel-based audio, and wherein the playback location of a channel-based monophonic audio stream comprises a location in the three-dimensional space;

decode the encoded audio data to provide the plurality of monophonic audio streams; and

render the plurality of monophonic audio streams to a plurality of speaker feeds corresponding to speakers in a playback environment, wherein the speakers are placed at specific positions within the playback environment, and wherein one or more additional metadata elements associated with each respective object-based monophonic audio stream indicate whether rendering the respective monophonic audio stream into one or more specific speaker feeds of the plurality of speaker feeds is prohibited, such that the respective object-based monophonic audio stream is not rendered into any of the one or more specific speaker feeds of the plurality of speaker feeds.

2. The system of claim 1 , wherein the metadata elements associated with each object-based monophonic audio stream further indicate spatial parameters controlling the playback of a corresponding sound component comprising one or more of: sound position, sound width, and sound velocity.

3. The system of claim 1 , wherein the playback location for each of the plurality of object-based monophonic audio streams is independently specified with respect to either an egocentric frame of reference or an allocentric frame of reference, wherein the egocentric frame of reference is taken in relation to a listener in the playback environment, and wherein the allocentric frame of reference is taken with respect to a characteristic of the playback environment.

4. A method for authoring audio content for rendering, comprising:

receiving a plurality of audio signals;

generating a plurality of monophonic audio streams and metadata associated with each of the audio streams and indicating a playback location of a respective monophonic audio stream, wherein at least some of the plurality of monophonic audio streams are identified as object-based audio, and wherein the playback location of the object-based audio comprises a location in a three-dimensional space, wherein at least some other of the plurality of monophonic audio streams are identified as channel-based audio, and wherein the playback location of a channel-based monophonic audio stream comprises a location in the three-dimensional space;

encoding the plurality of monophonic audio streams to provide encoded audio data; and

encapsulating the encoded audio data and the metadata in a bitstream for transmission to a rendering system configured to render the plurality of monophonic audio streams to a plurality of speaker feeds corresponding to speakers in a playback environment, wherein the speakers are placed at specific positions within the playback environment, and wherein one or more additional metadata elements associated with each respective object-based monophonic audio stream indicate whether rendering the respective monophonic audio stream into one or more specific speaker feeds of the plurality of speaker feeds is prohibited, such that the respective object-based monophonic audio stream is not rendered into any of the one or more specific speaker feeds of the plurality of speaker feeds.

5. A system for authoring audio content for rendering configured to perform the method of claim 4 .

6. A non-transitory computer readable storage medium comprising a sequence of instructions, wherein, when executed by a system for processing audio signals, the sequence of instructions causes the system to perform the method of claim 4 .

7. A method for rendering audio signals, comprising:

receiving a bitstream comprising encoded audio data representing a plurality of monophonic audio streams, and further comprising metadata associated with each of the audio streams and indicating a playback location of a respective monophonic audio stream, wherein at least some of the plurality of monophonic audio streams are identified as object-based audio, and wherein the playback location of an object-based monophonic audio stream comprises a location in three-dimensional space, wherein at least some other of the plurality of monophonic audio streams are identified as channel-based audio, and wherein the playback location of a channel-based monophonic audio stream comprises a location in the three-dimensional space;

decoding the encoded audio data to provide the plurality of monophonic audio bitstreams; and

rendering the plurality of monophonic audio streams to a plurality of speaker feeds corresponding to speakers in a playback environment, wherein the speakers are placed at specific positions within the playback environment, and wherein one or more additional metadata elements associated with each respective object-based monophonic audio stream indicate whether rendering the respective monophonic audio stream into one or more specific speaker feeds of the plurality of speaker feeds is prohibited, such that the respective object-based monophonic audio stream is not rendered into any of the one or more specific speaker feeds of the plurality of speaker feeds.

8. The method of claim 7 , wherein the metadata elements associated with each object-based monophonic audio stream further indicate spatial parameters controlling the playback of a corresponding sound component comprising one or more of: sound position, sound width, and sound velocity.

9. The method of claim 7 , wherein the playback location for each of the plurality of object-based monophonic audio streams comprises a spatial position relative to a screen within a playback environment, or a surface that encloses the playback environment, and wherein the surface comprises a front plane, a back plane, a left plane, right plane, an upper plane, and a lower plane, and/or is independently specified with respect to either an egocentric frame of reference or an allocentric frame of reference, wherein the egocentric frame of reference is taken in relation to a listener in the playback environment, and wherein the allocentric frame of reference is taken with respect to a characteristic of the playback environment.

10. A non-transitory computer readable storage medium comprising a sequence of instructions, wherein, when executed by a system for processing audio signals, the sequence of instructions causes the system to perform the method of claim 7 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: ROBINSON, CHARLES; TSINGOS, NICOLAS; CHABANNE, CHRISTOPHE
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 068239/0728 →
Continuity (15)
Continuation 17883440 · Aug 8, 2022
Continuation 17156459 · Jan 22, 2021
Continuation 16679945 · Nov 11, 2019
Continuation 16443268 · Jun 17, 2019
Continuation 16207006 · Nov 30, 2018
Continuation 16035262 · Jul 13, 2018
Continuation 15905536 · Feb 26, 2018
Continuation 15672656 · Aug 9, 2017
Continuation 15483806 · Apr 10, 2017
Continuation 15263279 · Sep 12, 2016
Continuation 14866350 · Sep 25, 2015
Continuation 14130386
Provisional Application 61636429 · Apr 20, 2012
Provisional Application 61504005 · Jul 1, 2011
Related Publication 20240349010A1 · Oct 17, 2024
References Cited (101)
US 5155510A · Beard · 1992 [cited by applicant]
US 5602923A · Ozaki · 1997 [cited by applicant]
US 5642423A · Embree · 1997 [cited by applicant]
US 5970152A · Klayman · 1999 [cited by applicant]
US 6164018A · Runge · 2000 [cited by applicant]
US 6229899B1 · Norris · 2001 [cited by applicant]
US 6624873B1 · Raymond · 2003 [cited by applicant]
US 6646800B2 · Choi · 2003 [cited by applicant]
US 6931370B1 · McDowell · 2005 [cited by applicant]
US 7106411B2 · Read · 2006 [cited by applicant]
US 7212872B1 · Smith et al. · 2007 [cited by applicant]
US 7333154B2 · Dean · 2008 [cited by applicant]
US 7782439B2 · Bogdanowicz · 2010 [cited by applicant]
US 7788395B2 · Bowra · 2010 [cited by applicant]
US 7796190B2 · Basso · 2010 [cited by applicant]
US 7911580B2 · Read · 2011 [cited by applicant]
US 8798776B2 · Schildbach · 2014 [cited by applicant]
US 8914137B2 · Crockett · 2014 [cited by applicant]
US 10085104B2 · Ertel · 2018 [cited by applicant]
US 20010055398A1 · Pachet et al. · 2001 [cited by applicant]
US 20030223603A1 · Beckman · 2003 [cited by applicant]
US 20050075882A1 · Fay · 2005 [cited by applicant]
US 20050222841A1 · McDowell · 2005 [cited by applicant]
US 20060206221A1 · Metcalf · 2006 [cited by applicant]
US 20060256985A1 · Pierre · 2006 [cited by applicant]
US 20070025559A1 · Mihelich · 2007 [cited by applicant]
US 20070101249A1 · Lee · 2007 [cited by applicant]
US 20070276656A1 · Solbach · 2007 [cited by applicant]
US 20080019534A1 · Reichelt et al. · 2008 [cited by applicant]
US 20080126461A1 · Christoph · 2008 [cited by applicant]
US 20090220107A1 · Every · 2009 [cited by applicant]
US 20100023544A1 · Shahraray · 2010 [cited by applicant]
US 20100050225A1 · Bennett · 2010 [cited by applicant]
US 20100135510A1 · Yoo · 2010 [cited by applicant]
US 20100198378A1 · Smithers · 2010 [cited by applicant]
US 20110004897A1 · Alexander · 2011 [cited by applicant]
US 20110013790A1 · Hilpert · 2011 [cited by applicant]
US 20110040395A1 · Kraemer · 2011 [cited by applicant]
US 20110040396A1 · Kraemer · 2011 [cited by applicant]
US 20110040397A1 · Kraemer · 2011 [cited by applicant]
US 20110072086A1 · Newsome · 2011 [cited by applicant]
US 20110088076A1 · Li · 2011 [cited by applicant]
US 20120078402A1 · Brown · 2012 [cited by applicant]
US 20130163794A1 · Groves · 2013 [cited by applicant]
CN 101001485 · 2007 [cited by applicant]
CN 101133454 · 2008 [cited by applicant]
EP 1584217 · 2005 [cited by applicant]
EP 1600035 · 2005 [cited by applicant]
EP 1695338 · 2006 [cited by applicant]
EP 1843635 · 2007 [cited by applicant]
JP 0951600 · 1997 [cited by applicant]
JP 2003531555 · 2003 [cited by applicant]
JP 2003348700A · 2003 [cited by applicant]
JP 2006507727 · 2006 [cited by applicant]
JP 2006304165 · 2006 [cited by applicant]
JP 2008532374 · 2008 [cited by applicant]
JP 2008537833 · 2008 [cited by applicant]
JP 2009501463 · 2009 [cited by applicant]
JP 2009278381A · 2009 [cited by applicant]
JP 2010505328A · 2010 [cited by applicant]
JP 2010507114 · 2010 [cited by applicant]
JP 2010511912 · 2010 [cited by applicant]
JP 2010521013 · 2010 [cited by applicant]
JP 2010525378 · 2010 [cited by applicant]
JP 2010154548A · 2010 [cited by applicant]
JP 2010529500 · 2010 [cited by applicant]
JP 2017215592 · 2017 [cited by applicant]
JP 6486995B2 · 2019 [cited by applicant]
JP 7348320B2 · 2022 [cited by applicant]
KR 1020100116223A · 2010 [cited by applicant]
RS 1332U · 2013 [cited by applicant]
RU 2347282 · 2009 [cited by applicant]
RU 2376654C2 · 2009 [cited by applicant]
RU 2008148961 · 2010 [cited by applicant]
TW 200921644A · 2009 [cited by applicant]
TW 201033938 · 2010 [cited by applicant]
WO 2009115299 · 2009 [cited by applicant]
WO 2010006719 · 2010 [cited by applicant]
WO 2010058518A1 · 2010 [cited by applicant]
WO 2011045813 · 2011 [cited by applicant]
WO 2011061174 · 2011 [cited by applicant]
WO 2011068490 · 2011 [cited by applicant]
Arumi, P. et al., “Remastering of Movie Soundtracks into Immersive 3D Audio”, Amsterdam Blender Conference, slides 1-25, Oct. 24, 2009. [cited by applicant]
Baelen, W. et al., “Auro-3D A New Dimension in Cinema Sound”, pp. 1-11, May 26, 2011. [cited by applicant]
Brandenburg, K. et al., “Wave Field Synthesis: New Possibilities for Large-Scale Immersive Sound Reinforcement”, Fraunhofer IDMT & Ilmenau Tech University, Mo5 E. 1, pp I-507-1-508, Apr. 2004. [cited by applicant]
Delancie P., “Dolby DP564”, <http://mixonline.com/mag/audio_dolby_dp_2/index.html>, Sep. 1, 2002. [cited by applicant]
Kim, S. et al., “A Novel Test-Bed for Immersive and Interactive Broadcasting Production Using Augmented Reality and Haptics”, vol. E89-D., No. 1, pp. 106-110, Jan. 2006. [cited by applicant]
Kyriakakis, C., “Fundamental and Technological Limitations of Immersive Audio Systems”, Proceedings of the IEEE 1998, vol. 86, No. 5, pp. 941-951, May 1998. [cited by applicant]
Neukom, Martin “Decoding Second Order Ambisonics to 5.1 Surround Systems” AES Convention, 121, Signal Processing, Oct. 1, 2006. [cited by applicant]
Stanojevic, T. “Some Technical Possibilities of Using the Total Surround Sound Concept in the Motion Picture Technology”, 133rd SMPTE Technical Conference and Equipment Exhibit, Los Angeles Convention Center, Los Angele… [cited by applicant]
Stanojevic, T. et al.“Designing of TSS Halls” 13th International Congress on Acoustics, Yugoslavia, 1989, pp. 327-330, 4 pages. [cited by applicant]
Stanojevic, T. et al.“The Total Surround Sound (TSS) Processor” SMPTE Journal, Nov. 1994, pp. 734-740, 8 pages. [cited by applicant]
Stanojevic, T. et al.“The Total Surround Sound System”, 86th AES Convention, Hamburg, Mar. 7-10, 1989, pp. 0-20, 21 pages. [cited by applicant]
Stanojevic, T. et al.“TSS System and Live Performance Sound” 88th AES Convention, Montreux, Mar. 13-16, 1990, pp. 1-27, 27 pages. [cited by applicant]
Stanojevic, T. et al. “TSS Processor” 135th SMPTE Technical Conference, Oct. 29-Nov. 2, 1993, Los Angeles Convention Center, Los Angeles, California, Society of Motion Picture and Television Engineers. [cited by applicant]
Stanojevic, Tomislav “3-D Sound in Future HDTV Projection Systems” presented at the 132nd SMPTE Technical Conference, Jacob K. Javits Convention Center, New York City, Oct. 13-17, 1990. [cited by applicant]
Stanojevic, Tomislav “Surround Sound for a New Generation of Theaters, Sound and Video Contractor” Dec. 20, 1995. [cited by applicant]
Stanojevic, Tomislav, “Virtual Sound Sources in the Total Surround Sound System” Proc. 137th SMPTE Technical Conference and World Media Expo, Sep. 6-9, 1995, New Orleans Convention Center, New Orleans, Louisiana. [cited by applicant]
Theile, G., “Wave Field Synthesis - A Promising Spatial Audio Rendering Concept”, Proc. of 7th Int. Conference on Digital Audio Effects, pp. 125-132, Oct. 5-8, 2004. [cited by applicant]
Van Beek, P. et al., “Metadata-Driven Multimedia Access”, IEEE Signal Processing Magazine, pp. 40-52, Mar. 2003. [cited by applicant]
Anonymous, Multichannel sound technology in home and broadcasting applications, Report ITU-R BS.2159-6 ; ITU-R, BS Series Broadcasting service (sound), Nov. 1, 2013, 84 pages. [cited by applicant]