IP Library › Granted Patent US 12,621,621
Granted Patent B2
US 12,621,621 · App. 18/535,192 · Granted May 5, 2026

Adaptive panner of audio objects

Inventors: Jun Wang (Beijing, CN); Giulio Cengarle (Barcelona, ES); Juan Felix Torres (Darlinghurst, AU); Daniel Arteaga (Barcelona, ES)
Assignees: DOLBY LABORATORIES LICENSING CORPORATION; DOLBY INTERNATIONAL AB
H04S3/002H04S7/30H04S7/302H04S7/308H04R5/02H04R5/04H04S2400/11H04S2400/13H04S2420/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,621
App. No.
18/535,192
Granted
May 5, 2026
Kind
B2
Abstract

An audio object including audio content and object metadata is received. The object metadata indicates an object spatial position of the audio object to be rendered by audio speakers in a playback environment. Based on the object spatial position and source spatial positions of the audio speakers, initial gain values for the audio speakers are determined. The initial gain values can be used to select a set of audio speakers from among the audio speakers. Based on the object spatial position and a set of source spatial positions at which the set of audio speakers are respectively located in the playback environment, a set of non-negative optimized gain values for the set of audio speakers is determined. The audio object at the object spatial position is rendered with the set of optimized gain values for the set of audio speakers.

Claims (20)

1 . A computer-implemented method, comprising:

receiving an audio object comprising audio content and object metadata, the object metadata of the audio object indicating an object spatial position of the audio object to be rendered by a plurality of audio speakers, each audio speaker in the plurality of audio speakers being located in a respective source spatial position in a plurality of source spatial positions, wherein the object spatial position is related to audio content in one or more audio frames, or one or more subdivisions of an audio frame;

determining, based on the object spatial position of the audio object and the plurality of source spatial positions of the plurality of audio speakers, a plurality of initial gain values for the plurality of audio speakers, each audio speaker in the plurality of audio speakers being assigned with a respective initial gain value in the plurality of initial gain values;

determining, for each of the plurality of audio speakers, whether a respective audio speaker is an active audio speaker or whether the respective audio speaker is not an active audio speaker, wherein an audio speaker is an active audio speaker if the initial gain value assigned to the audio speaker is above a threshold value, and wherein the audio speaker is not an active audio speaker if the initial gain value assigned to the audio speaker is below or at a threshold value;

determining, based on the object spatial position of the audio object and a set of source spatial positions at which the set of active audio speakers are respectively located, a set of optimized non-negative gain values for the set of active audio speakers, wherein the set of optimized gain values are yielded using the initial gain values of the active speakers as input; and

outputting, for each audio speaker in the set of active audio speakers a respective optimized gain value in the plurality of optimized gain values.

2 . A system comprising:

one or more processors; and

a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations of:

receiving an audio object comprising audio content and object metadata, the object metadata of the audio object indicating an object spatial position of the audio object to be rendered by a plurality of audio speakers, each audio speaker in the plurality of audio speakers being located in a respective source spatial position in a plurality of source spatial positions, wherein the object spatial position is related to audio content in one or more audio frames, or one or more subdivisions of an audio frame;

determining, based on the object spatial position of the audio object and the plurality of source spatial positions of the plurality of audio speakers, a plurality of initial gain values for the plurality of audio speakers, each audio speaker in the plurality of audio speakers being assigned with a respective initial gain value in the plurality of initial gain values;

determining, for each of the plurality of audio speakers, whether a respective audio speaker is an active audio speaker or whether the respective audio speaker is not an active audio speaker, wherein an audio speaker is an active audio speaker if the initial gain value assigned to the audio speaker is above a threshold value, and wherein the audio speaker is not an active audio speaker if the initial gain value assigned to the audio speaker is below or at a threshold value;

determining, based on the object spatial position of the audio object and a set of source spatial positions at which the set of active audio speakers are respectively located, a set of optimized non-negative gain values for the set of active audio speakers, wherein the set of optimized gain values are yielded using the initial gain values of the active speakers as input; and

outputting, for each audio speaker in the set of active audio speakers a respective optimized gain value in the plurality of optimized gain values.

3 . A non-transitory computer-readable medium storing instructions that, when exceed by a processors, cause the one or more processors to perform the operations of:

receiving an audio object comprising audio content and object metadata, the object metadata of the audio object indicating an object spatial position of the audio object to be rendered by a plurality of audio speakers, each audio speaker in the plurality of audio speakers being located in a respective source spatial position in a plurality of source spatial positions, wherein the object spatial position is related to audio content in one or more audio frames, or one or more subdivisions of an audio frame;

determining, based on the object spatial position of the audio object and the plurality of source spatial positions of the plurality of audio speakers, a plurality of initial gain values for the plurality of audio speakers, each audio speaker in the plurality of audio speakers being assigned with a respective initial gain value in the plurality of initial gain values;

determining, for each of the plurality of audio speakers, whether a respective audio speaker is an active audio speaker or whether the respective audio speaker is not an active audio speaker, wherein an audio speaker is an active audio speaker if the initial gain value assigned to the audio speaker is above a threshold value, and wherein the audio speaker is not an active audio speaker if the initial gain value assigned to the audio speaker is below or at a threshold value;

determining, based on the object spatial position of the audio object and a set of source spatial positions at which the set of active audio speakers are respectively located, a set of optimized non-negative gain values for the set of active audio speakers, wherein the set of optimized gain values are yielded using the initial gain values of the active speakers as input; and

outputting, for each audio speaker in the set of active audio speakers a respective optimized gain value in the plurality of optimized gain values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2024
From: WANG, JUN; CENGARLE, GIULIO; TORRES, JUAN FELIX; ARTEAGA, DANIEL
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 066204/0225 →
Priority Claims (2)
ES ES201630341 · Mar 22, 2016 · national
EP 16181436 · Jul 27, 2016 · regional
Continuity (7)
Continuation 17833761 · Jun 6, 2022
Continuation 17149683 · Jan 14, 2021
Continuation 16555126 · Aug 29, 2019
Continuation 15647121 · Jul 11, 2017
Continuation 15451241 · Mar 6, 2017
Provisional Application 62345602 · Jun 3, 2016
Related Publication 20240179485A1 · May 30, 2024
References Cited (34)
US 9949052B2 · Wang · 2018 [cited by applicant]
US 10405120B2 · Wang · 2019 [cited by applicant]
US 10897682B2 · Wang · 2021 [cited by applicant]
US 11356787B2 · Wang · 2022 [cited by applicant]
US 11843930B2 · Wang · 2023 [cited by examiner]
US 20080013746A1 · Reichelt · 2008 [cited by applicant]
US 20110013790A1 · Hilpert et al. · 2011 [cited by applicant]
US 20110081023A1 · Raghuvanshi · 2011 [cited by applicant]
US 20120057715A1 · Johnston · 2012 [cited by applicant]
US 20130142341A1 · Del Galdo · 2013 [cited by applicant]
US 20140016802A1 · Sen · 2014 [cited by applicant]
US 20140050325A1 · Norris · 2014 [cited by applicant]
US 20150146873A1 · Chabanne · 2015 [cited by applicant]
US 20150221313A1 · Purnhagen et al. · 2015 [cited by applicant]
US 20150319530A1 · Kalevi · 2015 [cited by applicant]
US 20150332663A1 · Jot · 2015 [cited by applicant]
US 20160127847A1 · Shi · 2016 [cited by examiner]
US 20160295343A1 · Tsingos · 2016 [cited by applicant]
US 20180160250A1 · Yamamoto · 2018 [cited by applicant]
JP 2015080119 · 2015 [cited by applicant]
WO 2013181272 · 2013 [cited by applicant]
WO 2014147442 · 2014 [cited by applicant]
WO 20140159272 · 2014 [cited by applicant]
WO 20150017037 · 2015 [cited by applicant]
WO 20150054033 · 2015 [cited by applicant]
WO 20150080967 · 2015 [cited by applicant]
WO 20150105748 · 2015 [cited by applicant]
WO 20150150480 · 2015 [cited by applicant]
WO 20170027308 · 2017 [cited by applicant]
Bucar, Dejan “Reducing Interrupt Latency Using the Cache” Master's Thesis in Electrical Engineering Stockholm, Jan. 31, 2001, pp. 1-43. [cited by applicant]
ITU-R BS.2051-0 “Advanced Sound System for Programme Production” Feb. 2014, pp. 1-14. [cited by applicant]
Jeon, Se-Woon et al “Virtual Source Panning Using Multiple-Wise Vector Base in the Multispeaker Stereo Format” 18th European Signal Processing Conference, Aalborg, Denmark, Aug. 23-27, 2010, pp. 1337-1341. [cited by applicant]
Lee, D.D. et al Algorithms for Non-Negative Matrix Factorization in Advances in Neural and Information Processing Systems 13, pp. 556-562, 2001. [cited by applicant]
Cichocki, A. et al “Nonnegative Matrix and Tensor Factorizations: Applications to Exploratory Multi-way Data Analysis and Blind Source Separation”, Wiley 2009 205 pages. [cited by applicant]