IP Library Granted Patent US 12,634,649
Granted Patent B2
US 12,634,649 · App. 18/547,006 · Granted May 19, 2026

Clustering audio objects

Inventors: Ziyu Yang (Beijing, CN); Lie Lu (Dublin, CA)
Assignee: Dolby Laboratories Licensing Corporation
H04S7/30H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,649
App. No.
18/547,006
Granted
May 19, 2026
Kind
B2
Abstract

A method for clustering audio objects may involve identifying a plurality of audio objects, wherein each audio object of the plurality of audio objects is associated with respective metadata that indicates respective spatial position information and respective rendering metadata. The method may involve assigning audio objects of the plurality of audio objects to categories of rendering metadata of a plurality of categories of rendering metadata, wherein at least one category of rendering metadata comprises a plurality of types of rendering metadata to be preserved. The method may involve determining an allocation of a plurality of audio object clusters to each category of rendering metadata. The method may involve rendering audio objects of the plurality of audio objects to an allocated plurality of audio object clusters based on the metadata that indicates spatial position information and based on the assignments of the audio objects to the categories of rendering metadata.

Claims (34)

1 . A method for clustering audio objects, comprising:

identifying a plurality of audio objects, wherein an audio object of the plurality of audio objects is associated with respective metadata that indicates respective spatial position information and respective rendering metadata;

assigning audio objects of the plurality of audio objects to categories of rendering metadata of a plurality of categories of rendering metadata, wherein at least one category of rendering metadata comprises a plurality of types of rendering metadata to be preserved;

determining an allocation of a plurality of audio object clusters to each category of rendering metadata, wherein an audio object cluster comprises one or more audio objects of the plurality of audio objects having similar attributes;

rendering audio objects of the plurality of audio objects to an allocated plurality of audio object clusters based on the metadata that indicates spatial position information and based on the assignments of the audio objects to the categories of rendering metadata.

2 . The method of claim 1 , wherein the categories of rendering metadata comprise a bypass mode category and a virtualization category.

3 . The method of claim 2 , wherein the plurality of types of rendering metadata included in the virtualization category comprise a plurality of types of virtualization, each representing a distance from a head center to the audio object.

4 . The method of claim 1 , wherein the categories of rendering metadata comprise one of a zone category or a snap category, or wherein an audio object assigned to a first category of rendering metadata is inhibited from being assigned to an audio object cluster of the plurality of audio object clusters allocated to a second category of rendering metadata.

5 . The method of claim 1 , further comprising transmitting an audio signal that comprises spatial information and gain information associated with each audio object cluster of the allocated plurality of audio object clusters, wherein the audio signal has less spatial distortion than an audio signal comprising spatial information and gain information associated with audio object clusters in which an audio object assigned to the first category of rendering metadata is assigned to an audio object cluster associated with the second category of rendering metadata.

6 . The method of claim 1 , wherein determining the allocation of the plurality of audio object clusters to each category of rendering metadata comprises:

(i) determining an initial allocation of an initial plurality of audio object clusters to each category of rendering metadata;

(ii) assigning the audio objects to the initial plurality of audio object clusters based on the metadata that indicates spatial position information and based on the assignments of the audio objects to the categories of rendering metadata;

(iii) for each category of rendering metadata, determining a category cost of the assignment of the audio objects to the initial plurality of audio object clusters;

(iv) determining an updated allocation of the initial plurality of audio object clusters to each category of rendering metadata based at least in part on the category cost for each category of rendering metadata; and

(v) repeating (ii)-(iv) until a stopping criterion is reached.

7 . The method of claim 6 , wherein determining the category cost of the assignment of the audio objects to the initial plurality of audio object clusters is based on positions of audio object clusters allocated to the category of rendering metadata and positions of audio objects assigned to the audio object clusters allocated to the category of rendering metadata.

8 . The method of claim 7 , wherein the category cost is based on a left versus right placement of an audio object relative to a left versus right placement of an audio object cluster the audio object has been assigned to.

9 . The method of claim 6 , wherein determining the category cost of the assignment of the audio objects to the initial plurality of audio object clusters is based on:

loudness of the audio objects; and/or

a distance of an audio object to an audio object cluster the audio object has been assigned to, and/or

a similarity of a type of rendering metadata of an audio object to a type of rendering metadata of an audio object cluster the audio object has been assigned to.

10 . The method of claim 6 , further comprising determining a global cost based on the category cost for each category of rendering metadata, wherein the updated allocation of the initial plurality of audio object clusters is based on the global cost.

11 . The method of claim 10 , wherein repeating (ii)-(iv) until the stopping criterion is reached comprises determining a minimum of the global cost has been achieved.

12 . The method of claim 6 , wherein determining the updated allocation comprises changing a number of audio object clusters allocated to at least one category of rendering metadata of the plurality of categories of rendering metadata.

13 . The method of claim 12 , further comprising determining a global cost based on the category cost for each category of rendering metadata, wherein the number of audio object clusters is determined based on the global cost.

14 . The method of claim 13 , wherein determining the number of audio object clusters comprises minimizing the global cost subject to a constraint on the number of audio object clusters that indicates a maximum number of audio object clusters that can be added.

15 . The method of claim 1 , wherein rendering audio objects of the plurality of audio objects to the allocated plurality of audio object clusters comprises determining an object-to-cluster gain for each audio object of the plurality of audio objects when rendered to one or more audio object clusters allocated to a category of rendering metadata to which the audio object is assigned.

16 . The method of claim 15 , wherein object-to-cluster gains for audio objects assigned to a first category of the plurality of categories of rendering metadata are determined either:

separately from object-to-cluster gains for audio objects assigned to a second category of the plurality of categories of rendering metadata; or

jointly with object-to-cluster gains for audio objects assigned to a second category of the plurality of categories of rendering metadata.

17 . The method of claim 1 , further comprising transmitting an audio signal that comprises spatial information and gain information associated with each audio object cluster of the allocated plurality of audio object clusters, wherein transmitting the audio signal requires less bandwidth than an audio signal that comprises spatial information and gain information associated with each audio object of the plurality of audio objects.

18 . An apparatus configured for implementing the method of claim 1 .

19 . A system configured for implementing the method of claim 1 .

20 . One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2023
From: YANG, ZIYU; LU, LIE
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 065584/0231 →
Priority Claims (2)
WO PCT/CN2021/077110 · Feb 20, 2021 · international
EP 21178179 · Jun 8, 2021 · regional
Continuity (3)
Provisional Application 63202227 · Jun 2, 2021
Provisional Application 63165220 · Mar 24, 2021
Related Publication 20240187807A1 · Jun 6, 2024
References Cited (39)
US 7904797B2 · Wong · 2011 [cited by applicant]
US 8824688B2 · Schreiner · 2014 [cited by applicant]
US 8892450B2 · Schildbach · 2014 [cited by applicant]
US 8902224B2 · Wyeld · 2014 [cited by applicant]
US 9031243B2 · Leboeuf · 2015 [cited by applicant]
US 9479886B2 · Xiang · 2016 [cited by applicant]
US 9489954B2 · Hooks · 2016 [cited by applicant]
US 9728181B2 · Jot · 2017 [cited by applicant]
US 9805725B2 · Crockett · 2017 [cited by applicant]
US 10257638B2 · Van Brandenburg · 2019 [cited by applicant]
US 10277997B2 · Chen · 2019 [cited by applicant]
US 10278000B2 · Breebaart · 2019 [cited by applicant]
US 10638246B2 · Chen · 2020 [cited by applicant]
US 20140025386A1 · Xiang · 2014 [cited by applicant]
US 20150235645A1 · Hooks · 2015 [cited by examiner]
US 20160125887A1 · Purnhagen · 2016 [cited by applicant]
US 20160337776A1 · Breebaart · 2016 [cited by applicant]
US 20170180905A1 · Purnhagen · 2017 [cited by examiner]
US 20170339506A1 · Chen · 2017 [cited by applicant]
US 20180098173A1 · Van Brandenburg · 2018 [cited by applicant]
US 20190278504A1 · Matsui · 2019 [cited by applicant]
CN 106385660B · 2020 [cited by applicant]
EP 3780661A2 · 2021 [cited by applicant]
JP 2016509249A · 2016 [cited by applicant]
JP 2016522911A · 2016 [cited by applicant]
JP 2017535905A · 2017 [cited by applicant]
KR 101507901B1 · 2015 [cited by applicant]
KR 20170081688A · 2017 [cited by applicant]
RU 2019100704A · 2019 [cited by applicant]
WO WO2014015299A1 · 2014 [cited by examiner]
WO 2014099285A1 · 2014 [cited by applicant]
WO 2014187990A1 · 2014 [cited by applicant]
WO 2015017037A1 · 2015 [cited by applicant]
WO 2016094674A1 · 2016 [cited by applicant]
WO 2017085562A2 · 2017 [cited by applicant]
WO 2018017394A1 · 2018 [cited by applicant]
WO 2018026828A1 · 2018 [cited by applicant]
WO 2020105423A1 · 2020 [cited by applicant]
Liu et al., Multiple Speaker Tracking in Spatial Audio via PHD Filtering and Depth-Audio Fusion, IEEE Transactions on Multimedia ( vol. 20, Issue: 7, Jul. 2018). [cited by applicant]