IP Library Granted Patent US 11,386,913
Granted Patent B2
US 11,386,913 · App. 16/636,241 · Granted Jul 12, 2022

Audio object classification based on location metadata

Inventor: Mark William Gerrard (Rozelle, AU)
Assignee: Dolby Laboratories Licensing Corporation
G10L21/0272G06F16/635G06F16/65G06F16/683G06F16/687
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,386,913
App. No.
16/636,241
Granted
Jul 12, 2022
Kind
B2
Abstract

Methods ( 700, 800, 900 ), systems ( 200, 300, 400, 500, 600 ) and computer program products are provided. Location metadata ( 620 ) associated with an audio object is received ( 801 ). The location metadata defines a position of the audio object in an audio scene. It is estimated ( 630, 802 ), based on the location metadata, whether the audio object includes dialog. A value representative of a result of the estimation is assigned ( 803 ) to an object type parameter ( 231 ). In some example embodiments, audio objects are selected ( 661, 662, 804 ) based on values of their respective of object type parameters. In some example embodiments, at least one of the selected audio objects is submitted to dialog enhancement ( 690, 807 ).

Claims (51)

1. A method comprising:

receiving location metadata associated with an audio object, wherein the location metadata defines a position of the audio object in an audio scene;

estimating, based on the location metadata, whether the audio object includes dialog; and

assigning a value to an object type parameter representative of a result of the estimation,

wherein estimating whether the audio object includes dialog comprises:

computing a speed of the audio object based on location metadata associated with different time frames; and

estimating, based on said speed, whether the audio object includes dialog.

2. The method of claim 1 , wherein the object type parameter:

indicates a level of confidence that the audio object includes dialog; or

is a Boolean type parameter indicating whether or not a level of confidence that the audio object includes dialog is above a threshold.

3. The method of claim 1 , wherein the estimation is performed based on a position of the audio object in a front-back direction of the audio scene, the position in the front-back direction being defined by the location metadata.

4. The method of claim 3 , wherein estimating whether the audio object includes dialog comprises:

associating a position at a front of the audio scene with a higher level of confidence that the audio object includes dialog than levels of confidence associated with positions further back in the audio scene.

5. The method of claim 1 , wherein estimating whether the audio object includes dialog comprises:

associating a first value of said speed with a higher level of confidence that the audio object includes dialog than a level of confidence associated with a second value of said speed, wherein the first value of said speed is lower than the second value of said speed.

6. The method of claim 1 , wherein the estimation is performed based on a level of elevation of the audio object defined by the location metadata.

7. The method of claim 6 , wherein estimating whether the audio object includes dialog comprises:

associating a first level of elevation of the audio object with a higher level of confidence that the audio object includes dialog than levels of confidence associated with other levels of elevation of the audio object, wherein the first level of elevation corresponds to a floor level of the audio scene or a vertical position of an intended listener.

8. The method of claim 1 , comprising:

receiving a plurality of audio objects, each of the received audio objects including audio content and location metadata, wherein the location metadata of an audio object defines a position of that audio object in an audio scene;

estimating, based on the location metadata of the respective audio objects, whether the respective audio objects include dialog;

assigning values to object type parameters representative of results of the respective estimations; and

selecting a subset of the plurality of audio objects based on the assigned values of the object type parameters, wherein the subset includes one or more audio objects.

9. The method of claim 8 , further comprising:

subjecting at least one audio object in the selected subset to dialog enhancement; and/or

performing clustering such that the audio content from those of the plurality of audio objects outside the selected subset is included in a collection of clusters and such that:

at least one audio object of the selected subset is excluded from the clustering; or

the audio content of at least one audio object of the selected subset is included in a cluster which does not include audio content from any of those of the plurality of audio objects outside the selected subset.

10. The method of claim 8 , further comprising, for each of the one or more audio objects in the selected subset:

analyzing the audio content of the audio object; and

determining, based on said analysis, a value indicating a level of confidence that the audio object includes dialog.

11. The method of claim 10 , comprising:

subjecting at least one audio object from the selected subset to dialog enhancement, wherein a degree of dialog enhancement to which said at least one audio object is subjected is determined based on the corresponding at least one determined value.

12. The method of claim 10 , wherein the selected subset includes multiple audio objects, the method comprising:

selecting at least one audio object from the selected subset based on the determined values; and

subjecting the selected at least one audio object to dialog enhancement and/or performing clustering such that the audio content from those of the plurality of audio objects outside the selected at least one audio object is included in a collection of clusters, wherein the clustering is performed such that:

the at least one selected audio object is excluded from the clustering; or

the audio content of the at least one selected audio object is included in a cluster which does not include audio content from any of those of the plurality of audio objects outside the at least one selected audio object.

13. A computer program product comprising a non-transitory computer-readable medium with instructions, which when executed by one or more processors of an electronic device, cause the device to perform the steps of:

receiving location metadata associated with an audio object, wherein the location metadata defines a position of the audio object in an audio scene;

estimating, based on the location metadata, whether the audio object includes dialog; and

assigning a value to an object type parameter representative of a result of the estimation, wherein estimating whether the audio object includes dialog comprises:

computing a speed of the audio object based on location metadata associated with different time frames; and

estimating, based on said speed, whether the audio object includes dialog.

14. A system configured to receive location metadata associated with an audio object, wherein the location metadata defines a position of the audio object in an audio scene, the system comprising:

a processor configured to perform the following:

estimating, based on the location metadata, whether the audio object includes dialog, and

assigning a value to an object type parameter representative of a result of the estimation,

wherein estimating whether the audio object includes dialog comprises:

computing a speed of the audio object based on location metadata associated with different time frames; and

estimating, based on said speed, whether the audio object includes dialog.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2020
From: GERRARD, MARK WILLIAM
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 051730/0073 →
Priority Claims (1)
EP 17184244 · Aug 1, 2017 · regional
Continuity (2)
Provisional Application 62539599 · Aug 1, 2017
Related Publication 20200381003A1 · Dec 3, 2020