IP Library › Granted Patent US 8,032,385
Granted Patent B2
US 8,032,385 · App. 12/566,621 · Granted Oct 4, 2011

Method for correcting metadata affecting the playback loudness of audio information

Assignee: Dolby Laboratories Licensing Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,032,385
App. No.
12/566,621
Granted
Oct 4, 2011
Kind
B2
Abstract

A coded signal conveys encoded audio information and metadata that may be used to control the loudness of the audio information during its playback. If the values for these metadata parameters are set incorrectly, annoying fluctuations in loudness during playback can result. The present invention overcomes this problem by detecting incorrect metadata parameter values in the signal and replacing the incorrect values with corrected values.

Claims (33)

1. A method for correcting playback loudness of audio information, wherein the method comprises steps that:

receive an input signal that conveys data representing a first loudness normalization level and first encoded audio information, wherein the data conveyed by the input signal was produced by an encoding process that generated the first encoded audio information according to psychoacoustic principles;

obtain segments of decoded audio information from an application of a decoding process to the input signal;

identify which of the segments of decoded audio information are predominantly speech;

obtain a respective measure of loudness for each of the segments of audio information from an analysis of the decoded audio information that accounts for presence or absence of speech and derive a second loudness normalization level for each segment from its respective measure of loudness;

generate an output signal that conveys data representing a third loudness normalization level and segments of third encoded audio information representing the segments of decoded audio information in an encoded form, wherein:

if a difference between the first and second loudness normalization levels does not exceed a threshold, the third loudness level represents the first loudness normalization level, and the third encoded audio information represents the first encoded audio information, and

if the difference between the first and second loudness normalization levels exceeds the threshold, the third loudness level is derived from the second loudness normalization level.

2. The method of claim 1 wherein, for each segment of decoded audio information that is predominantly speech, the respective measure of loudness represents loudness of the speech in the segment, and for each segment of decoded audio information that is not predominantly speech, the respective measure of loudness represents an average loudness of the audio information.

3. The method of claim 1 wherein, if the difference between the first and second loudness normalization levels exceeds the threshold, the third encoded audio information is generated by encoding the decoded audio information according to psychoacoustic principles.

4. The method of claim 1 wherein, if the difference between the first and second loudness normalization levels exceeds the threshold, the third encoded audio information represents the first encoded audio information.

5. An apparatus for correcting playback loudness of audio information, wherein the apparatus comprises:

means for receiving an input signal that conveys data representing a first loudness normalization level and first encoded audio information, wherein the data conveyed by the input signal was produced by an encoding process that generated the first encoded audio information according to psychoacoustic principles;

means for obtaining segments of decoded audio information from an application of a decoding process to the input signal;

means for identifying which of the segments of decoded audio information are predominantly speech;

means for obtaining a respective measure of loudness for each of the segments of audio information from an analysis of the decoded audio information that accounts for presence or absence of speech and derive a second loudness normalization level for each segment from its respective measure of loudness;

means for generating an output signal that conveys data representing a third loudness normalization level and segments of third encoded audio information representing the segments of decoded audio information in an encoded form, wherein:

if a difference between the first and second loudness normalization levels does not exceed a threshold, the third loudness level represents the first loudness normalization level, and the third encoded audio information represents the first encoded audio information, and

if the difference between the first and second loudness normalization levels exceeds the threshold, the third loudness level is derived from the second loudness normalization level.

6. The apparatus of claim 5 wherein, for each segment of decoded audio information that is predominantly speech, the respective measure of loudness represents loudness of the speech in the segment, and for each segment of decoded audio information that is not predominantly speech, the respective measure of loudness represents an average loudness of the audio information.

7. The apparatus of claim 5 wherein, if the difference between the first and second loudness normalization levels exceeds the threshold, the third encoded audio information is generated by encoding the decoded audio information according to psychoacoustic principles.

8. The apparatus of claim 5 wherein, if the difference between the first and second loudness normalization levels exceeds the threshold, the third encoded audio information represents the first encoded audio information.

9. A non-transitory storage medium recording a program of instructions that is executable by device to perform a method for correcting playback loudness of audio information, wherein the method comprises steps that:

receive an input signal that conveys data representing a first loudness normalization level and first encoded audio information, wherein the data conveyed by the input signal was produced by an encoding process that generated the first encoded audio information according to psychoacoustic principles;

obtain segments of decoded audio information from an application of a decoding process to the input signal;

identify which of the segments of decoded audio information are predominantly speech;

obtain a respective measure of loudness for each of the segments of audio information from an analysis of the decoded audio information that accounts for presence or absence of speech and derive a second loudness normalization level for each segment from its respective measure of loudness;

generate an output signal that conveys data representing a third loudness normalization level and segments of third encoded audio information representing the segments of decoded audio information in an encoded form, wherein:

if a difference between the first and second loudness normalization levels does not exceed a threshold, the third loudness level represents the first loudness normalization level, and the third encoded audio information represents the first encoded audio information, and

if the difference between the first and second loudness normalization levels exceeds the threshold, the third loudness level is derived from the second loudness normalization level.

10. The medium of claim 9 wherein, for each segment of decoded audio information that is predominantly speech, the respective measure of loudness represents loudness of the speech in the segment, and for each segment of decoded audio information that is not predominantly speech, the respective measure of loudness represents an average loudness of the audio information.

11. The medium of claim 9 wherein, if the difference between the first and second loudness normalization levels exceeds the threshold, the third encoded audio information is generated by encoding the decoded audio information according to psychoacoustic principles.

12. The medium of claim 9 wherein, if the difference between the first and second loudness normalization levels exceeds the threshold, the third encoded audio information represents the first encoded audio information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2009
From: SMITHERS, MICHAEL; RIEDMILLER, JEFFREY; ROBINSON, CHARLES; CROCKETT, BRETT
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 023392/0663 →
Continuity (2)
Continuation 10884177 · Jul 1, 2004
Related Publication 20100250258A1 · Sep 30, 2010