IP Library › Granted Patent US 10,958,229
Granted Patent B2
US 10,958,229 · App. 16/776,297 · Granted Mar 23, 2021

Metadata for loudness and dynamic range control

Inventors: Frank Baumgarte (Sunnyvale, CA); Eric A. Allamanche (Sunnyvale, CA); Stefan K. O. Strommer (Mountain View, CA)
Assignee: APPLE INC.
H03G3/20G10L19/008G10L21/0316H03G7/007
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,958,229
App. No.
16/776,297
Granted
Mar 23, 2021
Kind
B2
Abstract

An audio normalization gain value is applied to an audio signal to produce a normalized signal. The normalized signal is processed to compute dynamic range control (DRC) gain values in accordance with a selected one of several pre-defined DRC characteristics. The audio signal is encoded, and the DRC gain values are provided as metadata associated with the encoded audio signal. Several other embodiments are also described and claimed.

Claims (40)

1. A method for decoding audio during playback processing, comprising:

receiving an encoded audio signal;

receiving metadata associated with the encoded audio signal, the metadata including a plurality of dynamic range control (DRC) gain values and an index of a previously selected DRC characteristic in accordance with which the DRC gain values were determined for the encoded audio signal;

decoding the encoded audio signal to produce a decoded audio signal;

selecting a current DRC characteristic, wherein the selected current DRC characteristic is associated with the index;

applying the plurality of DRC gain values from the metadata to the current DRC characteristic to obtain a plurality of input levels; and

applying the plurality of input levels to a target characteristic to produce gain values and applying the gain values to the decoded audio signal to produce an adjusted audio signal during playback processing.

2. The method of claim 1 wherein the received metadata further includes a plurality of values selected from the group consisting of: program loudness, true peak, loudness range, maximum momentary loudness, and short-term loudness values.

3. The method of claim 1 wherein applying the plurality of input levels to a target characteristic is based on a playback condition being late night, noisy environment, inside a moving car, walking, running, or speaker dynamic range.

4. The method of claim 1 wherein applying the plurality of input levels to a target characteristic to produce gain values and applying the gain values to the decoded audio signal comprises

performing DRC compression for downmixing only if playback volume is above a threshold.

5. The method of claim 4 further comprising

extracting a true peak value of a stereo downmix from the metadata; and

using the true peak value to estimate how much DRC compression is needed.

6. A digital audio decoder apparatus, comprising:

a digital media player having a decoder, a processor, the decoder to receive an encoded audio signal and produce a decoded audio signal;

the processor to receive metadata which is associated with the encoded audio signal, wherein the metadata includes a plurality of dynamic range control (DRC) gain values and an index of a previously selected DRC characteristic in accordance with which the DRC gain values were determined for the encoded audio signal, the processor to select a current DRC characteristic that is associated with the index, and apply the plurality of DRC gain values from the metadata to the current DRC characteristic to obtain a plurality of input levels; and

the digital media player to apply the plurality of input levels to a target characteristic to produce gain values and apply the gain values to the decoded audio signal to produce an adjusted audio signal.

7. The apparatus of claim 6 wherein the digital media player is part of an end user device that further comprises a digital to analog converter (DAC) to convert the adjusted audio signal into analog form during playback of the encoded audio signal.

8. The apparatus of claim 7 further comprising a downmix processor to perform a downmix conversion upon the adjusted audio signal, prior to conversion into analog form, based on downmix gain values extracted from the metadata.

9. The apparatus of claim 8 wherein the digital media player is to apply the plurality of input levels to the target characteristic based on a playback condition being one of: playback volume setting; user context including late night, walking, running, or car; DAC dynamic range; or speaker dynamic range.

10. The apparatus of claim 8 wherein the digital media player is to apply the plurality of input levels to the target characteristic to produce gain values and apply the gain values to the decoded audio signal by

performing DRC compression for downmixing only if playback volume is above a threshold.

11. The apparatus of claim 10 wherein the processor is to extract from the metadata a true peak value of a stereo downmix of the encoded audio signal, and use the true peak value to estimate how much DRC compression is to be applied to the decoded signal prior to a downmix conversion.

12. The apparatus of claim 7 wherein the processor is to select the target characteristic from amongst a plurality of stored characteristics based on one or more of the following: playback volume setting; user context including late night, walking, running, or car; DAC dynamic range, speaker dynamic range.

13. An audio system comprising:

a processor; and

memory having stored therein instructions that when executed by the processor

receive metadata associated with the encoded audio signal, the metadata including a plurality of dynamic range control (DRC) gain values and an index of a previously selected DRC characteristic in accordance with which the DRC gain values were computed for the encoded audio signal;

decode the encoded audio signal to produce a decoded audio signal,

select a current DRC characteristic, wherein the selected current DRC characteristic is associated with the index in the received metadata,

apply the plurality of DRC gain values from the metadata to the current DRC characteristic to obtain a plurality of input levels, and

apply the plurality of input levels to a target characteristic to produce gain values and apply the gain values to the decoded audio signal to produce an adjusted audio signal.

14. The audio system of claim 13 wherein the received metadata further includes a plurality of values selected from the group consisting of: program loudness, true peak, loudness range, maximum momentary loudness, and short-term loudness values.

15. The audio system of claim 13 wherein applying the plurality of input levels to a target characteristic is based on a playback condition being late night, noisy environment, inside a moving car, walking, running, or speaker dynamic range.

16. The audio system of claim 13 wherein applying the plurality of input levels to the target characteristic to produce gain values and applying the gain values to the decoded audio signal comprises

performing DRC compression for downmixing only if playback volume is above a threshold.

17. The audio system of claim 16 wherein the memory has further instructions that when executed by the processor

extract a true peak value of a stereo downmix from the metadata, and

use the true peak value to estimate how much DRC compression is needed.

Continuity (4)
Continuation 15417424 · Jan 27, 2017
Continuation 14225950 · Mar 26, 2014
Provisional Application 61806570 · Mar 29, 2013
Related Publication 20200169233A1 · May 28, 2020
Cited By (2)
US 12,563,339 US 12,694,879