IP Library Granted Patent US 10,643,626
Granted Patent B2
US 10,643,626 · App. 16/436,835 · Granted May 5, 2020

Methods for parametric multi-channel encoding

Inventors: Tobias Friedrich (Furth, DE); Alexander Mueller (Nuremberg, DE); Karsten Linzmeier (Nuremberg, DE); Claus-Christian Spenger (Nuremberg, DE); Tobias R. Wagenblass (Nuremberg, DE)
Assignee: Dolby International AB
G10L19/008G10L19/167H04S3/008H04S2400/01H04S2400/03H04S2420/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,643,626
App. No.
16/436,835
Granted
May 5, 2020
Kind
B2
Abstract

The present document relates to audio coding systems. In particular, the present document relates to efficient methods and systems for parametric multi-channel audio coding. An audio encoding system configured to generate a bitstream indicative of a downmix signal and spatial metadata for generating a multi-channel upmix signal from the downmix signal is described. The system comprises a downmix processing unit configured to generate the downmix signal from a multi-channel input signal; wherein the downmix signal comprises m channels and wherein the multi-channel input signal comprises n channels; n, m being integers with m<n. Furthermore, the system comprises a parameter processing unit configured to determine the spatial metadata from the multi-channel input signal. In addition, the system comprises a configuration unit configured to determine one or more control settings for the parameter processing unit based on one or more external settings; wherein the one or more external settings comprise a target data-rate for the bitstream and wherein the one or more control settings comprise a maximum data-rate for the spatial metadata.

Claims (30)

1. A method comprising:

obtaining, by an audio decoder of a playback equipment, an encoded bitstream generated by an audio encoding system;

extracting, from the encoded bitstream by the audio decoder, an audio signal;

extracting, from the encoded bitstream by the audio decoder, a first set of dynamic range control (DRC) values configured for controlling a dynamic range of the audio signal during playback by the playback equipment, wherein the first set of DRC values are generated and encoded into the encoded bitstream by the audio encoding system;

extracting, from the encoded bitstream, a second set of DRC values configured for preventing the audio signal from clipping during playback by the playback equipment, wherein the second set of DRC values are generated and encoded into the encoded bitstream by the audio encoding system, wherein each DRC value in the second set of DRC values represents a clipping protection gain indicating an attenuation to be applied to a corresponding frame of the audio signal to prevent clipping;

extracting, from the encoded bitstream, first metadata indicating how to apply the first and second sets of DRC values to the audio signal;

interpolating one or more DRC values from the first and second sets of DRC values to generate interpolated DRC values;

applying the interpolated DRC values to the audio signal during playback by the playback equipment according to the first metadata; and

rendering the audio signal with the playback equipment.

2. The method of claim 1 , wherein the audio signal is a m-channel downmix audio signal, the method further comprises:

applying the second set of DRC values to the m-channel downmix audio signal;

extracting, from the encoded bitstream, spatial metadata; and

upmixing the m-channel downmix audio signal into an n-channel audio signal using the spatial metadata, where m and n are positive integers and m is less than n.

3. The method of claim 1 , wherein the first set of DRC values are configured to dynamically compress the audio signal.

4. An apparatus comprising:

one or more processors;

memory storing instructions, which, when executed by the one or more processors, causes the one or more processors to perform operations comprising:

obtaining, by an audio decoder of a playback equipment, an encoded bitstream generated by an audio encoding system;

extracting, from the encoded bitstream by the audio decoder, an audio signal;

extracting, from the encoded bitstream by the audio decoder, a first set of dynamic range control (DRC) values configured for controlling a dynamic range of the audio signal during playback by the playback equipment, wherein the first set of DRC values are generated and encoded into the encoded bitstream by the audio encoding system;

extracting, from the encoded bitstream, a second set of DRC values configured for preventing the audio signal from clipping during playback by the playback equipment, wherein the second set of DRC values are generated and encoded into the encoded bitstream by the audio encoding system, wherein each DRC value in the second set of DRC values represents a clipping protection gain indicating an attenuation to be applied to a corresponding frame of the audio signal to prevent clipping;

extracting, from the encoded bitstream, first metadata indicating how to apply the first and second sets of DRC values to the audio signal; and

interpolating one or more DRC values from the first and second sets of DRC values to generate interpolated DRC values;

applying the interpolated DRC values to the audio signal during playback by the playback equipment according to the first metadata;

rendering the audio signal with the playback equipment.

5. The apparatus of claim 4 , wherein the audio signal is a m-channel downmix audio signal, the operations further comprising:

applying the second set of DRC values to the m-channel downmix audio signal; extracting, from the encoded bitstream, spatial metadata; and

upmixing the m-channel downmix audio signal into an n-channel audio signal using the spatial metadata, where m and n are positive integers and m is less than n.

6. The apparatus of claim 4 , wherein the first set of DRC values are configured to dynamically compress the audio signal.

7. A non-transitory computer-readable storage medium comprising a sequence of instructions, wherein, when executed by one or more processors, the sequence of instructions causes the one or more processors to perform the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2019
From: FRIEDRICH, TOBIAS; MUELLER, ALEXANDER; LINZMEIER, KARSTEN; SPENGER, CLAUS-CHRISTIAN; WAGENBLASS, TOBIAS R.
To: DOLBY INTERNATIONAL AB
Reel/Frame 049429/0429 →
Continuity (4)
Continuation 15646482 · Jul 11, 2017
Continuation 14767883
Provisional Application 61767673 · Feb 21, 2013
Related Publication 20190348052A1 · Nov 14, 2019