IP Library › Granted Patent US 10,170,131
Granted Patent B2
US 10,170,131 · App. 15/513,543 · Granted Jan 1, 2019

Decoding method and decoder for dialog enhancement

Inventors: Jeroen Koppens (Sodertalje, SE); Per Ekstrand (Saltsjobaden, SE)
Assignee: Dolby International AB
G10L21/0205G10L21/0316H04S3/008G10L19/008H04S2400/01H04S2400/03H04S2420/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,170,131
App. No.
15/513,543
Granted
Jan 1, 2019
Kind
B2
Abstract

There is provided a method for enhancing dialog in a decoder of an audio system. The method comprises receiving a plurality of downmix signals being a downmix of a larger plurality of channels; receiving parameters for dialog enhancement being defined with respect to a subset of the plurality of channels that is downmixed into a subset of the plurality of downmix signals; upmixing the subset of downmix signals parametrically in order to reconstruct the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined; applying dialog enhancement to the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined using the parameters for dialog enhancement to provide at least one dialog enhanced signal; and subjecting the at least one dialog enhanced signal to mixing to provide dialog enhanced versions of the subset of downmix signals.

Claims (46)

1. A method for enhancing dialog in a decoder of an audio system, the method comprising the steps of:

receiving a plurality of downmix signals being a downmix of a larger plurality of channels;

receiving parameters for dialog enhancement, wherein the parameters are defined with respect to a subset of the plurality of channels including channels comprising dialog, wherein the subset of the plurality of channels is downmixed into a subset of the plurality of downmix signals, wherein the subset of the plurality of downmix signals contains fewer downmix signals than the plurality of downmix signals;

receiving reconstruction parameters allowing parametric reconstruction of channels that are downmixed into the subset of the plurality of downmix signals;

upmixing only the subset of the plurality of downmix signals parametrically based on the reconstruction parameters in order to reconstruct only a subset of the plurality of channels including the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined;

applying dialog enhancement to the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined using the parameters for dialog enhancement so as to provide at least one dialog enhanced signal; and

providing dialog enhanced versions of the subset of the plurality of downmix signals by mixing the at least one dialog enhanced signal with at least one other signal.

2. The method of claim 1 , wherein, in the step of upmixing only the subset of the plurality of downmix signals parametrically, no decorrelated signals are used in order to reconstruct only a subset of the plurality of channels including the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined.

3. The method of claim 1 , wherein the mixing is made in accordance with mixing parameters describing a contribution of the at least one dialog enhanced signal to the dialog enhanced versions of the subset of the plurality of downmix signals.

4. The method of claim 1 , wherein the step of upmixing only the subset of the plurality of downmix signals parametrically comprises reconstructing only the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined,

wherein the step of applying dialog enhancement comprises predicting and enhancing a dialog component from the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined using the parameters for dialog enhancement so as to provide the at least one dialog enhanced signal, and

wherein the mixing comprises mixing the at least one dialog enhanced signal with the subset of the plurality of downmix signals.

5. The method of claim 1 , further comprising: receiving an audio signal representing dialog, wherein the step of applying dialog enhancement comprises applying dialog enhancement to the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined further using the audio signal representing dialog.

6. The method of claim 1 , further comprising receiving mixing parameters for mixing the at least one dialog enhanced signal with at least one other signal.

7. The method of claim 1 , wherein the steps of upmixing only the subset of the plurality of downmix signals, applying dialog enhancement, and mixing are performed as matrix operations defined by the reconstruction parameters, the parameters for dialog enhancement, and the mixing parameters, respectively, and optionally, further comprising combining, by matrix multiplication, the matrix operations corresponding to the steps of upmixing only the subset of the plurality of downmix signals, applying dialog enhancement, and mixing, into a single matrix operation before application to the subset of the plurality of downmix signals.

8. The method of claim 1 , wherein the dialog enhancement parameters and the reconstruction parameters are frequency dependent.

9. The method of claim 8 , wherein the parameters for dialog enhancement are defined with respect to a first set of frequency bands and the reconstruction parameters are defined with respect to a second set of frequency bands, the second set of frequency bands being different than the first set of frequency bands.

10. The method of claim 1 , wherein:

values of the parameters for dialog enhancement are received repeatedly and are associated with a first set of time instants (T1={t11, t12, t13, . . . }), at which respective values apply exactly, wherein a predefined first interpolation pattern (I1) is to be performed between consecutive time instants; and

values of the reconstruction parameters are received repeatedly and are associated with a second set of time instants (T2={t21, t22, t23, . . . }), at which respective values apply exactly, wherein a predefined second interpolation pattern (I2) is to be performed between consecutive time instants,

the method further comprising:

selecting a parameter type being either parameters for dialog enhancement or reconstruction parameters and in such manner that the set of time instants associated with the selected type comprises at least one prediction instant being a time instant (t p ) that is absent from the set associated with the not-selected type;

predicting a value of the parameters of the not-selected type at the prediction instant (t p );

computing, based on at least the predicted value of the parameters of the not-selected type and a received value of the parameters of the selected type, a joint processing operation representing at least upmixing of only the subset of the downmix signals followed by dialog enhancement at the prediction instant (t p ); and

computing, based on at least a value of the parameters of the selected type and a value of the parameters of the not-selected type, at least either being a received value, said joint processing operation at an adjacent time instant (t a ) in the set associated with the selected or the not-selected type,

wherein said steps of upmixing only the subset of the plurality of downmix signals and applying dialog enhancement are performed between the prediction instant (t p ) and the adjacent time instant (t a ) by way of an interpolated value of the computed joint processing operation.

11. The method of claim 10 , wherein the selected type of parameters is the reconstruction parameters.

12. The method of claim 10 , wherein said joint processing operation at the adjacent time instant (t a ) is computed based on a received value of the parameters of the selected type and a received value of the parameters of the not-selected type.

13. The method of claim 10 ,

further comprising selecting, on the basis of the first and second interpolation patterns, a joint interpolation pattern (I3) according to a predefined selection rule,

wherein said interpolation of the computed respective joint processing operations is in accordance with the joint interpolation pattern.

14. The method of claim 13 , wherein the predefined selection rule is defined for the case where the first and second interpolation patterns are different.

15. The method of claim 14 , wherein, in response to the first interpolation pattern (I1) being linear and the second interpolation pattern (I2) being piecewise constant, linear interpolation is selected as the joint interpolation pattern.

16. The method of claim 10 , wherein the prediction of the value of the parameters of the not-selected type at the prediction instant (t p ) is made in accordance with the interpolation pattern for the parameters of the not-selected type.

17. The method of claim 10 , wherein the joint processing operation is computed as a single matrix operation before it is applied to the subset of the plurality of downmix signals, and optionally, wherein:

linear interpolation is selected as the joint interpolation pattern; and

the interpolated value of the computed respective joint processing operations is computed by linear matrix interpolation.

18. The method of claim 1 , wherein the mixing of the at least one dialog enhanced signal with at least one other signal is restricted to a non-complete selection of the plurality of downmix signals.

19. A non-transitory computer-readable storage medium comprising a sequence of instructions, which, when performed by one or more processing devices, cause the one or more processing devices to perform the method of claim 1 .

20. A decoder for enhancing dialog in an audio system, wherein the decoder:

receives a plurality of downmix signals being a downmix of a larger plurality of channels;

receives parameters for dialog enhancement, wherein the parameters are defined with respect to a subset of the plurality of channels including channels comprising dialog, wherein the subset of the plurality of channels is downmixed into a subset of the plurality of downmix signals, wherein the subset of the plurality of downmix signals contains fewer downmix signals than the plurality of downmix signals;

receives reconstruction parameters allowing parametric reconstruction of channels that are downmixed into the subset of the plurality of downmix signals;

upmixes only the subset of the plurality of downmix signals parametrically based on the reconstruction parameters in order to reconstruct only a subset of the plurality of channels including the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined; and

applies dialog enhancement to the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined using the parameters for dialog enhancement so as to provide at least one dialog enhanced signal; and

provides dialog enhanced versions of the subset of the plurality of downmix signals by mixing the at least one dialog enhanced signal with at least one other signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2017
From: KOPPENS, JEROEN; EKSTRAND, PER
To: DOLBY INTERNATIONAL AB
Reel/Frame 042334/0102 →
Continuity (3)
Provisional Application 62128331 · Mar 4, 2015
Provisional Application 62059015 · Oct 2, 2014
Related Publication 20170309288A1 · Oct 26, 2017