IP Library › Granted Patent US 12,431,145
Granted Patent B2
US 12,431,145 · App. 18/327,623 · Granted Sep 30, 2025

Immersive voice and audio services (IVAS) with adaptive downmix strategies

Inventors: Harald Mundt (Fürth, DE); David S. McGrath (Rose Bay, AU); Rishabh Tyagi (Sydney, AU)
Assignees: Dolby Laboratories Licensing Corporation; Dolby International AB
G10L19/008G10L19/083H04S7/00H04S2400/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,431,145
App. No.
18/327,623
Granted
Sep 30, 2025
Kind
B2
Abstract

Disclosed is an audio signal encoding/decoding method that uses an encoding downmix strategy applied at an encoder that is different than a decoding re-mix/upmix strategy applied at a decoder. Based on the type of downmix coding scheme, the method comprises: computing input downmixing gains to be applied to the input audio signal to construct a primary downmix channel; determining downmix scaling gains to scale the primary downmix channel; generating prediction gains based on the input audio signal, the input downmixing gains and the downmix scaling gains; determining residual channel(s) from the side channels by using the primary downmix channel and the prediction gains to generate side channel predictions and subtracting the side channel predictions from the side channels; determining decorrelation gains based on energy in the residual channels; encoding the primary downmix channel, the residual channel(s), the prediction gains and the decorrelation gains; and sending the bitstream to a decoder.

Claims (11)

1. An audio signal encoding method comprising:

obtaining, with at least one processor, an input audio signal, the input audio signal representing an input audio scene and comprising a primary input audio channel and side channels;

determining, with the at least one processor, a type of downmix coding scheme based on the input audio signal;

based on the type of downmix coding scheme:

computing, with the at least one processor, one or more input downmixing gains to be applied to the input audio signal to construct a primary downmix channel, wherein the input downmixing gains are determined to minimize an overall prediction error on the side channels;

determining, with the at least one processor, one or more downmix scaling gains to scale the primary downmix channel, wherein the downmix scaling gains are determined by minimizing an energy difference between a reconstructed representation of the input audio scene from the primary downmix channel and the input audio signal;

generating, with the at least one processor, prediction gains based on the input audio signal, the input downmixing gains and the downmix scaling gains;

determining, with the at least one processor, one or more residual channels from the side channels in the input audio signal by using the primary downmix channel and the prediction gains to generate side channel predictions and then subtracting the side channel predictions from the side channels;

determining, with the at least one processor, decorrelation gains based on energy in the residual channels;

encoding, with the at least one processor, the primary downmix channel, zero or more of the residual channels and side information into a bitstream, the side information comprising the prediction gains and the decorrelation gains; and

outputting, with the at least one processor, the bitstream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2024
From: MUNDT, HARALD; MCGRATH, DAVID S.; TYAGI, RISHABH
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 066025/0334 →
Continuity (4)
Provisional Application 63228732 · Aug 3, 2021
Provisional Application 63171404 · Apr 6, 2021
Provisional Application 63120365 · Dec 2, 2020
Related Publication 20240135937A1 · Apr 25, 2024
References Cited (37)
US 8249883B2 · Mehrotra et al. · 2012 [cited by applicant]
US 8290783B2 · Schnell · 2012 [cited by applicant]
US 8325929B2 · Koppens · 2012 [cited by applicant]
US 8972270B2 · Oh · 2015 [cited by applicant]
US 9137603B2 · Breebaart · 2015 [cited by applicant]
US 9584912B2 · Koppens · 2017 [cited by applicant]
US 9761229B2 · Xiang · 2017 [cited by applicant]
US 9786285B2 · Herre · 2017 [cited by applicant]
US 9812136B2 · Kristofer · 2017 [cited by applicant]
US 9848272B2 · Lars · 2017 [cited by applicant]
US 10448185B2 · Disch · 2019 [cited by applicant]
US 10986456B2 · Song et al. · 2021 [cited by applicant]
US 20100014679A1 · Kim · 2010 [cited by applicant]
US 20140211947A1 · Wu · 2014 [cited by examiner]
US 20150086022A1 · Engdegard · 2015 [cited by applicant]
US 20160155448A1 · Purnhagen · 2016 [cited by applicant]
US 20170365264A1 · Disch · 2017 [cited by examiner]
US 20190110147A1 · Song · 2019 [cited by applicant]
US 20190156841A1 · Fatus · 2019 [cited by applicant]
US 20190272833A1 · Borss · 2019 [cited by examiner]
US 20190287542A1 · Fueg · 2019 [cited by applicant]
US 20200302943A1 · Lars · 2020 [cited by applicant]
US 20200395023A1 · Purnhagen · 2020 [cited by applicant]
US 20220036911A1 · Reutelhuber · 2022 [cited by examiner]
US 20220108707A1 · Bouthéon · 2022 [cited by examiner]
US 20230051420A1 · Eksler · 2023 [cited by examiner]
US 20230215444A1 · McGrath · 2023 [cited by applicant]
US 20230298602A1 · Eichenseer · 2023 [cited by examiner]
EP 3079379A1 · 2016 [cited by applicant]
EP 3550561A1 · 2019 [cited by applicant]
EP 3079379B1 · 2020 [cited by applicant]
KR 20140003619A · 2014 [cited by applicant]
RU 2666640C2 · 2018 [cited by applicant]
WO 2024097485A1 · 2024 [cited by applicant]
Adrien, “Spatial auditory blurring and applications to multichannel audio coding.” Acoustics [physics.class-ph]. PhD diss., Universite Pierre et Marie Curie—Paris VI, Sep. 14, 2011, pp. 1-173, 173 pages. [cited by applicant]
Bleidt et al., “Development of the MPEG-H TV Audio System for ATSC 3.0”, IEEE Transactions On Broadcasting., vol. 63, No. 1, Mar. 1, 2017 (Mar. 1, 2017), pp. 202-236, 35 pages. [cited by applicant]
McGrath et al., “Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec”, ICASSP 2019—2019 IEEE International Conference On Acoustics, Speech and Signal Processing (ICASSP),… [cited by applicant]