IP Library › Granted Patent US 12,380,898
Granted Patent B2
US 12,380,898 · App. 18/000,841 · Granted Aug 5, 2025

Encoding of multi-channel audio signals comprising downmixing of a primary and two or more scaled non-primary input channels

Inventor: David S. McGrath (New South Wales, AU)
Assignee: Dolby Laboratories Licensing Corporation
G10L19/008G10L19/002G10L19/0204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,898
App. No.
18/000,841
Granted
Aug 5, 2025
Kind
B2
Abstract

Systems, methods, and computer program products are disclosed for adaptive downmixing of audio signals with improved continuity. An audio encoding system receives an input multi-channel audio signal including a primary input audio channel and L non-primary input audio channels. The system determines a set of L input gains. For each of the channels and gains, the system forms a respective scaled non-primary input audio channel. The system forms a primary output audio channel from the sum of the primary input audio channel and the scaled non-primary input audio channels. The system determines a set of L prediction gains. The system forms a prediction channel from the primary output audio channel. The system forms L non-primary output audio channels. The system forms an output multi-channel audio signal from the primary output audio channel and the L non-primary output audio channels.

Claims (12)

1. An audio encoding method comprising:

receiving, with at least one processor, an input multi-channel audio signal comprising a primary input audio channel and L non-primary input audio channels;

determining, with the at least one processor, a set of L input gains, wherein L is a positive integer greater than one, and wherein the set of L input gains are determined by scaling a set of L mixing coefficients by an input mixture strength coefficient;

for each of the L non-primary input audio channels and L input gains, forming a respective scaled non-primary input audio channel from the respective non-primary input audio channel scaled according to the input gain;

forming a primary output audio channel from a sum of the primary input audio channel and the scaled non-primary input audio channels;

determining, with the at least one processor, a set of L prediction gains, wherein the set of L prediction gains is determined by scaling the set of L mixing coefficients by a prediction mixture strength coefficient;

for each of the L prediction gains, forming, with the at least one processor, a prediction channel from the primary output audio channel scaled according to the prediction gain;

forming, with the at least one processor, L non-primary output audio channels from a difference of the respective non-primary input audio channel and the respective prediction channel;

forming, with the at least one processor, an output multi-channel audio signal from the primary output audio channel and the L non-primary output audio channels;

encoding, with an audio encoder, the output multi-channel audio signal; and

transmitting or storing, with the at least one processor, the encoded output multi-channel audio signal.

2. A non-transitory computer-readable medium storing instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform operations of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2023
From: MCGRATH, DAVID S.
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 064436/0754 →
Continuity (3)
Provisional Application 63037635 · Jun 11, 2020
Provisional Application 63193926 · May 27, 2021
Related Publication 20230215444A1 · Jul 6, 2023
References Cited (31)
US 8340325B2 · Oh · 2012 [cited by applicant]
US 8538766B2 · Hellmuth · 2013 [cited by applicant]
US 8817991B2 · Jaillet · 2014 [cited by applicant]
US 9129593B2 · Ojala · 2015 [cited by applicant]
US 9514759B2 · Virette · 2016 [cited by applicant]
US 9584235B2 · Ojala · 2017 [cited by applicant]
US 9875746B2 · Honma · 2018 [cited by applicant]
US 20070033056A1 · Herre · 2007 [cited by applicant]
US 20080192941A1 · Oh · 2008 [cited by applicant]
US 20160241982A1 · Seefeldt · 2016 [cited by examiner]
US 20160247507A1 · Disch · 2016 [cited by applicant]
US 20180122384A1 · Melkote · 2018 [cited by examiner]
US 20190259398A1 · Buethe · 2019 [cited by applicant]
US 20200152209A1 · Büthe · 2020 [cited by applicant]
US 20220122619A1 · Shlomot · 2022 [cited by examiner]
CN 105723453A · 2019 [cited by applicant]
CN 110544484B · 2021 [cited by applicant]
EA 035064B1 · 2020 [cited by applicant]
JP H0870252A · 1996 [cited by applicant]
RU 2609097C2 · 2017 [cited by applicant]
TW 201214416A · 2012 [cited by applicant]
TW 201830379A · 2018 [cited by applicant]
WO 2007009548A1 · 2007 [cited by applicant]
WO 2016001357A1 · 2016 [cited by applicant]
WO 2020010072A1 · 2020 [cited by applicant]
WO 2021022087A1 · 2021 [cited by applicant]
Mahe Pierre, et al “First-Order Ambisonic Coding with PCA Matrixing and Quaternion-Based Interpolation”, Proceedings of the 22 nd International Conference on Digital Audio Effects, Sep. 2, 2019 (Sep. 2, 2019), pp. 1-8, … [cited by applicant]
McGrath D et al: Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec11 , ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP… [cited by applicant]
Text of ISO/IEC 23008-3:201x 3D Audio, Second Edition11 , 117. MPEG Meeting; Jan. 16, 2017-Jan. 20, 2017; Geneva; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), Mar. 9, 2017. ⋅ No. n16582. [cited by applicant]
Wu, B. et al “Downmix and Coding of Multichannel Signals Based on Spatial Correlation” IEEE 2015 8th International Congress on Image and Signal Processing (CISP 2015) pp. 1142-1146. [cited by applicant]
Choi et al., “Multi-Channel Audio CODEC with Channel Interface Suppression,” Journal of Semiconductor Technology and Science, vol. 15, No. 6, Dec. 30, 2015, pp. 1-7, 7 pages. [cited by applicant]