IP Library › Granted Patent US 12,283,281
Granted Patent B2
US 12,283,281 · App. 17/772,497 · Granted Apr 22, 2025

Bitrate distribution in immersive voice and audio services

Inventors: Rishabh Tyagi (Sydney, AU); Juan Felix Torres (Darlinghurst, AU); Stefanie Brown (Lewisham, AU)
Assignee: Dolby Laboratories Licensing Corporation
G10L19/032G10L19/008G10L19/167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,281
App. No.
17/772,497
Granted
Apr 22, 2025
Kind
B2
Abstract

Embodiments are disclosed for bitrate distribution in immersive voice and audio services. In an embodiment, a method of encoding an IVAS bitstream comprises: receiving an input audio signal; downmixing the input audio signal into one or more downmix channels and spatial metadata; reading a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata from a bitrate distribution control table; determining a combination of the one or more bitrates for the downmix channels; determining a metadata quantization level from the set of metadata quantization levels using a bitrate distribution process; quantizing and coding the spatial metadata using the metadata quantization level; generating, using the combination of one or more bitrates, a downmix bitstream for the one or more downmix channels; combining the downmix bitstream, the quantized and coded spatial metadata and the set of quantization levels into the IVAS bitstream.

Claims (26)

1. A method of encoding an immersive voice and audio services (IVAS) bitstream, the method comprising:

receiving, using one or more processors, an input audio signal;

downmixing, using the one or more processors, the input audio signal into one or more downmix channels and spatial metadata associated with one or more channels of the input audio signal;

obtaining using the one or more processors, a set of one or more target bitrates for the one or more downmix channels and a set of metadata quantization levels for the spatial metadata from a bitrate distribution control table;

determining, using the one or more processors, a combination of the one or more target bitrates for the one or more downmix channels;

determining, using the one or more processors, a metadata quantization level from the set of metadata quantization levels using a bitrate distribution process, wherein the bitrate distribution process adjusts at least one of the target bitrates or at least one of the metadata quantization levels of the spatial metadata based at least in part on a bitrate budget for the IVAS bitstream;

quantizing and coding, using the one or more processors, the spatial metadata using the metadata quantization level;

generating, using the one or more processors and the combination of one or more target bitrates, a downmix bitstream for the one or more downmix channels;

combining, using the one or more processors, the downmix bitstream, the quantized and coded spatial metadata and the coded set of metadata quantization levels into the IVAS bitstream; and

outputting, streaming or storing the IVAS bitstream for playback on an IVAS- enabled device.

2. The method of claim 1 , wherein the input audio signal is a four-channel first order Ambisonic (FoA) audio signal, three-channel planar FoA signal or a two-channel stereo audio signal.

3. The method of claim 1 , wherein the one or more target bitrates are bitrates of one or more instances of a mono audio coder/decoder (codec).

4. The method of claim 1 , wherein the mono audio codec is an enhanced voice services (EVS) codec and the downmix bitstream is an EVS bitstream.

5. The method of claim 1 , wherein obtaining, using the one or more processors, the set of one or more target bitrates for the one or more downmix channels and the set of metadata quantization levels for the spatial metadata using the bitrate distribution control table, further comprises:

identifying a row in the bitrate distribution control table using a table index that includes one or more of a format of the input audio signal, a bandwidth of the input audio signal, an allowed spatial coding tool, a transition mode and a mono downmix backward compatible mode; and

extracting from the identified row of the bitrate distribution control table, one or more of a target bitrate, a bitrate ratio, a minimum bitrate and bitrate deviation steps, wherein the bitrate ratio indicates a ratio in which a total bitrate is to be distributed between the downmix audio signal channels, the minimum bitrate is a value below which the total bitrate is not allowed to go and the bitrate deviation steps are target bitrate reduction steps when a first priority for the downmix signals is higher than or equal to, or lower, than a second priority of the spatial metadata; and

wherein determining the combination of the one or more bitrates for the one or more downmix channels and the spatial metadata is based on one or more of the target bitrate, the bitrate ratio, the minimum bitrate and the bitrate deviation steps.

6. The method of claim 1 , wherein quantizing and coding the spatial metadata for the one or more channels of the input audio signal using a the set of metadata quantization levels is performed in a quantization loop that applies increasingly coarse quantization strategies based on a difference between a target metadata bit rate and an actual metadata bitrate.

7. The method of claim 1 , wherein the quantization is determined in accordance with a mono codec priority and a spatial metadata priority based on properties extracted from the input audio signal and channel banded co-variance values.

8. The method of claim 1 , wherein the input audio signal is a stereo signal and the downmix signals include a representation of a mid-signal, residuals from the stereo signal and the spatial metadata.

9. The method of claim 1 , wherein the spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C) and decorrelation coefficients (P) for a spatial reconstructor (SPAR) format and prediction coefficients (PR) or decorrelation coefficients (PR) for complex advanced coupling (CACPL) format.

10. The method of claim 1 , wherein obtaining, using the one or more processors, the set of one or more target bitrates for the one or more downmix channels using the bitrate distribution control table, further comprises:

identifying a row in the bitrate distribution control table using a table index that includes one or more of a format of the input audio signal, a bandwidth of the input audio signal and a IVAS bitrate; and

extracting, from the identified row of the bitrate distribution control table, one or more of a target bitrate, a minimum bitrate, and a maximum bitrate for each of the one or more downmix channels, wherein the minimum bitrate and maximum bitrate define a bitrate range for the bitrate of the downmix channel, and wherein the target bitrate is a preferred bitrate for the downmix channel; and

computing, a total downmix bitrate by subtracting the metadata bitrate and IVAS header bitrate from the total IVAS bitrate; and

determining, the combination of the one or more bitrates for the one or more downmix channels based on one or more of the target bitrate, the minimum bitrate, the maximum bitrate, the total downmix bitrate and a priority assigned to the one or more downmix channels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2022
From: TYAGI, RISHABH; TORRES, JUAN FELIX; BROWN, STEFANIE
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 061097/0902 →
Continuity (3)
Provisional Application 63092830 · Oct 16, 2020
Provisional Application 62927772 · Oct 30, 2019
Related Publication 20220406318A1 · Dec 22, 2022
References Cited (34)
US 7573912B2 · Lindblom · 2009 [cited by applicant]
US 8442836B2 · Li · 2013 [cited by applicant]
US 8918636B2 · Kiefer · 2014 [cited by applicant]
US 9398337B2 · Lee · 2016 [cited by applicant]
US 9530422B2 · Klejsa · 2016 [cited by applicant]
US 10395664B2 · Tsingos · 2019 [cited by applicant]
US 10937435B2 · Fueg · 2021 [cited by examiner]
US 11096002B2 · Pihlajakuja · 2021 [cited by examiner]
US 20150340044A1 · Kim · 2015 [cited by applicant]
US 20170236521A1 · Chebiyyam et al. · 2017 [cited by applicant]
US 20190013028A1 · Atti · 2019 [cited by applicant]
US 20190103118A1 · Atti · 2019 [cited by applicant]
US 20190251986A1 · Niedermeier · 2019 [cited by applicant]
US 20190295559A1 · Atti · 2019 [cited by applicant]
US 20220279299A1 · Vasilache · 2022 [cited by examiner]
GB 2595891A · 2021 [cited by examiner]
JP 2008529056A · 2008 [cited by applicant]
JP 2016509260A · 2016 [cited by applicant]
RU 2616774C1 · 2017 [cited by applicant]
TW 201134135A · 2011 [cited by applicant]
TW 201907392A · 2019 [cited by applicant]
TW 201923744A · 2019 [cited by applicant]
WO WO2007016107A2 · 2007 [cited by examiner]
WO WO2013186345A1 · 2013 [cited by examiner]
WO WO2019023488A1 · 2019 [cited by examiner]
WO 2019056107A1 · 2019 [cited by applicant]
WO 2019068638A1 · 2019 [cited by applicant]
WO 2019105575A1 · 2019 [cited by applicant]
WO 2019106221A1 · 2019 [cited by applicant]
Dolby Laboratories Inc. “Dolby VRStream audio profile candidate - Description of Bitstream, Decoder, and Renderer plus informative Encoder Description,” Jul. 9-13, 2018, Rome, Italy, Jul. 2018 (Year: 2018). [cited by examiner]
Jürgen Herre, Senior Member, IEEE, Johannes Hilpert, Achim Kuntz, and Jan Plogsties, “Mpeg-H 3D Audio—The New Standard for Coding of Immersive Spatial Audio” IEEE Journal of Selected Topics in Signal Processing, vol. 9,… [cited by examiner]
Breebaart, J. et al.“MPEG Spatial Audio Coding/MPEG Surround: Overview and Current Status” presented at the 119th Convention, Oct. 7-10, 2005, New York, USA, pp. 1-17. [cited by applicant]
Dolby Laboratories Inc: Dolby VRStream audio profile candidate—Description of Bitstream, Decoder, and Renderer plus informative Encoder Description11 , 3gpp Draft; S4-180806—Dolby Vrstream Audio Candidate—Description of… [cited by applicant]
McGrath D. et al: Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec11 , ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASS… [cited by applicant]