IP Library Granted Patent US 9,530,421
Granted Patent B2
US 9,530,421 · App. 14/026,984 · Granted Dec 27, 2016

Encoding and reproduction of three dimensional audio soundtracks

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,530,421
App. No.
14/026,984
Granted
Dec 27, 2016
Kind
B2
Abstract

The present invention provides a novel end-to-end solution for creating, encoding, transmitting, decoding and reproducing spatial audio soundtracks. The provided soundtrack encoding format is compatible with legacy surround-sound encoding formats, so that soundtracks encoded in the new format may be decoded and reproduced on legacy playback equipment with no loss of quality compared to legacy formats.

Claims (62)

1. A method of encoding an audio soundtrack, comprising the steps of:

receiving a base mix signal representing a physical sound;

receiving at least one object audio signal, each object audio signal having at least one audio object component of the audio soundtrack;

receiving at least one object mix cue stream, the object mix cue streams defining mixing parameters of the object audio signals;

receiving at least one object render cue stream, the object render cue streams defining rendering parameters for rendering the object audio signals in a target spatial audio format;

encoding the object audio signals by a first audio encoding processor to obtain encoded object audio signals that contain encoded audio objects;

decoding the encoded object audio signals by a first audio decoding processor;

utilizing the decoded object audio signals and the object mix cue streams to combine the audio object components with the base mix signal, thereby obtaining a downmix signal; and

multiplexing the downmix signal, the encoded object audio signals, the object render cue streams, and the object mix cue streams to form a soundtrack data stream.

2. The method of claim 1 , wherein the downmix signal is encoded by a second audio encoding processor before being multiplexed.

3. The method of claim 2 , wherein the second audio encoding processor is a lossy digital encoding processor.

4. A method of decoding an audio soundtrack, representing a physical sound, comprising the steps of:

receiving a soundtrack data stream, having:

a downmix signal representing an audio scene;

at least one object audio signal, the object audio signals having at least one audio object component of the audio soundtrack;

at least one object mix cue stream, the object mix cue streams defining mixing parameters of the object audio signals; and

at least one object render cue stream, the object render cue streams defining rendering parameters for rendering the object audio signals in a target spatial audio format;

utilizing the object audio signals and the object mix cue streams to substantially remove at least one audio object component from the downmix signal, thereby obtaining a residual downmix signal;

applying a spatial format conversion to the residual downmix signal, thereby outputting a converted residual downmix signal, wherein the spatial format conversion utilizes spatial parameters determined by the target spatial audio format;

utilizing the object audio signals and the object render cue streams to derive at least one object rendering signal; and

combining the converted residual downmix signal and the object rendering signal to obtain a soundtrack rendering signal.

5. The method of claim 4 , wherein the audio object component is subtracted from the downmix signal.

6. The method of claim 4 , wherein the audio object component is substantially removed from the downmix signal such that the audio object component is unnoticeable in the downmix signal.

7. The method of claim 4 , wherein the downmix signal is an encoded audio signal.

8. The method of claim 7 , wherein the downmix signal is decoded by an audio decoder.

9. The method of claim 4 , wherein the object audio signals are mono audio signals.

10. The method of claim 4 , wherein the object audio signals are multi-channel audio signals having at least 2 channels.

11. The method of claim 4 , wherein the object audio signals are discrete loudspeaker-feed audio channels.

12. The method of claim 4 , wherein the audio object components are voices, instruments, or sound effects of the audio scene.

13. The method of claim 4 , wherein the spatial audio format represents a listening environment.

14. An audio encoding processor, comprising:

a receiver processor for receiving:

a base mix signal representing a physical sound;

at least one object audio signal, each object audio signal having at least one audio object component of the audio soundtrack;

at least one object mix cue stream, the object mix cue streams defining mixing parameters of the object audio signals; and

at least one object render cue stream, the object render cue streams defining rendering parameters for rendering the object audio signals in a target spatial audio format;

a first audio encoding processor for encoding the object audio signals to obtain encoded object audio signals that contain encoded audio objects;

a first audio decoding processor for decoding the encoded object audio signals;

a combining processor for combining the audio object components with the base mix signal based on the decoded object audio signals and the object mix cue streams, the combining processor outputting a downmix signal; and

a multiplexer processor for multiplexing the downmix signal, the encoded object audio signals, the object render cue streams, and the object mix cue streams to form a soundtrack data stream.

15. The audio encoding processor of claim 14 , wherein the downmix signal is encoded by a second audio encoding processor before being multiplexed.

16. An audio decoding processor, comprising:

a receiving processor for receiving:

a downmix signal representing an audio scene;

at least one object audio signal, the object audio signal having at least one audio object component of the audio scene;

at least one object mix cue stream, the object mix cue streams defining mixing parameters of the object audio signals; and

at least one object render cue stream, the object render cue stream defining rendering parameters for rendering the object audio signals in a target spatial format;

an object audio processor for substantially removing at least one audio object component from the downmix signal based on the object audio signals and the object mix cue streams, and outputting a residual downmix signal;

a spatial format converter for applying a spatial format conversion to the residual downmix signal, thereby outputting a converted residual downmix signal, wherein the spatial format converter utilizes spatial parameters determined by the target spatial audio format;

a rendering processor for processing the object audio signals and the object render cue streams to derive at least one object rendering signal; and

a combining processor for combining the converted residual downmix signal and the object rendering signal to obtain a soundtrack rendering signal.

17. The audio decoding processor of claim 16 , wherein the audio object component is subtracted from the downmix signal.

18. The audio decoding processor of claim 16 , wherein the audio object component is partially removed from the downmix signal such that the audio object component is unnoticeable in the downmix signal.

19. A method of decoding an audio soundtrack, representing a physical sound, comprising the steps of:

receiving a soundtrack data stream, having:

a downmix signal representing an audio scene;

at least one object audio signal, the object audio signal having at least one audio object component of the audio soundtrack; and

at least one object render cue stream, the object render cue stream defining rendering parameters for rendering the object audio signals in a target spatial format;

utilizing the object audio signals and the object render cue streams to substantially remove at least one audio object component from the downmix signal, thereby obtaining a residual downmix signal;

applying a spatial format conversion to the residual downmix signal, thereby outputting a converted residual downmix signal, wherein the spatial format converter utilizes spatial parameters determined by the target spatial audio format;

utilizing the object audio signals and the object render cue streams to derive at least one object rendering signal; and

combining the converted residual downmix signal and the object rendering signal to obtain a soundtrack rendering signal.

Assignments (6)
PARTIAL RELEASE OF SECURITY INTEREST IN PATENTS Recorded Oct 27, 2022
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: VEVEO LLC (F.K.A. VEVEO, INC.); DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
Reel/Frame 061786/0675 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2020
From: ROYAL BANK OF CANADA
To: TESSERA, INC.; INVENSAS BONDING TECHNOLOGIES, INC. (F/K/A ZIPTRONIX, INC.); FOTONATION CORPORATION (F/K/A DIGITALOPTICS CORPORATION AND F/K/A DIGITALOPTICS CORPORATION MEMS); INVENSAS CORPORATION; TESSERA ADVANCED TECHNOLOGIES, INC; DTS, INC.; DTS LLC; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
Reel/Frame 052920/0001 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
RELEASE OF SECURITY INTEREST Recorded Dec 6, 2016
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: DTS, INC.
Reel/Frame 040821/0083 →
SECURITY INTEREST Recorded Dec 2, 2016
From: INVENSAS CORPORATION; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; ZIPTRONIX, INC.; DIGITALOPTICS CORPORATION; DIGITALOPTICS CORPORATION MEMS; DTS, LLC; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 040797/0001 →
SECURITY INTEREST Recorded Nov 2, 2015
From: DTS, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS ADMINISTRATIVE AGENT
Reel/Frame 037032/0109 →