IP Library Granted Patent US 12706102
Granted Patent B2
US 12706102 · App. 18/554,234 · Granted Aug 11, 2026

Separating spatial audio objects

Inventors: Mikko-Ville Laitinen (Espoo, FI); Anssi Sakari Rämö (Tampere, FI)
Assignee: NOKIA TECHNOLOGIES OY
G10L19/008G10L25/21
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12706102
App. No.
18/554,234
Granted
Aug 11, 2026
Kind
B2
Abstract

There is inter alia disclosed an apparatus for spatial audio encoding configured to: determine an audio object for separation ( 306 ) from a plurality of audio objects of an audio frame ( 1281 ); separate the audio object for separation ( 308 ) from the plurality of audio objects to provide a separated audio object ( 126 ) and at least one remaining audio object ( 124 ); encode the separated audio object with an audio object encoder; and encode the plurality of remaining audio objects together with another input audio format.

Claims (51)

1 . A method for spatial audio signal encoding comprising:

determining an energy of each of a plurality of audio object signals over an audio frame to give a plurality of audio object signal energies, wherein each audio object signal is associated with an audio object of a plurality of audio objects of the audio frame;

determining an energy of at least one audio signal of another input audio format over the audio frame;

determining a loudest audio object by selecting an audio object with an audio object signal energy which is a largest audio object signal energy from the plurality of audio object signal energies;

determining an energy proportion factor;

determining a threshold value for the audio frame according to the energy proportion factor;

determining a ratio of the audio object signal energy of the loudest audio object to an energy of a separated audio object signal for a previous audio frame, wherein the energy of the separated audio object signal for the previous audio frame is calculated over the audio frame, and wherein the separated audio object signal for the previous audio frame is associated with an audio object separated in the previous frame;

comparing the ratio of the audio object signal energy of the loudest audio object to the energy of the separated audio object signal for the previous audio frame against the threshold value;

depending on the comparison identifying for the audio frame, either the loudest audio object as an audio object for separation, or the audio object separated in the previous frame as the audio object for separation;

separating the audio object for separation from the plurality of audio objects of the audio frame to provide a separated audio object and at least one remaining audio object for the audio frame;

encoding the separated audio object with an audio object encoder; and

combining the encoding of the at least one remaining audio object together with the encoding of the another input audio format.

2 . The method as claimed in claim 1 , wherein determining the energy proportion factor comprises:

determining a total energy by summing the energy of each of the plurality of audio object signals over the audio frame, the energy of each of a plurality of audio object signals over the previous audio frame, the energy of the at least one audio signal of the other audio input format over the audio frame and the energy of the at least one audio signal of the other audio input format over the previous audio frame; and

determining the ratio of a sum energy to the total energy, wherein the sum energy is the sum of the loudest audio object signal energy of the loudest audio object, an audio object signal energy of a loudest audio object from the previous audio frame, the energy of the separated audio object signal for the previous audio frame and an energy of the separated audio object signal for the previous audio frame calculated over the previous audio frame.

3 . The method as claimed in claim 1 , wherein the method further comprises determining a manner of transition by which a change from the audio object separated in the previous audio frame to the separated audio object for the audio frame is performed.

4 . The method as claimed in claim 3 , wherein determining the manner of transition comprises:

comparing the energy proportion factor against a further threshold;

determining that the manner of transition from the audio object separated in the previous audio frame to a separated audio object for the audio frame is performed using a hard transition when the energy proportion factor is less than the threshold; and

determining that the manner of transition from the audio object separated in the previous audio frame to the separated audio object for the audio frame is performed using a fade out fade in transition when the energy proportion factor is greater than or equal to the threshold.

5 . An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to:

determine an energy of each of a plurality of audio object signals over an audio frame to give a plurality of audio object signal energies, wherein each audio object signal is associated with an audio object of a plurality of audio objects of the audio frame;

determine an energy of at least one audio signal of another input audio format over the audio frame;

determine a loudest audio object by selecting an audio object with a largest audio object signal energy which is a largest audio object signal energy from the plurality of audio object signal energies;

determine an energy proportion factor;

determine a threshold value for the audio frame according to the energy proportion factor;

determine a ratio of the audio object signal energy of the loudest audio object to an energy of a separated audio object signal for a previous audio frame, wherein the energy of the separated audio object signal for the previous audio frame is calculated over the audio frame, and wherein the separated audio object signal for the previous audio frame is associated with an audio object separated in the previous frame;

compare the ratio of the loudest audio object signal energy of the loudest audio object to the energy of the separated audio object signal for the previous audio frame against the threshold value;

depending on the comparison identify for the audio frame, either the loudest audio object as an audio object for separation, or the audio object separated in the previous frame as the audio object for separation;

separate the audio object for separation from the plurality of audio objects of the audio frame to provide a separated audio object and at least one remaining audio object for the audio frame;

encode the separated audio object with an audio object encoder; and

combine the encoding of the at least one remaining audio object together with the encoding of the another input audio format.

6 . The apparatus as claimed in claim 5 , wherein the apparatus caused to determine the energy proportion factor is caused to:

determine a total energy by summing the energy of each of the plurality of audio object signals over the audio frame, the energy of each of a plurality of audio object signals over the previous audio frame, the energy of the at least one audio signal of the other audio input format over the audio frame and the energy of the at least one audio signal of the other audio input format over the previous audio frame; and

determine the ratio of a sum energy to the total energy, wherein the sum energy is the sum of the audio object signal energy of the loudest audio object, an audio object signal energy of a loudest audio object from the previous audio frame, the energy of the separated audio object signal for the previous audio frame calculated over the audio frame and an energy of the separated audio object for the previous audio frame calculated over the previous audio frame.

7 . The apparatus as claimed in claim 5 , wherein the apparatus is further caused to determine a manner of transition by which a change from the audio object separated in the previous audio frame to the separated audio object for the audio frame is performed.

8 . The apparatus as claimed in claim 7 , wherein the apparatus caused to determine the manner of transition is caused to:

compare the energy proportion factor against a further threshold;

determine that the manner of transition from the audio object separated in the previous audio frame to a separated audio object for the audio frame is performed using a hard transition when the energy proportion factor is less than the threshold; and

determine that the manner of transition from the audio object separated in the previous audio frame to the separated audio object for the audio frame is performed using a fade out fade in transition when the energy proportion factor is greater than or equal to the threshold.

9 . The apparatus as claimed in claim 5 , wherein the apparatus is further caused to:

provide a separated audio object for at least one following audio frame and a plurality of remaining audio objects for the at least one following audio frame, wherein that least one following audio frame follows the audio frame; set the audio object signal of the separated audio object for the audio frame as the audio object signal of the audio frame of the separated audio object for the previous audio frame multiplied by a fading out window function;

set an audio object signal of the separated audio object for the at least one following audio frame as the audio object signal of the at least one following audio frame of the audio object for separation multiplied by a fading in window function;

set an audio object signal corresponding to the separated audio object for the previous audio frame within the at least one remaining audio object for the audio frame as the audio object signal for the audio frame of the separated audio object from the previous audio multiplied by a fading in window function; and

set an audio object signal corresponding to the separated audio object for the audio frame within the at least one remaining audio object for the at least one following audio frame as the audio object signal of the audio object for separation multiplied by a fading out window function.

10 . The apparatus as claimed in claim 9 , wherein the manner of transition from the separated audio object signal for the previous audio frame to a separated audio object signal for the audio frame is performed using the fade in fade out transition.

11 . The apparatus as claimed in claim 9 , wherein the fading out window function is a latter half of a Hann window function and wherein the fading in window function is one minus the latter half of the Hann window function.

12 . The apparatus as claimed in claim 5 , wherein the apparatus caused to determine the energy of each of the plurality of audio object signals over an audio frame is further caused to smooth the energy of each of the plurality of audio object signals by using an energy of a corresponding audio object signal from a previous audio frame.

13 . The apparatus as claimed in claim 5 , wherein the other input audio format comprises at least one of:

at least one audio signal and an input audio format metadata set; and

at least two audio signals.