IP Library Granted Patent US 12,666,218
Granted Patent B2
US 12,666,218 · App. 18/490,442 · Granted Jun 23, 2026

Generating parametric spatial audio representations

Inventors: Mikko-Ville Laitinen (Espoo, FI); Juha Tapio Vilkamo (Helsinki, FI); Jussi Kalevi Virolainen (Espoo, FI)
Assignee: Nokia Technologies Oy
H04S7/303G10L19/008H04S3/008H04S7/307H04S2400/01H04S2400/11H04S2400/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,666,218
App. No.
18/490,442
Granted
Jun 23, 2026
Kind
B2
Abstract

A method for generating a spatial audio stream, the method including: obtaining at least two audio signals from at least two microphones; extracting from the at least two audio signals a first audio signal, the first audio signal including at least partially speech of a user; extracting from the at least two audio signals a second audio signal, wherein speech of the user is substantially not present within the second audio signal; and encoding the first audio signal and the second audio signal to generate the spatial audio stream such that a rendering of speech of the user to a controllable direction and/or distance is enabled.

Claims (50)

1 . A method for generating a spatial audio stream, the method comprising:

obtaining at least two audio signals from at least two microphones, wherein the at least two microphones are configured to capture binaural audio, wherein the at least two microphones are mounted on earphones or headphones worn by a user;

extracting from the at least two audio signals a first audio signal, the first audio signal comprising at least partially speech of the user;

extracting from the at least two audio signals a second audio signal, wherein the speech of the user is substantially not present within the second audio signal; and

encoding the first audio signal and the second audio signal to generate the spatial audio stream such that a rendering of the speech of the user to at least one of a controllable direction or distance is enabled.

2 . The method as claimed in claim 1 , wherein the spatial audio stream further enables a controllable rendering of captured ambience audio content.

3 . The method as claimed in claim 1 , wherein extracting from the at least two audio signals the first audio signal further comprises applying a machine learning model to the at least two audio signals or at least one audio signal based on the at least two audio signals to generate the first audio signal.

4 . The method as claimed in claim 3 , wherein applying the machine learning model to the at least two audio signals or at least one audio signal based on the at least two audio signals to generate the first audio signal further comprises:

generating a first speech mask based on the at least two audio signals; and

separating the at least two audio signals into a mask processed speech audio signal and a mask processed remainder audio signal based on application of the first speech mask to the at least two audio signals or at least one audio signal based on the at least two audio signals.

5 . The method as claimed in claim 4 , wherein extracting from the at least two audio signals the first audio signal further comprises beamforming the at least two audio signals to generate a speech audio signal.

6 . The method as claimed in claim 5 , wherein beamforming the at least two audio signals to generate the speech audio signal comprises:

determining steering vectors for the beamforming based on the mask processed speech audio signal;

determining a remainder covariance matrix based on the mask processed remainder audio signal; and

applying a beamformer configured based on the steering vectors and the remainder covariance matrix to generate a beam audio signal.

7 . The method as claimed in claim 6 , wherein applying the machine learning model to the at least two audio signals or at least one audio signal based on the at least two audio signals to generate the first audio signal further comprises:

generating a second speech mask based on the beam audio signal; and

applying a gain processing to the beam audio signal based on the second speech mask to generate the speech audio signal.

8 . The method as claimed in claim 7 , wherein extracting from the at least two audio signals the second audio signal comprises:

generating a positioned speech audio signal from the speech audio signal; and

subtracting from the at least two audio signals the positioned speech audio signal to generate the at least one remainder audio signal.

9 . The method as claimed in claim 3 , wherein applying the machine learning model to the at least two audio signals or at least one signal based on the at least two audio signals to generate the first audio signal further comprises equalizing the first audio signal.

10 . The method as claimed in claim 1 , wherein extracting from the at least two audio signals the first audio signal comprising the speech of the user comprises:

generating the first audio signal based on the at least two audio signals; and

generating an audio object representation, the audio object representation comprising the first audio signal.

11 . The method as claimed in claim 10 , wherein extracting from the at least two audio signals the first audio signal further comprises:

analysing the at least two audio signals to determine at least one of a direction or position relative to the at least two microphones associated with the speech of the user, wherein the audio object representation further comprises at least one of the direction or position relative to the at least two microphones.

12 . The method as claimed in claim 10 , wherein generating the second audio signal further comprises generating binaural audio signals.

13 . The method as claimed in claim 1 , wherein encoding the first audio signal and the second audio signal to generate the spatial audio stream comprises:

mixing the first audio signal and the second audio signal to generate at least one transport audio signal;

determining at least one directional or positional spatial parameter associated with a desired direction or position of the speech of the user; and

encoding the at least one transport audio signal and the at least one directional or positional spatial parameter to generate the spatial audio stream.

14 . The method as claimed in claim 13 , further comprising obtaining an energy ratio parameter, and wherein encoding the at least one transport audio signal and the at least one directional or positional spatial parameter comprises further encoding the energy ratio parameter.

15 . The method as claimed in claim 1 , wherein the first audio signal is a single channel audio signal.

16 . The method as claimed in claim 1 , wherein the at least two microphones are located in an audio scene comprising the user as a first audio source and a second audio source, and the method further comprises:

extracting from the at least two audio signals the first audio signal, the first audio signal comprising at least partially the second audio source; and

extracting from the at least two audio signals at the second audio signal, wherein sound of the first audio source is further first substantially not present within the second audio signal, or the sound of the first audio source is within the second audio signal.

17 . The method as claimed in claim 16 , wherein the first audio source is a talker and the second audio source is a further talker.

18 . An apparatus for generating a spatial audio stream, the apparatus comprising:

at least one processor; and

at least one memory storing instructions that, when executed with the at least one processor, cause the apparatus at least to:

obtain at least two audio signals from at least two microphones, wherein the at least two microphones are configured to capture binaural audio, wherein the at least two microphones are mounted on earphones or headphones worn by a user;

extract from the at least two audio signals a first audio signal, the first audio signal comprising at least partially speech of the user;

extract from the at least two audio signals a second audio signal, wherein the speech of the user is substantially not present within the second audio signal; and

encode the first audio signal and the second audio signal to generate the spatial audio stream such that a rendering of the speech of the user to at least one of a controllable direction or distance is enabled.

19 . A non-transitory computer readable medium comprising instructions that, when executed by an apparatus, cause the apparatus to perform at least the following:

obtaining at least two audio signals from at least two microphones, wherein the at least two microphones are configured to capture binaural audio, wherein the at least two microphones are mounted on earphones or headphones worn by a user;

extracting from the at least two audio signals a first audio signal, the first audio signal comprising at least partially speech of the user;

extracting from the at least two audio signals a second audio signal, wherein the speech of the user is substantially not present within the second audio signal; and

encoding the first audio signal and the second audio signal to generate a spatial audio stream such that a rendering of the speech of the user to at least one of a controllable direction or distance is enabled.