IP Library Granted Patent US 12,243,540
Granted Patent B2
US 12,243,540 · App. 17/786,088 · Granted Mar 4, 2025

Merging of spatial audio parameters

Inventors: Mikko-Ville Laitinen (Espoo, FI); Lasse Laaksonen (Tampere, FI); Adriana Vasilache (Tampere, FI); Tapani Pihlajakuja (Kellokoski, FI); Anssi Rämö (Tampere, FI)
Assignee: NOKIA TECHNOLOGIES OY
G10L19/008H04S7/302H04S2420/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,540
App. No.
17/786,088
Granted
Mar 4, 2025
Kind
B2
Abstract

There is inter alia disclosed an apparatus for spatial audio encoding comprising: means for determining at least two of a type of spatial audio parameter for one or more audio signals, wherein a first of the type of spatial audio parameter is associated with a first group of samples in a domain of the one or more audio signals and a second of the type of spatial audio parameter is associated with a second group of samples in the domain of the one or more audio signals; and means for merging the first of the type of spatial audio parameter and the second of the type of spatial audio parameter into a merged spatial audio parameter.

Claims (78)

1. An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to:

determine or receive at least two of a type of spatial audio parameter for one or more audio signals, wherein a first of the type of spatial audio parameter is associated with a first group of samples in a domain of the one or more audio signals and a second of the type of spatial audio parameter is associated with a second group of samples in the domain of the one or more audio signals;

merge the first of the type of spatial audio parameter and the second of the type of spatial audio parameter into a merged spatial audio parameter; and

at least one of store or transmit an encoded representation of at least one of the first of the type of spatial audio parameter, the second of the type of spatial audio parameter or the merged spatial audio parameter.

2. The apparatus as claimed in claim 1 , wherein the apparatus is further caused to:

determine whether the merged spatial audio parameter is encoded for at least one of storage or transmission; or

determine whether the at least two of the type of spatial audio parameter is encoded for at least one of storage or transmission.

3. The apparatus as claimed in claim 2 , wherein the apparatus is further caused to:

determine a metric for the first group of samples and the second group of samples; and

compare the metric against a threshold value;

wherein when the metric is above the threshold value the apparatus is caused to determine that the at least two of the type of spatial audio parameter is encoded for at least one storage or transmission; and

wherein when the metric is below or equal to the threshold value the apparatus is caused to determine that the merged spatial audio parameter band is encoded for at least one of storage or transmission.

4. The apparatus as claimed in claim 1 , wherein the apparatus is further caused to:

determine a metric for the first group of samples and the second group of samples;

determine a further at least two of a type of spatial audio parameter for the one or more audio signals, wherein a further first of the type of spatial audio parameter is associated with a first further group of samples in the domain of the one or more audio signals and a further second of the type of spatial audio parameter is associated with a second further group of samples in the domain of the one or more audio signals;

merge the further first of the type of spatial audio parameter and the further second of the type of spatial audio parameter into a further merged spatial audio parameter;

determine a metric for the first further group of samples and second further group of samples; and

determine that the further first of the type of spatial audio parameter and the further second of the type of spatial audio parameter are encoded for at least one of storage or transmission and the merged spatial audio parameter is encoded for at least one of storage or transmission when the metric for the first further group of samples and second further group of samples is higher than the metric for the first group of samples and the second group of samples.

5. The apparatus as claimed in claim 1 , wherein the apparatus is further caused to determine an energy of the first group of samples of the one or more audio signals and an energy of the second group of samples of the one or more audio signals, wherein the value of the merged spatial audio parameter is based on the energy of the first group of samples and the energy of the second group of samples.

6. The apparatus as claimed in claim 5 , wherein the type of spatial audio parameter comprises a spherical direction vector and wherein the merged spatial audio parameter comprises a merged spherical direction vector, and wherein to merge the first of the type of spatial audio parameter and the second of the type of spatial audio parameter into the merged spatial audio parameter, the apparatus is caused to:

convert a first spherical direction vector into a first cartesian vector converting a second spherical direction vector into a second cartesian vector, wherein the first cartesian direction vector and second cartesian direction vector each comprise an x-axis component, y-axis component and a z-axis component, and wherein for each component the apparatus is caused to;

weight the component of the first cartesian vector by the energy of the first group of samples of the one or more audio signals and a direct to total energy ratio calculated for the first group of samples of the one or more audio signals;

weight the component of the second cartesian vector by the energy of the second group of samples of the one or more audio signals and a direct to total energy ratio calculated for the second group of samples of the one or more audio signals;

sum, the weighted component of the first cartesian vector and the weighted respective component of the second cartesian vector to give a merged respective cartesian component vector; and

convert the merged cartesian x-axis component value, the merged cartesian y-axis component value and the merged cartesian z-axis component value into the merged spherical direction vector.

7. The apparatus as claimed in claim 6 , wherein the apparatus is further caused to merge the direct to total energy ratio for the first group of samples of the one or more audio signals and the direct to total energy ratio of the second group of samples of the one or more audio signals into a merged direct to total energy ratio, by being caused to determine the length of the merged cartesian vector; and

normalize the length of the merged cartesian vector by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of the one or more audio signals.

8. The apparatus as claimed in claim 6 , wherein the apparatus caused to determine a metric, is caused to:

determine a sum of the length of the first cartesian vector and the length of the second cartesian vector; and

determine a difference between the length of the merged cartesian vector and the sum.

9. The apparatus as claimed in claim 5 , wherein the apparatus is further caused to:

determine a first spread coherence parameter associated with the first group of samples in the domain of the one or more audio signals and a second spread coherence parameter associated with the second group of samples in the domain of the one or more audio signals; and

merge the first spread coherence parameter and the second spread coherence parameter into a merged spread coherence parameter, and wherein to merge the first spread coherence parameter and the second spread coherence parameter into a merged spread coherence parameter, the apparatus is caused to:

weight a first spread coherence value by the energy of the first group of samples of the one or more audio signals;

weight a second spread coherence value by the energy of the second group of samples of the one or more audio;

sum the weighted first spread coherence value and the weighted second spread coherence value to give a merged spread coherence value; and

normalise the merged spread coherence value by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of the one or more audio signals.

10. The apparatus as claimed in claim 5 , wherein the apparatus is further caused to:

determine a first surround coherence parameter associated with the first group of samples in the domain of the one or more audio signals and a second surround coherence parameter associated with the second group of samples in the domain of the one or more audio signals; and

merge the first surround coherence parameter and the second surround coherence parameter into a merged surround coherence parameter, and

wherein to merge the first surround coherence parameter and the second surround coherence parameter into a merged surround coherence parameter, the apparatus is caused to:

weight the first surround coherence value by the energy of the first group of samples of the one or more audio signals;

weight the second surround coherence value by the energy of the second group of samples of the one or more audio;

sum, the weighted first surround coherence value and the weighted second surround coherence value to give the merged spread coherence value; and

normalise the merged surround coherence value by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of the one or more audio signals.

11. The apparatus as claimed in claim 1 , wherein the apparatus is further caused to:

determine a first spread coherence parameter associated with the first group of samples in the domain of the one or more audio signals and a second spread coherence parameter associated with the second group of samples in the domain of the one or more audio signals; and

merge the first spread coherence parameter and the second spread coherence parameter into a merged spread coherence parameter.

12. The apparatus as claimed in claim 1 , wherein the apparatus is further caused to:

determine a first surround coherence parameter associated with the first group of samples in the domain of the one or more audio signals and a second surround coherence parameter associated with the second group of samples in the domain of the one or more audio signals; and

merge the first surround coherence parameter and the second surround coherence parameter into a merged surround coherence parameter.

13. The apparatus as claimed in claim 1 , wherein the first group of samples is a first subframe in the time domain and the second group of samples is a second subframe in the time domain.

14. The apparatus as claimed in claim 1 , wherein the first group of samples is a first sub band in the frequency domain and the second group of samples is a second sub band in the frequency domain.

15. A method comprising:

determining or receiving at least two of a type of spatial audio parameter for one or more audio signals, wherein a first of the type of spatial audio parameter is associated with a first group of samples in a domain of the one or more audio signals and a second of the type of spatial audio parameter is associated with a second group of samples in the domain of the one or more audio signals;

merging the first of the type of spatial audio parameter and the second of the type of spatial audio parameter into a merged spatial audio parameter; and

at least one of storing or transmitting an encoded representation of at least one of the first of the type of spatial audio parameter, the second of the type of spatial audio parameter or the merged spatial audio parameter.

16. The method as claimed in claim 15 , wherein the method further comprises:

determining whether the merged spatial audio parameter is encoded for at least one of storage or transmission; or

determining whether the at least two of the type of spatial audio parameter is encoded for at least one of storage or transmission.

17. The method as claimed in claim 16 , wherein the method further comprises:

determining a metric for the first group of samples and the second group of samples; and

comparing the metric against a threshold value,

wherein when the metric is above the threshold value the method comprises determining that the at least two of the type of spatial audio parameter is encoded for at least one of storage or transmission; and

wherein when the metric is below or equal to the threshold value then determining that the merged spatial audio parameter band is encoded for at least one of storage or transmission.

18. The method as claimed in claim 15 , wherein the method further comprises:

determining a metric for the first group of samples and the second group of samples;

determining a further at least two of a type of spatial audio parameter for the one or more audio signals, wherein a further first of the type of spatial audio parameter is associated with a first further group of samples in the domain of the one or more audio signals and a further second of the type of spatial audio parameter is associated with a second further group of samples in the domain of the one or more audio signals;

merging the further first of the type of spatial audio parameter and the further second of the type of spatial audio parameter into a further merged spatial audio parameter;

determining a metric for the first further group of samples and second further group of samples; and

determining that the further first of the type of spatial audio parameter and the further second of the type of spatial audio parameter are encoded for at least one of storage or transmission and the merged spatial audio parameter is encoded for at least one of storage or transmission when the metric for the first further group of samples and second further group of samples is higher than the metric for the first group of samples and the second group of samples.

19. The method as claimed in claim 15 , wherein the method further comprises determining an energy of the first group of samples of the one or more audio signals and an energy of the second group of samples of the one or more audio signals, wherein the value of the merged spatial audio parameter is based on the energy of the first group of samples and the energy of the second group of samples.

20. The method as claimed in claim 19 , wherein the type of spatial audio parameter comprises a spherical direction vector and wherein the merged spatial audio parameter comprises a merged spherical direction vector, and wherein merging the first of the type of spatial audio parameter and the second of the type of spatial audio parameter into the merged spatial audio parameter comprises:

converting a first spherical direction vector into a first cartesian vector converting a second spherical direction vector into a second cartesian vector, wherein the first cartesian direction vector and second cartesian direction vector each comprise an x-axis component, y-axis component and a z-axis component, and wherein for each component in turn the method comprises:

weighting the component of the first cartesian vector by the energy of the first group of samples of the one or more audio signals and a direct to total energy ratio calculated for the first group of samples of the one or more audio signals;

weighting the component of the second cartesian vector by the energy of the second group of samples of the one or more audio signals and a direct to total energy ratio calculated for the second group of samples of the one or more audio signals;

summing, the weighted component of the first cartesian vector and the weighted respective component of the second cartesian vector to give a merged respective cartesian component vector; and

converting the merged cartesian x-axis component value, the merged cartesian y-axis component value and the merged cartesian z-axis component value into the merged spherical direction vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2022
From: LAITINEN, MIKKO-VILLE ILARI; VASILACHE, ADRIANA; RÄMÖ, ANSSI SAKARI; LAAKSONEN, LASSE JUHANI; PIHLAJAKUJA, TAPANI JOHANNES
To: NOKIA TECHNOLOGIES OY
Reel/Frame 060225/0165 →
Priority Claims (1)
GB 1919130 · Dec 23, 2019 · national
Continuity (1)
Related Publication 20230197086A1 · Jun 22, 2023
References Cited (35)
US 9489955B2 · Peters et al. · 2016 [cited by applicant]
US 20090234657A1 · Takagi · 2009 [cited by examiner]
US 20120314877A1 · Ojala · 2012 [cited by applicant]
US 20130177887A1 · Hoehne · 2013 [cited by examiner]
US 20140025386A1 · Xiang et al. · 2014 [cited by applicant]
US 20140297296A1 · Koppens · 2014 [cited by examiner]
US 20150170657A1 · Thompson · 2015 [cited by examiner]
US 20160078877A1 · Vasilache et al. · 2016 [cited by applicant]
EP 2600343A1 · 2013 [cited by applicant]
GB 2574238A · 2019 [cited by applicant]
WO 2014099285A1 · 2014 [cited by applicant]
WO 2014187991A1 · 2014 [cited by applicant]
WO 2016133785A1 · 2016 [cited by applicant]
WO 2017005978A1 · 2017 [cited by applicant]
WO 2018091776A1 · 2018 [cited by applicant]
WO 2019086757A1 · 2019 [cited by applicant]
WO 2019097018A1 · 2019 [cited by applicant]
WO 2019234290A1 · 2019 [cited by applicant]
WO 2020008105A1 · 2020 [cited by applicant]
WO 2020070377A1 · 2020 [cited by applicant]
WO 2020089510A1 · 2020 [cited by applicant]
WO 2020193865A1 · 2020 [cited by applicant]
WO 2021048468A1 · 2021 [cited by applicant]
B. Wu and L. Gao, “Downmix and coding of multichannel signals based on spatial correlation,” 2015 8th International Congress on Image and Signal Processing (CISP), Shenyang, China, 2015, pp. 1142-1146, doi: 10.1109/CISP… [cited by examiner]
B. Wu and L. Gao, “Downmix and coding of multichannel signals based on spatial correlation,” 2015 8th International Congress on Image and Signal Processing (CISP), Shenyang, China, 2015, pp. 1142-1146, doi: 10.1109/CISP… [cited by examiner]
J. Capobianco, G. Pallone and L. Daudet, “Dynamic strategy for window splitting, parameters estimation and interpolation in spatial parametric audio coders,” 2012 IEEE International Conference on Acoustics, Speech and S… [cited by examiner]
Extended European Search Report received for corresponding European Patent Application No. 20907123.2, dated Dec. 1, 18, 2023, 10 pages. [cited by applicant]
Office action received for corresponding Indian Patent Application No. 202247041316, dated Oct. 18, 2022, 5 pages. [cited by applicant]
“Proposal for MASA common metadata and metadata structure”, 3GPP TSG-SA4#101 meeting, S4-181353, Agenda: 7.5, Nokia Corporation, Nov. 19-23, 2018, pp. 1-4. [cited by applicant]
Search Report received for corresponding United Kingdom Patent Application No. 1919130.3, dated Jun. 18, 2020, 5 pages. [cited by applicant]
International Search Report and Written Opinion received for corresponding Patent Cooperation Treaty Application No. PCT/FI2020/050750, dated Feb. 23, 2021, 15 pages. [cited by applicant]
Tsingos et al., “Perceptual audio rendering of complex virtual environments”, ACM Transactions on Graphics, vol. 23, No. 3, Aug. 2004, pp. 249-258. [cited by applicant]
Yang et al., “Multi-channel Object-Based Spatial Parameter Compression Approach for 3D Audio”, Advances in Multimedia Information Processing, 2015, pp. 354-364. [cited by applicant]
Kazakova et al., “Iterative weighted 2D orientation averaging that minimizes arc-length between vectors”, IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 24-28, 2017, pp. 2499-2504. [cited by applicant]
“Proposal for MASA format”, 3GPP TSG-SA4#102 meeting, S4-190121, Agenda: 7.5, Nokia Corporation, Jan. 28-Feb. 1, 2019, pp. 1-10. [cited by applicant]