IP Library Granted Patent US 12,483,849
Granted Patent B2
US 12,483,849 · App. 17/910,990 · Granted Nov 25, 2025

Rendering of audio objects with a complex shape

Inventors: Tommy Falk (Spånga, SE); Werner De Bruijn (Stockholm, SE)
Assignee: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
H04S7/302G06T19/006H04S2400/11H04S2420/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,483,849
App. No.
17/910,990
Granted
Nov 25, 2025
Kind
B2
Abstract

A method ( 900 ) for representing an audio object in an extended reality scene. The method includes obtaining (s 902 ) first metadata describing a first three-dimensional (3D) shape associated with the audio object. The method also includes transforming (s 904 ) the obtained first metadata to produce transformed metadata describing a two-dimensional (2D) plane or a one-dimensional (1D) line, wherein the 2D plane or the 1D line represent at least a portion of the audio object.

Claims (74)

1 . A method for representing an audio object with respect to a listening position of a listener in an extended reality scene, the method comprising:

obtaining first metadata describing a first three-dimensional (3D) shape associated with the audio object; and

transforming the obtained first metadata, which describes the first 3D shape associated with the audio object, to produce transformed metadata describing a two-dimensional (2D) plane or a one-dimensional (1D) line, wherein

the 2D plane or the 1D line represent at least a portion of the audio object,

transforming the obtained first metadata to produce the transformed metadata comprises determining an anchor point, the anchor point being a point on the 3D shape that is deemed closest to the listening position of the listener in the extended reality scene,

the 2D plane or 1D lines passes through the anchor point,

transforming the obtained first metadata to produce the transformed metadata further comprises:

determining a first point on the first 3D shape that represents a first edge of the first 3D shape with respect to the listening position of the listener,

determining a second point on the first 3D shape that represents a second edge of the first 3D shape with respect to the listening position of the listener, and

determining a first property of the 2D plane or 1D line based on the first and second points, and

the method further comprises:

prior to determining the first point on the first 3D shape and the second point on the first 3D shape, determining a line between the listening position of the listener and the anchor point; and

using the determined line to determine the first point on the first 3D shape and the second point on the first 3D shape, wherein

the first point on the first 3D shape has a highest azimuth angle relative to the line between the listening position of the listener and the anchor point, and

the second point on the first 3D shape has a lowest azimuth angle relative to the line between the listening position of the listener and the anchor point.

2 . The method of claim 1 , wherein

obtaining the first metadata comprises selecting the first metadata from a set of metadata comprising the first metadata that describes the first 3D shape associated with the audio object and second metadata that describes a second 3D shape associated with the audio object, wherein the second 3D shape is different than the first 3D shape.

3 . The method of claim 2 , wherein the second 3D shape is in the form of one of the following:

a mesh structure with fewer vertices than the first 3D shape described by the first metadata,

a box shape,

a sphere shape,

an ellipsoid shape,

a cylinder shape.

4 . The method of claim 2 , wherein the set of metadata further comprises third metadata that describes a single point in the extended reality scene.

5 . The method of claim 2 , wherein the selection is based on one or more of the following:

a distance between the audio object and the listening position of the listener in the extended reality scene and at least one threshold distance parameter,

a size of the audio object,

the number of currently active audio objects in the extended reality scene,

a current load of a renderer that will be used to render the audio object, or

a current audio energy level of the audio object.

6 . The method of claim 1 , wherein determining the first property of the 2D plane or 1D line based on the first and second points comprises determining a first dimension of the 2D plane or 1D line based on the first and second points.

7 . The method of claim 6 , wherein transforming the obtained first metadata to produce the transformed metadata further comprises:

determining a third point on the first 3D shape that represents a third edge of the first 3D shape with respect to the listening position of the listener,

determining a fourth point on the first 3D shape that represents a fourth edge of the first 3D shape with respect to the listening position, and

determining a second dimension of the 2D plane or 1D line based on the third and fourth points.

8 . The method of claim 1 , wherein

determining the first property of the 2D plane or 1D line based on the first and second points comprises determining a horizontal angle of the 2D plane or 1D line based on the first and second points.

9 . The method of claim 8 , wherein transforming the obtained first metadata to produce the transformed metadata further comprises:

determining a third point on the first 3D shape that represents a third edge of the first 3D shape with respect to the listening position of the listener,

determining a fourth point on the first 3D shape that represents a fourth edge of the first 3D shape with respect to the listening position, and

determining a vertical angle of the 2D plane or 1D line based on the third and fourth points.

10 . The method of claim 1 , wherein a position of the first point is smoothed using a technique where the magnitude of the position change between two time instances is limited to be, at most, proportional to the magnitude of the change in relative distance between the audio object and the listener between the same two time instances.

11 . The method of claim 10 , wherein a change of rotation of the extent of the audio object is taken into account by calculating an expected position change of the anchor point, and where the position change of the first point is limited to be, at most, proportional to the sum of the magnitude of the expected position change of the anchor point and the magnitude of the change in relative position between the audio object and the listener.

12 . The method of claim 1 , wherein

the audio object is represented by a multi-channel audio signal, and

the method further comprises rendering the multi-channel audio signal using the transformed metadata.

13 . The method of claim 12 , wherein the multi-channel audio signal is rendered using a virtual source that represents an edge of the 2D shape or 1D line.

14 . The method of claim 12 , wherein the multi-channel audio signal is rendered using a virtual sound source that has an associated head related transfer function (HRTF) that represents a sub-area of the 2D shape or 1D line.

15 . A non-transitory computer readable storage medium storing a computer program comprising instructions which, when executed by processing circuitry of an apparatus causes the apparatus to perform the method of claim 1 .

16 . The method of claim 1 , wherein determining the anchor point comprises:

evaluating a set of points to identify which point in the set of points is closest to the listening position, wherein

the anchor point is the point in the set of points that is identified as being closest to the listening position, and

each point in the set of points represents a vertex of the first 3D shape or is a point on a surface of the first 3D shape.

17 . The method of claim 1 , wherein

determining the anchor point comprises determining a point on the first 3D shape that is closest to the listening position, and

the anchor point is the point on the first 3D shape that is determined to be closest to the listening position.

18 . An apparatus for representing an audio object in an extended reality scene, the apparatus comprising:

a storage unit; and

processing circuitry coupled to the storage unit, wherein the apparatus is configured to perform a method comprising:

obtaining first metadata describing a first three-dimensional (3D) shape associated with the audio object; and

transforming the obtained first metadata to produce transformed metadata describing a two-dimensional (2D) plane or a one-dimensional (1D) line, wherein

the 2D plane or the 1D line represent at least a portion of the audio object,

transforming the obtained first metadata to produce the transformed metadata comprises determining an anchor point, the anchor point being a point on the surface of the 3D shape that is deemed closest to the listening position of the listener in the extended reality scene,

the 2D plane or 1D lines passes through the anchor point,

transforming the obtained first metadata to produce the transformed metadata further comprises:

determining a first point on the first 3D shape that represents a first edge of the first 3D shape with respect to the listening position of the listener,

determining a second point on the first 3D shape that represents a second edge of the first 3D shape with respect to the listening position of the listener, and

determining a first property of the 2D plane or 1D line based on the first and second points, and

the method further comprises:

prior to determining the first point on the first 3D shape and the second point on the first 3D shape, determining a line between the listening position of the listener and the anchor point; and

using the determined line to determine the first point on the first 3D shape and the second point on the first 3D shape, wherein

the first point on the first 3D shape has a highest azimuth angle relative to the line between the listening position of the listener and the anchor point, and

the second point on the first 3D shape has a lowest azimuth angle relative to the line between the listening position of the listener and the anchor point.

19 . The apparatus of claim 18 , wherein the storage unit comprises memory that stores instructions for configuring the apparatus to obtain the first metadata by selecting the first metadata from a set of metadata comprising the first metadata and second metadata that describes a second 3D shape associated with the audio object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2022
From: FALK, TOMMY; DE BRUIJN, WERNER
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 061154/0176 →
Continuity (2)
Provisional Application 62988983 · Mar 13, 2020
Related Publication 20230132745A1 · May 4, 2023
References Cited (13)
US 10425762B1 · Schissler · 2019 [cited by applicant]
US 20200128347A1 · Schissler · 2020 [cited by examiner]
US 20200275230A1 · Laaksonen · 2020 [cited by examiner]
US 20210092546A1 · Terentiv · 2021 [cited by examiner]
US 20210289309A1 · Herre · 2021 [cited by examiner]
US 20210306792A1 · De Bruijn · 2021 [cited by examiner]
WO 2019121773A1 · 2019 [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/EP2021/056112 dated Jun. 16, 2021 (13 pages). [cited by applicant]
Schissler, C. et al., “Efficient HRTF-based Spatial Audio for Area and Volumetric Sources”, IEEE Transactions on Visualization and Computer Graphics, vol. 22, No. 4, Apr. 2016 (11 pages). [cited by applicant]
Garland, M. et al., “Surface Simplification Using Quadric Error Metrics”, Carnegie Mellon University, 1997 (8 pages). [cited by applicant]
Huebner, K. et al., “Minimum Volume Bounding Box Decomposition for Shape Approximation in Robot Grasping”, 2008 IEEE International Conference on Robotics and Automation, May 2008 (6 pages). [cited by applicant]
ISO/IEC 23008-3:201x(E), “Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio”, ISO/IEC JTC 1/SC 29, Oct. 12, 2016 (23 pages). [cited by applicant]
EBU Operating Eurovision and Euroradio Tech 3388, “ADM Renderer for Use in Next Generation Audio Broadcasting”, Source: BTF Renderer Group, Specification Version 1.0, Geneva, Mar. 2018 (57 pages). [cited by applicant]