IP Library Granted Patent US 12,380,903
Granted Patent B2
US 12,380,903 · App. 18/416,775 · Granted Aug 5, 2025

Apparatus and method for screen related audio object remapping

Inventors: Simone Fueg (Kalchreuth, DE); Jan Plogsties (Fuerth, DE); Sascha Dick (Nuremberg, DE); Johannes Hilpert (Nuremberg, DE); Julien Robilliard (Nuremberg, DE); Achim Kuntz (Hemhofen, DE); Andreas Hoelzer (Erlangen, DE)
Assignee: Fraunhofer-Gesellschaft zur Foerderung der angewandten Forschung e.V.
G10L19/20G10L19/008H04N21/233H04N21/4318H04N21/439H04N21/4516H04N21/8106H04N21/84H04S3/008H04S7/00H04S7/30H04S7/308G10L19/167H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,903
App. No.
18/416,775
Granted
Aug 5, 2025
Kind
B2
Abstract

An apparatus for generating loudspeaker signals includes an object metadata processor configured to receive metadata, to calculate a second position of the audio object depending on the first position of the audio object and on a size of a screen if the audio object is indicated in the metadata as being screen-related, to feed the first position of the audio object as the position information into the object renderer if the audio object is indicated in the metadata as being not screen-related, and to feed the second position of the audio object as the position information into the object renderer if the audio object is indicated in the metadata as being screen-related. The apparatus further includes an object renderer configured to receive an audio object and to generate the loudspeaker signals depending on the audio object and on position information.

Claims (289)

1. An apparatus for generating loudspeaker signals, comprising:

an object metadata processor, and

an object renderer,

wherein,

metadata comprises an indication on whether an audio object is screen-related, and further comprises a first position of the audio object,

wherein the object metadata processor is configured to calculate a second position of the audio object depending on the first position of the audio object and depending on a size of a screen if the audio object is indicated in the metadata as being screen-related,

wherein the object renderer is configured to generate the loudspeaker signals depending on the audio object and depending on position information,

wherein the object metadata processor is configured to feed the first position of the audio object as the position information into the object renderer if the audio object is indicated in the metadata as being not screen-related, and

wherein the object metadata processor is configured to feed the second position of the audio object as the position information into the object renderer if the audio object is indicated in the metadata as being screen-related.

2. The apparatus according to claim 1 , wherein the object metadata processor is configured to not calculate the second position of the audio object if the audio object is indicated in the metadata as being not screen-related.

3. The apparatus according to claim 1 , wherein the object renderer is configured to not determine whether the position information is the first position of the audio object or the second position of the audio object.

4. The apparatus according to claim 1 , wherein the object renderer is configured to generate the loudspeaker signals further depending on the number of the loudspeakers of a playback environment.

5. The apparatus according to claim 4 , wherein the object renderer is configured to generate the loudspeaker signals further depending on a loudspeaker position of each of the loudspeakers of the playback environment.

6. The apparatus according to claim 1 , wherein the object metadata processor is configured to calculate the second position of the audio object depending on the first position of the audio object and depending on the size of the screen if the audio object is indicated in the metadata as being screen-related, wherein the first position indicates the first position in a three-dimensional space, and wherein the second position indicates the second position in the three-dimensional space.

7. The apparatus according to claim 6 , wherein the object metadata processor is configured to calculate the second position of the audio object depending on the first position of the audio object and depending on the size of the screen if the audio object is indicated in the metadata as being screen-related, wherein the first position indicates a first azimuth, a first elevation and a first distance, and wherein the second position indicates a second azimuth, a second elevation and a second distance.

8. The apparatus according to claim 1 ,

wherein the object metadata processor is configured to receive the metadata, comprising the indication on whether the audio object is screen-related as a first indication, and further comprising a second indication if the audio object is screen-related, said second indication indicating whether the audio object is an on-screen object, and

wherein the object metadata processor is configured to calculate the second position of the audio object depending on the first position of the audio object and depending on the size of the screen, such that the second position takes a first value on a screen area of the screen if the second indication indicates that the audio object is an on-screen object.

9. The apparatus according to claim 8 , wherein the object metadata processor is configured to calculate the second position of the audio object depending on the first position of the audio object and depending on the size of the screen, such that the second position takes a second value, which is either on the screen area or not on the screen area if the second indication indicates that the audio object is not an on-screen object.

10. The apparatus according to claim 1 ,

wherein the object metadata processor is configured to receive the metadata, comprising the indication on whether the audio object is screen-related as a first indication, and further comprising a second indication if the audio object is screen-related, said second indication indicating whether the audio object is an on-screen object,

wherein the object metadata processor is configured to calculate the second position of the audio object depending on the first position of the audio object, depending on the size of the screen, and depending on a first mapping curve as the mapping curve if the second indication indicates that the audio object is an on-screen object, wherein the first mapping curve defines a mapping of original object positions in a first value interval to remapped object positions in a second value interval, and

wherein the object metadata processor is configured to calculate the second position of the audio object depending on the first position of the audio object, depending on the size of the screen, and depending on a second mapping curve as the mapping curve if the second indication indicates that the audio object is not an on-screen object, wherein the second mapping curve defines a mapping of original object positions in the first value interval to remapped object positions in a third value interval, and wherein said second value interval is comprised by the third value interval, and wherein said second value interval is smaller than said third value interval.

11. The apparatus according to claim 10 ,

wherein each of the first value interval and the second value interval and the third value interval is a value interval of azimuth angles, or

wherein each of the first value interval and the second value interval and the third value interval is a value interval of elevation angles.

12. The apparatus according to claim 1 ,

wherein the object metadata processor is configured to calculate the second position of the audio object depending on at least one of a first linear mapping function and a second linear mapping function,

wherein the first linear mapping function is defined to map a first azimuth value to a second azimuth value,

wherein the second linear mapping function is defined to map a first elevation value to a second elevation value,

wherein φ left nominal indicates a left azimuth screen edge reference,

wherein φ right nominal right indicates a right azimuth screen edge reference,

wherein θ top nominal indicates a top elevation screen edge reference,

wherein θ bottom nominal indicates a bottom elevation screen edge reference,

wherein φ left repro indicates a left azimuth screen edge of the screen,

wherein φ right repro indicates a right azimuth screen edge of the screen,

wherein θ top repro indicates a top elevation screen edge of the screen,

wherein θ bottom repro indicates a bottom elevation screen edge of the screen,

wherein φ indicates the first azimuth value,

wherein φ′ indicates the second azimuth value,

wherein θ indicates the first elevation value,

wherein θ′ indicates the second elevation value,

wherein the second azimuth value φ′ results from a first mapping of the first azimuth value φ according to the first linear mapping function according to

φ

=

{

φ

right

repro

+

180

°

φ

right

nominal

+

180

°

·

(

φ

+

180

°

)

-

180

°

for

-1

80

°

φ

<

φ

right

nominal

φ

left

repro

-

φ

right

repro

φ

left

nominal

-

φ

right

nominal

·

(

φ

-

φ

right

nominal

)

+

φ

right

repro

for

φ

right

nominal

φ

<

φ

left

nominal

180

°

-

φ

left

repro

180

°

-

φ

left

nominal

·

(

φ

-

φ

left

nominal

)

+

φ

left

repro

for

φ

left

nominal

φ

<

180

°

,

and

wherein the second elevation value θ′ results from a second mapping of the first elevation value θ according to the second linear mapping function according to

θ

=

{

θ

bottom

repro

+

90

°

θ

bottom

nominal

+

90

°

·

(

θ

+

90

°

)

-

90

°

for

-9

0

°

θ

<

θ

bottom

nominal

θ

top

repro

-

θ

bottom

repro

θ

top

nominal

-

θ

bottom

nominal

·

(

θ

-

θ

bottom

nominal

)

+

θ

bottom

repro

for

θ

bottom

nominal

θ

<

θ

top

nominal

90

°

-

θ

top

repro

90

°

-

θ

top

nominal

·

(

θ

-

θ

top

nominal

)

+

θ

top

repro

for

θ

top

nominal

θ

<

90

°

.

13. A method for generating loudspeaker signals, wherein metadata comprises an indication on whether an audio object is screen-related, and further comprises a first position of the audio object, wherein the method comprises:

calculating a second position of the audio object depending on the first position of the audio object and depending on a size of a screen if the audio object is indicated in the metadata as being screen-related,

generating the loudspeaker signals depending on the audio object and depending on position information,

wherein the position information is the first position of the audio object if the audio object is indicated in the metadata as being not screen-related, and

wherein the position information is the second position of the audio object if the audio object is indicated in the metadata as being screen-related.

14. A non-transitory digital storage medium having a computer program stored thereon to perform the method of claim 13 when said computer program is run by a computer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2024
From: FUEG, SIMONE; PLOGSTIES, JAN; DICK, SASCHA; HILPERT, JOHANNES; ROBILLIARD, JULIEN; KUNTZ, ACHIM; HOELZER, ANDREAS
To: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 066172/0802 →
Priority Claims (2)
EP 14161819 · Mar 26, 2014 · regional
EP 14196769 · Dec 8, 2014 · regional
Continuity (6)
Continuation 18057188 · Nov 18, 2022
Continuation 16950768 · Nov 17, 2020
Continuation 16236079 · Dec 28, 2018
Continuation 15274310 · Sep 23, 2016
Continuation PCTEP2015056417 · Mar 25, 2015
Related Publication 20240265930A1 · Aug 8, 2024
References Cited (42)
US 7606372B2 · Melchior et al. · 2009 [cited by applicant]
US 9179236B2 · Robinson · 2015 [cited by examiner]
US 9451363B2 · Jax · 2016 [cited by examiner]
US 10299062B2 · Jax et al. · 2019 [cited by applicant]
US 10504527B2 · Herre et al. · 2019 [cited by applicant]
US 20030007648A1 · Currell · 2003 [cited by applicant]
US 20050147257A1 · Melchior et al. · 2005 [cited by applicant]
US 20060294125A1 · Deaven et al. · 2006 [cited by applicant]
US 20100017002A1 · Oh et al. · 2010 [cited by applicant]
US 20120078642A1 · Seo et al. · 2012 [cited by applicant]
US 20120183162A1 · Chabanne et al. · 2012 [cited by applicant]
US 20130202129A1 · Kraemer et al. · 2013 [cited by applicant]
US 20130216070A1 · Keiler et al. · 2013 [cited by applicant]
US 20130236033A1 · Ayres et al. · 2013 [cited by applicant]
US 20130236039A1 · Jax et al. · 2013 [cited by applicant]
US 20140016785A1 · Neuendorf et al. · 2014 [cited by applicant]
US 20140016786A1 · Sen et al. · 2014 [cited by applicant]
US 20140019146A1 · Neuendorf et al. · 2014 [cited by applicant]
US 20140023196A1 · Xiang et al. · 2014 [cited by applicant]
US 20140023197A1 · Xiang et al. · 2014 [cited by applicant]
US 20140133683A1 · Robinson et al. · 2014 [cited by applicant]
CN 102099854A · 2011 [cited by applicant]
CN 102460571A · 2012 [cited by applicant]
CN 102667919A · 2012 [cited by applicant]
CN 103250207A · 2013 [cited by applicant]
CN 103313182A · 2013 [cited by applicant]
CN 103650539A · 2014 [cited by applicant]
EP 1318502A2 · 2003 [cited by applicant]
EP 2637427A1 · 2013 [cited by applicant]
EP 2637428A1 · 2013 [cited by applicant]
JP 2009278381A · 2009 [cited by applicant]
JP 2013521725A · 2013 [cited by applicant]
JP 2013187903A · 2013 [cited by applicant]
JP 2013187908A · 2013 [cited by applicant]
JP 6422995B2 · 2018 [cited by applicant]
TW 201325269A · 2013 [cited by applicant]
WO 2004073352A1 · 2004 [cited by applicant]
WO 2013006330A2 · 2013 [cited by applicant]
WO 2013006338A2 · 2013 [cited by applicant]
WO 2014032709A1 · 2014 [cited by applicant]
Neuendorf, Max, et al., “The ISO/MPEG Unified Speech and Audio Coding Standard Consistent High Quality for all Content Types and at all Bit Rates”, J. Audio Eng. Soc., vol. 61, No. 12,, pp. 956-977. [cited by applicant]
Schultz-Amling, et al., “Acoustical Zooming Based on a Parametric Sound Field Representation”, AES Convention Paper 8120; Presented at the 128th Convention; London, UK, May 22-25, 2010, 9 pages. [cited by applicant]