IP Library Granted Patent US 12,308,034
Granted Patent B2
US 12,308,034 · App. 16/907,771 · Granted May 20, 2025

Performing psychoacoustic audio coding based on operating conditions

Inventors: Ferdinando Olivieri (San Diego, CA); Taher Shahbazi Mirzahasanloo (San Diego, CA); Nils Günther Peters (San Diego, CA)
Assignee: QUALCOMM Incorporated
G10L19/008G10L19/032G10L25/48H04L65/75H04L65/752
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,308,034
App. No.
16/907,771
Granted
May 20, 2025
Kind
B2
Abstract

A device comprising a memory and one or more processors may be configured to perform the techniques described in this disclosure. The memory may be configured to store the encoded scene-based audio data. The one or more processors may be configured to obtain an operating condition of the device for decoding the encoded scene-based audio data and perform, based on the operating condition, psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain ambisonic transport format audio data. The one or more processors may also be configured to perform spatial audio decoding with respect to the ambisonic transport format audio data to obtain scene-based audio data.

Claims (63)

1. A device configured to decode encoded scene-based audio data, the device comprising:

a memory configured to store the encoded scene-based audio data; and

one or more processors configured to:

monitor one or more internal components of the device to obtain an operating condition of the device for decoding the encoded scene-based audio data;

perform, based on the operating condition, psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain ambisonic transport format audio data; and

perform spatial audio decoding with respect to the ambisonic transport format audio data to obtain scene-based audio data.

2. The device of claim 1 , wherein the one or more processors perform a single instance of the psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain the ambisonic transport format audio data.

3. The device of claim 1 , wherein the one or more processors are configured to perform, based on the operating conditions, psychoacoustic audio decoding according to a compression algorithm with respect to the encoded scene-based audio data to obtain the ambisonic transport format audio data.

4. The device of claim 1 ,

wherein the one or more processors are further configured to obtain a latency requirement for decoding the encoded scene-based audio data; and

wherein the one or more processors are configured to perform, based on the operating condition and the latency requirement, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain the ambisonic transport format audio data.

5. The device of claim 1 ,

wherein the one or more processors are further configured to establish a communication link over which the encoded scene-based audio data is received, and

wherein the one or more processors are configured to monitor an internal communication component that established the communication link to obtain a status of the communication link as the operating condition.

6. The device of claim 1 , wherein the one or more processors are configured to perform, when the operating condition is below a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode a subset of bits representative of the ambisonic transport format audio data, the subset of bits less than a total number of bits used to represent the ambisonic transport format audio data.

7. The device of claim 6 , wherein the one or more processors are configured to:

determine a number of course bits used during quantization of an energy of the ambisonic transport format audio data; and

perform the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the number of course bits representative of the ambisonic transport format audio data.

8. The device of claim 1 , wherein the one or more processors are configured to perform, when the operating condition is greater than or equal to a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode a subset of bits representative of the ambisonic transport format audio data and additional bits representative of the ambisonic transport format audio data, the subset of bits less than a total number of bits used to represent the ambisonic transport format audio data.

9. The device of claim 8 , wherein the one or more processors are configured to perform the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the ambisonic transport format audio data corresponding to all of the orders up to a defined order either when reconstructing the subset of bits or both the subset of bits and the additional bits.

10. The device of claim 1 , wherein the one or more processors are configured to:

determine a number of course bits used during quantization of an energy of the ambisonic transport format audio data;

determine a number of fine bits used during quantization of the energy of the ambisonic transport format audio data; and

perform, when the operating condition is greater than or equal to a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the number of course bits and the number of fine bits representative of the ambisonic transport format audio data.

11. The device of claim 1 , wherein the one or more processors are configured to perform, when the operating condition is below a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the ambisonic transport format audio data corresponding to spherical basis functions having a subset of orders, the subset of orders being less than all of the orders of the spherical basis functions to which the scene-based audio data corresponds.

12. The device of claim 1 , wherein the one or more processors are configured to perform, when the operating condition is above a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to reconstruct the ambisonic transport format audio data corresponding to spherical basis functions having a subset of orders and at least one additional order, the subset of orders being less than all of the orders of the spherical basis functions to which the scene-based audio data corresponds.

13. The device of claim 12 , wherein the one or more processors are configured to perform the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode all of the bits representative of the ambisonic transport format audio data corresponding to the spherical basis functions having either the subset of orders or both the subset of orders and the at least one additional order.

14. The device of claim 1 , wherein the one or more processors are further configured to:

render the scene-based audio data to one or more speaker feeds; and

output the speaker feeds to one or more speakers to reproduce, based on the speaker feeds, a soundfield represented by the scene-based audio data.

15. The device of claim 1 ,

wherein the one or more processors are further configured to render the scene-based audio data to one or more speaker feeds, and

wherein the device comprises one or more speakers configured to reproduce, based on the speaker feeds, a soundfield represented by the scene-based audio data.

16. The device of claim 1 , wherein the one or more processors are configured to monitor an internal battery of the device to obtain a battery power as the operating condition.

17. A method of decoding encoded scene-based audio data, the method comprising:

monitoring one or more internal components of a device configured to decode the scene-based audio data to obtain an operating condition of the device;

performing, based on the operating condition, psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain ambisonic transport format audio data; and

performing spatial audio decoding with respect to the ambisonic transport format audio data to obtain scene-based audio data.

18. The method of claim 17 , wherein performing the psychoacoustic audio decoding comprises performing a single instance of the psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain the ambisonic transport format audio data.

19. The method of claim 17 , wherein performing the psychoacoustic audio decoding comprises performing, based on the operating conditions, the psychoacoustic audio decoding according to a compression algorithm with respect to the encoded scene-based audio data to obtain the ambisonic transport format audio data.

20. The method of claim 17 , further comprising obtaining a latency requirement for decoding the encoded scene-based audio data, and

wherein performing the psychoacoustic audio decoding comprises performing, based on the operating condition and the latency requirement, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain the ambisonic transport format audio data.

21. The method of claim 17 , further comprising establishing a communication link over which the encoded scene-based audio data is received, and

wherein monitoring the one or more internal components includes monitoring a communication unit that established the communication link to obtain a status of the communication link as the operating condition.

22. The method of claim 17 , wherein performing the psychoacoustic audio decoding comprises performing, when the operating condition is below a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode a subset of bits representative of the ambisonic transport format audio data, the subset of bits less than a total number of bits used to represent the ambisonic transport format audio data.

23. The method of claim 22 , wherein obtaining the psychoacoustic audio decoding comprises:

determining a number of course bits used during quantization of an energy of the ambisonic transport format audio data; and

performing the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the number of course bits representative of the ambisonic transport format audio data.

24. The method of claim 17 , wherein performing the psychoacoustic audio decoding comprises performing, when the operating condition is greater than or equal to a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode a subset of bits representative of the ambisonic transport format audio data and additional bits representative of the ambisonic transport format audio data, the subset of bits less than a total number of bits used to represent the ambisonic transport format audio data.

25. The method of claim 17 , where performing the psychoacoustic audio decoding comprises:

determining a number of course bits used during quantization of an energy of the ambisonic transport format audio data;

determining a number of fine bits used during quantization of the energy of the ambisonic transport format audio data; and

performing, based on the operating condition is greater than or equal to a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the number of course bits and the number of fine bits representative of the ambisonic transport format audio data.

26. The method of claim 22 , wherein performing the psychoacoustic audio decoding comprises performing the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the ambisonic transport format audio data corresponding to all of the orders up to a defined order either when reconstructing the subset of bits or both the subset of bits and additional bits.

27. The method of claim 17 , wherein performing the psychoacoustic audio decoding comprises performing, when the operating condition is below a threshold, the psychoacoustic audio decoding with respect to the encoded scene-based audio data to decode the ambisonic transport format audio data corresponding to spherical basis functions having a subset of orders, the subset of orders being less than all of the orders of the spherical basis functions to which the scene-based audio data corresponds.

28. A device configured to decode encoded scene-based audio data, the device comprising:

means for monitoring one or more internal components of the device to obtain an operating condition of the device;

means for performing, based on the operating condition, psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain ambisonic transport format audio data; and

means for performing spatial audio decoding with respect to the ambisonic transport format audio data to obtain scene-based audio data.

29. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors of a device configured to decode encoded scene-based audio data to:

monitor one or more internal components of the device to obtain an operating condition of the device;

perform, based on the operating condition, psychoacoustic audio decoding with respect to the encoded scene-based audio data to obtain ambisonic transport format audio data; and

perform spatial audio decoding with respect to the ambisonic transport format audio data to obtain scene-based audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2021
From: OLIVIERI, FERDINANDO; SHAHBAZI MIRZAHASANLOO, TAHER; PETERS, NILS GÜNTHER
To: QUALCOMM INCORPORATED
Reel/Frame 055460/0294 →
Continuity (2)
Provisional Application 62865848 · Jun 24, 2019
Related Publication 20200402521A1 · Dec 24, 2020
References Cited (90)
US 5651090A · Moriya et al. · 1997 [cited by applicant]
US 9626973B2 · Taleb · 2017 [cited by examiner]
US 9747910B2 · Kim · 2017 [cited by examiner]
US 9852737B2 · Kim et al. · 2017 [cited by applicant]
US 9959880B2 · Peters et al. · 2018 [cited by applicant]
US 10075802B1 · Kim et al. · 2018 [cited by applicant]
US 10657974B2 · Kim et al. · 2020 [cited by applicant]
US 10770087B2 · Kim et al. · 2020 [cited by applicant]
US 20030006916A1 · Takamizawa · 2003 [cited by applicant]
US 20070168197A1 · Vasilache · 2007 [cited by applicant]
US 20070269063A1 · Goodwin et al. · 2007 [cited by applicant]
US 20070286441A1 · Harsch · 2007 [cited by examiner]
US 20080027709A1 · Baumgarte · 2008 [cited by applicant]
US 20080136686A1 · Feiten · 2008 [cited by applicant]
US 20080252510A1 · Jung et al. · 2008 [cited by applicant]
US 20100017204A1 · Oshikiri et al. · 2010 [cited by applicant]
US 20110170711A1 · Rettelbach et al. · 2011 [cited by applicant]
US 20110249821A1 · Jaillet et al. · 2011 [cited by applicant]
US 20120042107A1 · Rabii · 2012 [cited by examiner]
US 20130275140A1 · Kim et al. · 2013 [cited by applicant]
US 20140219459A1 · Daniel et al. · 2014 [cited by applicant]
US 20140249827A1 · Sen · 2014 [cited by examiner]
US 20140358557A1 · Sen et al. · 2014 [cited by applicant]
US 20140358565A1 · Peters et al. · 2014 [cited by applicant]
US 20150025895A1 · Schildbach · 2015 [cited by applicant]
US 20150255076A1 · Fejzo · 2015 [cited by applicant]
US 20150271621A1 · Sen et al. · 2015 [cited by applicant]
US 20150332681A1 · Kim et al. · 2015 [cited by applicant]
US 20150340044A1 · Kim · 2015 [cited by applicant]
US 20150358810A1 · Chao · 2015 [cited by examiner]
US 20160005407A1 · Friedrich et al. · 2016 [cited by applicant]
US 20160007132A1 · Peters · 2016 [cited by examiner]
US 20160064005A1 · Peters · 2016 [cited by examiner]
US 20160093311A1 · Kim et al. · 2016 [cited by applicant]
US 20160104493A1 · Kim et al. · 2016 [cited by applicant]
US 20160104494A1 · Kim · 2016 [cited by examiner]
US 20180013446A1 · Milicevic · 2018 [cited by examiner]
US 20180082694A1 · Kim · 2018 [cited by applicant]
US 20180278962A1 · Johnston · 2018 [cited by examiner]
US 20190007781A1 · Peters et al. · 2019 [cited by applicant]
US 20190103118A1 · Atti et al. · 2019 [cited by applicant]
US 20190259398A1 · Buethe et al. · 2019 [cited by applicant]
US 20190371348A1 · Shahbazi Mirzahasanloo · 2019 [cited by examiner]
US 20200013414A1 · Thagadur Shivappa et al. · 2020 [cited by applicant]
US 20200402519A1 · Olivieri · 2020 [cited by applicant]
CN 102089808A · 2011 [cited by applicant]
CN 102124517A · 2011 [cited by applicant]
CN 104285451A · 2015 [cited by applicant]
CN 104471641A · 2015 [cited by applicant]
CN 105593931A · 2016 [cited by applicant]
EP 3067885A1 · 2016 [cited by applicant]
WO 2015000819A1 · 2015 [cited by applicant]
WO 2015176003A1 · 2015 [cited by applicant]
WO 2017066312A1 · 2017 [cited by applicant]
International Preliminary Report on Patentability—PCT/US2020/039158, The International Bureau of WIPO—Geneva, Switzerland, Jan. 6, 2022 8 Pages. [cited by applicant]
Final Office Action from U.S. Appl. No. 116/907,934 mailed Jan. 28, 2022, 17 pages. [cited by applicant]
Notice of Allowance from U.S. Appl. No. 16/907,969, mailed Feb. 16, 2022, 12 pp. [cited by applicant]
Peters N., et al., “Scene-Based Audio Implemented with Higher Order Ambisonics (HOA)”, SMPTE Motion Imaging Journal, Nov.-Dec. 2016, 13 pages. [cited by applicant]
Shankar Shivappa, et al., “Efficient, Compelling and Immersive VR Audio Experience Using Scene Based Audo/Higher Order Ambisonics”, AES conference paper, presented on the conference on Audio for Virtual and Augmented Re… [cited by applicant]
Advisory Action for U.S. Appl. No. 16/907,934, mailed Apr. 5, 2022, 4 pp. [cited by applicant]
Appeal Brief for U.S. Appl. No. 16/907,934, filed Jun. 21, 2022, 27 pp. [cited by applicant]
Ma N., et al., “Combining Speech Fragment Decoding and Adaptive Noise Floor Modeling”, in IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, No. 3, pp. 818-827, Mar. 2012. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/908,032, mailed May 18, 2022, 41 pp. [cited by applicant]
Zamani S., et al., “Spatial Audio Coding without Recourse to Background Signal Compression”, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 720-724, DOI: 10… [cited by applicant]
“3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Virtual Reality (VR) Media Services Over 3GPP (Release 15)”,3GPP Draft, S4-170494 TR 26.918 Virtual Reality (VR) Media Serv… [cited by applicant]
“Advanced Audio Distribution Profile Specification,” version 1.3.1, published Jul. 14, 2015, 35 pp. [cited by applicant]
AUDIO: “Call for Proposals for 3D Audio”, International Organisation for Standardisation Organisation Internationale De Normalisation ISO/IEC JTC1/SC29/WG11 Coding of Moving Pictures and Audio, ISO/IEC JTC1/SC29/WG11/N1… [cited by applicant]
“Bluetooth Core Specification v 5.0,” published Dec. 6, 2016 accessed from https://www.bluetooth.com/specifications, pp. 1-5. [cited by applicant]
Boehm J., et al., “Scalable Decoding Mode for MPEG-H 3D Audio HOA”, 108. MPEG Meeting; Mar. 31, 2014-Apr. 4, 2014; Valencia; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. m33195, Mar. 26, 2014 (Mar. 26, 2… [cited by applicant]
ETSI TS 103 589 V1.1.1, “Higher Order Ambisonics (HOA) Transport Format”, Jun. 2018, 33 pages. [cited by applicant]
Herre J., et al., “MPEG-H 3D Audio—The New Standard for Coding of Immersive Spatial Audio”, IEEE Journal of Selected Topics in Signal Processing, vol. 9, No. 5, Aug. 1, 2015 (Aug. 1, 2015), pp. 770-779, XP055243182, US … [cited by applicant]
Hollerweger F., “An Introduction to Higher Order Ambisonic,” Oct. 2008, pp. 13, Accessed online [Jul. 8, 2013]. [cited by applicant]
“Information technology—High Efficiency Coding and Media Delivery in Heterogeneous Environments—Part 3: 3D Audio,” ISO/IEC JTC 1/SC 29, ISO/IEC DIS 23008-3, Jul. 25, 2014, 433 Pages. [cited by applicant]
“Information technology—High Efficiency Coding and Media Delivery in Heterogeneous Environments—Part 3: 3D Audio”, ISO/IEC JTC 1/SC 29/WG11, ISO/IEC 23008-3, 201x(E), Oct. 12, 2016, 797 Pages. [cited by applicant]
“Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: Part 3: BD Audio, Amendment 3: MPEG-H 3D Audio Phase 2,” ISO/IEC JTC 1/SC 29N, ISO/IEC 23008-3:2015/PDAM 3, Jul. 25… [cited by applicant]
International Search Report and Written Opinion—PCT/US2020/039158—ISA/EPO—Oct. 9, 2020. [cited by applicant]
ISO/IEC/JTC: “ISO/IEC JTC 1/SC 29 N ISO/IEC CD 23008-3 Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio,” Apr. 4, 2014 (Apr. 4, 2014), 337 Pages, XP05520637… [cited by applicant]
Poletti M.A., “Three-Dimensional Surround Sound Systems Based on Spherical Harmonics”, The Journal of the Audio Engineering Society, vol. 53, No. 11, Nov. 2005, pp. 1004-1025. [cited by applicant]
Schonefeld V., “Spherical Harmonics”, Jul. 1, 2005, XP002599101, 25 Pages, Accessed online [Jul. 9, 2013] at URL:http://heim.c-otto.de/˜volker/prosem_paper.pdf. [cited by applicant]
Sen D., et al., “Efficient Compression and Transportation of Scene Based Audio for Television Broadcast”, Jul. 18, 2016 (Jul. 18, 2016), XP055327771, 8 pages, Retrieved from the Internet: URL: http://www.aes.org/tmpFile… [cited by applicant]
Sen D., et al., “RM1-HOA Working Draft Text”, 107. MPEG Meeting, Jan. 13, 2014-Jan. 17, 2014, San Jose, (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. m31827, Jan. 11, 2014 (Jan. 11, 2014), 83 Pages, XP030… [cited by applicant]
Sen D., et al., “Technical Description of the Qualcomm's HoA Coding Technology for Phase II”, 109. MPEG Meeting; Jul. 7, 2014-Jul. 11, 2014; Sapporo; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. m34104, … [cited by applicant]
Sen D (Qualcomm)., et al., “Thoughts on Layered/Scalable Coding for HOA the Signal”, 110. MPEG Meeting, Oct. 20, 2014-Oct. 24, 2014; Strasbourg, (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. m35160, Oct. … [cited by applicant]
Yang D., “High Fidelity Multichannel Audio Compression”, Jan. 1, 2002 (Jan. 1, 2002), 211 Pages, XP055139610. Retrieved from the Internet : URL: http://search.proquest.com/docview/305523844. section 5.1. [cited by applicant]
U.S. Appl. No. 16/907,934, filed Jun. 22, 2020, 65 Pages. [cited by applicant]
U.S. Appl. No. 16/907,969, filed Jun. 22, 2020, 73 Pages. [cited by applicant]
U.S. Appl. No. 16/908,032, filed Jun. 22, 2020, 74 Pages. [cited by applicant]
Patent Board Decision for U.S. Appl. No. 16/907,934, mailed Dec. 11, 2023, 10 pp. [cited by applicant]
Response to Decision on Appeal dated Dec. 11, 2023 from U.S. Appl. No. 16/907,934, filed Feb. 2, 2024, 10 pp. [cited by applicant]
Notice of Allowance from U.S. Appl. No. 16/907,934 dated Mar. 1, 2024, 12 pp. [cited by applicant]