IP Library › Granted Patent US 12,361,956
Granted Patent B2
US 12,361,956 · App. 18/507,824 · Granted Jul 15, 2025

Perceptually-based loss functions for audio encoding and decoding based on machine learning

Inventors: Roy M. Fejgin (San Francisco, CA); Grant A. Davidson (Burlingame, CA); Chih-Wei Wu (San Francisco, CA); Vivek Kumar (Foster City, CA)
Assignee: Dolby Laboratories Licensing Corporation
G10L19/022G06F3/16G06N3/048G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,956
App. No.
18/507,824
Granted
Jul 15, 2025
Kind
B2
Abstract

Computer-implemented methods for training a neural network, as well as for implementing audio encoders and decoders via trained neural networks, are provided. The neural network may receive an input audio signal, generate an encoded audio signal and decode the encoded audio signal. A loss function generating module may receive the decoded audio signal and a ground truth audio signal, and may generate a loss function value corresponding to the decoded audio signal. Generating the loss function value may involve applying a psychoacoustic model. The neural network may be trained based on the loss function value. The training may involve updating at least one weight of the neural network.

Claims (34)

1. A method of encoding an audio signal using a neural network, the method comprising:

receiving the audio signal; and

encoding the audio signal via the neural network by using at least one weighted parameter of the neural network, wherein the neural network is trained based on a perceptually-based loss function, the perceptually-based loss function modeling an acoustic response of a human ear canal, and wherein the at least one weighted parameter is determined by the perceptually-based loss function.

2. The method of claim 1 , wherein encoding the audio signal via the neural network comprises:

generating, by the neural network and based on the audio signal, the encoded audio signal.

3. The method of claim 1 , further comprising:

prior to the encoding, transforming the audio signal from a time domain representation to a frequency domain representation.

4. The method of claim 1 , wherein the perceptually-based loss function is based on at least one selected from a group consisting of a psychoacoustic model, a psychoacoustic masking threshold, an average noise-to-masking ratio, minimizing an average noise-to-masking ratio, modeling an outer ear transfer function, grouping into critical bands, frequency-domain masking, level dependent spreading, and modeling of a frequency-dependent hearing threshold.

5. The method of claim 1 , wherein the audio signal comprises at least one selected from a group consisting of speech, music, and a mixture of sounds.

6. The method of claim 1 , wherein the neural network comprises an autoencoder.

7. The method of claim 1 , further comprising:

outputting the encoded audio signal in a compressed audio format.

8. A method of decoding an audio signal using a neural network, the method comprising:

receiving an encoded audio signal; and

decoding the encoded audio signal via the neural network to produce at least one selected from a group consisting of a decoded audio signal and at least one decoded transform coefficient by using at least one weighted parameter of the neural network, wherein the neural network is trained based on a perceptually-based loss function, the perceptually-based loss function modeling an acoustic response of a human ear canal, and wherein the at least one weighted parameter is determined by the perceptually-based loss function.

9. The method of claim 8 , wherein the decoding comprises:

generating, by the neural network and based on the encoded audio signal, the at least one selected from the group consisting of the decoded audio signal and the at least one decoded transform coefficient.

10. The method of claim 8 , wherein the perceptually-based loss function is based on at least one selected from a group consisting of a psychoacoustic model, a psychoacoustic masking threshold, an average noise-to-masking ratio, minimizing an average noise-to-masking ratio, modeling an outer ear transfer function, grouping into critical bands, frequency-domain masking, level dependent spreading, and modeling of a frequency dependent hearing threshold.

11. The method of claim 8 , wherein the decoded audio signal comprises at least one selected from a group consisting of speech, music, and a mixture of sounds.

12. An audio decoder comprising:

a processor; and

non-transitory storage media coupled to the processor, the non-transitory storage media storing instructions executable by the processor to:

receive an encoded audio signal; and

generate, via a neural network, at least one selected from a group consisting of a decoded audio signal and at least one decoded transform coefficient based on the received encoded audio signal by using at least one weighted parameter of the neural network, wherein the neural network is trained based on a perceptually-based loss function, the perceptually-based loss function modeling an acoustic response of a human ear canal, and wherein the at least one weighted parameter is determined by the perceptually-based loss function.

13. The audio decoder of claim 12 , wherein the neural network comprises:

a plurality of neuron layers comprising:

an input neuron layer;

a plurality of hidden neuron layers; and

an output neuron layer.

14. The audio decoder of claim 13 , wherein each neuron layer of the plurality of neuron layers are configured with at least one selected from a group consisting of a rectified linear unit (ReLU) activation function, a sigmoidal activation function, a tanh activation function, and an exponential linear unit (ELU) activation function.

15. The audio decoder of claim 12 , wherein the neural network is configured based on a perceptually-based loss function.

16. The audio decoder of claim 15 , wherein the perceptually-based loss function is based on at least one selected from a group consisting of a psychoacoustic model, a psychoacoustic masking threshold, an average noise-to-masking ratio, minimizing an average noise-to-masking ratio, modeling an outer ear transfer function, grouping into critical bands, frequency-domain masking, level dependent spreading, and modeling of a frequency-dependent hearing threshold.

17. The audio decoder of claim 12 , wherein the decoded audio signal comprises at least one selected from a group consisting of speech, music, and a mixture of sounds.

18. The audio decoder of claim 12 , wherein the encoded audio signal is configured in a compressed audio format.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2024
From: FEJGIN, ROY M.; DAVIDSON, GRANT A.; WU, CHIH-WEI; KUMAR, VIVEK
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 066064/0062 →
Priority Claims (1)
EP 18173673 · May 22, 2018 · regional
Continuity (4)
Continuation 17046284
Provisional Application 62829552 · Apr 4, 2019
Provisional Application 62656275 · Apr 11, 2018
Related Publication 20240079019A1 · Mar 7, 2024
References Cited (60)
US 8484022B1 · Vanhoucke · 2013 [cited by examiner]
US 9294060B2 · Myllyla · 2016 [cited by examiner]
US 9640194B1 · Nemala · 2017 [cited by examiner]
US 9679258B2 · Mnih · 2017 [cited by examiner]
US 9779727B2 · Yu · 2017 [cited by examiner]
US 11316767B2 · Hussa · 2022 [cited by examiner]
US 11373672B2 · Mesgarani · 2022 [cited by examiner]
US 11538455B2 · Zhou · 2022 [cited by examiner]
US 11687778B2 · Ciftci · 2023 [cited by examiner]
US 11817111B2 · Fejgin · 2023 [cited by examiner]
US 11948066B2 · van den Oord · 2024 [cited by examiner]
US 20030115041A1 · Chen · 2003 [cited by examiner]
US 20040044533A1 · Najaf-Zadeh · 2004 [cited by examiner]
US 20130144614A1 · Myllyla · 2013 [cited by examiner]
US 20140175534A1 · Kofuji · 2014 [cited by examiner]
US 20140372112A1 · Xue · 2014 [cited by examiner]
US 20150149165A1 · Saon · 2015 [cited by examiner]
US 20160307095A1 · Li · 2016 [cited by examiner]
US 20170040798A1 · Mazuelas · 2017 [cited by examiner]
US 20170177993A1 · Draelos · 2017 [cited by examiner]
US 20180018990A1 · Kim · 2018 [cited by examiner]
US 20180075343A1 · van den Oord · 2018 [cited by examiner]
US 20180108363A1 · Kim · 2018 [cited by examiner]
US 20180174046A1 · Xiao · 2018 [cited by examiner]
US 20190066713A1 · Mesgarani · 2019 [cited by examiner]
US 20200410976A1 · Zhou · 2020 [cited by examiner]
US 20210082444A1 · Fejgin · 2021 [cited by examiner]
US 20220392482A1 · Mesgarani · 2022 [cited by examiner]
US 20230229892A1 · Biswas · 2023 [cited by examiner]
US 20230395086A1 · Vinton · 2023 [cited by examiner]
US 20240079019A1 · Fejgin · 2024 [cited by examiner]
CA 2482427C · 2010 [cited by applicant]
CN 101790757B · 2012 [cited by applicant]
CN 101501759B · 2012 [cited by applicant]
CN 101872618B · 2012 [cited by applicant]
CN 103026407B · 2015 [cited by applicant]
CN 107516527A · 2017 [cited by applicant]
CN 105070293B · 2018 [cited by applicant]
CN 106782575B · 2020 [cited by applicant]
TW 201812744A · 2018 [cited by applicant]
WO 2017168870A1 · 2017 [cited by applicant]
WO 2018036972A1 · 2018 [cited by applicant]
WO WO2019199995A1 · 2019 [cited by examiner]
Atreya, A. et al “Novel Lossy Compression Algorithms with Stacked Autoencoders” Internet Citation, Dec. 11, 2009, pp. 1-5. [cited by applicant]
Balle, J. et al (2016). “End-to-end optimization of nonlinear transform codes for perceptual quality”. In Proceedings of the Picture Coding Symposium (PCS). [cited by applicant]
Beerends, J. G. et al., 2013, “Perceptual Objective Listening Quality Assessment (POLQA), the third generation ITU-T standard for end-to-end speech quality measurement Part I-temporal alignment”. AES: Journal of the Aud… [cited by applicant]
Fastl, H. et al . (2007), Psychoacoustics: Facts and Models (3rd ed., Springer). [cited by applicant]
Goodfellow, I. et al “Deep Learning” MIT Press Book, 2016. [cited by applicant]
Johnson, J. et al “Perceptual losses for real-time style transfer and super-resolution” in Proceedings of the European Conference on Computer Vision (ECCV) (vol. 9906 LNCS, pp. 694-711), 2016. [cited by applicant]
Kankanahalli, Srihari “End-to-End Optimized Speech Coding with Deep Neural Networks” IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 15, 2018, pp. 2521-2525. [cited by applicant]
Kingma, D.P. et al “Adam: a Method for Stochastic Optimization,” in Proceedings of the International Conference on Learning Representations (ICLR), 2015, pp. 1-15. [cited by applicant]
Ledig, C. et al “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network” Sep. 16, 2016, pp. 4681-4690. [cited by applicant]
Marchi, E. et al “Deep Recurrent Neural Network-Based Autoencoders for Acoustic Novelty Detection” Hindawi Computational Intelligence and Neuroscience, vol. 2017, published Jan. 15, 2017, 14 pages. [cited by applicant]
Martinez, A.M. et al., “Should deep neural nets have ears? the role of auditory features in deep learning approaches”, in Proceedings of the Annual Conference of the International Speech Communication Association (pp. 2… [cited by applicant]
Moore, B. et al chapter 3 (Frequency Selectivity, Masking, and the Critical Band) of . (2012), an Introduction to the Psychology of Hearing (Emerald Group Publishing),. [cited by applicant]
Nikunen, J. et al “Noise-to-mask ratio minimization by weighted non-negative matrix factorization,” in Proceedings of the International Conference on Acoustics Speech and Signal Processing (ICASSP), 2010, pp. 25-28. [cited by applicant]
Tripathi, A. et al “Asymmetric Stacked Autoencoder” International Joint Conference on Neural Networks, May 14, 2017, pp. 911-918. [cited by applicant]
Davidson, et al. “ATSC Video and Audio Coding”, in Proceedings of the IEEE, vol. 94, No. 1, Jun. 28, 2005, pp. 60-76,, 17 Pages. [cited by applicant]
Yuqing, et al. “Dialog generation based on hierarchical encoding and deep reinforcement learning”, in Journal of Computer Applications, vol. 37, Issue 10, 2313-2818, Oct. 10, 2017, 7 Pages. [cited by applicant]
Compression of LSP parameters using layered neural networks, The Institute of Electronics, Information and Communication Engineers, Technical Report of IEICE, Oct. 22, 1992, vol. 92 No. 275, 9 Pages. [cited by applicant]