IP Library › Granted Patent US 12,198,715
Granted Patent B1
US 12,198,715 · App. 18/388,694 · Granted Jan 14, 2025

System and method for generating impulse responses using neural networks

Inventor: Martin Eineborg (Gardabaer, IS)
Assignee: TREBLE TECHNOLOGIES
G10L25/30G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,715
App. No.
18/388,694
Filed
Nov 10, 2023
Granted
Jan 14, 2025
Kind
B1
Art Unit
2658
USPC
704/232
Abstract

A method for generating an impulse response representing a sound wave propagation from at least one sound source received at a listening point in a room includes obtaining the generated impulse response at the listening point in the room from a neural network architecture by providing at least the position of the listening point as input. The generated impulse response is generated using a neural network architecture. The network is trained by obtaining a 3D model of the room including the at least one sound source emitting sound in the room and obtaining a training group of simulated impulse responses, wherein each simulated impulse response is generated for a respective predefined listening point in the 3D model of the virtual room. An autoencoder is trained by training an encoder of the autoencoder by using the training group of simulated impulse responses as input in order to obtain a corresponding training group of compressed simulated impulse response as outputs and training a decoder of the autoencoder by using the training group of compressed impulse responses as input in order to obtain a corresponding training group of uncompressed simulated impulse response as outputs. An IR neural network is trained using the training group of compressed simulated impulse responses of the autoencoder and the corresponding position of the predefined listening points as input.

Claims (50)

1. A computer-implemented method for generating an impulse response (IR) representing a sound wave propagation from at least one sound source received at a listening point in a room, the method comprising:

obtaining a generated impulse response at the listening point in the room from a neural network architecture by providing to the neural network architecture at least a position of the listening point as input, wherein the generated impulse response is generated using the neural network architecture trained according to:

obtaining a 3D model of the room comprising the at least one sound source virtually emitting sound in the room; and

obtaining a training group of simulated impulse responses, wherein each simulated impulse response is generated for a respective predefined listening point in the 3D model of the room;

training an autoencoder, the training comprising:

training an encoder of the autoencoder by using the training group of simulated impulse responses as input in order to obtain a corresponding training group of compressed simulated impulse responses as outputs; and

training a decoder of the autoencoder by using the training group of compressed impulse responses as input in order to obtain a corresponding training group of uncompressed simulated impulse responses as outputs; and

training an IR neural network using the training group of compressed simulated impulse responses of the autoencoder and the corresponding positions of the predefined listening points as input.

2. The method according to claim 1 , wherein obtaining the generated impulse response further comprises:

generating a compressed generated impulse response using the IR neural network; and

generating the generated impulse response by using the decoder of the trained autoencoder to decompress the compressed generated impulse response.

3. The method according to claim 1 , wherein training the neural network architecture further comprises:

obtaining a validation group of simulated impulse responses, wherein each simulated impulse response is generated for a respective predefined listening point in the 3D model of the room, wherein each of the simulated impulse responses is generated using at least a wave-based solver; and

validating the autoencoder and the neural network using the validation group of simulated impulse responses, wherein the validation group of simulated impulse responses is different from the training group of simulated impulse responses.

4. The method according to claim 1 , wherein training the IR neural network comprises using the 3D model of the room as input.

5. The method according to claim 1 , wherein training the IR neural network comprises using at least one of a position and a directivity of at least one sound source as input.

6. The method according to claim 1 , wherein the simulated impulse responses are generated using at least a wave-based solver.

7. The method according to claim 1 , further comprising:

generating a reverberating audio signal received at the listening point in the room by:

obtaining the generated impulse response;

obtaining an anechoic audio signal; and

generating the reverberating audio signal received at the listening point by convolving the anechoic audio signal and the generated impulse response.

8. The method according to claim 1 , wherein the method further comprises generating a first part of the generated impulse response using a first neural network architecture, wherein a first part of the simulated impulse responses is obtained, and wherein the first part includes a predetermined set of first data points.

9. The method according to claim 8 , wherein the predetermined set of first data points corresponds to at least one of a first, second, and third reverberation of the simulated impulse response.

10. The method according to claim 8 , wherein the predetermined set of first data points includes at least 2000 data points.

11. The method according to claim 8 , wherein the method further comprises generating a second part of the generated impulse response using a second neural network architecture, wherein a second part of the simulate impulse response is obtained, and wherein the second part includes a predetermined set of second data points.

12. The method according to claim 11 , wherein the predetermined set of second data points corresponds to a second reverberation following a first reverberation corresponding to the set of first data points.

13. The method according to claim 11 , wherein the first part of the generated impulse response and the second part of the generated impulse response are combined into a combined generated impulse response.

14. A computer implemented method for training a neural network architecture to generate an impulse response signal for a position in a 3D model of a room, the method comprising:

obtaining a 3D model of the room comprising at least one sound source virtually emitting sound in the room;

obtaining a training group of simulated impulse responses, wherein each simulated impulse response is generated for a respective predefined listening point in the 3D model of the room;

training an autoencoder, the training comprising:

training an encoder of the autoencoder by using the training group of simulated impulse responses as input in order to obtain a corresponding training group of compressed simulated impulse response as outputs; and

training a decoder of the autoencoder by using the training group of compressed impulse responses as input in order to obtain a corresponding training group of uncompressed simulated impulse response as outputs; and

training an IR neural network using the training group of compressed simulated impulse responses of the autoencoder and the corresponding position of the predefined listening points as input.

15. A system for generating an impulse response representing a sound wave propagation from at least one sound source received at a listening point in a room, the system comprising a computer system having processing circuitry coupled to a memory, and a neural network architecture coupled to the computer system, wherein the processing circuitry is configured to:

obtain a generated impulse response at the listening point in the room from the neural network architecture by providing to the neural network architecture at least a position of the listening point as input, wherein the generated impulse response is generated using the neural network architecture trained according to:

obtaining a 3D model of the room comprising the at least one sound source virtually emitting sound in the room; and

obtaining a training group of simulated impulse responses, wherein each simulated impulse response is generated for a respective predefined listening point in the 3D model of the room;

training an autoencoder, the training comprising:

training an encoder of the autoencoder by using the training group of simulated impulse responses as input in order to obtain a corresponding training group of compressed simulated impulse responses as outputs; and

training a decoder of the autoencoder by using the training group of compressed simulated impulse responses as input in order to obtain a corresponding training group of uncompressed simulated impulse responses as outputs; and

training an IR neural network using the training group of compressed simulated impulse responses of the autoencoder and the corresponding positions of the predefined listening points as input.

16. The system according to claim 15 , wherein the processing circuitry is further configured to:

generate a compressed generated impulse response using the IR neural network; and

generate the generated impulse response by using the decoder of the trained autoencoder to decompress the compressed generated impulse response.

17. The system according to claim 15 , wherein the processing circuitry is further configured to obtain an anechoic audio signal and to generate a reverberating audio signal received at the listening point by convolving the anechoic audio signal and the generated impulse response.

18. The system according to claim 15 , wherein the neural network architecture includes a first neural network architecture and wherein processing circuitry is configured to generate a first part of the generated impulse response using the first neural network architecture, wherein a first part of the simulated impulse responses is obtained, and wherein the first part includes a predetermined set of first data points.

19. The system according to claim 18 , wherein the neural network architecture further includes a second neural network architecture, and wherein the processing circuitry is further configured to generate a second part of the generated impulse response using the second neural network architecture, and wherein a second part of the simulated impulse responses is obtained and wherein the second part includes a predetermined set of second data points.

20. The system according to claim 19 , wherein the processing circuitry is further configured to combine the first part of the generated impulse responses and the second part of the generated impulse responses into a combined generated impulse response.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: EINEBORG, MARTIN
To: TREBLE TECHNOLOGIES
Reel/Frame 067448/0754 →
Priority Claims (1)
EP 23196535 · Sep 11, 2023 · regional
References Cited (75)
US 6826483B1 · Anderson et al. · 2004 [cited by applicant]
US 9383464B2 · Shin · 2016 [cited by applicant]
US 9560467B2 · Gorzel et al. · 2017 [cited by applicant]
US 9711126B2 · Mehra et al. · 2017 [cited by applicant]
US 10440498B1 · Gari et al. · 2019 [cited by applicant]
US 10559295B1 · Abel · 2020 [cited by applicant]
US 10777214B1 · Shi et al. · 2020 [cited by applicant]
US 10897570B1 · Robinson et al. · 2021 [cited by applicant]
US 10986444B2 · Mansour et al. · 2021 [cited by applicant]
US 11830471B1 · Mansour et al. · 2023 [cited by applicant]
US 20110015924A1 · Gunel Hacihabiboglu · 2011 [cited by examiner]
US 20150110310A1 · Minnaar · 2015 [cited by applicant]
US 20200214559A1 · Krueger et al. · 2020 [cited by applicant]
US 20200395028A1 · Kameoka · 2020 [cited by examiner]
US 20210074282A1 · Borgstrom · 2021 [cited by examiner]
US 20210074308A1 · Skordilis · 2021 [cited by examiner]
US 20210136510A1 · Tang et al. · 2021 [cited by applicant]
US 20220051479A1 · Agarwal et al. · 2022 [cited by applicant]
US 20220079499A1 · Doron · 2022 [cited by examiner]
US 20220101126A1 · Bharitkar · 2022 [cited by examiner]
US 20220327316A1 · Grauman et al. · 2022 [cited by applicant]
US 20220405602A1 · Yoo · 2022 [cited by examiner]
US 20230164509A1 · Sporer · 2023 [cited by examiner]
US 20230197043A1 · Martinez Ramirez · 2023 [cited by examiner]
US 20230362572A1 · Jang et al. · 2023 [cited by applicant]
WO WO2022167720A1 · 2022 [cited by applicant]
Abadi, M. et al, TensorFlow: A system for large-scale machine learning, Proceedings of the 12th USENIX conference on Operating Systems Design and ImplementationNov. 2016, pp. 265-283. [cited by applicant]
Dozat, T. Incorporating Nesterov Momentum into Adam, Proceedings of the 4th International Conference on Learning Representations, Workshop Track, San Juan, Puerto Rico, May 2-4, 2016, pp. 1-4. [cited by applicant]
Käser, M. et al, An arbitrary high-order discontinuous Galerkin method for elastic waves on unstructured meshes—I. The two-dimensional isotropic case with external source terms, Geophys. J. Int. (2006) 166, 855-877. [cited by applicant]
Majumder, S. et al, Few-Shot Audio-Visual Learning of Environment Acoustics, 36th Conference on Neural Information Processing Systems (NeurIPS 2022). [cited by applicant]
Melander, A. et al, Massively parallel nodal discontinous Galerkin finite element method simulator for room acoustics, The International Journal of High Performance Computing Applications. [cited by applicant]
Pind Jörgensson, F. K, Wave-Based Virtual Acoustics. Technical University of Denmark, (2020) , 194 pages. [cited by applicant]
Pind, F. et al, Time-domain room acoustic simulations with extendedreacting porous absorbers using the discontinuous Galerkin method, J. Acoust. Soc. Am. 148, 2851-2863 (2020). [cited by applicant]
Ratnarajah, A. et al, IR-GAN: Room impulse response generator for far-field speech recognition, INTERSPEECH 2021, 22nd Annual Conference of the International Speech Communication Association, Brno, Czechia, Aug. 30-Sep.… [cited by applicant]
Reed, W. H. and Hill, T. R. 1973. “Triangular mesh methods for the neutron transport equation”, Los Alamos Scientific Laboratory, pp. 1-23 Oct. 31, 1973. [cited by applicant]
Richard, A. et al, Deep Impulse Responses: Estimating And Parameterizing Filters With Deep Networks, arXiv:2202.03416v1 [cs. SD] Feb. 7, 2022. [cited by applicant]
Singh, S. et al, Image2Reverb: Cross-Modal Reverb Impulse Response Synthesis, arXiv:2103.14201v2 [cs.SD] Aug. 13, 2021. [cited by applicant]
N. Ketkar and N. Ketkar, “Introduction to keras,” Deep learning with python: a hands-on introduction, pp. 97-111, 2017. [cited by applicant]
Ahrens, J. et al., “Computation of Spherical Harmonics Based Sound Source Directivity Models from Sparse Measurement Data”, Forum Acusticum, Dec. 7-11, 2020, pp. 2019-2026,HAL open science. [cited by applicant]
Anonymous, “Hybrid Model for Acoustic Simulation” May 15, 2021, pp. 1-6, XP93044320, obtained from Internet: https://reuk.github.io/wayverb/hybrid.html. [cited by applicant]
Aretz, M., “Combined Wave And Ray Based Room Acoustic Simulations Of Small Rooms”, Logos Verlag Berling GmbH, Sep. 2012, pp. 1-211. [cited by applicant]
Atkins, H.L et al, “Quadrature-Free Implementation of Discontinuous Galerkin Method for Hyperbolic Equations”, AIAA Journal vol. 36, No. 5, May 1998, pp. 775-782, Downloaded by North Dakota State University. [cited by applicant]
Bank, D. et al., “Autoencoders”, Version 2, Submitted Apr. 3, 2021, pp. 1-22, Obtained from Internet: https://arxiv.org/abs/2003.05991v2. [cited by applicant]
Bansal, M. et al., “First Approach to Combine Particle Model Algorithms with Modal Analysis using FEM”, Conventional Paper 6392, AES Convention 118, May 28-31, 2005, Barcelona, Spain, pp. 1-9, AES. [cited by applicant]
Berland, J. et al, “Low-dissipation and low-dispersion fourth-order Runge-Kutta algorithm”, Computers & Fluids 35.10, 2006, pp. 1459-1463. [cited by applicant]
Bilbao, S. et al., “Local time-domain spherical harmonic spatial encoding for wave-based acoustic simulation”, IEEE Signal Processing Letters, 26.4, Mar. 1, 2019, pp. 617-621, obtained from Internet: https://www.researc… [cited by applicant]
Cosnefroy, M. “Propagation of impulsive sounds in the atmosphere: numerical simulations and comparison with experiments”, Partly in French, PhD thesisPHD thesis, École Centrale de Lyon, Submitted Dec. 18, 2019, pp. 1-22… [cited by applicant]
Denk, F. et al, “Equalization filter design for achieving acoustic transparency in a semi-open fit hearing device”, Speech Communication; 13th ITG-Symposium, Oct. 10-12, 2018, Oldenburg, Germany, pp. 226-230. [cited by applicant]
Pind, F. et al., “A novel wave-based virtual acoustics and spatial audio framework”, Audio Engineering Society Conference Paper, AVAR Conference, Richmond, VA, Aug. 15-17, 2022, pp. 1-10, AES. [cited by applicant]
Dragna, D. et al “A generalized recursive convolution method for time-domain propagation in porous media”, The Journal of the Acoustical Society of America 138.2, published online Aug. 20, 2015, pp. 1030-1042, https://d… [cited by applicant]
Funkhouser, T., “Survey of Methods for Modeling Sound Propagation in Interactive Virtual Environment Systems”, Department of Computer Science of Princeton University, Jan. 1, 2003, pp. 1-53, XP055746257. [cited by applicant]
Gabard, G. et al.“A full discrete dispersion analysis of time-domain simulations of acoustic liners with flow”, Manuscript, Journal of Computational Physics 273, Received date Nov. 25, 2013, Accepted date May 2, 2014, p… [cited by applicant]
Hart, C. et al., “Machine-learning of long-range sound propagation through simulated atmospheric turbulence”, article, The Journal of the Acoustical Society of America, American Institute of Physics, vol. 149, No. 6, pu… [cited by applicant]
Hesthaven, J.S. et al, “Nodal Discontinuous Galerkin Methods, Algorithms, Analysis, and Applications”, Texts in Applied Mathematics, Chapter 3, pp. 1-507, Springer, New York, 2008. [cited by applicant]
Hu, F.Q. et al., “Low-dissipation and low-dispersion Runge-Kutta schemes for computational acoustics”, Article No. 0052, Journal of Computational Physics 124, 1996, received Dec. 23, 1994, Revised Jul. 1995, pp. 177-191… [cited by applicant]
Jameson, A. et al “Solution of the Euler equations for complex configurations”, 6th Computational Fluid Dynamics Conference, paper No. 83-1929, pp. 1-11, 1983, American Institute of Aeronautics and Astronautics (AIAA), … [cited by applicant]
Sakamoto, S. et al., “Directional sound source modeling by using spherical harmonic functions for finite-difference time-domain analysis”, Proceedings of Meetings on Acoustics, vol. 19, 2013, ICA 2013 Montreal, Jun. 2-7… [cited by applicant]
Yeh C-Y. et al., “Wave-ray coupling for interactive sound propagation in large complex scenes”, ACM Transactions on Graphics, ACM, NY, US, vol. 32, No. 6, Article 165, Nov. 2013, pp. 1-11, XP058033914. [cited by applicant]
Kuttruff, H., “Room Acoustics: 6th edition”, Dec. 10, 2019, pp. 1-302, CRC Press. [cited by applicant]
Yeh C-Y. et al., “Using Machine Learning to Predict Indoor Acoustic Indicators of Multi-Functional Activity Centers”, Article, Applied Sciences, vol. 11, No. 12, Submitted May 28, 2021; Published Jun. 18, 2021, pp. 1-24… [cited by applicant]
Xu, Z. et al, “Simulating room transfer functions between transducers mounted on audio devices using a modified image source method”, J. Coust. Soc. Am., Sep. 8, 2023, Submitted Sep. 7, 2023, arXiv:2309.03486 [eess.AS],… [cited by applicant]
Miccini, R. et al., “A hybrid approach to structural modeling of individualized HRTFs”, 2021 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), Mar. 27, 2021, pp. 80-85, IEEE. [cited by applicant]
Milo, A. et al., “Treble Auralizer: a real time Web Audio Engine enabling 3DoF auralization of simulated room acoustics designs”, Presented at conference 2023 Immersive and 3D Audio: from Architecture to Automotive (I3D… [cited by applicant]
Moreau, S. et al. ,“Study of Higher Order Ambisonic Microphone”, CFA/DAGA'04, Strasbourg, Mar. 24-25, 2004, pp. 215-216. [cited by applicant]
Wang, H. et al., “An arbitrary high-order discontinuous Galerkin method with local time-stepping for linear acoustic wave propagation”, The Journal of the Acoustical Society of America 149.1, publication date Jan. 25, 2… [cited by applicant]
Pind, F. et al, “A phenomenological extended-reaction boundary model for time-domain wave-based acoustic simulations under sparse reflection conditions using a wave splitting method”, preprint submitted to Applied Acous… [cited by applicant]
Strøm, E. et al., “Massively Parallel Nodal Discontinous Galerkin Finite Element Method Simulator for Room Acoustics”, Master thesis, Apr. 2020, pp. 1-133, Technical University of Denmark. [cited by applicant]
Pind, F. et al, “Time-domain room acoustic simulations with extended-reacting porous absorbers using the discontinuous Galerkin method”, The Journal of the Acoustical Society of America 148.5, Nov. 24, 2020, pp. 2851-28… [cited by applicant]
Wang, H. et al., “Time-domain impedance boundary condition modeling with the discontinuous Galerkin method for room acoustics simulations”, The Journal of the Acoustical Society of America 147.4, 2020, pp. 2534-2546, Ac… [cited by applicant]
Savioja, L. et al, “Overview of geometrical room acoustic modeling techniques”, J. Acoust. Soc. Am., 138, published online Aug. 10, 2015, pp. 708-730. [cited by applicant]
Thomas, M.R.P., “Practical Concentric Open Sphere Cardioid Microphone Array Design For Higher Order Sound Field Capture”, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP… [cited by applicant]
Rumelhart, D.E. et al, “Learning Internal Representations by Error Propagation”, Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Foundations, Chapter 8, 1987, pp. 318-362, MIT Press. [cited by applicant]
Sakamoto, S. et al, “Calculation of impulse responses and acoustic parameters in a hall by the finite-difference time-domain method”, Acoust. Sci. & Tech. 29, 4, 2008, accepted Feb. 1, 2008, pp. 256-265, The Acoustical … [cited by applicant]
Sanaguano-Moreno, D.A. et al., “A Deep Learning approach for the Generation of Room Impulse Responses”, 2022 Third International Conference of Information Systems and Software Technologies (ICI2ST), IEEE, Nov. 8, 2022, … [cited by applicant]
Rocchesso, D., “Maximally Diffusive Yet Efficient Feedback Delay Networks for Artificial Reverberation”, Sep. 1997, pp. 1-4, vol. 4, No. 9, IEEE. [cited by applicant]