IP Library › Granted Patent US 12,444,398
Granted Patent B1
US 12,444,398 · App. 18/476,197 · Granted Oct 14, 2025

Manifold learning for sound field estimation

Inventors: Karim Helwani (San Mateo, CA); Michael Mark Goodwin (Scotts Valley, CA); Paris Smaragdis (Urbana, IL)
Assignee: Amazon Technologies, Inc.
G10K11/17823G10K11/17873G10K2210/12G10K2210/3027G10K2210/3028G10K2210/3035G10K2210/3038G10K2210/505
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,398
App. No.
18/476,197
Granted
Oct 14, 2025
Kind
B1
Abstract

System and methods are provided for estimating the sound field from partial observations. Estimating an acoustic environment for virtual reality and augmented reality applications is a step in the creation of simulated acoustic sound scenes. In particular, the impulse responses of room can be estimated with a generative model. In a teleconferencing scenario with remote participants and a group of participants in a common physical space, giving the remote participants the impression that all other participants are sitting is in the same room acoustically requires filtering the speech of the remote participants with impulse responses estimated at the desired rendering position in the conference room.

Claims (92)

1. A computer-implemented method for estimating a sound field for virtual reality or augmented reality, comprising:

receiving, for a first position associated with a near end room, room data comprising (i) input audio data and (ii) target audio data;

generating measurement vector data from the input audio data;

generating initial input vector data for a second position associated with the near end room;

generating input data from (i) the measurement vector data, (ii) the first position, (iii) the initial input vector data, and (iv) the second position;

applying initial filter parameters to the input data that results in filtered data;

generating target data from the target audio data;

determining an estimated loss from the filtered data and the target data;

determining a matrix from a decoder of a trained variational autoencoder;

combining the matrix, the estimated loss, and a step value that results in a tangent vector;

applying the decoder to a point in a tangent space indicated by the tangent vector, wherein the decoder outputs updated filter parameters;

receiving far end audio data;

in response to receiving the far end audio data, substantially in real-time:

generating near end audio data from (i) the far end audio data, (ii) the updated filter parameters, and (iii) the second position; and

outputting the near end audio data.

2. The computer-implemented method of claim 1 , further comprising:

training a machine learning model with training data comprising a plurality of impulse responses for a second room as input training data and a training label,

wherein the training data further comprises, for each impulse response in the plurality of impulse responses, a position relative to a source in the second room, and

wherein training the machine learning model further comprises:

determining a loss and a gradient of a neural network; and

updating, based on the loss and the gradient, a weight of the neural network that results in the trained variational autoencoder.

3. The computer-implemented method of claim 2 , wherein the training data further comprises a room type for the second room, and the input data further comprises a near end room type.

4. The computer-implemented method of claim 1 , wherein generating the near end audio data further comprises:

applying a machine learning model to the far end audio data, wherein the machine learning model outputs de-reverbed audio data.

5. The computer-implemented method of claim 4 , wherein generating the near end audio data further comprises:

determining a near end impulse response from the updated filter parameters at the second position; and

applying the near end impulse response at the second position to the de-reverbed audio data that results in the near end audio data as reverbed.

6. The computer-implemented method of claim 1 , further comprising:

iteratively determining filter parameters until a threshold is satisfied.

7. One or more non-transitory computer-readable storage media storing computer executable instructions that when executed by a computing system perform operations comprising:

receiving, for a first position associated with a near end room, room data comprising (i) input audio data and (ii) target audio data;

generating measurement vector data from the input audio data;

generating initial input vector data for a second position associated with the near end room;

generating input data from (i) the measurement vector data, (ii) the first position, (iii) the initial input vector data, and (iv) the second position;

applying initial filter parameters to the input data that results in filtered data;

generating target data from the target audio data;

determining an estimated loss from the filtered data and the target data;

determining a matrix from a decoder of a trained generative model;

combining the matrix, the estimated loss, and a step value that results in a tangent vector;

applying the decoder to a point in a tangent space indicated by the tangent vector, wherein the decoder outputs updated filter parameters;

receiving far end audio data;

in response to receiving the far end audio data, substantially in real-time:

generating near end audio data from (i) the far end audio data, (ii) the updated filter parameters, and (iii) the second position; and

transmitting the near end audio data.

8. The one or more non-transitory computer-readable storage media of claim 7 storing further computer-executable instructions that when executed by the computing system perform further operations comprising:

training a machine learning model with training data comprising a plurality of impulse responses for a second room as input training data and a training label,

wherein the training data further comprises, for each impulse response in the plurality of impulse responses, a position relative to a source in the second room, and

wherein training the machine learning model further comprises:

determining a loss and a gradient of a neural network; and

updating, based on the loss and the gradient, a weight of the neural network that results in the trained generative model.

9. The one or more non-transitory computer-readable storage media of claim 8 , wherein determining the loss of the neural network further comprises:

applying a loss function with a regularization term, wherein the regularization term causes a representation of a Hessian matrix of an adaptive filter cost function in a latent space to be approximately diagonal.

10. The one or more non-transitory computer-readable storage media of claim 7 , wherein combining the matrix, the estimated loss, and the step value further comprises:

calculating an inverse Hessian matrix from the matrix; and

calculating the tangent vector from the matrix, the inverse Hessian matrix, the estimated loss, and the step value.

11. The one or more non-transitory computer-readable storage media of claim 7 , wherein generating the near end audio data further comprises:

applying a machine learning model to the far end audio data, wherein the machine learning model outputs de-reverbed audio data.

12. The one or more non-transitory computer-readable storage media of claim 11 , wherein generating the near end audio data further comprises:

determining a near end impulse response from the updated filter parameters at the second position; and

applying the near end impulse response at the second position to the de-reverbed audio data that results in the near end audio data as reverbed.

13. The one or more non-transitory computer-readable storage media of claim 7 , wherein the trained generative model comprises a variational autoencoder.

14. A system comprising:

a non-transitory data storage medium; and

a computer hardware processor in communication with the non-transitory data storage medium, wherein the computer hardware processor is configured to execute computer-executable instructions to at least:

receive, for a first position associated with a near end room, room data comprising (i) input audio data and (ii) target audio data;

generate measurement vector data from the input audio data;

generate initial input vector data for a second position associated with the near end room;

generate input data from (i) the measurement vector data, (ii) the first position, (iii) the initial input vector data, and (iv) the second position;

apply initial filter parameters to the input data that results in filtered data;

generate target data from the target audio data;

determine an estimated loss from the filtered data and the target data;

determine a matrix from a decoder of a trained generative model;

combine the matrix, the estimated loss, and a step value that results in a tangent vector;

apply the decoder to a point in a tangent space indicated by the tangent vector, wherein the decoder outputs updated filter parameters;

receive far end audio data;

generate near end audio data from (i) the far end audio data, (ii) the updated filter parameters, and (iii) the second position; and

transmit the near end audio data.

15. The system of claim 14 , wherein the computer hardware processor executes additional computer-executable instructions to at least:

train a machine learning model with training data comprising a plurality of impulse responses for a second room as input training data and a training label,

wherein the training data further comprises, for each impulse response in the plurality of impulse responses, a position relative to a source in the second room, and

wherein to train the machine learning model, the computer hardware processor executes the additional computer-executable instructions to at least:

determine a loss and a gradient of a neural network; and

update, based on the loss and the gradient, a weight of the neural network that results in the trained generative model.

16. The system of claim 15 , wherein to train the machine learning model with the training data, the computer hardware processor executes further computer-executable instructions to at least:

apply a Kirchhoff-Helmholtz integral to the plurality of impulse responses at a respective position relative to the source in the second room.

17. The system of claim 15 , wherein the training data further comprises a room type for the second room, and the input data further comprises a near end room type.

18. The system of claim 17 , wherein the room type comprises at least one of a small room type, a medium room type, or a large room type.

19. The system of claim 14 , wherein to generate the near end audio data, the computer hardware processor executes additional computer-executable instructions to at least:

apply a machine learning model to the far end audio data, wherein the machine learning model outputs de-reverbed audio data.

20. The system of claim 19 , wherein to generate the near end audio data, the computer hardware processor executes further computer-executable instructions to at least:

determine a near end impulse response from the updated filter parameters at the second position; and

apply the near end impulse response at the second position to the de-reverbed audio data that results in the near end audio data as reverbed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2023
From: HELWANI, KARIM; GOODWIN, MICHAEL MARK; SMARAGDIS, PARIS
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 065141/0312 →
References Cited (29)
US 20020071573A1 · Finn · 2002 [cited by examiner]
US 20030235312A1 · Pessoa · 2003 [cited by examiner]
US 20140307882A1 · LeBlanc · 2014 [cited by examiner]
US 20230224635A1 · Tian · 2023 [cited by examiner]
Belkin et al., “Laplacian eigenmaps and spectral techniques for embedding and clustering,” Advances in neural information processing systems, vol. 14, 2001. [cited by applicant]
Berkhout et al., “A holographic approach to acoustic control,” Journal of The Audio Engineering Society, vol. 36, pp. 977-995, 1988. [cited by applicant]
Benesty et al., “A robust fast recursive least squares adaptive algorithm,” in 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 01CH37221), 2001, vol. 6, pp. 3785-378… [cited by applicant]
Boche et al., “Limitations of deep learning for inverse problems on digital hardware,” arXiv preprint arXiv:202.13490, 2022. [cited by applicant]
Buchner et al., “Adaptive dynamical systems in compressive domains as a manifold learning framework,” in SPARS Workshop, 2015. [cited by applicant]
Buchner et al., “A systematic approach to incorporate deterministic prior knowledge in broadband adaptive mimo systems,” in 2010 Conference Record of the Forty Fourth Asilomar Conference on Signals, Systems and Computer… [cited by applicant]
Buchner et al., “Unsupervised bayesian estimation and tracking of time-varying convolutive multichannel systems,” in 2019 22th International Conference on Information Fusion (Fusion). IEEE, 2019, pp. 1-8. [cited by applicant]
Casebeer et al., “Meta-af: Meta-learning for adaptive filters,” arXiv preprint arXiv:2204.11942, 2022. [cited by applicant]
Coifman et al., Applied and Computational Harmonic Analysis, vol. 21, No. 1, pp. 5-30, 2006, Special Issue: Diffusion Maps and Wavelets. [cited by applicant]
Donoho et al., “Hessian eigenmaps: Locally linear embedding techniques for high-dimensional data,” Proceedings of the National Academy of Sciences, vol. 100, No. 10, pp. 5591-5596, 2003. [cited by applicant]
Edelman et al. “The geometry of algorithms with orthogonality constraints,” 1998. [cited by applicant]
Griffiths et al., “An alternative approach to linearly constrained adaptive beamforming,” IEEE Transactions on Antennas and Propagation, vol. 30, No. 1, pp. 27-34, 1982. [cited by applicant]
Helwani et al., “Multichannel adaptive filtering in compressive domains,” in 2014 14th International Workshop on Acoustic Signal Enhancement (IWAENC). IEEE, 2014, pp. 174-177. [cited by applicant]
Helwani et al., “Multichannel adaptive filtering with sparseness constraints,” in IWAENC 2012; International Workshop on Acoustic Signal Enhancement, 2012, pp. 1-4. [cited by applicant]
Kumar et al., “Variational inference of disentangled latent concepts from unlabeled observations,” 2017. [cited by applicant]
Luo et al., “Gaussian process models for hrtf based sound-source localization and active-learning,” arXiv preprint arXiv:1502.03163, 2015. [cited by applicant]
Moor et al., “Topological autoencoders,” 2021. [cited by applicant]
Posada et al., “Simplicial autoencoders: A connection between algebraic topology and probabilistic modelling,” 2018. [cited by applicant]
Plumbley, “Geometry and manifolds for independent component analysis,” in 2007 IEEE International Conference on Acoustics, Speech and Signal Processing—ICASSP '07, 2007, vol. 4, pp. IV-1397-IV-1400. [cited by applicant]
Roweis et al., “Nonlinear dimensionality reduction by locally linear embedding,” science, vol. 290, No. 5500, pp. 2323-2326, 2000. [cited by applicant]
Scheibler et al., Eric Bezzam, and Ivan Dokmanic, “Pyroomacoustics: A python package for audio room simulation and array processing algorithms,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Proc… [cited by applicant]
Shun-Ichi et al., “Natural Gradient Works Efficiently in Learning,” Neural Computation, vol. 10, No. 2, pp. 251-276, Feb. 1998. [cited by applicant]
Talmon et al., “Diffusion maps for signal processing: A deeper look at manifold-learning techniques based on kernels and graphs,” IEEE signal processing magazine, vol. 30, No. 4, pp. 75-86, 2013. [cited by applicant]
Valin et al., “A hybrid dsp/deep learning approach to realtime full-band speech enhancement,” 2017. [cited by applicant]
Valin et al., “Low-complexity, real-time joint neural echo control and speech enhancement based on percepnet,” 2021. [cited by applicant]