IP Library › Granted Patent US 11,678,111
Granted Patent B1
US 11,678,111 · App. 17/348,336 · Granted Jun 13, 2023

Deep-learning based beam forming synthesis for spatial audio

Inventors: Shai Messingher Lang (Santa Clara, CA); Symeon Delikaris Manias (Los Angeles, CA)
Assignee: Apple Inc.
H04R1/406G06F3/012G06N3/08G10L25/30H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,678,111
App. No.
17/348,336
Granted
Jun 13, 2023
Kind
B1
Abstract

A machine learning model can determine output frequency response at different directions relative to a target audio output format, based on input including frequency response at directions relative to microphones of a capture device. A spatial filter determined based on the output frequency responses is applied to one or more of the microphone signals to map the spatial information from the microphone signals to the target audio output.

Claims (33)

1. A method for spatial audio reproduction comprising:

obtaining a plurality of microphone signals representing sounds sensed by a plurality of microphones;

providing, as input to a machine learning model, a frequency response for each of a plurality of directions around each of the plurality of microphones;

obtaining, from the machine learning model, an output frequency response for each of a second plurality of directions associated with audio channels of a target audio output format; and

applying spatial filter parameters, determined based on the output frequency response, to one or more microphone signals selected from the plurality of microphone signals, resulting in output audio signals for each of the audio channels of the target audio output format.

2. The method of claim 1 , wherein the target output audio output format is one of the following: a binaural output, a 3D speaker layout, or surround loudspeaker layout.

3. The method of claim 1 , wherein the machine learning model performs a non-linear least-squares optimization to determine the output frequency response for each of the second plurality of directions associated with the audio channels of the target audio output format.

4. The method of claim 1 , wherein during training of the machine learning model, a cost function that represents a difference between the output audio signals and a sample recording of the sensed sounds, the sample recording using microphones placed at locations corresponding to the target audio output format, is used to adjust the machine learning model to minimize the difference.

5. The method of claim 4 , wherein the cost function includes a speech intelligibility cost term.

6. The method of claim 4 , wherein the cost function includes a signal distortion ratio as a cost term.

7. The method of claim 1 , wherein the spatial filter parameters are updated based on the output frequency response and a tracked position of a user's head.

8. The method of claim 1 , wherein the machine learning model includes a trained neural network that is trained using a plurality of audio recordings recorded with a second plurality of microphones having a geometrical arrangement resembling that of the plurality of microphones.

9. The method of claim 1 , wherein speakers are driven with the output audio signals to generate sound that, when perceived by a listener, spatially resemble the sounds as sensed by the plurality of microphones.

10. A spatial audio reproduction system comprising

a processor, configured to perform the following:

obtaining a plurality of microphone signals representing sounds sensed by a plurality of microphones;

providing, as input to a machine learning model, a frequency response for each of a plurality of directions around each of the plurality of microphones;

obtaining, from the machine learning model, an output frequency response for each of a second plurality of directions associated with audio channels of a target audio output format; and

applying spatial filter parameters, determined based on the output frequency response, to one or more microphone signals selected from the plurality of microphone signals, resulting in output audio signals for each of the audio channels of the target audio output format.

11. The spatial audio reproduction system of claim 10 , wherein the target output audio output format is one of the following: a binaural output, a 3D speaker layout, or surround loudspeaker layout.

12. The spatial audio reproduction system of claim 10 , wherein the machine learning model performs a non-linear least-squares optimization to determine the output frequency response for each of the second plurality of directions associated with the audio channels of the target audio output format.

13. The spatial audio reproduction system of claim 10 , wherein during training of the machine learning model, a cost function represents a difference between the output audio signals and a sample recording of the sensed sounds, the sample recording using microphones placed at locations corresponding to the target audio output format, is used to adjust the machine learning model to minimize the difference.

14. The spatial audio reproduction system of claim 13 , wherein the cost function includes a speech intelligibility cost term.

15. The spatial audio reproduction system of claim 13 , wherein the cost function includes a signal distortion ratio as a cost term.

16. The spatial audio reproduction system of claim 10 , wherein the spatial filter parameters are updated based on the output frequency response and a tracked position of a user's head.

17. The spatial audio reproduction system of claim 10 , wherein the machine learning model includes a trained neural network that is trained using a plurality of audio recordings recorded with a second plurality of microphones having a geometrical arrangement resembling that of the plurality of microphones.

18. The spatial audio reproduction system of claim 10 , wherein speakers are driven with the output audio signals to generate sound that, when perceived by a listener, spatially resemble the sounds as sensed by the plurality of microphones.

19. A non-transitory computer readable medium having stored therein instructions that, when executed by a processor, causes performance of the following:

obtaining a plurality of microphone signals representing sounds sensed by a plurality of microphones;

providing, as input to a machine learning model, a frequency response for each of a plurality of directions around each of the plurality of microphones;

obtaining, from the machine learning model, an output frequency response for each of a second plurality of directions associated with audio channels of a target audio output format; and

applying spatial filter parameters, determined based on the output frequency response, to one or more microphone signals selected from the plurality of microphone signals, resulting in output audio signals for each of the audio channels of the target audio output format.

20. The computer readable medium of claim 19 , wherein the target output audio output format is one of the following: a binaural output, a 3D speaker layout, or surround loudspeaker layout.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2021
From: MESSINGHER LANG, SHAI; DELIKARIS MANIAS, SYMEON
To: APPLE INC.
Reel/Frame 056551/0887 →
Continuity (1)
Provisional Application 63054924 · Jul 22, 2020
Cited By (8)
US 12,425,782 US 12,470,880 US 12,574,691 US 12,581,250 US 12,610,200 US 12,634,642 US 12,713,188 US 12,750,623