IP Library › Granted Patent US 11,508,388
Granted Patent B1
US 11,508,388 · App. 17/100,802 · Granted Nov 22, 2022

Microphone array based deep learning for time-domain speech signal extraction

Inventors: Mehrez Souden (Los Angeles, CA); Symeon Delikaris Manias (Cupertino, CA); Joshua D. Atkins (Los Angeles, CA); Ante Jukic (Los Angeles, CA); Ramin Pishehvar (Cupertino, CA)
Assignee: Apple Inc.
G10L21/0232G06N3/08G10L25/30H04R1/406H04R3/005G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,508,388
App. No.
17/100,802
Filed
Nov 20, 2020
Granted
Nov 22, 2022
Kind
B1
Art Unit
2651
USPC
704/232
Abstract

A device for processing audio signals in a time-domain includes a processor configured to receive multiple audio signals corresponding to respective microphones of at least two or more microphones of the device, at least one of the multiple audio signals comprising speech of a user of the device. The processor is configured to provide the multiple audio signals to a machine learning model, the machine learning model having been trained based at least in part on an expected position of the user of the device and expected positions of the respective microphones on the device. The processor is configured to provide an audio signal that is enhanced with respect to the speech of the user relative to the multiple audio signals, wherein the audio signal is a waveform output from the machine learning model.

Claims (33)

1. A method comprising:

receiving multiple audio signals corresponding to respective microphones of a device, at least one of the multiple audio signals comprising speech of a user of the device;

providing the multiple audio signals to a machine learning model, the machine learning model having been trained based at least in part on an expected position of the user of the device and expected positions of the respective microphones on the device; and

providing, responsive to the providing of the multiple audio signals to the machine learning model, an audio signal that is enhanced with respect to the speech of the user relative to the multiple audio signals, wherein the audio signal is a waveform output from the machine learning model.

2. The method of claim 1 , wherein the machine learning model having been further trained to optimize an application-dependent cost function with respect to the waveform.

3. The method of claim 2 , further comprising:

providing the audio signal to an application related to the application-dependent cost function.

4. The method of claim 1 , wherein the waveform comprises a voice of the user exclusive of other audio data present in the received multiple audio signals.

5. The method of claim 1 , wherein the machine learning model is a deep neural network (DNN).

6. The method of claim 1 , wherein the waveform is a time-domain waveform.

7. The method of claim 1 , wherein each of the multiple audio signals comprise a time-domain waveform.

8. The method of claim 1 , wherein the machine learning model having been further trained to transform audio data of the multiple audio signals into a different domain from time-domain.

9. The method of claim 8 , wherein the machine learning model having been further trained to combine the transformed audio data with one or more estimated filter masks to enhance the audio data with respect to the speech of the user relative to the multiple audio signals.

10. The method of claim 9 , wherein the machine learning model having been further trained to transform the audio data enhanced with respect to the speech of the user relative to the multiple audio signals into time-domain and output the waveform comprising the audio data enhanced with respect to the speech of the user relative to the multiple audio signals in the time-domain.

11. The method of claim 1 , wherein the multiple received audio signals are beamformed signals based on signals of the microphones of the device.

12. A device comprising:

at least two or more microphones;

a processor; and

a memory including instructions that, when executed by the processor, causes the processor to:

receive multiple audio signals corresponding to respective microphones of the at least two or more microphones, at least one of the multiple audio signals comprising speech of a user of the device;

provide the multiple audio signals to a machine learning model, the machine learning model having been trained based at least in part on an expected position of the user of the device and expected positions of the respective microphones on the device; and

provide an audio signal that is enhanced with respect to the speech of the user relative to the multiple audio signals, wherein the audio signal is a waveform output from the machine learning model.

13. The device of claim 12 , wherein the machine learning model having been further trained to optimize an application-dependent cost function with respect to the waveform.

14. The device of claim 12 , wherein the waveform comprises a voice of the user exclusive of other audio data present in the received multiple audio signals.

15. The device of claim 12 , wherein the machine learning model is a deep neural network (DNN).

16. The device of claim 12 , wherein the waveform is a time-domain waveform.

17. The device of claim 12 , wherein each of the multiple audio signals comprise a time-domain waveform.

18. The device of claim 12 , wherein the machine learning model having been further trained to transform audio data of the multiple audio signals into a different domain from time-domain.

19. The device of claim 18 , wherein the machine learning model having been further trained to combine the transformed audio data with one or more estimated filter masks to enhance the audio data with respect to the speech of the user relative to the multiple audio signals.

20. A computer program product comprising code, stored in a non-transitory computer-readable storage medium, the code comprising:

code to receive multiple audio signals corresponding to respective microphones of a device, at least one of the multiple audio signals comprising speech of a user of the device;

code to provide the multiple audio signals to a machine learning model, the machine learning model having been trained based at least in part on an expected position of the user of the device, expected positions of the respective microphones on the device; and

code to provide, to an application, an audio signal that is enhanced with respect to the speech of the user relative to the multiple audio signals, wherein the audio signal is a waveform output from the machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2021
From: SOUDEN, MEHREZ; DELIKARIS MANIAS, SYMEON; ATKINS, JOSHUA D.; JUKIC, ANTE; PISHEHVAR, RAMIN
To: APPLE INC.
Reel/Frame 054969/0808 →
Continuity (1)
Provisional Application 62939528 · Nov 22, 2019
Cited By (4)
US 12,230,259 US 12,417,777 US 12,587,804 US 12,711,961