IP Library Granted Patent US 10,715,909
Granted Patent B1
US 10,715,909 · App. 16/215,473 · Granted Jul 14, 2020

Direct path acoustic signal selection using a soft mask

Inventors: Vladimir Tourbabin (Sammamish, WA); Ravish Mehra (Tacoma, WA)
Assignee: FACEBOOK TECHNOLOGIES, LLC
H04R3/005G10L25/18H04R1/406
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,715,909
App. No.
16/215,473
Granted
Jul 14, 2020
Kind
B1
Abstract

One embodiment of the present application sets forth a computer-implemented method that includes receiving, from a first microphone, a first input acoustic signal, generating a first audio spectrum from at least the first input acoustic signal, where the first audio spectrum includes a set of time-frequency bins, for each time-frequency bin included in the set of time-frequency bins, computing a weighted local space-domain distance (LSDD) spectrum value based on a portion of the first audio spectrum that is included in the time-frequency bin, generating a combined spectrum value based on a set of the weighted LSDD spectrum values computed for the set of time-frequency bins, and determining a first estimated direction of the first input acoustic signal based on the combined spectrum value.

Claims (53)

1. A computer-implemented method, comprising:

receiving, from a first microphone, a first input acoustic signal;

generating a first audio spectrum from at least the first input acoustic signal, wherein the first audio spectrum includes a set of time-frequency bins;

for each time-frequency bin included in the set of time-frequency bins, computing a weighted local space-domain distance (LSDD) spectrum value based on a portion of the first audio spectrum that is included in the time-frequency bin;

generating a combined spectrum value based on a set of the weighted LSDD spectrum values computed for the set of time-frequency bins; and

determining a first estimated direction of the first input acoustic signal based on the combined spectrum value.

2. The computer implemented method of claim 1 , wherein computing the weighted LSDD spectrum value comprises:

computing an LSDD spectrum value based on the portion of the first audio spectrum;

computing a weight value associated with the portion of the first audio spectrum; and

combining the LSDD spectrum value with the weight value to generate the weighted LSDD spectrum value.

3. The computer-implemented method of claim 2 , wherein computing the weight value comprises:

computing a first metric associated with the portion of the first audio spectrum; and

computing the weight value based on the first metric and the LSDD spectrum value.

4. The computer-implemented method of claim 3 , wherein the first metric comprises a direct-to-reverberant ratio (DRR) metric that is based on a ratio of a maximum peak value of the LSDD spectrum value relative to an average peak value of the LSDD spectrum value.

5. The computer-implemented method of claim 4 , wherein the weight value is based on an inverse of the DRR metric.

6. The computer-implemented method of claim 1 , wherein generating the first audio spectrum from the first input acoustic signal comprises generating a short-time Fourier transform (STFT) from the first input acoustic signal.

7. The computer-implemented method of claim 1 , wherein the first microphone is included in a wearable headset.

8. One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

receiving, from a first microphone, a first input acoustic signal;

generating a first audio spectrum from at least the first input acoustic signal, wherein the first audio spectrum includes a set of time-frequency bins;

for each time-frequency bin included in the set of time-frequency bins, computing a weighted local space-domain distance (LSDD) spectrum value based on a portion of the first audio spectrum that is included in the time-frequency bin;

generating a combined spectrum value based on a set of the weighted LSDD spectrum values computed for the set of time-frequency bins; and

determining a first estimated direction of the first input acoustic signal based on the combined spectrum value.

9. The non-transitory computer-readable storage media of claim 8 , wherein computing the weighted LSDD spectrum value comprises:

computing an LSDD spectrum value based on the portion of the first audio spectrum;

computing a weight value associated with the portion of the first audio spectrum; and

combining the LSDD spectrum value with the weight value to generate the weighted LSDD spectrum value.

10. The non-transitory computer-readable storage media of claim 9 , wherein computing the weight value comprises:

computing a first metric associated with the portion of the first audio spectrum; and

computing the weight value based on the first metric and the LSDD spectrum value.

11. The non-transitory computer-readable storage media of claim 10 , wherein the first metric comprises a direct-to-reverberant ratio (DRR) metric that is based on a ratio of a maximum peak value of the LSDD spectrum value relative to an average peak value of the LSDD spectrum value.

12. The non-transitory computer-readable storage media of claim 11 , wherein the weight value is based on an inverse of the DRR metric.

13. The non-transitory computer-readable storage media of claim 8 , wherein generating the first audio spectrum from the first input acoustic signal comprises generating a short-time Fourier transform (STFT) from the first input acoustic signal.

14. A wearable device, comprising:

a microphone array that receives a first input acoustic signal; and

a controller that:

generates a first audio spectrum from at least the first input acoustic signal, wherein the first audio spectrum includes a set of time-frequency bins,

for each time-frequency bin included in the set of time-frequency bins, computes a weighted local space-domain distance (LSDD) spectrum value based on a portion of the first audio spectrum that is included in the time-frequency bin,

generates a combined spectrum value based on a set of the weighted LSDD spectrum values computed for the set of time-frequency bins, and

determines a first estimated direction of the first input acoustic signal based on the combined spectrum value.

15. The wearable device of claim 14 , wherein the microphone array comprises two or more distinct microphones at different locations on the wearable device.

16. The wearable device of claim 15 , wherein:

the two or more distinct microphones receive the first input acoustic signal as least two or more acoustic signals; and

the controller adds the two or more acoustic signals to generate a combined input acoustic signal, wherein the first audio spectrum is generated from the combined input acoustic signal.

17. The wearable device of claim 14 , wherein the controller computes the weighted LSDD spectrum value by:

computing an LSDD spectrum value based on the portion of the first audio spectrum;

computing a weight value associated with the portion of the first audio spectrum; and

combining the LSDD spectrum value with the weight value to generate the weighted LSDD spectrum value.

18. The wearable device of claim 17 , wherein the controller computes the weight value by:

computing a first metric associated with the portion of the first audio spectrum; and

computing the weight value based on the first metric and the LSDD spectrum value.

19. The wearable device of claim 18 , wherein the first metric comprises a direct-to-reverberant ratio (DRR) metric that is based on a ratio of a maximum peak value of the LSDD spectrum value relative to an average peak value of the LSDD spectrum value.

20. The wearable device of claim 14 , wherein the controller generates the first audio spectrum from the first input acoustic signal by generating a short-time Fourier transform (STFT) from the first input acoustic signal.

Assignments (2)
CHANGE OF NAME Recorded Jul 12, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060637/0858 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2019
From: TOURBABIN, VLADIMIR; MEHRA, RAVISH
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 048852/0242 →
Continuity (1)
Continuation In Part 15947502 · Apr 6, 2018