IP Library Granted Patent US 10,536,775
Granted Patent B1
US 10,536,775 · App. 16/448,259 · Granted Jan 14, 2020

Auditory signal processor using spiking neural network and stimulus reconstruction with top-down attention control

Inventors: Kamal Sen (Boston, MA); Harry Steven Colburn (Jamaica Plain, MA); Junzi Dong (Cambridge, MA); Kenny Feng-Hsu Chou (Brookline, MA)
Assignee: Trustees of Boston University
H04R3/04G10L21/0208H04R3/005H04R5/04H04S7/307H04S2420/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,536,775
App. No.
16/448,259
Granted
Jan 14, 2020
Kind
B1
Abstract

An auditory signal processor includes a filter bank generating frequency components of a source audio signal; a spatial localization network operative in response to the frequency components to generate spike trains for respective spatially separated components of the source audio signal; a cortical network operative in response to the spike trains to generate a resultant spike train for selected spatially separated components of the source audio signal; and a stimulus reconstruction circuit that processes the resultant spike train to generate a reconstructed audio output signal for a target component of the source audio signal. The cortical network incorporates top-down attentional inhibitory modulation of respective spatial channels to produce the resultant spike train for the selected spatially separate components of the source audio signal, and the stimulus reconstruction circuit employs convolution of a reconstruction kernel with the resultant spike train to generate the reconstructed audio output.

Claims (37)

1. An auditory signal processor, comprising:

a filter bank configured and operative to generate a plurality of frequency components of a source audio signal;

a spatial localization network configured and operative in response to the frequency components to generate a plurality of spike trains for respective spatially separated components of the source audio signal;

a cortical network configured and operative in response to the spike trains to generate a resultant spike train for selected spatially separated components of the source audio signal; and

a stimulus reconstruction circuit configured and operative to process the resultant spike train to generate a reconstructed audio signal for a target component of the source audio signal,

wherein (1) the cortical network incorporates top-down attentional inhibitory modulation of respective spatial channels to produce the resultant spike train for the selected spatially separate components of the source audio signal, and (2) the stimulus reconstruction circuit employs convolution of a reconstruction kernel with the resultant spike train to generate the reconstructed audio signal.

2. The auditory signal processor of claim 1 , wherein the convolution produces a time-frequency spike mask that is further processed to produce the reconstructed audio signal.

3. The auditory signal processor of claim 2 , wherein the further processing includes direct voice coding of the time-frequency spike mask to produce the reconstructed audio signal.

4. The auditory signal processor of claim 2 , wherein the further processing includes applying the time-frequency spike mask to the frequency components from the filter bank to produce the reconstructed audio signal.

5. The auditory signal processor of claim 2 , wherein the further processing includes applying the time-frequency spike mask to an extracted envelope of the frequency components from the filter bank to produce the reconstructed audio signal.

6. The auditory signal processor of claim 1 , wherein the top-down attentional inhibitory modulation is performed by structure including:

relay neuron elements selectively relaying corresponding spatial-channel inputs to an output element in the absence of respective first inhibitory inputs;

bottom-up inhibitory neuron elements selectively generating the first inhibitory inputs to corresponding relay neuron elements based on the absence of respective second inhibitory inputs; and

top-down inhibitory neuron elements selectively generating the second inhibitory inputs to corresponding sets of the bottom-up inhibitory neuron elements to cause selection of a corresponding set of spatial-channel inputs to the output element.

7. An auditory device, comprising:

audio input circuitry for receiving or generating a source audio signal;

an auditory signal processor to produce a reconstructed audio signal for a target component of the source audio signal; and

output circuitry to produce an output from the reconstructed audio signal,

wherein the auditory signal processor includes:

a filter bank configured and operative to generate a plurality of frequency components of the source audio signal;

a spatial localization network configured and operative in response to the frequency components to generate a plurality of spike trains for respective spatially separated components of the source audio signal;

a cortical network configured and operative in response to the spike trains to generate a resultant spike train for selected spatially separated components of the source audio signal; and

a stimulus reconstruction circuit configured and operative to process the resultant spike train to generate the reconstructed audio signal for a target component of the source audio signal,

wherein (1) the cortical network incorporates top-down attentional inhibitory modulation of respective spatial channels to produce the resultant spike train for the selected spatially separate components of the source audio signal, and (2) the stimulus reconstruction circuit employs convolution of a reconstruction kernel with the resultant spike train to generate the reconstructed audio signal.

8. The auditory device of claim 7 , wherein the auditory signal processor is configured to receive user input to generate corresponding control for the top-down attentional inhibitory modulation.

9. The auditory device of claim 7 , wherein the convolution of the auditory signal processor produces a time-frequency spike mask that is further processed to produce the reconstructed audio signal.

10. The auditory device of claim 9 , wherein the further processing of the auditory signal processor includes direct voice coding of the time-frequency spike mask to produce the reconstructed audio signal.

11. The auditory device of claim 9 , wherein the further processing of the auditory signal processor includes applying the time-frequency spike mask to the frequency components from the filter bank to produce the reconstructed audio signal.

12. The auditory device of claim 9 , wherein the further processing of the auditory signal processor includes applying the time-frequency spike mask to an extracted envelope of the frequency components from the filter bank to produce the reconstructed audio signal.

13. The auditory device of claim 7 , wherein the top-down attentional inhibitory modulation of the auditory signal processor is performed by structure including:

relay neuron elements selectively relaying corresponding spatial-channel inputs to an output element in the absence of respective first inhibitory inputs;

bottom-up inhibitory neuron elements selectively generating the first inhibitory inputs to corresponding relay neuron elements based on the absence of respective second inhibitory inputs; and

top-down inhibitory neuron elements selectively generating the second inhibitory inputs to corresponding sets of the bottom-up inhibitory neuron elements to cause selection of a corresponding set of spatial-channel inputs to the output element.

14. The auditory device of claim 7 , wherein audio input circuitry include one or more pairs of microphones.

15. The auditory device of claim 7 , wherein the output circuitry includes an output transducer operative to generate an acoustic audio output signal.

16. The auditory device of claim 7 , wherein the source audio signal includes multiple-channel audio or video recordings for which a spatial configuration of recording microphones is known.

17. The auditory device of claim 7 , wherein the output is transcribed into text using voice-to-text technology.

Assignments (2)
CONFIRMATORY LICENSE Recorded Feb 7, 2020
From: BOSTON UNIVERSITY, CHARLES RIVER CAMPUS
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 051845/0915 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2019
From: SEN, KAMAL; COLBURN, HARRY STEVEN; DONG, JUNZI; CHOU, KENNY FENG-HSU
To: TRUSTEES OF BOSTON UNIVERSITY
Reel/Frame 050037/0968 →
Continuity (1)
Provisional Application 62688181 · Jun 21, 2018
Cited By (10)
US 12,356,153 US 12,356,154 US 12,356,156 US 12,363,489 US 12,413,929 US 12,418,756 US 12,574,691 US 12,610,200 US 12,634,642 US 12,713,188