IP Library Granted Patent US 9,640,193
Granted Patent B2
US 9,640,193 · App. 14/355,458 · Granted May 2, 2017

Systems and methods for enhancing place-of-articulation features in frequency-lowered speech

Inventor: Ying-Yee Kong (Somerville, MA)
Assignee: Northeastern University
G10L21/02G10L21/003G10L21/0364H04R1/08H04R25/353G10L13/02G10L25/18G10L25/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,640,193
App. No.
14/355,458
Granted
May 2, 2017
Kind
B2
Abstract

To improve the intelligibility of speech for users with high-frequency hearing loss, the present systems and methods provide an improved frequency lowering system with enhancement of spectral features responsive to place-of-articulation of the input speech. High frequency components of speech, such as fricatives, may be classified based on one or more features that distinguish place of articulation, including spectral slope, peak location, relative amplitudes in various frequency bands, or a combination of these or other such features. Responsive to the classification of the input speech, a signal or signals may be added to the input speech in a frequency band audible to the hearing-impaired listener, said signal or signals having predetermined distinct spectral features corresponding to the classification, and allowing a listener to easily distinguish various consonants in the input.

Claims (44)

1. A method for frequency-lowering of audio signals for improved speech perception, comprising:

receiving, by an analysis module of a device, a first audio signal;

detecting, by the analysis module, one or more spectral characteristics of the first audio signal, the detected one or more spectral characteristics corresponding to one or more respective non-sonorant sounds;

classifying, by the analysis module, the one or more respective non-sonorant sounds, based on the detected one or more spectral characteristics of the first audio signal;

selecting, by a synthesis module of the device, a second audio signal from a plurality of audio signals, responsive to at least the classification of the one or more respective non-sonorant sounds; and

combining, by the synthesis module of the device, at least a portion of the first audio signal with the second audio signal for output to form a combined audio signal with frequency characteristics audible to the user.

2. The method of claim 1 , wherein detecting one or more spectral characteristics of the first audio signal comprises detecting a spectral slope or a peak location of the first audio signal.

3. The method of claim 1 , wherein detecting the one or more spectral characteristics comprises detecting the one or more spectral characteristics corresponding to the one or more non-sonorant sounds based on identifying that the first audio signal comprises an aperiodic signal above a predetermined frequency.

4. The method of claim 1 , wherein detecting the one or more spectral characteristics comprises detecting the one or more spectral characteristics corresponding to the one or more non-sonorant sounds based on analyzing amplitudes of energy of the first audio signal in one or more predetermined frequency bands.

5. The method of claim 1 further comprising:

classifying the one or more non-sonorant sounds in the first audio signal as belonging to a first group of one of a predetermined plurality of groups having distinct spectral characteristics, based on a spectral slope of the first audio signal not exceeding a threshold.

6. The method of claim 1 further comprising:

classifying the one or more non-sonorant sounds in the first audio signal as belonging to a second group of one of a predetermined plurality of groups having distinct spectral characteristics, based on a spectral slope of the first audio signal exceeding a threshold and a spectral peak location of the first audio signal not exceeding a second threshold.

7. The method of claim 1 further comprising:

classifying the one or more non-sonorant sounds in the first audio signal as belonging to a third group of one of a predetermined plurality of groups having distinct spectral characteristics, based on a spectral slope of the first audio signal exceeding a threshold and a spectral peak location of the first audio signal above a predetermined frequency exceeding a second threshold.

8. The method of claim 1 further comprising:

classifying the one or more non-sonorant sounds in the first audio signal as belonging to a first, second, or third group of one of a predetermined plurality of groups having distinct spectral characteristics, based on amplitudes of energy of the first audio signal in one or more predetermined frequency bands.

9. The method of claim 1 wherein selecting the second audio signal further comprises:

selecting the second audio signal from the plurality of audio signals responsive to the classification of the one or more non-sonorant sounds in the first audio signal, each of the plurality of audio signals comprising a plurality of noise signals and each having a different spectral shape, and wherein the spectral shape of each of the plurality of audio signals is based on the relative amplitudes of each of the plurality of noise signals at a plurality of predetermined frequencies.

10. The method of claim 1 wherein each audio signal of the plurality of audio signals has a different shape, and wherein selecting the second audio signal further comprises:

selecting a given audio signal of the plurality of audio signals having a spectral shape corresponding to spectral features of a given one of the one or more non-sonorant sounds in the first audio signal, responsive to the classification of the given one of the one or more non-sonorant sounds in the first audio signal.

11. The method of claim 1 , wherein combining the first audio signal with the second audio signal comprises combining at least a portion of the one or more non-sonorant sounds in the first audio signal with the second audio signal for output, the second audio signal having an amplitude proportional to a portion of the first audio signal above a predetermined frequency and wherein a portion of the second audio signal includes spectral content below a portion of the first audio signal above a predetermined frequency.

12. The method of claim 1 , further comprising:

receiving, by the analysis module, a third audio signal;

detecting, by the analysis module, one or more spectral characteristics of the third audio signal;

classifying, by the analysis module, the third audio signal as a sonorant sound, based on the detected one or more spectral characteristics of the third audio signal; and

outputting the third audio signal without performing a frequency lowering process.

13. A system for improving speech perception, comprising:

a first transducer for receiving a first audio signal;

an analysis module configured for:

detecting one or more spectral characteristics of the first audio signal, the detected one or more spectral characteristics corresponding to one or more respective non-sonorant sounds; and

classifying the one or more respective non-sonorant sounds, based on the detected one or more spectral characteristics of the first audio signal;

a synthesis module configured for:

selecting a second audio signal from a plurality of audio signals, responsive to at least the classification of the one or more respective non-sonorant sounds; and

combining at least a portion of the first audio signal with the second audio signal for output to form a combined audio signal with frequency characteristics audible to the user; and

a second transducer for outputting the combined audio signal.

14. The system of claim 13 , wherein the analysis module is further configured to detect the one or more spectral characteristics by detecting the one or more spectral characteristics corresponding to the one or more non-sonorant sounds based on identifying that the first audio signal comprises an aperiodic signal above a predetermined frequency.

15. The system of claim 13 , wherein the analysis module is further configured to detect the one or more spectral characteristics by detecting the one or more spectral characteristics corresponding to the one or more non-sonorant sounds based on analyzing amplitudes of energy of the first audio signal in one or more predetermined frequency bands.

16. The system of claim 13 , wherein the analysis module is further configured for classifying the one or more non-sonorant sounds in the first audio signal as belonging to a first group of one of a predetermined plurality of groups having distinct spectral characteristics, based on a spectral slope of the first audio signal not exceeding a threshold.

17. The system of claim 13 , wherein the analysis module is further configured for classifying the one or more non-sonorant sounds in the first audio signal as belonging to a second group of one of a predetermined plurality of groups having distinct spectral characteristics, based on a spectral slope of the first audio signal exceeding a threshold and a spectral peak location of the first audio signal not exceeding a second threshold.

18. The system of claim 13 , wherein the analysis module is further configured for classifying the one or more non-sonorant sounds in the first audio signal as belonging to a third group of one of a predetermined plurality of groups having distinct spectral characteristics, based on a spectral slope of the first audio signal exceeding a threshold and a spectral peak location of the first audio signal above a predetermined frequency exceeding a second threshold.

19. The system of claim 13 , wherein the analysis module is further configured for classifying the one or more non-sonorant sounds in the first audio signal as belonging to a first, second, or third group of one of a predetermined plurality of groups having distinct spectral characteristics, based on amplitudes of energy of the first audio signal in one or more predetermined frequency bands.

20. The system of claim 13 , wherein the synthesis module is further configured for selecting the second audio signal from the plurality of audio signals responsive to the classification of the one or more non-sonorant sounds in the first audio signal, each of the plurality of audio signals comprising a plurality of noise signals and each having a different spectral shape, and wherein the spectral shape of each of the plurality of audio signals is based on the relative amplitudes of each of the plurality of noise signals at a plurality of predetermined frequencies.

21. The system of claim 13 , wherein the synthesis module is further configured for combining at least a portion of the one or more non-sonorant sounds in the first audio signal with the second audio signal, the second audio signal having an amplitude proportional to a portion of the first audio signal above a predetermined frequency and wherein a portion of the second audio signal includes spectral content below a portion of the first audio signal above a predetermined frequency.

Assignments (5)
CONFIRMATORY LICENSE Recorded Apr 27, 2017
From: NORTHEASTERN UNIVERSITY
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 042352/0780 →
CONFIRMATORY LICENSE Recorded Apr 24, 2017
From: NORTHEASTERN UNVIERSITY
To: NATIONAL INSTITUTES OF HEALTH-DIRECTOR DEITR NIH
Reel/Frame 042320/0733 →
CONFIRMATORY LICENSE Recorded Feb 16, 2017
From: NORTHEASTERN UNVIERSITY
To: NATIONAL INSTITUTES OF HEALTH-DIRECTOR DEITR NIH
Reel/Frame 041736/0259 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2017
From: NORTHEASTERN UNIVERSITY
To: KONG, YING-YEE
Reel/Frame 041070/0068 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2014
From: KONG, YING-YEE
To: NORTHEASTERN UNIVERSITY
Reel/Frame 032803/0330 →
Continuity (2)
Provisional Application 61555720 · Nov 4, 2011
Related Publication 20140288938A1 · Sep 25, 2014