IP Library › Granted Patent US 11,363,390
Granted Patent B2
US 11,363,390 · App. 17/074,144 · Granted Jun 14, 2022

Perceptually guided speech enhancement using deep neural networks

Inventors: Joao Felipe Santos (Montreal, CA); Tao Zhang (Eden Prairie, MN); Yan Zhao (Columbus, OH); Buye Xu (Eden Prairie, MN); Ritwik Giri (Eden Prairie, MN)
Assignee: Starkey Laboratories, Inc.
H04R25/50G10L15/16G10L21/0208G10L25/30G10L25/48G10L25/84H04R25/507A61N1/36039H04R25/407H04R2225/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,363,390
App. No.
17/074,144
Granted
Jun 14, 2022
Kind
B2
Abstract

A method, comprising receiving at least one sound at an electronic device. The at least one sound is enhanced for the at least one user based on a compound metric. The compound metric is calculated using at least two sound metrics selected from an engineering metric, a perceptual metric, and a physiological metric. The engineering metric comprises a difference between an output signal and a desired signal. At least one of the perceptual metric and the physiological metric is based at least in part on input sensed from the at least one user in response to the received at least one sound.

Claims (25)

1. A system, comprising:

an electronic device configured to receive at least one sound;

at least one sensor communicatively coupled to the electronic device, the at least one sensor configured to sense an input from at least one user of the electronic device in response to the received at least one sound; and

a processor communicatively coupled to the electronic device configured to enhance the at least one sound for the at least one user based on a compound metric using a neural network, the compound metric calculated using a physiological metric and at least one sound metric selected from an engineering metric and a perceptual metric, the engineering metric comprising a difference between an output signal and a desired signal.

2. The system of claim 1 , wherein the processor is further configured to train the neural network using training data to build a compound metric model.

3. The system of claim 2 , wherein the training data is based on a at least one user characteristic.

4. The system of claim 3 , wherein the at least one user characteristic comprises a user age range.

5. The system of claim 1 , wherein the processor is further configured to train the neural network based on the user input.

6. The system of claim 1 , wherein the user input comprises one or more of an audible input, a gesture input, and a tactile input.

7. The system of claim 1 , wherein the perceptual metric is calculated by using at least one of a short term objective intelligibility metric (STOI), a hearing-aid speech quality index (HASQI), a hearing-aid speech perception index (HASPI), a perceptual evaluation of speech quality (PESQ), and a perceptual evaluation of audio quality (PEAQ).

8. The system of claim 1 , wherein the compound metric is calculated during a predetermined testing time period.

9. The system of claim 8 , wherein the compound metric is calculated at a plurality of time periods subsequent to the predetermined testing time period.

10. The system of claim 1 , wherein the electronic device is an ear-worn electronic device configured to be worn by the at least one user.

11. The system of claim 1 , wherein the electronic device is a mobile communication device.

12. The system of claim 1 , wherein the at least one sound comprises one or more of speech, music, and an alarm.

13. The system of claim 1 , further comprising calculating the compound metric in response to the user input.

14. The system of claim 1 , further comprising calculating the compound metric in response to the difference between the output signal and the clean signal being greater than a predetermined threshold.

15. The system of claim 1 , wherein the physiological metric comprises at least one of an electroencephalogram (EEG) signal, a skin conductance signal, and a heart rate signal.

16. A method, comprising:

receiving at least one sound at an electronic device; and

using a neural network, enhancing the at least one sound for at least one user based on a compound metric, the compound metric calculated using a physiological metric and at least one of an engineering metric and a perceptual metric, the perceptual metric and the physiological metric based at least in part on input sensed from the at least one user in response to the received at least one sound.

17. The method of claim 16 , wherein enhancing the at least one sound for the at least one user comprises training the neural network using training data to build a compound metric model.

18. The method of claim 17 , wherein the training data is based on a at least one user characteristic.

19. The method of claim 16 , wherein enhancing the at least one sound for the at least one user comprises training the neural network based on the user input.

20. The method of claim 16 , wherein the user input comprises one or more of an audible input, a gesture input, and a tactile input.

Continuity (3)
Continuation 16170858 · Oct 25, 2018
Provisional Application 62577903 · Oct 27, 2017
Related Publication 20210037323A1 · Feb 4, 2021
Cited By (1)
US 12,363,488