IP Library › Granted Patent US 12,363,488
Granted Patent B2
US 12,363,488 · App. 18/367,017 · Granted Jul 15, 2025

Electronic device using a compound metric for sound enhancement

Inventors: Joao Felipe Santos (Montreal, CA); Tao Zhang (Eden Prairie, MN); Yan Zhao (Columbus, OH); Buye Xu (Eden Prairie, MN); Ritwik Giri (Eden Prairie, MN)
Assignee: Starkey Laboratories, Inc.
H04R25/50G10L15/16G10L21/0208G10L25/30G10L25/48G10L25/84H04R25/507A61N1/36039H04R25/407H04R2225/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,363,488
App. No.
18/367,017
Granted
Jul 15, 2025
Kind
B2
Abstract

A method, comprising receiving at least one sound at an electronic device. The at least one sound is enhanced for the at least one user based on a compound metric. The compound metric is calculated using at least two sound metrics selected from an engineering metric, a perceptual metric, and a physiological metric. The engineering metric comprises a difference between an output signal and a desired signal. At least one of the perceptual metric and the physiological metric is based at least in part on input sensed from the at least one user in response to the received at least one sound.

Claims (39)

1. A system, comprising:

an electronic device configured to receive at least one sound, the electronic device comprising an ear-worn electronic device or a mobile communication device;

at least one sensor communicatively coupled to the electronic device, the at least one sensor configured to sense an input from at least one user of the electronic device in response to the received at least one sound; and

a processor communicatively coupled to the electronic device and configured to enhance the at least one sound for the at least one user based on a compound metric using a neural network defining a compound metric model and a loss function.

2. The system of claim 1 , wherein the processor is configured to calculate the compound metric using two or more of a perceptual metric, an engineering metric, and a physiological metric.

3. The system of claim 1 , wherein the neural network is configured as a denoising system that provides noisy speech enhancement.

4. The system of claim 1 , wherein a perceptual metric is incorporated in the loss function.

5. The system of claim 1 , wherein a short term objective intelligibility (STOI) metric is incorporated in the loss function.

6. The system of claim 1 , wherein the neural network comprises one or more deep neural networks.

7. The system of claim 1 , wherein the processor is configured to calculate the compound metric in response to a user input.

8. The system of claim 1 , wherein:

the at least one sensor comprises one or more of an electroencephalogram (EEG) sensor configured to produce an EEG signal, a skin conductance sensor configured to produce a skin conductance signal, and a heart rate sensor configured to produce a heart rate signal; and

the processor is configured to calculate a physiological metric of the compound metric using at least one of the EEG signal, the skin conductance signal, and the heart rate signal.

9. The system of claim 1 , wherein:

the processor is further configured to train the neural network using training data to build the compound metric model; and

the training data is based on at least one user characteristic.

10. The system of claim 1 , wherein:

the processor is further configured to train the neural network based on a user input; and

the user input comprises one or more of an audible input, a gesture input, and a tactile input.

11. The system of claim 1 , wherein the processor is configured to calculate the compound metric during a predetermined testing time period.

12. The system of claim 11 , wherein the processor is configured to calculate the compound metric at a plurality of time periods subsequent to the predetermined testing time period.

13. The system of claim 1 , wherein the at least one sound comprises one or more of speech, music, and an alarm.

14. A method, comprising:

receiving at least one sound at an electronic device, the electronic device comprising an ear-worn electronic device or a mobile communication device;

sensing an input from at least one user of the electronic device in response to the received at least one sound; and

using a neural network defining a compound metric model and a loss function, enhancing the at least one sound for the at least one user based on a compound metric.

15. The method of claim 14 , comprising calculating the compound metric using two or more of a perceptual metric, an engineering metric, and a physiological metric.

16. The method of claim 14 , wherein a perceptual metric is incorporated in the loss function.

17. The method of claim 14 , wherein a short term objective intelligibility (STOI) metric is incorporated in the loss function.

18. The method of claim 14 , comprising calculating the compound metric in response to the user input.

19. The method of claim 14 , comprising calculating the compound metric using at least one of an EEG signal, a skin conductance signal, and a heart rate signal sensed from the at least one user.

20. The method of claim 14 , comprising:

training the neural network using training data to build the compound metric model;

wherein the training data is based on at least one user characteristic.

21. The method of claim 14 , comprising:

training the neural network based on a user input;

wherein the user input comprises one or more of an audible input, a gesture input, and a tactile input.

22. The method of claim 14 , comprising calculating the compound metric during a predetermined testing time period.

23. The method of claim 22 , comprising calculating the compound metric at a plurality of time periods subsequent to the predetermined testing time period.

Continuity (5)
Continuation 17834288 · Jun 7, 2022
Continuation 17074144 · Oct 19, 2020
Continuation 16170858 · Oct 25, 2018
Provisional Application 62577903 · Oct 27, 2017
Related Publication 20230421973A1 · Dec 28, 2023
References Cited (27)
US 9633671B2 · Giacobello et al. · 2017 [cited by applicant]
US 10812915B2 · Santos · 2020 [cited by examiner]
US 11363390B2 · Santos · 2022 [cited by examiner]
US 11812223B2 · Santos · 2023 [cited by examiner]
US 20160055420A1 · Karanam et al. · 2016 [cited by applicant]
US 20170061978A1 · Wang et al. · 2017 [cited by applicant]
US 20170123824A1 · Franck · 2017 [cited by examiner]
EP 3229496 · 2017 [cited by applicant]
Clevert et al., “Fast and accurate deep network learning by exponential linear units (ELUs),” arXiv preprint arXiv:1511.07289, 2015. [cited by applicant]
Emiya et al., “Subjective and objective quality assessment of audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, pp. 2046-2057, 2011. [cited by applicant]
International Search Report and Written Opinion dated Feb. 18, 2019 from PCT Application No. PCT/US2018/057713, 11 pages. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014. [cited by applicant]
Koizumi et al., “Dnn-based source enhancement self-optimized by reinforcement learning using sound quality measurements”, IEEE International Conference on Acoustics, Speech and Signal Processing, 2017, pp. 81-85. [cited by applicant]
Li et al., “An overview of noise-robust automatic speech recognition”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 22, 2014, pp. 745-777. [cited by applicant]
Ming et al., “Robust speaker recognition in noisy conditions,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 15, pp. 1711-1723, 2007. [cited by applicant]
Rix et al., “Perceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs,” in IEEE International Conference on Acoustics, Speech, and Signal Processing, 200… [cited by applicant]
Rothauser et al., “IEEE recommended practice for speech quality measurements”, IEEE Transactions on Audio Electroacoust, vol. 17, 1969, pp. 225-246. [cited by applicant]
Srivastava et al., “Dropout: a simple way to prevent neural networks from overfitting.,” Journal of Machine Learning Research, vol. 15, pp. 1929-1958, 2014. [cited by applicant]
Taal et al., “An algorithm for intelligibility prediction of time-frequency weighted noisy speech,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, pp. 2125-2136, 2011. [cited by applicant]
Varga et al., “Assessment for automatic speech recognition: li noisex-92: A database and an experiment to study the effect of additive noise on speech recognition systems”, Speech Communication, vol. 12, 1993, pp. 247-2… [cited by applicant]
Vincent et al., “Performance measurement in blind audio source separation,” IEEE transactions on audio, speech, and language processing, vol. 14, pp. 1462-1469, 2006. [cited by applicant]
Wang et al., “A deep neural network for time-domain signal reconstruction”, IEEE International Conference on Acoustics, Speech and Signal Processing, 2015, pp. 4390-4394. [cited by applicant]
Wang et al., “On training targets for supervised speech separation”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 22, 2014, pp. 1849-1858. [cited by applicant]
Weninger et al., “Discriminatively trained recurrent neural networks for singlechannel speech separation,” in IEEE Global Conference on Signal and Information Processing (GlobalSIP), pp. 577-581, 2014. [cited by applicant]
Williamson et al., “Complex ratio masking for monaural speech separation,” IEEE/ACM transactions on audio, speech, and language processing, vol. 24, pp. 483-492, 2016. [cited by applicant]
Xu et al., “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 23, pp. 7-19, 2015. [cited by applicant]
Zhao et al., “A two-stage algorithm for noisy and reverberant speech enhancement,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5580-5584, 2017. [cited by applicant]