IP Library › Granted Patent US 12,340,789
Granted Patent B2
US 12,340,789 · App. 17/509,892 · Granted Jun 24, 2025

Hearing apparatus with bone conduction sensor

Inventors: Andreas Tiefenau (Gammel Holte, DK); Brian Dam Pedersen (Ringsted, DK); Antonie Johannes Hendrikse (Eindhoven, NL); Anuj Dev (Amsterdam, NL)
Assignee: GN HEARING A/S
G10L13/047G10L13/02H04R25/507H04R25/606H04R25/554H04R2225/55
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,789
App. No.
17/509,892
Granted
Jun 24, 2025
Kind
B2
Abstract

The present disclosure relates to a hearing apparatus comprising: a bone conduction sensor configured to convert bone vibrations of voice sound information into a bone conduction signal; a signal processing unit configured to implement a synthetic speech generation process, the synthetic speech generation process implementing a speech model; wherein the synthetic speech generation process receives the bone conduction signal as a control input and outputs a synthetic speech signal.

Claims (39)

1. A hearing apparatus comprising:

a bone conduction sensor configured to provide a bone conduction signal indicative of bone-conducted vibration conducted by a bone of a wearer of the hearing apparatus; and

a signal processing unit comprising a speech model, the speech model comprising a neural network model configured to obtain a representation of the bone conduction signal as a control input, wherein the signal processing unit is configured to provide a synthetic speech signal;

wherein the signal processing unit is configured to predict a current sample of a time series from one or more previous samples of the time series, the time series representing a speech waveform, wherein the signal processing unit is configured to predict the current sample of the time series based on the representation of the bone conduction signal;

wherein the neural network model comprises a layer implemented as a part of the neural network model of the signal processing unit that provides the synthetic speech signal;

wherein the neural network model is trained based on a plurality of training speech samples; and

wherein at least one of the training speech samples comprises a training bone conduction data representing a speech and a corresponding training microphone data representing airborne sound of the speech, the training microphone data and the training bone conduction data corresponding with each other temporally.

2. The hearing apparatus according to claim 1 , wherein the speech model defines an internal state that evolves over time.

3. The hearing apparatus according to claim 1 , wherein the neural network comprises a recurrent neural network.

4. The hearing apparatus according to claim 3 , wherein the recurrent neural network has a density estimation mode during operation.

5. The hearing apparatus according to claim 1 , wherein the neural network comprises a layered neural network comprising two or more layers, at least one of the two or more layers being a softmax layer.

6. The hearing apparatus according to claim 1 , wherein the speech model comprises an autoregressive speech model.

7. The hearing apparatus according to claim 1 , wherein the speech model is configured to compute a probability distribution over a plurality of output classes, at least one of the output classes representing a sample value of a sample of a sampled audio waveform.

8. The hearing apparatus according to claim 1 , further comprising a head-worn hearing device, the head-worn hearing device comprising the bone conduction sensor and a first communication interface.

9. The hearing apparatus according to claim 8 , wherein the head-worn hearing device further comprises the signal processing unit, and wherein the head-worn hearing device is configured to communicate the synthetic speech signal via the first communication interface to a handheld communication device.

10. The hearing apparatus according to claim 8 , further comprising a signal processing device, the signal processing device comprising the signal processing unit and a second communication interface;

wherein the first communication interface of the head-worn hearing device is configured to communicate the bone conduction signal to the second communication interface of the signal processing device.

11. The hearing apparatus according to claim 1 , further comprising an ambient microphone configured to detect air-borne speech spoken by the wearer of the hearing apparatus, and to provide an ambient microphone signal indicative of the detected air-borne speech.

12. The hearing apparatus according to claim 1 , further comprising a memory configured to store training data, the training data comprising one or more signal pairs, at least one of the signal pairs comprising the training bone conduction data and the training microphone data.

13. The hearing apparatus according to claim 1 , wherein the signal processing unit is configured to generate a synthetic filtered signal corresponding to a speech signal filtered by a first filter, after receiving the representation of the bone conduction signal as the control input.

14. The hearing apparatus according to claim 13 , wherein the synthetic filtered signal is the synthetic speech signal.

15. The hearing apparatus according to claim 1 , wherein the hearing apparatus is a hearing aid.

16. The hearing apparatus according to claim 1 , wherein the hearing apparatus is a BTE, RIE, ITE, ITC or CIC hearing instrument.

17. A hearing apparatus comprising:

a bone conduction sensor configured to provide a bone conduction signal indicative of bone-conducted vibration conducted by a bone of a wearer of the hearing apparatus; and

a signal processing unit comprising a speech model, the speech model comprising a neural network model configured to obtain a representation of the bone conduction signal as a control input, wherein the signal processing unit is configured to provide a synthetic speech signal;

wherein the signal processing unit is configured to generate a synthetic filtered signal corresponding to a speech signal filtered by a first filter; and

wherein the signal processing unit is configured to receive an ambient microphone signal associated with an ambient microphone, the ambient microphone signal and the bone conduction signal corresponding with each other temporally, and

wherein the signal processing unit is configured to create a filtered version of the received ambient microphone signal using a second filter, and to combine the generated synthetic filtered signal with the created filtered version of the received ambient microphone signal to create the synthetic speech signal.

18. The hearing apparatus according to claim 17 , wherein the speech model is a machine learning model, and wherein the machine learning model is trained based on a plurality of training speech samples.

19. The hearing apparatus according to claim 18 , wherein at least one of the training speech samples comprises a training bone conduction data and a corresponding training microphone data, the training microphone data and the training bone conduction data corresponding with each other temporally.

20. The hearing apparatus according to claim 17 , wherein the signal processing unit, when in a training mode, is configured to adapt one or more model parameters of the speech model.

21. The hearing apparatus according to claim 20 , wherein the adapted one or more model parameters are configured to allow the speech model to provide an improved match between a model output representing the synthetic speech and a corresponding training ambient microphone signal.

22. A processor-implemented method of obtaining a synthetic speech signal, comprising:

receiving, by a processing unit of an apparatus, a bone conduction signal from a bone conduction sensor, the bone conduction sensor configured to detect a bone-conducted vibration conducted by a bone of a person; and

using the signal processing unit to predict a current sample of a time series from one or more previous samples of the time series, the time series representing a speech waveform, wherein the signal processing unit is configured to predict the current sample of the time series based on the bone conduction signal, wherein the signal processing unit comprises a neural network model configured to receive the bone conduction signal as a control input;

wherein the neural network model comprises a layer implemented as a part of the neural network model of the signal processing unit;

wherein the neural network model is trained based on a plurality of training speech samples; and

wherein at least one of the training speech samples comprises a training bone conduction data representing a speech and a corresponding training microphone data representing airborne sound of the speech, the training microphone data and the training bone conduction data corresponding with each other temporally.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2025
From: TIEFENAU, ANDREAS; PEDERSEN, BRIAN DAM; HENDRIKSE, ANTONIE JOHANNES; DEV, ANUJ
To: GN HEARING A/S
Reel/Frame 070487/0431 →
Priority Claims (1)
EP 19172713 · May 6, 2019 · regional
Continuity (2)
Continuation PCTEP2020062561 · May 6, 2020
Related Publication 20230290333A1 · Sep 14, 2023
References Cited (60)
US 4588867A · Konomi · 1986 [cited by applicant]
US 5280524A · Norris · 1994 [cited by applicant]
US 5692059A · Kruger · 1997 [cited by applicant]
US 5794191A · Chen · 1998 [cited by examiner]
US 6354299B1 · Fischell et al. · 2002 [cited by applicant]
US 7676372B1 · Oba · 2010 [cited by examiner]
US 8200486B1 · Jorgensen · 2012 [cited by examiner]
US 20050049856A1 · Baraff · 2005 [cited by examiner]
US 20050114137A1 · Saito · 2005 [cited by examiner]
US 20080220718A1 · Sakamoto · 2008 [cited by examiner]
US 20100183161A1 · Boretzki · 2010 [cited by examiner]
US 20120278070A1 · Herve · 2012 [cited by examiner]
US 20130195302A1 · Meincke · 2013 [cited by examiner]
US 20150139459A1 · Olsen · 2015 [cited by examiner]
US 20150228271A1 · Morita · 2015 [cited by examiner]
US 20150242180A1 · Boulanger-Lewandowski · 2015 [cited by examiner]
US 20160030744A1 · Hubert-Brierre · 2016 [cited by examiner]
US 20160343387A1 · Kamamoto · 2016 [cited by examiner]
US 20170295439A1 · Xu · 2017 [cited by examiner]
US 20170323636A1 · Xiao · 2017 [cited by examiner]
US 20180113673A1 · Sheynblat · 2018 [cited by examiner]
US 20180310159A1 · Katz · 2018 [cited by examiner]
US 20180331668A1 · Yuzuriha · 2018 [cited by examiner]
US 20180343525A1 · Karlsen · 2018 [cited by examiner]
US 20180367882A1 · Watts · 2018 [cited by examiner]
US 20190222943A1 · Andersen · 2019 [cited by examiner]
US 20200135171A1 · Tachibana · 2020 [cited by examiner]
US 20210020161A1 · Gao · 2021 [cited by examiner]
US 20210327407A1 · Chae · 2021 [cited by examiner]
CN 105185371 · 2015 [cited by applicant]
CN 106782577 · 2017 [cited by applicant]
CN 109120790 · 2019 [cited by applicant]
EP 3188507 · 2017 [cited by applicant]
EP 3229496 · 2017 [cited by applicant]
WO WO0069215 · 2000 [cited by applicant]
Valin, J. M., & Skoglund, J. (May 2019). LPCNet: Improving neural speech synthesis through linear prediction. In ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 58… [cited by examiner]
Srinivasan, S., & Kechichian, P. (Sep. 2012). Robustness analysis of speech enhancement using a bone conduction microphone—preliminary results. In IWAENC 2012; International Workshop on Acoustic Signal Enhancement (pp. … [cited by examiner]
Kapur, A., Kapur, S., & Maes, P. (Mar. 2018). Alterego: A personalized wearable silent speech interface. In 23rd International conference on intelligent user interfaces (pp. 43-53). (Year: 2018). [cited by examiner]
Shahina, A., & Yegnanarayana, B. (2007). Mapping speech spectra from throat microphone to close-speaking microphone: A neural network approach. EURASIP Journal on Advances in Signal Processing, 2007, 1-10. (Year: 2007). [cited by examiner]
Shan, D., Zhang, X., Zhang, C., & Li, L. (2018). A novel encoder-decoder model via NS-LSTM used for bone-conducted speech enhancement. IEEE Access. Oct. 4, 2018; 6: 62638-44. (Year: 2018). [cited by examiner]
Liu, H. P., Tsao, Y., & Fuh, C. S. (2018). Bone-conducted speech enhancement using deep denoising autoencoder. Speech Communication, 104, 106-112. (Year: 2018). [cited by examiner]
Liu, R., Cornelius, C., Rawassizadeh, R., Peterson, R., & Kotz, D. (2018). Vocal resonance: Using internal body voice for wearable authentication. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous T… [cited by examiner]
Huang, B., Gong, Y., Sun, J., & Shen, Y. (Aug. 2017). A wearable bone-conducted speech enhancement system for strong background noises. In 2017 18th International Conference on Electronic Packaging Technology (ICEPT) (p… [cited by examiner]
Maruri, H. A. C., Lopez-Meyer, P., Huang, J., Beltman, W. M., Nachman, L., & Lu, H. (2018). V-Speech: noise-robust speech capturing glasses using vibration sensors. Proceedings of the ACM on Interactive, Mobile, Wearabl… [cited by examiner]
Turan, M. A. T. (2018). Enhancement of throat microphone recordings using gaussian mixture model probabilistic estimator. arXiv preprint arXiv:1804.05937. (Year: 2018). [cited by examiner]
Turan, M. T., & Erzin, E. (2015). Source and filter estimation for throat-microphone speech enhancement. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 24(2), 265-275. (Year: 2015). [cited by examiner]
Extended European Search Report for EP Patent Appln. No. 19172713.0 dated Jun. 25, 2019. [cited by applicant]
Lee, C., et al., “Bone-Conduction Sensor Assisted Noise Estimation for Improved Speech Enhancement,” Department of Electrical and Computer Engineering University of California, San Diego, Interspeech 2018. [cited by applicant]
“Reconstruction Filter Design for Bone-Conducted Speech”_T. Tamiya and T. Shimamura, Interspeech 2004—ICSLP 8 [cited by applicant]
Kalchbrenner, N., et al., “Efficient Neural Audio Synthesis,” dated Jun. 2018. [cited by applicant]
Ping, W., et al., “ClariNet: ParallelWave Generation in End-to-End Text-to-Speech,” Baidu Research, Feb. 2019. [cited by applicant]
Foreign OA for CN Patent Appln. No. 202080044974.3 dated Aug. 8, 2023. [cited by applicant]
“LPCNET: Improving Neural Speech Synthesis Through Linear Prediction” (Jaen-Marc Valin & Jan Skoglund), dated Feb. 19, 2019. [cited by applicant]
Written Opinion for Foreign Patent Appln. No. PCT/EP2020/062561 dated Aug. 31, 2020. [cited by applicant]
International Search Report for Foreign Patent Appln. No. PCT/EP2020/062561 dated Aug. 31, 2020. [cited by applicant]
English Translation for Foreign OA for CN Patent Appln. No. 20208004974.3 dated Aug. 8, 2023. [cited by applicant]
Foreign Exam Report for EP Patent Appln. No. 20722603.6 dated Jan. 30, 2024. [cited by applicant]
Translation of foreign office action dated Apr. 25, 2024 for Chinese Appln. No. 202080044974.3. [cited by applicant]
Foreign OA for CN Patent Appln. No. 202080044974.3 dated Aug. 13, 2024. [cited by applicant]
Translation of foreign office action dated Aug. 13, 2024 for Chinese Appln. No. 202080044974.3. [cited by applicant]