IP Library Granted Patent US 12,567,428
Granted Patent B1
US 12,567,428 · App. 17/962,311 · Granted Mar 3, 2026

Contact transducer based audio enhancement

Inventors: Lin Li (Seattle, WA); Tetsuro Oishi (Bothell, WA); Gongqiang Yu (Redmond, WA)
Assignee: Meta Platforms Technologies, LLC
G10L21/0232H04R5/033H04R5/04G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,428
App. No.
17/962,311
Granted
Mar 3, 2026
Kind
B1
Abstract

Contact transducer based audio enhancement (e.g., of sounds from a local area, speech of a user, etc.) is described. An audio system includes a microphone array, a contact transducer, and a controller. The contact transducer is in contact with tissue of the user and can detect tissue-based vibrations generated by the speech of the user. The sounds and the detected vibrations are pre-processed. Input parameters are determined using the pre-processed sounds and vibrations. The audio system analyzes the input parameters to determine one or more signal characteristics. The audio system adjusts one or more sound filters to enhance a signal corresponding to the speech of the user based in part on status of the signal characteristics. The audio system performs an action associated with the enhanced signal corresponding to the speech.

Claims (36)

1 . A method comprising:

detecting, via a microphone array of an audio system, sounds from a local area, the sounds from the local area including speech of a user of the audio system;

detecting, via a contact transducer that is in contact with tissue of the user, tissue-based vibrations generated by the speech of the user;

determine, in real-time and based on the sounds detected via the microphone array and the tissue-based vibrations detected by the contact transducer, a state associated with wind noise in the sounds and tissue-based vibrations;

determine, in real-time and based on the sounds detected via the microphone array and the tissue-based vibrations detected by the contact transducer, a state associated with a voice of the user in the sounds and tissue-based vibrations; and

combining, based on the determinations of the states associated with wind noise and the voice of the user, low-frequency components of the tissue-based vibrations detected by the contact transducer with high-frequency components of the sounds detected by the microphone array.

2 . The method of claim 1 , wherein the state associated with wind noise is determined via an inter-channel coherence analysis of the sounds and the tissue-based vibrations.

3 . The method of claim 1 , wherein the state associated with the voice of the user is determined using a spectrum centroid analysis of the tissue-based vibrations.

4 . The method of claim 1 , wherein the combining further comprises attenuating frequency components below a threshold in the sounds detected by the microphone array.

5 . The method of claim 4 , wherein the combining further comprises augmenting missing frequency components in the tissue-based vibrations with corresponding components from the sounds detected by the microphone array.

6 . The method of claim 1 , wherein determining the state associated with wind noise comprises calculating a root-mean-square difference between a first signal representing the sounds detected by the microphone array and a second signal representing the tissue-based vibrations detected by contact transducer.

7 . The method of claim 1 , wherein the combining is performed in response to the state of the voice of the user indicating that speech is present.

8 . The method of claim 1 , wherein the combining comprises adjusting one or more sound filters based on the states associated with wind noise and the voice of the user in the sounds and tissue-based vibrations.

9 . The method of claim 1 , further comprising analyzing an output of the combining to determine a user command.

10 . The method of claim 1 , wherein the contact transducer is configured to be in contact with a head of a user.

11 . An audio system comprising:

a microphone array configured to detect sounds from a local area, the sounds from the local area including a voice of a user of the audio system;

a contact transducer configured to be in contact with a portion of a head of the user, and detect tissue-based vibrations that are generated by the voice of the user; and

a controller configured to:

determine, in real-time and based on the sounds detected via the microphone array and the tissue-based vibrations detected by the contact transducer, a state associated with wind noise in the sounds and tissue-based vibrations;

determine, in real-time and based on the sounds detected via the microphone array and the tissue-based vibrations detected by the contact transducer, a state associated with a voice of the user in the sounds and tissue-based vibrations; and

combine, based on the determinations of the states associated with wind noise and the voice of the user, low-frequency components of the tissue-based vibrations detected by the contact transducer with high-frequency components of the sounds detected by the microphone array.

12 . The audio system of claim 11 , wherein the controller is further configured to determine ed with the wind noise by performing an inter-channel coherence analysis of the sounds and the tissue-based vibrations.

13 . The audio system of claim 11 , wherein the controller is further configured to determine the state associated with the voice of the user via a spectrum centroid analysis of the tissue-based vibrations.

14 . The audio system of claim 11 , wherein the controller is configured to perform the combining of the low-frequency components and the high-frequency components by attenuating frequency components below a threshold in the sounds detected by the microphone array.

15 . The audio system of claim 11 , wherein controller is configured to perform the combining of the low-frequency components and the high-frequency components by augmenting missing frequency components in the tissue-based vibrations with corresponding components from the sounds detected by the microphone array.

16 . A non-transitory computer-readable storage medium comprising memory with executable computer instructions encoded thereon that, when executed by one or more processors of an audio system, cause the audio system to:

detect, via a microphone array of the audio system, sounds from a local area, the sounds from the local area including speech of a user of the audio system;

detect, via a contact transducer that is in contact with tissue of the user, tissue-based vibrations generated by the speech of the user;

determine, in real-time and based on the sounds detected via the microphone array and the tissue-based vibrations detected by the contact transducer, a state associated with wind noise in the sounds and tissue-based vibrations;

determine, in real-time and based on the sounds detected via the microphone array and the tissue-based vibrations detected by the contact transducer, a state associated with a voice of the user in the sounds and tissue-based vibrations; and

combine, based on the determinations of the states associated with wind noise and the voice of the user, low-frequency components of the tissue-based vibrations detected by the contact transducer with high-frequency components of the sounds detected by the microphone array.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the executable computer instructions cause the audio system to determine the state associated with the wind noise by performing an inter-channel coherence analysis of the sounds and the tissue-based vibrations.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the executable computer instructions cause the audio system to determine the state associated with the voice of the user by using a spectrum centroid analysis of the tissue-based vibrations.

19 . The non-transitory computer-readable storage medium of claim 16 , wherein the executable computer instructions cause the audio system to combine the low-frequency components and the high-frequency components by attenuating frequency components below a threshold in the sounds detected by the microphone array.

20 . The non-transitory computer-readable storage medium of claim 1 , wherein the executable computer instructions cause the audio system to selectively combine the low-frequency components and the high-frequency components by augmenting missing frequency components in the tissue-based vibrations with corresponding components from the sounds detected by the microphone array.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2022
From: LI, LIN; OISHI, TETSURO; YU, GONGQIANG
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 061741/0104 →
Continuity (1)
Provisional Application 63254493 · Oct 11, 2021
References Cited (57)
US 8165875B2 · Hetherington et al. · 2012 [cited by applicant]
US 9330675B2 · Zhang et al. · 2016 [cited by applicant]
US 9916841B2 · Hetherington et al. · 2018 [cited by applicant]
US 10297245B1 · Chen · 2019 [cited by examiner]
US 10728649B1 · Holman · 2020 [cited by examiner]
US 10791389B1 · Khaleghimeybodi · 2020 [cited by examiner]
US 10966043B1 · Khaleghimeybodi · 2021 [cited by examiner]
US 11069365B2 · Kar et al. · 2021 [cited by applicant]
US 11490092B2 · Zhao · 2022 [cited by examiner]
US 20090097674A1 · Watson · 2009 [cited by examiner]
US 20090202084A1 · Joeng · 2009 [cited by examiner]
US 20110123031A1 · Ojala · 2011 [cited by examiner]
US 20110224481A1 · Lee · 2011 [cited by examiner]
US 20120191447A1 · Joshi · 2012 [cited by examiner]
US 20130051585A1 · Karkkainen · 2013 [cited by examiner]
US 20150222989A1 · Labrosse · 2015 [cited by examiner]
US 20160381453A1 · Ushakov · 2016 [cited by examiner]
US 20180233127A1 · Visser · 2018 [cited by examiner]
US 20180278967A1 · Kerofsky · 2018 [cited by examiner]
US 20180367882A1 · Watts · 2018 [cited by examiner]
US 20190089550A1 · Rexach · 2019 [cited by examiner]
US 20190124454A1 · Aschbacher · 2019 [cited by examiner]
US 20190227766A1 · Nahman · 2019 [cited by examiner]
US 20200013395A1 · Jeong · 2020 [cited by examiner]
US 20200184996A1 · Steele · 2020 [cited by examiner]
US 20200357406A1 · York · 2020 [cited by examiner]
US 20210056984A1 · Zhang · 2021 [cited by examiner]
US 20210074310A1 · Bryan · 2021 [cited by examiner]
US 20210089863A1 · Yang · 2021 [cited by examiner]
US 20210110841A1 · Weber · 2021 [cited by examiner]
US 20210125625A1 · Huang · 2021 [cited by examiner]
US 20210134312A1 · Koishida · 2021 [cited by examiner]
US 20210168554A1 · Zhang · 2021 [cited by examiner]
US 20210256988A1 · Gallart Mauri · 2021 [cited by examiner]
US 20210266683A1 · Hersbach · 2021 [cited by examiner]
US 20210321205A1 · Akers · 2021 [cited by examiner]
US 20210368255A1 · Lee · 2021 [cited by examiner]
US 20210377649A1 · Lewis · 2021 [cited by examiner]
US 20210386320A1 · Lesso · 2021 [cited by examiner]
US 20210390941A1 · Lakshminarayanan · 2021 [cited by examiner]
US 20220141561A1 · Wang · 2022 [cited by examiner]
US 20220180767A1 · Aharonson · 2022 [cited by examiner]
US 20220208209A1 · Zheng · 2022 [cited by examiner]
US 20220284913A1 · Yu · 2022 [cited by examiner]
US 20220343887A1 · Xiao · 2022 [cited by examiner]
US 20240153518A1 · Vondersaar · 2024 [cited by examiner]
US 20240155290A1 · Hiroe · 2024 [cited by examiner]
US 20240323616A1 · Nakamura · 2024 [cited by examiner]
CA 3045600A1 · 2018 [cited by examiner]
CN 110010143A · 2019 [cited by examiner]
CN 113270106A · 2021 [cited by applicant]
WO WO2021068120A1 · 2021 [cited by examiner]
WO WO2022089563A1 · 2022 [cited by examiner]
WO WO2022226696A1 · 2022 [cited by examiner]
WO WO2022226792A1 · 2022 [cited by examiner]
Zheng, Yanli, et al. “Air-and bone-conductive integrated microphones for robust speech detection and enhancement.” 2003 IEEE Workshop on Automatic Speech Recognition and Understanding (IEEE Cat. No. 03EX721). (Year: 200… [cited by examiner]
Zhou, Yi, et al. “A real-time dual-microphone speech enhancement algorithm assisted by bone conduction sensor.” Sensors 20.18 (2020): 5050. (Year: 2020). [cited by examiner]