IP Library › Granted Patent US 12,300,241
Granted Patent B2
US 12,300,241 · App. 17/994,968 · Granted May 13, 2025

Speech signal processing method and related device thereof

Inventors: Libin Zhang (Beijing, CN); Hui Yang (Beijing, CN); Shu Fang (Beijing, CN); Siwei Dong (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L15/24G06V40/171G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,300,241
App. No.
17/994,968
Granted
May 13, 2025
Kind
B2
Abstract

A speech signal processing method and a related device thereof are provided. The method may be applied to the audio field and includes: obtaining a user speech signal captured by a sensor; obtaining a corresponding vibration signal when a user generates a speech, where the vibration signal indicates a vibration feature of a body part of the user, and the body part is a part that vibrates correspondingly based on sound-making behavior when the user is making a sound; and obtaining target speech information based on the vibration signal and the user speech signal captured by the sensor. In this application, the vibration signal is used as a basis for speech recognition.

Claims (65)

1. A speech signal processing method, wherein the method comprises:

obtaining a user speech signal captured by a sensor;

obtaining a corresponding vibration signal when a user generates the speech signal, wherein the vibration signal indicates a vibration feature of a body part of the user, and the body part is a part that vibrates based on sound-making behavior when the user is making a sound;

obtaining a corresponding brain wave signal of the user when the user generates the speech signal; and

determining target speech information based on the vibration signal, the user speech signal captured by the sensor, and the brain wave signal, or performing voiceprint recognition based on the user speech signal captured by the sensor, the vibration signal, and the brain wave signal.

2. The method according to claim 1 , wherein the vibration signal indicates a vibration feature corresponding to a vibration generated when the user generates the speech signal.

3. The method according to claim 1 , wherein the body part comprises at least one of the following: a calvarium, a face, a larynx, or a neck.

4. The method according to claim 1 , wherein the obtaining a corresponding vibration signal when a user generates the speech signal comprises:

obtaining a video frame comprising the user; and

extracting, based on the video frame, the corresponding vibration signal when the user generates the speech signal.

5. The method according to claim 4 , wherein the video frame is captured using a dynamic vision sensor or a high-speed camera.

6. The method according to claim 1 , wherein the determining target speech information based on the vibration signal and the user speech signal captured by the sensor comprises:

obtaining a corresponding target audio signal based on the vibration signal;

filtering the target audio signal from the user speech signal captured by the sensor to obtain a to-be-filtered signal; and

filtering the to-be-filtered signal from the user speech signal captured by the sensor to obtain the target speech information.

7. The method according to claim 1 , wherein the determining target speech information based on the vibration signal and the user speech signal captured by the sensor comprises:

determining, based on the vibration signal and the user speech signal captured by the sensor, the target speech information by using a cyclic neural network model; or

determining a corresponding target audio signal based on the vibration signal, and determining, based on the target audio signal and the user speech signal captured by the sensor, the target speech information by using a cyclic neural network model.

8. The method according to claim 1 , wherein the method further comprises:

obtaining, based on the brain wave signal, a motion signal of a vocal tract occlusion part when the user generates the speech; and correspondingly, the determining the target speech information based on the vibration signal, the brain wave signal, and the user speech signal captured by the sensor comprises:

determining the target speech information based on the vibration signal, the motion signal, and the user speech signal captured by the sensor.

9. The method according to claim 1 , wherein the determining the target speech information based on the vibration signal, the brain wave signal, and the user speech signal captured by the sensor comprises:

determining, based on the vibration signal, the brain wave signal, and the user speech signal captured by the sensor, the target speech information by using a cyclic neural network model; or

determining a corresponding first target audio signal based on the vibration signal; and determining a corresponding second target audio signal based on the brain wave signal, and determining, based on the first target audio signal, the second target audio signal, and the user speech signal captured by the sensor, the target speech information by using a cyclic neural network model.

10. The method according to claim 1 , wherein the performing voiceprint recognition based on the user speech signal captured by the sensor, the vibration signal, and the brain wave signal comprises:

performing voiceprint recognition based on the user speech signal captured by the sensor to obtain a first confidence level that the user speech signal captured by the sensor belongs to the user;

performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user speech signal captured by the sensor belongs to the user;

performing voiceprint recognition based on the brain wave signal to obtain a third confidence level that the user speech signal captured by the sensor belongs to the user; and

determining the voiceprint recognition result based on the first confidence level, the second confidence level, and the third confidence level.

11. The method according to claim 1 , wherein the performing voiceprint recognition based on the user speech signal captured by the sensor and the vibration signal comprises:

performing voiceprint recognition based on the user speech signal captured by the sensor to obtain a first confidence level that the user speech signal captured by the sensor belongs to a target user;

performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user speech signal captured by the sensor belongs to the target user; and

determining a voiceprint recognition result based on the first confidence level and the second confidence level.

12. A speech signal processing apparatus, comprising:

a memory storing executable instructions; and

a processor configured to execute the executable instructions to perform operations of:

obtaining a user speech signal captured by a sensor;

obtaining a corresponding vibration signal when a user generates the speech signal, wherein the vibration signal indicates a vibration feature of a body part of the user, and the body part is a part that vibrates based on sound-making behavior when the user is making a sound;

obtaining a corresponding brain wave signal of the user when the user generates the speech signal; and

determining target speech information based on the vibration signal, the user speech signal captured by the sensor, and the brain wave signal, or performing voiceprint recognition based on the user speech signal captured by the sensor, the vibration signal, and the brain wave signal.

13. The apparatus according to claim 12 , wherein the body part comprises at least one of the following: a calvarium, a face, a larynx, or a neck.

14. The apparatus according to claim 12 , wherein the processor is further configured to execute the executable instructions to perform operations of:

obtaining a video frame comprising the user; and

extracting, based on the video frame, the corresponding vibration signal when the user generates the speech signal.

15. The apparatus according to claim 12 , wherein the processor is further configured to execute the executable instructions to perform operations of:

obtaining a corresponding target audio signal based on the vibration signal;

filtering the target audio signal from the user speech signal captured by the sensor to obtain a to-be-filtered signal; and

filtering the to-be-filtered signal from the user speech signal captured by the sensor to obtain the target speech information.

16. The apparatus according to claim 12 , wherein the processor is further configured to execute the executable instructions to perform operations of:

determining, based on the vibration signal and the user speech signal captured by the sensor, the target speech information by using a cyclic neural network model; or

determining a corresponding target audio signal based on the vibration signal, and determining, based on the target audio signal and the user speech signal captured by the sensor, the target speech information by using a cyclic neural network model.

17. The apparatus according to claim 12 , wherein the processor is further configured to execute the executable instructions to perform operations of:

obtaining, based on the brain wave signal, a motion signal of a vocal tract occlusion part when the user generates the speech; and correspondingly, the determining the target speech information based on the vibration signal, the brain wave signal, and the user speech signal captured by the sensor comprises:

determining the target speech information based on the vibration signal, the motion signal, and the user speech signal captured by the sensor.

18. A non-transitory computer-readable storage medium, comprising a program, wherein when the program runs on a computer, the computer is enabled to perform:

obtaining a user speech signal captured by a sensor;

obtaining a corresponding vibration signal when a user generates the speech signal, wherein the vibration signal indicates a vibration feature of a body part of the user, and the body part is a part that vibrates based on sound-making behavior when the user is making a sound;

obtaining a corresponding brain wave signal of the user when the user generates the speech signal; and

determining target speech information based on the vibration signal, the user speech signal captured by the sensor, and the brain wave signal, or performing voiceprint recognition based on the user speech signal captured by the sensor, the vibration signal, and the brain wave signal.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the obtaining a corresponding vibration signal when a user generates the speech signal comprises:

obtaining a video frame comprising the user; and

extracting, based on the video frame, the corresponding vibration signal when the user generates the speech signal.

20. The non-transitory computer-readable storage medium of claim 18 , wherein the method further comprises:

obtaining, based on the brain wave signal, a motion signal of a vocal tract occlusion part when the user generates the speech; and correspondingly, the determining the target speech information based on the vibration signal, the brain wave signal, and the user speech signal captured by the sensor comprises:

determining the target speech information based on the vibration signal, the motion signal, and the user speech signal captured by the sensor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2025
From: ZHANG, LIBIN; YANG, HUI; FANG, SHU; DONG, SIWEI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 069830/0781 →
Continuity (2)
Continuation PCTCN2020093523 · May 29, 2020
Related Publication 20230098678A1 · Mar 30, 2023
References Cited (39)
US 11282523B2 · Gross · 2022 [cited by examiner]
US 20100204987A1 · Miyauchi · 2010 [cited by examiner]
US 20150045007A1 · Cash · 2015 [cited by applicant]
US 20160011063A1 · Zhang · 2016 [cited by examiner]
US 20160203002A1 · Kannan · 2016 [cited by examiner]
US 20160267911A1 · Koetje · 2016 [cited by examiner]
US 20160285793A1 · Anderson · 2016 [cited by examiner]
US 20170011210A1 · Cheong · 2017 [cited by examiner]
US 20170235364A1 · Nakamura · 2017 [cited by examiner]
US 20170238144A1 · Chatani · 2017 [cited by examiner]
US 20180124458A1 · Knox · 2018 [cited by examiner]
US 20180232511A1 · Bakish · 2018 [cited by applicant]
US 20190043512A1 · Huang · 2019 [cited by examiner]
US 20190182415A1 · Sivan · 2019 [cited by examiner]
US 20190272325A1 · Korn · 2019 [cited by examiner]
US 20190348041A1 · Cella · 2019 [cited by examiner]
US 20200005770A1 · Lunner et al. · 2020 [cited by applicant]
US 20200077206A1 · Jensen et al. · 2020 [cited by applicant]
CN 1622200A · 2005 [cited by applicant]
CN 101947152A · 2011 [cited by applicant]
CN 102027536A · 2011 [cited by applicant]
CN 103871419A · 2014 [cited by applicant]
CN 105632512A · 2016 [cited by applicant]
CN 106778186A · 2017 [cited by applicant]
CN 106872011A · 2017 [cited by applicant]
CN 108805087A · 2018 [cited by applicant]
CN 109104209A · 2018 [cited by applicant]
CN 109151530A · 2019 [cited by applicant]
CN 109241912A · 2019 [cited by applicant]
CN 110248281A · 2019 [cited by applicant]
CN 110366086A · 2019 [cited by applicant]
CN 110931031A · 2020 [cited by applicant]
JP 2010217453A · 2010 [cited by applicant]
Intelligent Noise Detection and Control—Active Noise Control (ANC) technology, 2011, with the English Translation, 26 pages. [cited by applicant]
James A. O Sullivan et al, Feature Article, Attentional Selection in a Cocktail Party Environment Can Be Decoded from Single-Trial EEG, Cerebral Cortex Jul. 2015;25:1697-1706, doi:10.1093/cercor/bht355, Advance Access p… [cited by applicant]
Xiaoya Li et al, Lip Reading Deep Network Exploiting Multi-modal Spiking Visual and Auditory Sensors, 2019 IEEE, 5 pages. [cited by applicant]
Cong Han et al, Speaker-independent auditory attention decoding without access to clean speech sources, Science Advances, Research Article, May 15, 2019, 12 pages. [cited by applicant]
Gopala K. Anumanchipalli et al, Speech synthesis from neural decoding of spoken sentences, Nature vol. 568, pp. 493-498 (2019), 20 pages. [cited by applicant]
Wang Dongxia, Study on Methods for Speech Enhancement based on Microphone Array, School of Electronic and Information Engineering Dalian University of Technology, Dalian, Liaoning, P.R.China, 116024 , Apr. 2007, with th… [cited by applicant]