IP Library › Granted Patent US 12,412,591
Granted Patent B2
US 12,412,591 · App. 18/279,475 · Granted Sep 9, 2025

Voice processing method and electronic device

Inventors: Haikuan Gao (Beijing, CN); Zhenyi Liu (Beijing, CN); Zhichao Wang (Beijing, CN); Jianyong Xuan (Beijing, CN); Risheng Xia (Beijing, CN)
Assignee: BEIJING HONOR DEVICE CO., LTD.
G10L21/0232G10L2021/02082G10L2021/02161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,591
App. No.
18/279,475
Granted
Sep 9, 2025
Kind
B2
Abstract

A voice processing method is provided. The method includes: An electronic device first performs de-reverberation processing on a first frequency domain signal to obtain a second frequency domain signal, performs noise reduction processing on the first frequency domain signal to obtain a third frequency domain signal, and then performs, based on a first voice feature of the second frequency domain signal and a second voice feature of the third frequency domain signal, fusion processing on the second frequency domain signal and the third frequency domain signal that belong to a same channel of first frequency domain signal, to obtain a fused frequency domain signal. In this case, background noise in the fused frequency domain signal is not damaged, thereby effectively ensuring stable background noise of a voice signal obtained after voice processing. In addition, an electronic device, a chip system, and a computer-readable storage medium are provided.

Claims (58)

1. A voice processing method, applied to an electronic device, wherein the electronic device comprises n microphones, n is greater than or equal to 2, and the method comprises:

performing Fourier transform on voice signals picked up by the n microphones to obtain n channels of corresponding first frequency domain signals S, wherein each channel of first frequency domain signal S has M frequencies, and M is a quantity of transform points used when the Fourier transform is performed;

performing de-reverberation processing on the n channels of first frequency domain signals S to obtain n channels of second frequency domain signals S E , and performing noise reduction processing on the n channels of first frequency domain signals S to obtain n channels of third frequency domain signals S s ;

determining a first voice feature corresponding to M frequencies of a second frequency domain signal S Ei corresponding to a first frequency domain signal S i and a second voice feature corresponding to M frequencies of a third frequency domain signal S Si corresponding to the first frequency domain signal S i , and obtaining M target amplitude values corresponding to the first frequency domain signal S i based on the first voice feature, the second voice feature, the second frequency domain signal S Ei , and the third frequency domain signal S Si , wherein i=1, 2, . . . , or n, the first voice feature is used to represent a de-reverberation degree of the second frequency domain signal S Ei , and the second voice feature is used to represent a noise reduction degree of the third frequency domain signal S Si ; and

determining a fused frequency domain signal corresponding to the first frequency domain signal S i based on the M target amplitude values.

2. The method according to claim 1 , wherein the obtaining M target amplitude values corresponding to the first frequency domain signal S i based on the first voice feature, the second voice feature, the second frequency domain signal S Ei , and the third frequency domain signal S Si specifically comprises:

when it is determined that the first voice feature and the second voice feature that correspond to a frequency A i in the M frequencies meet a first preset condition, determining a first amplitude value corresponding to a frequency A i in the second frequency domain signal S Ei as a target amplitude value corresponding to the frequency A i , or determining the target amplitude value corresponding to the frequency A i based on the first amplitude value and a second amplitude value corresponding to a frequency A i in the third frequency domain signal S Si , wherein i=1, 2, . . . , or M; or

when it is determined that the first voice feature and the second voice feature that correspond to the frequency A i do not meet the first preset condition, determining the second amplitude value as the target amplitude value corresponding to the frequency A i .

3. The method according to claim 2 , wherein the determining the target amplitude value corresponding to the frequency A i based on the first amplitude value and a second amplitude value corresponding to a frequency A i in the third frequency domain signal S Si specifically comprises:

determining a first weighted amplitude value based on the first amplitude value corresponding to the frequency A i and a corresponding first weight, and determining a second weighted amplitude value based on the second amplitude value corresponding to the frequency A i and a corresponding second weight; and

determining a sum of the first weighted amplitude value and the second weighted amplitude value as the target amplitude value corresponding to the frequency A i .

4. The method according to claim 2 , wherein the first voice feature comprises a first dual-microphone correlation coefficient and a first frequency energy value, and the second voice feature comprises a second dual-microphone correlation coefficient and a second frequency energy value; and

the first dual-microphone correlation coefficient is used to represent a signal correlation degree between the second frequency domain signal S Ei and a second frequency domain signal S Et at corresponding frequencies, the second frequency domain signal S Et is any channel of second frequency domain signal S E other than the second frequency domain signal S Ei in the n channels of second frequency domain signals S E , the second dual-microphone correlation coefficient is used to represent a signal correlation degree between the third frequency domain signal S Si and a third frequency domain signal S St at corresponding frequencies, and the third frequency domain signal S St is a third frequency domain signal S s that is in the n channels of third frequency domain signals S s and that corresponds to a same first frequency domain signal as the second frequency domain signal S Et .

5. The method according to claim 4 , wherein the first preset condition is that the first dual-microphone correlation coefficient and the second dual-microphone correlation coefficient of the frequency A i meet a second preset condition, and the first frequency energy value and the second frequency energy value of the frequency A i meet a third preset condition.

6. The method according to claim 5 , wherein the second preset condition is that a first difference of the first dual-microphone correlation coefficient of the frequency A i minus the second dual-microphone correlation coefficient of the frequency A i is greater than a first threshold; and the third preset condition is that a second difference of the first frequency energy value of the frequency A i minus the second frequency energy value of the frequency A i is less than a second threshold.

7. The method according to claim 1 , wherein a de-reverberation processing method comprises a de-reverberation method based on a coherent-to-diffuse power ratio or a de-reverberation method based on a weighted prediction error.

8. The method according to claim 1 , wherein the method further comprises:

performing inverse Fourier transform on the fused frequency domain signal to obtain a fused voice signal.

9. The method according to claim 1 , wherein before the Fourier transform is performed on the voice signals, the method further comprises:

displaying a shooting interface, wherein the shooting interface comprises a first control;

detecting a first operation performed on the first control; and

in response to the first operation, performing, by the electronic device, video shooting to obtain a video that comprises the voice signals.

10. An electronic device, wherein the electronic device comprises:

n microphones, n is greater than or equal to 2;

one or more processors and

one or more memories; and the one or more memories are coupled to the one or more processors, the one or more memories are configured to store computer program code, the computer program code comprises computer instructions, and when the one or more processors execute the computer instructions, the electronic device is enabled to perform the following steps:

performing Fourier transform on voice signals picked up by the n microphones to obtain n channels of corresponding first frequency domain signals S, wherein each channel of first frequency domain signal S has M frequencies, and M is a quantity of transform points used when the Fourier transform is performed;

performing de-reverberation processing on the n channels of first frequency domain signals S to obtain n channels of second frequency domain signals S E , and performing noise reduction processing on the n channels of first frequency domain signals S to obtain n channels of third frequency domain signals S s ;

determining a first voice feature corresponding to M frequencies of a second frequency domain signal S Ei corresponding to a first frequency domain signal S i and a second voice feature corresponding to M frequencies of a third frequency domain signal S Si corresponding to the first frequency domain signal S i , and obtaining M target amplitude values corresponding to the first frequency domain signal S i based on the first voice feature, the second voice feature, the second frequency domain signal S Ei , and the third frequency domain signal S Si , wherein i=1, 2, . . . , or n, the first voice feature is used to represent a de-reverberation degree of the second frequency domain signal S Ei , and the second voice feature is used to represent a noise reduction degree of the third frequency domain signal S Si ; and

determining a fused frequency domain signal corresponding to the first frequency domain signal S i based on the M target amplitude values.

11. The electronic device according to claim 10 , wherein the obtaining M target amplitude values corresponding to the first frequency domain signal S i based on the first voice feature, the second voice feature, the second frequency domain signal S Ei , and the third frequency domain signal S Si specifically comprises:

when it is determined that the first voice feature and the second voice feature that correspond to a frequency A i in the M frequencies meet a first preset condition, determining a first amplitude value corresponding to a frequency A i in the second frequency domain signal S Ei as a target amplitude value corresponding to the frequency A i , or determining the target amplitude value corresponding to the frequency A i based on the first amplitude value and a second amplitude value corresponding to a frequency A i in the third frequency domain signal S Si , wherein i=1, 2, . . . , or M; or

when it is determined that the first voice feature and the second voice feature that correspond to the frequency A i do not meet the first preset condition, determining the second amplitude value as the target amplitude value corresponding to the frequency A i .

12. The electronic device according to claim 11 , wherein the determining the target amplitude value corresponding to the frequency A i based on the first amplitude value and a second amplitude value corresponding to a frequency A i in the third frequency domain signal S Si specifically comprises:

determining a first weighted amplitude value based on the first amplitude value corresponding to the frequency A i and a corresponding first weight, and determining a second weighted amplitude value based on the second amplitude value corresponding to the frequency A i and a corresponding second weight; and

determining a sum of the first weighted amplitude value and the second weighted amplitude value as the target amplitude value corresponding to the frequency A i .

13. The electronic device according to claim 11 , wherein the first voice feature comprises a first dual-microphone correlation coefficient and a first frequency energy value, and the second voice feature comprises a second dual-microphone correlation coefficient and a second frequency energy value; and

the first dual-microphone correlation coefficient is used to represent a signal correlation degree between the second frequency domain signal S Ei and a second frequency domain signal S Et at corresponding frequencies, the second frequency domain signal S Et is any channel of second frequency domain signal S E other than the second frequency domain signal S Ei in the n channels of second frequency domain signals S E , the second dual-microphone correlation coefficient is used to represent a signal correlation degree between the third frequency domain signal S Si and a third frequency domain signal S St at corresponding frequencies, and the third frequency domain signal S St is a third frequency domain signal S s that is in the n channels of third frequency domain signals S s and that corresponds to a same first frequency domain signal as the second frequency domain signal S Et .

14. The electronic device according to claim 13 , wherein the first preset condition is that the first dual-microphone correlation coefficient and the second dual-microphone correlation coefficient of the frequency A i meet a second preset condition, and the first frequency energy value and the second frequency energy value of the frequency A i meet a third preset condition.

15. The electronic device according to claim 14 , wherein the second preset condition is that a first difference of the first dual-microphone correlation coefficient of the frequency A i minus the second dual-microphone correlation coefficient of the frequency A i is greater than a first threshold; and the third preset condition is that a second difference of the first frequency energy value of the frequency A i minus the second frequency energy value of the frequency A i is less than a second threshold.

16. The electronic device according to claim 10 wherein a de-reverberation processing method comprises a de-reverberation method based on a coherent-to-diffuse power ratio or a de-reverberation method based on a weighted prediction error.

17. The electronic device according to claim 10 , wherein when the one or more processors execute the computer instructions, the electronic device is enabled to further perform the following steps:

performing inverse Fourier transform on the fused frequency domain signal to obtain a fused voice signal.

18. The electronic device according to claim 10 , wherein when the one or more processors execute the computer instructions, the electronic device is enabled to further perform the following steps:

before the Fourier transform is performed on the voice signals,

displaying a shooting interface, wherein the shooting interface comprises a first control;

detecting a first operation performed on the first control; and

in response to the first operation, performing, by the electronic device, video shooting to obtain a video that comprises the voice signals.

19. The electronic device according to claim 10 , wherein when the one or more processors execute the computer instructions, the electronic device is enabled to further perform the following steps:

before the Fourier transform is performed on the voice signals,

displaying a recording interface, wherein the recording interface comprises a second control;

detecting a second operation performed on the second control; and

in response to the second operation, performing, by the electronic device, recording to obtain the voice signals.

20. A non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when executed on an electronic device, causes the electronic device to perform following operations:

performing Fourier transform on voice signals picked up by the n microphones to obtain n channels of corresponding first frequency domain signals S, wherein each channel of first frequency domain signal S has M frequencies, and M is a quantity of transform points used when the Fourier transform is performed;

performing de-reverberation processing on the n channels of first frequency domain signals S to obtain n channels of second frequency domain signals S E , and performing noise reduction processing on the n channels of first frequency domain signals S to obtain n channels of third frequency domain signals S s ;

determining a first voice feature corresponding to M frequencies of a second frequency domain signal S Ei corresponding to a first frequency domain signal S i and a second voice feature corresponding to M frequencies of a third frequency domain signal S Si corresponding to the first frequency domain signal S i , and obtaining M target amplitude values corresponding to the first frequency domain signal S i based on the first voice feature, the second voice feature, the second frequency domain signal S Ei , and the third frequency domain signal S Si , wherein i=1, 2, . . . , or n, the first voice feature is used to represent a de-reverberation degree of the second frequency domain signal S Ei , and the second voice feature is used to represent a noise reduction degree of the third frequency domain signal S Si ; and

determining a fused frequency domain signal corresponding to the first frequency domain signal S i based on the M target amplitude values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2024
From: GAO, HAIKUAN; LIU, ZHENYI; WANG, ZHICHAO; XUAN, JIANYONG; XIA, RISHENG
To: BEIJING HONOR DEVICE CO., LTD.
Reel/Frame 067435/0196 →
Priority Claims (1)
CN 202110925923.8 · Aug 12, 2021 · national
Continuity (1)
Related Publication 20240144951A1 · May 2, 2024
References Cited (45)
US 9008329B1 · Mandel · 2015 [cited by examiner]
US 9401158B1 · Yen et al. · 2016 [cited by applicant]
US 10629194B2 · Song · 2020 [cited by applicant]
US 12272369B1 · Chhetri · 2025 [cited by examiner]
US 20120051548A1 · Visser · 2012 [cited by examiner]
US 20120185247A1 · Tzirkel-Hancock et al. · 2012 [cited by applicant]
US 20150334489A1 · Iyengar et al. · 2015 [cited by applicant]
US 20180330726A1 · Song · 2018 [cited by applicant]
US 20190318757A1 · Chen · 2019 [cited by examiner]
US 20200075012A1 · Tian et al. · 2020 [cited by applicant]
US 20200286501A1 · Xiao · 2020 [cited by examiner]
US 20210134312A1 · Koishida et al. · 2021 [cited by applicant]
US 20210176558A1 · Li et al. · 2021 [cited by applicant]
US 20210319802A1 · Bai · 2021 [cited by applicant]
US 20220230651A1 · Zhu et al. · 2022 [cited by applicant]
US 20230403505A1 · Yu · 2023 [cited by examiner]
US 20240144948A1 · Xuan · 2024 [cited by examiner]
US 20240290338A1 · Huang · 2024 [cited by examiner]
CA 2386653A1 · 2001 [cited by applicant]
CN 105427861A · 2016 [cited by applicant]
CN 105635500A · 2016 [cited by applicant]
CN 105825865A · 2016 [cited by applicant]
CN 107316648A · 2017 [cited by applicant]
CN 107316649A · 2017 [cited by applicant]
CN 109195043A · 2019 [cited by applicant]
CN 109979476A · 2019 [cited by applicant]
CN 110197669A · 2019 [cited by applicant]
CN 110211602A · 2019 [cited by applicant]
CN 110310655A · 2019 [cited by applicant]
CN 110648684A · 2020 [cited by applicant]
CN 110827791A · 2020 [cited by applicant]
CN 111161751A · 2020 [cited by applicant]
CN 111223493A · 2020 [cited by applicant]
CN 111312273A · 2020 [cited by applicant]
CN 111345047A · 2020 [cited by applicant]
CN 111489760A · 2020 [cited by applicant]
CN 111599372A · 2020 [cited by applicant]
CN 112420073A · 2021 [cited by applicant]
CN 113823314A · 2021 [cited by applicant]
O. Schwartz, S. Gannot and E. A. P. Habets, “Multi-Microphone Speech Dereverberation and Noise Reduction Using Relative Early Transfer Functions,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol.… [cited by applicant]
I. Kodrasi and S. Doclo, “Joint Dereverberation and Noise Reduction Based on Acoustic Multi-Channel Equalization,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, No. 4, pp. 680-693, Apr. 20… [cited by applicant]
B. J. Borgström and M. S. Brandstein, “Speech Enhancement via Attention Masking Network (SEAMNET): An End-to-End System for Joint Suppression of Noise and Reverberation,” in IEEE/ACM Transactions on Audio, Speech, and L… [cited by applicant]
Lan Tian et al: “An Overview of Monaural Speech Denoising and Dereverberation Research”, Computer Research and Development . 2020,57(05), total 26 pages. [cited by applicant]
Liu et al: “A Research to Speech Dereverberation Method Based on BISTM Recurrent Neural Networks and Non-negative Matrix Factorization”, Signal processing . 2017,33(03), total 5 pages. [cited by applicant]
H. Li, X. Zhang and G. Gao, “Robust Speech Dereverberation Based on WPE and Deep Learning,” 2020 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Auckland, New Zealan… [cited by applicant]