IP Library Granted Patent US 12,603,096
Granted Patent B2
US 12,603,096 · App. 18/330,472 · Granted Apr 14, 2026

Voice enhancement methods and systems

Inventors: Le Xiao (Shenzhen, CN); Chengqian Zhang (Shenzhen, CN); Fengyun Liao (Shenzhen, CN); Xin Qi (Shenzhen, CN)
Assignee: SHENZHEN SHOKZ CO., LTD.
G10L21/0232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,603,096
App. No.
18/330,472
Granted
Apr 14, 2026
Kind
B2
Abstract

The embodiments of the present disclosure provide a method and system for voice enhancement, including: obtaining a first signal and a second signal of a target voice, the first signal and the second signal being voice signals of the target voice at different voice collection positions; determining a target signal-to-noise ratio (SNR) of the target voice based on the first signal or the second signal; determining a processing mode for the first signal and the second signal based on the target SNR; and processing the first signal and the second signal based on the determined processing mode to obtain a voice-enhanced output voice signal corresponding to the target voice.

Claims (77)

1 . A voice enhancement method applied to a voice enhancement system, comprising:

obtaining a first signal and a second signal of a target voice, the first signal and the second signal being voice signals of the target voice collected by different collection devices at different voice collection positions;

determining a target signal-to-noise ratio (SNR) of the target voice based on the first signal or the second signal;

determining a processing mode for the first signal and the second signal based on the target SNR; and

obtaining a voice-enhanced output voice signal corresponding to the target voice by processing the first signal and the second signal based on the determined processing mode;

wherein the determining a processing mode for the first signal and the second signal based on the target SNR comprises:

in response to determining that the target SNR is smaller than a first threshold, processing the first signal and the second signal in a first mode; and

in response to determining that the target SNR is greater than a second threshold, processing the first signal and the second signal in a second mode,

wherein

the first threshold is not larger than the second threshold, and

computing resources allocated to the first mode is more than computing resources allocated to the second mode;

wherein the processing the first signal and the second signal in a first mode comprises:

obtaining a first output voice signal with a low frequency part of the target voice enhanced by processing a low frequency part of the first signal and a low frequency part of the second signal using a first processing technique, wherein the first processing technique includes:

obtaining a first downsampling signal by performing a downsampling on the first signal, and

obtaining a second downsampling signal by performing a downsampling on the second signal;

obtaining a frequency domain signal of the first downsampling signal by translating the first downsampling signal from a time domain to a frequency domain, and obtaining a frequency domain signal of the second downsampling signal by translating the second downsampling signal from the time domain to the frequency domain;

obtaining an enhanced frequency domain signal corresponding to the target voice by processing the frequency domain signal of the first downsampling signal and the frequency domain signal of the second downsampling signal;

determining an enhanced voice signal based on the enhanced frequency domain signal;

obtaining the first output voice signal with the low frequency part of the target voice enhanced by upsampling a part of the enhanced voice signal;

obtaining a second output voice signal with a high frequency part of the target voice enhanced by processing a high frequency part of the first signal and a high frequency part of the second signal using a second processing technique; and

obtaining the voice-enhanced output voice signal by combining the first output voice signal and the second output voice signal.

2 . The method of claim 1 , wherein the determining a target SNR of the target voice based on the first signal or the second signal comprises:

obtaining current frame data of the first signal and the second signal, respectively;

determining estimated SNR corresponding to the current frame data of the first signal and the second signal;

determining, based on frame data of at least one of the first signal and the second signal before the current frame data, a verification SNR of the target voice; and

determining the target SNR corresponding to the current frame data of the first signal and the second signal based on the verification SNR and the estimated SNR.

3 . The method of claim 2 , wherein the determining, based on frame data of at least one of the first signal and the second signal before the current frame data, a verification SNR of the target voice; and determining the target SNR corresponding to the current frame data of the first signal and the second signal based on the verification SNR and the estimated SNR comprises:

obtaining at least one voice-enhanced frame data of the first signal and the second signal before the current frame data;

determining at least one verification SNR corresponding to the at least one voice-enhanced frame data; and

determining the target SNR corresponding to the current frame data of the first signal and the second signal based on the at least one verification SNR and the estimated SNR.

4 . The method of claim 1 , wherein the first processing technique further comprises:

supplementing the first downsampling signal and the second downsampling signal so that their signal lengths and sampling frequencies meet a preset condition.

5 . The method of claim 1 , wherein the obtaining an enhanced frequency domain signal corresponding to the target voice by processing the frequency domain signal of the first downsampling signal and the frequency domain signal of the second downsampling signal comprises:

obtaining the enhanced frequency domain signal by performing a differential operation on the frequency domain signal of the first downsampling signal and the frequency domain signal of the second downsampling signal based on a difference factor between a noise signal of the first downsampling signal and a noise signal of the second downsampling signal, wherein the difference factor is determined based on signal energies of the first downsampling signal and the second downsampling signal.

6 . The method of claim 1 , wherein the obtaining an enhanced frequency domain signal corresponding to the target voice by processing the frequency domain signal of the first downsampling signal and the frequency domain signal of the second downsampling signal comprises:

obtaining a preliminary enhanced frequency domain signal by performing a differential operation on the frequency domain signal of the first downsampling signal and the frequency domain signal of the second downsampling signal based on a difference factor between a noise signal of the first downsampling signal and a noise signal of the second downsampling signal; and

obtaining the enhanced frequency domain signal by performing the differential operation based on the preliminary enhanced frequency domain signal, the frequency domain signal of the first downsampling signal, and the frequency domain signal of the second downsampling signal.

7 . The method of claim 6 , wherein the preliminary enhanced frequency domain signal, the frequency domain signal of the first downsampling signal, or the frequency domain signal of the second downsampling signal corresponds to a first weight coefficient, the first weight coefficient being related to a voice existence probability of a currently processed signal.

8 . The method of claim 1 , wherein the first processing technique further comprises:

updating signal values of signal points in the enhanced frequency domain signal whose signal values are smaller than a preset parameter.

9 . The method of claim 1 , wherein the second processing technique comprises:

obtaining a first high frequency band signal corresponding to the high frequency part of the first signal and a second high frequency band signal corresponding to the high frequency part of the second signal; and

obtaining the second output voice signal with the high frequency part of the target voice enhanced by performing a differential operation based on the first high frequency band signal and the second high frequency band signal.

10 . The method of claim 9 , wherein the performing a differential operation based on the first high frequency band signal and the second high frequency band signal comprises:

obtaining a first upsampling signal and a second upsampling signal by upsampling the first high frequency band signal and the second high frequency band signal, respectively; and

obtaining the second output voice signal with the high frequency part of the target voice enhanced by performing the differential operation on the first upsampling signal and the second upsampling signal.

11 . The method of claim 9 , wherein the differential operation comprises:

performing the differential operation based on a first timing signal of the first high frequency band signal and at least one timing signal of the second high frequency band signal before the timing of the first timing signal.

12 . The method of claim 11 , wherein in the at least one timing signal before the timing of the first timing signal, each timing signal corresponds to a second weight coefficient, and the method comprises:

performing the differential operation based on the first timing signal of the first high frequency band signal, the at least one timing signal of the second high frequency band signal before the timing of the first timing signal, and the second weight coefficient corresponding to the at least one timing signal.

13 . A voice enhancement device, comprising:

collection devices located at different voice collection positions;

at least one terminal;

at least one storage medium; and

at least one processor,

wherein the at least one storage medium is configured to store a computer instruction; and the at least one processor is configured to execute the computer instruction to implement operations including:

obtaining a first signal and a second signal of a target voice, the first signal and the second signal being voice signals of the target voice by the collection devices located at the different voice collection positions;

determining a target signal-to-noise ratio (SNR) of the target voice based on the first signal or the second signal;

determining a processing mode for the first signal and the second signal based on the target SNR; and

obtaining a voice-enhanced output voice signal corresponding to the target voice by processing the first signal and the second signal based on the determined processing mode;

the at least one terminal is configured to receive the voice-enhanced output voice signal corresponding to the target voice;

wherein the determining a processing mode for the first signal and the second signal based on the target SNR comprises:

in response to determining that the target SNR is smaller than a first threshold, processing the first signal and the second signal in a first mode; and

in response to determining that the target SNR is greater than a second threshold, processing the first signal and the second signal in a second mode,

wherein

the first threshold is not larger than the second threshold, and

computing resources allocated to the first mode is more than computing resources allocated to the second mode;

wherein the processing the first signal and the second signal in a first mode comprises:

obtaining a first output voice signal with a low frequency part of the target voice enhanced by processing a low frequency part of the first signal and a low frequency part of the second signal using a first processing technique, wherein the first processing technique includes:

obtaining a first downsampling signal by performing a downsampling on the first signal, and

obtaining a second downsampling signal by performing a downsampling on the second signal;

obtaining a frequency domain signal of the first downsampling signal by translating the first downsampling signal from a time domain to a frequency domain, and obtaining a frequency domain signal of the second downsampling signal by translating the second downsampling signal from the time domain to the frequency domain;

obtaining an enhanced frequency domain signal corresponding to the target voice by processing the frequency domain signal of the first downsampling signal and the frequency domain signal of the second downsampling signal;

determining an enhanced voice signal based on the enhanced frequency domain signal;

obtaining the first output voice signal with the low frequency part of the target voice enhanced by upsampling a part of the enhanced voice signal;

obtaining a second output voice signal with a high frequency part of the target voice enhanced by processing a high frequency part of the first signal and a high frequency part of the second signal using a second processing technique; and

obtaining the voice-enhanced output voice signal by combining the first output voice signal and the second output voice signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: XIAO, LE; ZHANG, CHENGQIAN; LIAO, FENGYUN; QI, XIN
To: SHENZHEN SHOKZ CO., LTD.
Reel/Frame 065314/0336 →
Continuity (2)
Continuation PCTCN2021085039 · Apr 1, 2021
Related Publication 20230317093A1 · Oct 5, 2023
References Cited (39)
US 10192567B1 · Kamdar · 2019 [cited by examiner]
US 11259117B1 · Chelliah · 2022 [cited by examiner]
US 20060013412A1 · Goldin · 2006 [cited by applicant]
US 20180012614A1 · Soleymani · 2018 [cited by examiner]
US 20180167747A1 · Kuriger et al. · 2018 [cited by applicant]
US 20200265857A1 · Zhu et al. · 2020 [cited by applicant]
US 20210287687A1 · Disch et al. · 2021 [cited by applicant]
CN 101894563A · 2010 [cited by applicant]
CN 102074246A · 2011 [cited by applicant]
CN 102623016 · 2012 [cited by applicant]
CN 102623016A · 2012 [cited by examiner]
CN 104464745 · 2015 [cited by applicant]
CN 104575511 · 2015 [cited by applicant]
CN 105427859A · 2016 [cited by applicant]
CN 107967918 · 2018 [cited by applicant]
CN 109410976 · 2019 [cited by applicant]
CN 110085246 · 2019 [cited by applicant]
CN 110310651 · 2019 [cited by applicant]
CN 112116918 · 2020 [cited by applicant]
JP 2013068919 · 2013 [cited by applicant]
WO 2010029247A1 · 2010 [cited by applicant]
WO 2013065010A1 · 2013 [cited by applicant]
Yao et al., A priori SNR estimation and noise estimation for speech enhancement, 2016, EURASIP Journal on Advances in Signal Processing (Year: 2016). [cited by examiner]
Stahl et al., A Simple and Effective Framework for A Priori SNR Estimation, 2018, ICASSP 2018 (Year: 2018). [cited by examiner]
Nelke et al., Dual Microphone Wind Noise Reduction by Exploiting the Complex Coherence, 2014, ITG-Fachbericht 252: Speech Communication, 24.-26. (Year: 2014). [cited by examiner]
Hilbert, FFT Zero Padding, 2013, Obtained from: https://www.bitweenie.com/listings/fft-zero-padding/ (Year: 2013). [cited by examiner]
Montana State University, Using a Fast Fourier Transform Algorithm, 2007 (Year: 2007). [cited by examiner]
Luo et al., Adaptive Null-Forming Scheme in Digital Hearing Aids, 2002, IEEE Transactions on Signal Processing, vol. 50, No. 7 ( Year: 2002). [cited by examiner]
International Search Report in PCT/CN2021/085039 mailed on Jan. 4, 2022, 7 pages. [cited by applicant]
Harald Gustafsson et al., Dual-Microphone Spectral Subtraction, ResearchGate, 2000, 38 pages. [cited by applicant]
“A Survey of Beamforming Algorithms”, Web page <https://www.cnblogs.com/LXP-Never/p/12239399.html>, Mar. 1, 2020. [cited by applicant]
Gong, Qin et al., Beamforming and Maximum Likelihood Estimation for Speech Enhancement Using Dual Closely-Spaced Microphones, Journal of Tsinghua University, Science and Technology, 58(6): 603-608, 2018. [cited by applicant]
Philipos C. Loizou, Speech Enhancement: Theory and Practice, University of Electronic Science and Technology Press, 2012, 4 pages. [cited by applicant]
Takuya Higuchi et al., Robust MVDR Beamforming Using Time-Frequency Masks for Online/Offline ASR in Noise, International Conference on Acoustics, Speech and Signal Processing , 2016, 5 pages. [cited by applicant]
Gong, Qin et al., Speech enhancement algorithm for cochlear implants to suppress multi-directional speech noise, Journal of Tsinghua University, Science and Technology, 60(2): 181-188, 2020, 8 pages. [cited by applicant]
Zhang, Xueliang et al., A Speech Enhancement Algorithm by Iterating Single-and Multi-Microphone Processing and Its Application to Robust ASR, International Conference on Acoustics, 2017, 5 pages. [cited by applicant]
Sophocles J.Orfanidis, Introduction to Signal Processing, Tsinghua University Press, pp. 475-520, 1999. [cited by applicant]
Zhang, Xiaomin, The Research and Implementation of Dual-channel Speech Enhancement System, China Master's Theses Full-text Database (CMFD)—Information Science and Technology, 2013, 74 pages. [cited by applicant]
First Office Action in Chinese Application No. 202180068601.4 mailed on Dec. 30, 2025, 33 pages. [cited by applicant]