IP Library Granted Patent US 12,658,199
Granted Patent B2
US 12,658,199 · App. 18/582,989 · Granted Jun 16, 2026

Training method and enhancement method for speech enhancement model, apparatus, electronic device, storage medium and program product

Inventors: Xuefei Fang (Shenzhen, CN); Dong Yang (Shenzhen, CN); Muyong Cao (Shenzhen, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G10L21/0324G10L15/22G10L21/0208G10L21/0316
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,658,199
App. No.
18/582,989
Granted
Jun 16, 2026
Kind
B2
Abstract

A training method and an enhancement method for a speech enhancement model, an apparatus, an electronic device, a storage medium, and a program product are described. A training method for the model may include: invoking the model based on a noisy speech feature of a speech signal, to obtain first predicted mask values in an auditory domain. A first amplitude and first phase of the noisy speech signal, and a second amplitude and second phase of a clean speech signal may then be obtained. A phase difference at each frequency point may be determined based on the phases. The second amplitude corresponding to each frequency point may be corrected based on the phase difference at each frequency point, to obtain a corrected second amplitude corresponding to each frequency point. A loss value is determined, and parameters of the speech enhancement model may be updated based on the loss value.

Claims (106)

1 . A training method for a speech enhancement model, performed by an electronic device, the method comprising:

invoking the speech enhancement model based on a noisy speech feature of a noisy speech signal, to perform speech enhancement on the noisy speech signal, to obtain a plurality of first predicted mask values, the plurality of first predicted mask values having a one-to-one correspondence with a plurality of frequency bands in an auditory domain;

obtaining a first amplitude and a first phase corresponding to each frequency point of the noisy speech signal, and a second amplitude and a second phase corresponding to each frequency point of a clean speech signal;

determining a phase difference between the clean speech signal and the noisy speech signal at each frequency point based on the first phase and the second phase corresponding to each frequency point;

correcting the second amplitude corresponding to the frequency point based on the phase difference, to obtain a corrected second amplitude corresponding to each frequency point;

determining a loss value based on the plurality of first predicted mask values, and the first amplitude and the corrected second amplitude corresponding to each frequency point, including:

mapping the plurality of first predicted mask values respectively to obtain second predicted mask values corresponding to each frequency point; and

determining the loss value based on the second predicted mask value, the first amplitude, and the corrected second amplitude corresponding to each frequency point; and

performing backpropagation in the speech enhancement model based on the loss value, to update parameters of the speech enhancement model.

2 . The method according to claim 1 , wherein

the mapping the plurality of first predicted mask values respectively to obtain second predicted mask values corresponding to each frequency point comprises:

determining the second predicted mask value corresponding to each frequency point in one of the following manners:

determining a first frequency band to which the frequency point belongs in the auditory domain, and determining a first predicted mask value corresponding to the first frequency band as the second predicted mask value corresponding to the frequency point; or

determining the first frequency band to which the frequency point belongs in the auditory domain, and determining a reference frequency band adjacent to the first frequency band in the auditory domain; and

performing weighted summation on the first predicted mask value corresponding to the first frequency band and a first predicted mask value corresponding to the reference frequency band, to obtain the second predicted mask value corresponding to the frequency point,

a weight corresponding to each of the first predicted mask values being positively correlated with a distance between the following two elements: the frequency point and a center frequency point of a frequency band corresponding to the first predicted mask value.

3 . The method according to claim 1 , wherein

the determining the loss value based on the second predicted mask value, the first amplitude, and the corrected second amplitude corresponding to each frequency point further comprises:

multiplying the second predicted mask value corresponding to each frequency point by the first amplitude corresponding to the frequency point, to obtain a third amplitude corresponding to each frequency point; and

substituting the third amplitude corresponding to each frequency point and the corrected second amplitude corresponding to the frequency point into a first target loss function for calculation, to obtain the loss value.

4 . The method according to claim 1 , wherein

the determining the loss value based on the second predicted mask value, the first amplitude, and the corrected second amplitude corresponding to each frequency point further comprises:

determining a ratio of the corrected second amplitude corresponding to each frequency point to the first amplitude corresponding to the frequency point as a first target mask value corresponding to each frequency point; and

substituting the second predicted mask value corresponding to each frequency point and the first target mask value corresponding to the frequency point into a second target loss function for calculation, to obtain the loss value.

5 . The method of claim 1 , further comprising:

invoking the speech enhancement model with the updated parameters based on a to-be-processed speech feature of a to-be-processed speech signal, to perform speech enhancement, to obtain a plurality of third predicted mask values, the plurality of third predicted mask values having a one-to-one correspondence with the plurality of frequency bands in the auditory domain;

performing mapping processing based on the plurality of third predicted mask values to obtain a mapping processing result; and

performing signal reconstruction based on the mapping processing result and a phase spectrum of the to-be-processed speech signal, to obtain an enhanced speech signal.

6 . The method according to claim 5 , wherein

the mapping processing result comprises a third predicted mask value corresponding to a predicted frequency point of the to-be-processed speech signal; and

the performing mapping processing based on the plurality of predicted mask values in the auditory domain to obtain a mapping processing result comprises:

for each frequency point: determining a frequency band to which the frequency point belongs in the auditory domain, and determining a third predicted mask value corresponding to the frequency band as a fourth predicted mask value corresponding to the frequency point; or

determining a frequency band to which the frequency point belongs in the auditory domain, and determining at least one reference frequency band adjacent to the frequency band in the auditory domain; and performing weighted summation on the third predicted mask value corresponding to the frequency band and a third predicted mask value corresponding to the at least one reference frequency band, to obtain a fourth predicted mask value corresponding to the frequency point.

7 . The method according to claim 6 , wherein the performing signal reconstruction processing based on the mapping processing result and a phase spectrum of the to-be-processed speech signal, to obtain an enhanced speech signal comprises:

multiplying mask values corresponding to a plurality of frequency points in a frequency domain by an amplitude spectrum of the to-be-processed speech signal after obtaining the mask values corresponding to the plurality of frequency points in the frequency domain, including multiplying the mask values of the frequency points by amplitudes of corresponding frequency points in the amplitude spectrum, and maintaining amplitudes of other frequency points in the amplitude spectrum unchanged, to obtain an enhanced amplitude spectrum; and

performing an inverse Fourier transform on the enhanced amplitude spectrum and the phase spectrum of the to-be-processed speech signal, to obtain an enhanced speech signal.

8 . A training method for a speech enhancement model, performed by an electronic device, the method comprising:

invoking the speech enhancement model based on a noisy speech feature of a noisy speech signal, to perform speech enhancement on the noisy speech signal, to obtain a plurality of first predicted mask values, the plurality of first predicted mask values having a one-to-one correspondence with a plurality of frequency bands in an auditory domain;

obtaining a first amplitude and a first phase corresponding to each frequency point of the noisy speech signal, and a second amplitude and a second phase corresponding to each frequency point of a clean speech signal;

determining a phase difference between the clean speech signal and the noisy speech signal at each frequency point based on the first phase and the second phase corresponding to each frequency point;

correcting the second amplitude corresponding to the frequency point based on the phase difference, to obtain a corrected second amplitude corresponding to each frequency point;

determining a loss value based on the plurality of first predicted mask values, and the first amplitude and the corrected second amplitude corresponding to each frequency point, including:

mapping the first amplitude and the corrected second amplitude corresponding to each frequency point to a frequency band corresponding to the auditory domain;

determining a first energy corresponding to each frequency band based on a first amplitude mapped to each frequency band, wherein the first energy is a weighted summation result of the following parameters: squares of first amplitudes mapped to each frequency band;

determining a second energy corresponding to each frequency band based on the corrected second amplitude mapped to each frequency band, the second energy being a weighted summation result of the following parameters: squares of corrected second amplitudes mapped to each frequency band; and

determining the loss value based on the first predicted mask value, the first energy, and the second energy corresponding to each frequency band; and

performing backpropagation in the speech enhancement model based on the loss value, to update parameters of the speech enhancement model.

9 . The method according to claim 8 , wherein

the mapping the first amplitude and the corrected second amplitude that are corresponding to each frequency point to a frequency band corresponding to the auditory domain comprises:

determining a second frequency band to which each frequency point belongs in the auditory domain; and

mapping the first amplitude and the corrected second amplitude corresponding to each frequency point to the second frequency band to which the frequency point belongs in the auditory domain.

10 . The method according to claim 8 , wherein

the determining the loss value based on the first predicted mask value, the first energy, and the second energy corresponding to each frequency band further comprises:

determining a second target mask value corresponding to each frequency band based on the first energy and the second energy corresponding to each frequency band; and

substituting the first predicted mask value corresponding to each frequency band and the second target mask value corresponding to the frequency band into a third target loss function for calculation, to obtain the loss value.

11 . The method according to claim 10 , wherein

the determining a second target mask value corresponding to each frequency band based on the first energy and the second energy corresponding to each frequency band comprises:

determining the second target mask value corresponding to each frequency band in one of the following manners:

determining a ratio of the second energy to the first energy corresponding to the frequency band as the second target mask value corresponding to the frequency band; or

determining a difference between the first energy and the second energy corresponding to the frequency band as third energy corresponding to the frequency band; and

summing a square of the second energy and a square of the third energy corresponding to the frequency band, to obtain a first summation result; and determining a ratio of the square of the second energy to the first summation result as the second target mask value corresponding to the frequency band.

12 . The method according to claim 8 , wherein

the determining the loss value based on the first predicted mask value, the first energy, and the second energy corresponding to each frequency band further comprises:

multiplying the first predicted mask value corresponding to each frequency band by the first energy corresponding to the frequency band, to obtain fourth energy corresponding to each frequency band; and

substituting the second energy corresponding to each frequency band and the fourth energy corresponding to the frequency band into a fourth target loss function for calculation, to obtain the loss value.

13 . A training apparatus for a speech enhancement model, the apparatus comprising:

an enhancement module, configured to invoke a speech enhancement model based on a noisy speech feature of a noisy speech signal, to perform speech enhancement, to obtain a plurality of first predicted mask values in an auditory domain, the plurality of first predicted mask values having a one-to-one correspondence with a plurality of frequency bands in the auditory domain;

an obtaining module, configured to obtain a first amplitude and a first phase corresponding to each frequency point of the noisy speech signal, and a second amplitude and a second phase corresponding to each frequency point of a clean speech signal;

a correction module, configured to determine a phase difference between the clean speech signal and the noisy speech signal at each frequency point based on the first phase and the second phase corresponding to each frequency point, and correct the second amplitude corresponding to the frequency point based on the phase difference at each frequency point, to obtain a corrected second amplitude corresponding to each frequency point;

a determining module, configured to determine a loss value based on the plurality of first predicted mask values, and the first amplitude and the corrected second amplitude corresponding to each frequency point, wherein determining the loss value includes:

mapping the plurality of first predicted mask values respectively to obtain second predicted mask values corresponding to each frequency point; and

determining the loss value based on the second predicted mask value, the first amplitude, and the corrected second amplitude corresponding to each frequency point; and

an update module, configured to perform backpropagation in the speech enhancement model based on the loss value to update parameters of the speech enhancement model.

14 . The training apparatus of claim 13 , wherein:

the enhancement module is further configured to invoke the speech enhancement model with the updated parameters based on a to-be-processed speech feature of a to-be-processed speech signal, to perform speech enhancement, to obtain a plurality of third predicted mask values, the plurality of third predicted mask values having a one-to-one correspondence with the plurality of frequency bands in the auditory domain; and

wherein the apparatus further comprises:

a mapping module, configured to perform mapping processing based on the plurality of mask values in the auditory domain, to obtain a mapping processing result; and

a reconstruction module, configured to perform signal reconstruction based on the mapping processing result and a phase spectrum of the to-be-processed speech signal to obtain an enhanced speech signal.

15 . The training apparatus of claim 13 , wherein

the mapping the plurality of first predicted mask values respectively to obtain second predicted mask values corresponding to each frequency point comprises:

determining the second predicted mask value corresponding to each frequency point in one of the following manners:

determining a first frequency band to which the frequency point belongs in the auditory domain, and determining a first predicted mask value corresponding to the first frequency band as the second predicted mask value corresponding to the frequency point; or

determining the first frequency band to which the frequency point belongs in the auditory domain, and determining a reference frequency band adjacent to the first frequency band in the auditory domain; and

performing weighted summation on the first predicted mask value corresponding to the first frequency band and a first predicted mask value corresponding to the reference frequency band, to obtain the second predicted mask value corresponding to the frequency point,

a weight corresponding to each of the first predicted mask values being positively correlated with a distance between the following two elements: the frequency point and a center frequency point of a frequency band corresponding to the first predicted mask value.

16 . The training apparatus of claim 13 , wherein

the determining the loss value based on the second predicted mask value, the first amplitude, and the corrected second amplitude corresponding to each frequency point further comprises:

multiplying the second predicted mask value corresponding to each frequency point by the first amplitude corresponding to the frequency point, to obtain a third amplitude corresponding to each frequency point; and

substituting the third amplitude corresponding to each frequency point and the corrected second amplitude corresponding to the frequency point into a first target loss function for calculation, to obtain the loss value.

17 . A non-transitory computer-readable storage medium, having executable instructions stored therein, the executable instructions, when executed by a processor, causing an apparatus to perform:

invoking a speech enhancement model based on a noisy speech feature of a noisy speech signal, to perform speech enhancement on the noisy speech signal, to obtain a plurality of first predicted mask values, the plurality of first predicted mask values having a one-to-one correspondence with a plurality of frequency bands in an auditory domain;

obtaining a first amplitude and a first phase corresponding to each frequency point of the noisy speech signal, and a second amplitude and a second phase corresponding to each frequency point of a clean speech signal;

determining a phase difference between the clean speech signal and the noisy speech signal at each frequency point based on the first phase and the second phase corresponding to each frequency point;

correcting the second amplitude corresponding to the frequency point based on the phase difference, to obtain a corrected second amplitude corresponding to each frequency point;

determining a loss value based on the plurality of first predicted mask values, and the first amplitude and the corrected second amplitude corresponding to each frequency point, including:

mapping the plurality of first predicted mask values respectively to obtain second predicted mask values corresponding to each frequency point; and

determining the loss value based on the second predicted mask value, the first amplitude, and the corrected second amplitude corresponding to each frequency point; and

performing backpropagation in the speech enhancement model based on the loss value, to update parameters of the speech enhancement model.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the apparatus is further caused to perform:

invoking the speech enhancement model with the updated parameters based on a to-be-processed speech feature of a to-be-processed speech signal, to perform speech enhancement, to obtain a plurality of third predicted mask values, the plurality of third predicted mask values having a one-to-one correspondence with the plurality of frequency bands in the auditory domain;

performing mapping processing based on the plurality of third predicted mask values to obtain a mapping processing result; and

performing signal reconstruction based on the mapping processing result and a phase spectrum of the to-be-processed speech signal, to obtain an enhanced speech signal.

19 . The non-transitory computer-readable storage medium of claim 17 , wherein

the determining the loss value based on the second predicted mask value, the first amplitude, and the corrected second amplitude corresponding to each frequency point further comprises:

multiplying the second predicted mask value corresponding to each frequency point by the first amplitude corresponding to the frequency point, to obtain a third amplitude corresponding to each frequency point; and

substituting the third amplitude corresponding to each frequency point and the corrected second amplitude corresponding to the frequency point into a first target loss function for calculation, to obtain the loss value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2024
From: FANG, XUEFEI; YANG, DONG; CAO, MUYONG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 066512/0894 →
Priority Claims (1)
CN 202210917051.5 · Aug 1, 2022 · national
Continuity (2)
Continuation PCTCN2023096246 · May 25, 2023
Related Publication 20240194214A1 · Jun 13, 2024
References Cited (26)
US 9626987B2 · Matsuo · 2017 [cited by examiner]
US 10510358B1 · Barra-Chicote · 2019 [cited by examiner]
US 11056130B2 · Zhu · 2021 [cited by examiner]
US 11894011B2 · McCallum · 2024 [cited by examiner]
US 12217766B2 · McCallum · 2025 [cited by examiner]
US 20100246851A1 · Buck · 2010 [cited by examiner]
US 20100323652A1 · Visser · 2010 [cited by examiner]
US 20130136271A1 · Buck · 2013 [cited by examiner]
US 20140149111A1 · Matsuo · 2014 [cited by examiner]
US 20200265857A1 · Zhu · 2020 [cited by examiner]
US 20210201928A1 · Rao · 2021 [cited by examiner]
US 20220150627A1 · Zhou · 2022 [cited by examiner]
US 20220291328A1 · Ozturk · 2022 [cited by examiner]
US 20230097520A1 · Xiao · 2023 [cited by examiner]
US 20240194214A1 · Fang · 2024 [cited by examiner]
CN 102169694A · 2011 [cited by applicant]
CN 113436643A · 2021 [cited by applicant]
CN 110600017B · 2022 [cited by applicant]
CN 114974299A · 2022 [cited by applicant]
WO 2022012195A1 · 2022 [cited by applicant]
Sep. 18, 2023—(WO) Search Report—App PCT/CN2023/096246. [cited by applicant]
Zheng, Li et al., “Two-stage speech enhancement algorithm based on time-frequency mask optimization”, Electronic Design Engineering, vol. 30, No. 4, Feb. 2022, p. 17-21. [cited by applicant]
Erdogan, H. et al., “Phase Sensitive and Recognition-Boosted Speech Separation Using Deep Recurrent Neural Networks,” Mitsubishi Electric Research Laboratories, Apr. 2015, 7 pages. [cited by applicant]
Jamal, Norezmi et al., “Binary Time-Frequency Mask for Improved Malay Speech Intelligibility at Low SNR Condition,” 2020 IOP Conference Series: Materials Science and Engineering, Apr. 18, 2020, 7 pages. [cited by applicant]
European Patent Office Office Action for EP Application No. 23849005.6—1207/4394769 PCT/CN2023096246, date Dec. 16, 2024, including search strategy details in p. 10. [cited by applicant]
Michelsanti, Daniel et al., “An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation”, Electronic Systems, dated Aug. 21, 2020. [cited by applicant]