IP Library › Granted Patent US 12,586,599
Granted Patent B2
US 12,586,599 · App. 18/312,688 · Granted Mar 24, 2026

Audio signal processing method and apparatus, electronic device, and storage medium with machine learning and for microphone mute state features in a multi person voice call

Inventors: Siyu Zhang (Shenzhen, CN); Yi Gao (Shenzhen, CN); Cheng Luo (Shenzhen, CN); Bin Li (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G10L21/0364G06F3/165G10L17/02G10L21/034G10L25/21G10L25/30G10L25/84H04M3/568G10L25/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,599
App. No.
18/312,688
Granted
Mar 24, 2026
Kind
B2
Abstract

Herein are disclosed an audio signal processing method and apparatus, an electronic device, and a storage medium. The method includes: obtaining an audio signal acquired by an application while an account logging into the application is in a microphone mute state in a multi-person voice call; obtaining gain parameters, for each of a plurality of audio frames in the audio signal, respectively on a plurality of bands in a first band range; and outputting a prompt message responsive to a determination, based on the gain parameters, that a target voice is contained in the audio signal, the prompt message providing a prompt to disable the microphone mute state of the account.

Claims (75)

1 . An audio signal processing method, performed by at least one processor on a terminal, the method comprising:

obtaining an audio signal acquired by an application while an account logging into the application is in a microphone mute state in a multi-person voice call;

obtaining gain parameters, for each of a plurality of audio frames in the audio signal, respectively on a plurality of bands in a first band range, and the gain parameters are obtained by inputting data of the audio signal into a noise suppression (NS) model of a recurrent neural network (RNN); and

outputting a prompt message responsive to a determination, based on the gain parameters, that a target voice is contained in the audio signal, the prompt message providing a prompt to disable the microphone mute state of the account,

wherein the gain parameters are obtained by:

preprocessing the audio signal to obtain a first signal, the first signal comprising the data of the audio signal; and

inputting a plurality of audio frames in the first signal into the noise suppression (NS) model, and processing each audio frame in the plurality of audio frames by the NS model to thereby output a gain parameter of each audio frame on each band in the first band range, the gain parameter of the audio frame on a human voice band being greater than the gain parameter on a noise band,

wherein the RNN comprises at least one hidden layer, each hidden layer comprises a plurality of neurons, and the number of neurons in each hidden layer is equal to the number of inputted audio frames, and

the NS model processes each audio frame by:

weighting, through a present neuron in a present hidden layer in the RNN, a first frequency feature outputted by a previous neuron in the present hidden layer and a second frequency feature outputted by a neuron at a corresponding position in a previous hidden layer, and

inputting the weighted first and second frequency features to a next neuron in the present hidden layer and a neuron at a corresponding position in a next hidden layer respectively.

2 . The method according to claim 1 , wherein the target voice is determined to be contained in the audio signal by:

determining voice state parameters of the plurality of audio frames based on the gain parameters of the plurality of audio frames on the plurality of bands, the voice state parameters being used for representing whether a corresponding audio frame contains a target voice; and

determining, based on the voice state parameters of the plurality of audio frames, that the target voice is contained in the audio signal.

3 . The method according to claim 2 , wherein the voice state parameters are determined by:

determining a gain parameter of each audio frame on each band in a second band range based on the gain parameter of the audio frame on each band in the first band range, the second band range being a subset of the first band range; and

determining a voice state parameter of each audio frame based on the gain parameter of the audio frame on each band in the second band range.

4 . The method according to claim 3 , wherein the voice state parameter of each audio frame is determined by:

multiplying the gain parameter of the audio frame on each band in the second band range by a weight coefficient of a corresponding band to obtain a weighted gain parameter of the audio frame on each band in the second band range;

adding the weighted gain parameters of the audio frame on the respective bands in the second band range to obtain a comprehensive gain parameter of the audio frame; and

determining the voice state parameter of the audio frame based on the comprehensive gain parameter of the audio frame.

5 . The method according to claim 4 , wherein the voice state parameter of each audio frame is determined based on the comprehensive gain parameter by:

determining that the voice state parameter indicates the target voice when the comprehensive gain parameter amplified by a target multiple is greater than an activation threshold; and

determining that the voice state parameter indicates lack of the target voice when the comprehensive gain parameter amplified by the target multiple is less than or equal to the activation threshold.

6 . The method according to claim 2 , further comprising obtaining energy parameters of the plurality of audio frames,

wherein the voice state parameters of the plurality of audio frames are determined based on the gain parameters of the plurality of audio frames on the plurality of bands and the energy parameters of the plurality of audio frames.

7 . The method according to claim 6 , wherein the voice state parameters of the plurality of audio frames are determined by:

determining a comprehensive gain parameter of each audio frame based on the gain parameters of the audio frame on the plurality of bands;

determining that the voice state parameter of each audio frame indicates that the audio frame contains the target voice when the comprehensive gain parameter of the audio frame amplified by a target multiple is greater than an activation threshold and the energy parameter of the audio frame is greater than an energy threshold; and

determining that the voice state parameter of each audio frame indicates that the audio frame does not contain the target voice when the comprehensive gain parameter of the audio frame amplified by the target multiple is less than or equal to the activation threshold or the energy parameter of the audio frame is less than or equal to the energy threshold.

8 . The method according to claim 2 , wherein the target voice is determined to be contained in the audio signal based on voice state parameters of an audio frame and a target number of audio frames preceding the audio frame satisfying a first condition.

9 . The method according to claim 8 , wherein the target voice is determined to be contained in the audio signal by:

determining, based on the voice state parameters of the audio frame and of a first target number of audio frames preceding the audio frame, an activation state of an audio frame group comprising the audio frame and the first target number of audio frames preceding the audio frame; and

determining that the target voice is contained in the audio signal when the activation states of the audio frame group and of a second target number of audio frame groups preceding the audio frame group satisfy a second condition, the target number of audio frames being determined based on the first target number and the second target number.

10 . The method according to claim 9 , wherein the activation state of the audio frame group comprising the audio frame is determined by:

determining that the activation state of the audio frame group is activated when the number of audio frames containing the target voice exceeds a number threshold; and

determining that the activation state of the audio frame group is unactivated when the number of audio frames containing the target voice does not exceed the number threshold.

11 . The method according to claim 1 , wherein the target voice is determined to be contained in the audio signal by:

performing noise suppression on the plurality of audio frames, based on the gain parameters of the plurality of audio frames, on the plurality of bands to obtain a plurality of target audio frames;

performing voice activity detection (VAD) based on energy parameters of the plurality of target audio frames to obtain VAD values of the plurality of target audio frames; and

determining that the target voice is contained in the audio signal when the VAD values of the plurality of target audio frames satisfy a third condition.

12 . The method according to claim 1 , wherein the target voice is one of: a speech of a target object in the multi-person voice call, or a sound of the target object.

13 . The method according to claim 1 , further comprising,

responsive to the determination that the target voice is contained in the audio signal, automatically disabling the microphone mute state of the account.

14 . An audio signal processing apparatus, disposed in a terminal, the apparatus comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code comprising:

first obtaining code, configured to cause the at least one processor to obtain an audio signal acquired by an application while an account logging into the application is in a microphone mute state in a multi-person voice call;

second obtaining code, configured to cause the at least one processor to obtain gain parameters, for each of a plurality of audio frames in the audio signal, respectively on a plurality of bands in a first band range, and the gain parameters are obtained by inputting data of the audio signal into a noise suppression (NS) model of a recurrent neural network (RNN)I; and

output code, configured to cause the at least one processor to output a prompt message responsive to a determination, based on the gain parameters, that a target voice is contained in the audio signal, the prompt message providing a prompt to disable the microphone mute state of the account,

wherein the gain parameters are obtained by:

preprocessing the audio signal to obtain a first signal, the first signal comprising the data of the audio signal; and

inputting a plurality of audio frames in the first signal into the noise suppression (NS) model, and processing each audio frame in the plurality of audio frames by the NS model to thereby output a gain parameter of each audio frame on each band in the first band range, the gain parameter of the audio frame on a human voice band being greater than the gain parameter on a noise band,

wherein the RNN comprises at least one hidden layer, each hidden layer comprises a plurality of neurons, and the number of neurons in each hidden layer is equal to the number of inputted audio frames, and

the NS model processes each audio frame by:

weighting, through a present neuron in a present hidden layer in the RNN, a first frequency feature outputted by a previous neuron in the present hidden layer and a second frequency feature outputted by a neuron at a corresponding position in a previous hidden layer, and

inputting the weighted first and second frequency features to a next neuron in the present hidden layer and a neuron at a corresponding position in a next hidden layer respectively.

15 . The apparatus according to claim 14 , wherein the target voice is determined to be contained in the audio signal by:

determining voice state parameters of the plurality of audio frames based on the gain parameters of the plurality of audio frames on the plurality of bands, the voice state parameters being used for representing whether a corresponding audio frame contains a target voice; and

determining, based on the voice state parameters of the plurality of audio frames, that the target voice is contained in the audio signal.

16 . The apparatus according to claim 14 , wherein the target voice is determined to be contained in the audio signal by:

performing noise suppression on the plurality of audio frames, based on the gain parameters of the plurality of audio frames, on the plurality of bands to obtain a plurality of target audio frames;

performing voice activity detection (VAD) based on energy parameters of the plurality of target audio frames to obtain VAD values of the plurality of target audio frames; and

determining that the target voice is contained in the audio signal when the VAD values of the plurality of target audio frames satisfy a third condition.

17 . A non-transitory computer-readable storage medium, the storage medium storing at least one computer program, the at least one computer program being executable by a processor to perform audio signal processing operations of:

obtaining an audio signal acquired by an application while an account logging into the application is in a microphone mute state in a multi-person voice call;

obtaining gain parameters, for each of a plurality of audio frames in the audio signal, respectively on a plurality of bands in a first band range, and the gain parameters are obtained by inputting data of the audio signal into a noise suppression (NS) model of a recurrent neural network (RNN); and

outputting a prompt message responsive to a determination, based on the gain parameters, that a target voice is contained in the audio signal, the prompt message providing a prompt to disable the microphone mute state of the account,

wherein the gain parameters are obtained by:

preprocessing the audio signal to obtain a first signal, the first signal comprising the data of the audio signal; and

inputting a plurality of audio frames in the first signal into the noise suppression (NS) model, and processing each audio frame in the plurality of audio frames by the NS model to thereby output a gain parameter of each audio frame on each band in the first band range, the gain parameter of the audio frame on a human voice band being greater than the gain parameter on a noise band,

wherein the RNN comprises at least one hidden layer, each hidden layer comprises a plurality of neurons, and the number of neurons in each hidden layer is equal to the number of inputted audio frames, and

the NS model processes each audio frame by:

weighting, through a present neuron in a present hidden layer in the RNN, a first frequency feature outputted by a previous neuron in the present hidden layer and a second frequency feature outputted by a neuron at a corresponding position in a previous hidden layer, and

inputting the weighted first and second frequency features to a next neuron in the present hidden layer and a neuron at a corresponding position in a next hidden layer respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2023
From: ZHANG, SIYU; GAO, YI; LUO, CHENG; LI, BIN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 063548/0773 →
Priority Claims (1)
CN 202111087468.5 · Sep 16, 2021 · national
Continuity (2)
Continuation PCTCN2022111474 · Aug 10, 2022
Related Publication 20230317096A1 · Oct 5, 2023
References Cited (36)
US 6671667B1 · Chandran · 2003 [cited by examiner]
US 6810273B1 · Mattila · 2004 [cited by examiner]
US 7359504B1 · Reuss · 2008 [cited by examiner]
US 8903721B1 · Cowan · 2014 [cited by examiner]
US 9473643B2 · Baran et al. · 2016 [cited by applicant]
US 10805740B1 · Snyder · 2020 [cited by examiner]
US 20040260550A1 · Burges et al. · 2004 [cited by applicant]
US 20060280295A1 · Runcie · 2006 [cited by examiner]
US 20080052074A1 · Gopinath · 2008 [cited by examiner]
US 20090022305A1 · Chavez · 2009 [cited by examiner]
US 20100080382A1 · Dresher · 2010 [cited by examiner]
US 20100324891A1 · Cutler · 2010 [cited by examiner]
US 20130321156A1 · Liu · 2013 [cited by examiner]
US 20150195411A1 · Krack · 2015 [cited by examiner]
US 20160182727A1 · Baran et al. · 2016 [cited by applicant]
US 20160295539A1 · Atti · 2016 [cited by examiner]
US 20170092288A1 · Dewasurendra · 2017 [cited by examiner]
US 20200043514A1 · Liu · 2020 [cited by examiner]
US 20200312343A1 · Hsiung · 2020 [cited by examiner]
US 20210125625A1 · Huang · 2021 [cited by examiner]
US 20240195916A1 · Schiøler · 2024 [cited by examiner]
CN 1565144A · 2005 [cited by applicant]
CN 107276777A · 2017 [cited by applicant]
CN 110085249A · 2019 [cited by applicant]
CN 110111805A · 2019 [cited by applicant]
CN 110769354A · 2020 [cited by applicant]
CN 111343410A · 2020 [cited by applicant]
CN 111415685A · 2020 [cited by applicant]
CN 111429932A · 2020 [cited by applicant]
CN 113113039A · 2021 [cited by applicant]
JP 2010102203A · 2010 [cited by applicant]
Office Action issued Sep. 17, 2024 in Japanese Application No. 2023-551247. [cited by applicant]
International Search Report of PCT/CN2022/111474 dated Oct. 25, 2022 [PCT/ISA/210]. [cited by applicant]
Written Opinion of PCT/CN2022/111474 dated Oct. 25, 2022 [PCT/ISA/237]. [cited by applicant]
Extended European Search Report dated Oct. 15, 2024 in application No. 22868900.6. [cited by applicant]
Communication from the Chinese Patent Office in Application No. 202111087468.5 dated Jul. 19, 2025. [cited by applicant]