IP Library Granted Patent US 12676165
Granted Patent B2
US 12676165 · App. 18/595,546 · Granted Jul 7, 2026

Method for mixing microphone inputs, apparatus, and computer program product

Inventor: Magdalena Kaniewska (Leuven, BE)
Assignee: GOODIX TECHNOLOGY (HK) COMPANY LIMITED
G10L25/21G10L21/0208H04R3/005G10L2021/02082G10L2021/02165H04R2499/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12676165
App. No.
18/595,546
Granted
Jul 7, 2026
Kind
B2
Abstract

A method for mixing a plurality of input signals, an apparatus and a computer program product are provided. The method comprises receiving a plurality of current power values associated to a current time interval and a plurality of previous smoothed power values associated to a previous time interval, when it is determined that at least one of the plurality of input signals contains speech, calculating the current smoothed power value for each input signal based on a current power value and a previous smoothed power value, when it is determined that none of the plurality of input signals contains speech, calculating the current smoothed power value for each input signal based on a determined value and the previous smoothed power value corresponding to each input signal and calculating a plurality of mixing gains based on the plurality of current smoothed power values.

Claims (48)

1 . A method for mixing a plurality of input signals associated respectively with a plurality of microphones wherein each of the plurality of input signals comprises sound events generated by a one or more sound sources, the method comprising:

receiving, by a processor, a plurality of previous smoothed power values associated to a previous time interval, and a plurality of current power values corresponds respectively to each of the plurality of input signals;

determining, by the processor, whether at least one of the plurality of input signals contains speech;

calculating, by the processor, a plurality of current smoothed power values respectively for the plurality of input signals at a current time interval based on the plurality of previous smoothed power values associated to the previous time interval; and

mixing the plurality of input signals based on the plurality of current smoothed power values,

wherein the calculating the plurality of current smoothed power values comprises calculating a current smoothed power value for each input signal of the plurality of input signals based on whether at least one of the plurality of input signals contains speech; and

wherein the mixing the plurality of input signals comprises calculating a plurality of mixing gains for the plurality of input signals and a mixing gain among the plurality of mixing gains for an input signal among the plurality of input signals is determined based on a current smoothed power value among the plurality of current smoothed power values corresponding to the input signal.

2 . The method according to claim 1 , wherein the calculating the plurality of current smoothed power values comprises:

when it is determined that at least one of the plurality of input signals contains speech and echo is not dominating over speech, calculating the current smoothed power value for each input signal based on a current power value among a plurality of current power values and a previous smoothed power value among the plurality of previous smoothed power values, wherein each of the plurality of previous smoothed power values corresponds respectively to each of the plurality of input signals, and the current power value and the previous smoothed power value correspond to each input signal; and

when it is determined that none of the plurality of input signals contains speech or it is determined that at least one of the plurality of input signals contains speech but the echo dominates over speech, calculating for each input signal current smoothed power value based on a determined value and the previous smoothed power value corresponding to each input signal.

3 . The method according to claim 2 , wherein the calculating for each input signal current smoothed power value based on the determined value and the previous smoothed power value corresponding to each input signal comprises that the current smoothed power value is determined by smoothing between the determined value and the previous smoothed power value corresponding to each input signal and wherein the determined value is an average of the plurality of previous smoothed power values.

4 . The method according to claim 1 , wherein one of the plurality of mixing gains are further determined based on a power ratio between the current smoothed power value of a input signal and an average of the plurality of current smoothed power values of the plurality of input signals.

5 . The method according to claim 4 , wherein one of the plurality of mixing gains is further determined by a first updated power ratio, wherein the first updated power ratio equals to the square root of the power ratio diving divided by a number of the plurality of input signals.

6 . The method according to claim 5 , wherein one of the plurality of mixing gains are further determined by a second updated power ratio, wherein if the first updated power ratio is greater than a high threshold the second updated power ratio is determined to be the high threshold, if the first updated power ratio is not greater than a low threshold, the second updated power ratio is determined to be the low threshold, and if the second updated power ratio is determined to be not greater than the high threshold and greater than the low threshold, the second updated power ratio is determined to be the first updated power ratio, wherein the high threshold is greater than the low threshold.

7 . The method according to claim 6 , wherein one of plurality of mixing gains is determined as a ratio between a third updated power ratio and a sum of the plurality of the third updated power ratios, wherein the third updated power ratio is equal to a division between the second updated power ratio minus the low threshold and the high threshold minus the low threshold.

8 . The method according to claim 1 , wherein the plurality of mixing gains are determined based on the probabilities of speech or Signal to Noise Ratio (SNR).

9 . The method according to claim 8 , the method further comprising at least one of operations (a) and (b),

wherein the operation (a) includes:

determining each mixing gain of the plurality of mixing gains based on Voice Activity Detection ranging between 0 to 1 and a value 1/K, wherein K is a number of the plurality of input signals; and,

the operation (b) includes:

determining each mixing gain of the plurality of mixing gains based on Signal-to-Noise Ratio (SNR) of the plurality of input signals.

10 . The method according to claim 1 , the method further comprising:

splitting a frequency range of an input signal among the plurality of inputs signals into a plurality of frequency subranges;

calculating a plurality of power weight values respectively for the plurality of frequency subranges based on the Signal to Noise Ratio, SNR, of the input signal in corresponding frequency subrange; and

calculating a current power value among the plurality of current power values based on the plurality of power weight values.

11 . The method according to claim 10 , wherein the plurality of power weight values of the plurality of frequency subranges is calculated as a ratio between the average SNR of the input signal in corresponding frequency subrange and a sum of the plurality of the average SNR of the input signal in the plurality of frequency subranges.

12 . The method according to claim 10 , wherein the calculating the current power value based on the plurality of power weight values comprises weighing power of each frequency subrange of the plurality of frequency subrange by applying corresponding power weight among the plurality of power weight values to the power of the input signal of the corresponding subband and calculating a sum of the weighted powers of the plurality of the subbands for the input signals.

13 . The method according to claim 1 , wherein the plurality of input signals comprise microphone signals from microphones which are located more than 25 centimetres from each other, and the plurality of input signals further comprise output of the beamformer whose input are microphone signals from microphones which are located no more than 25 centimetres from each other.

14 . An apparatus for mixing a plurality of input signals associated respectively with a plurality of microphones wherein each of the plurality of input signals comprises sound events generated by a one or more sound sources, the apparatus comprising a memory and a processor communicatively connected to the memory and configured to execute instructions to perform a method for mixing the plurality of input signals, the method comprising:

receiving, by a processor, a plurality of previous smoothed power values associated to a previous time interval, and a plurality of current power values corresponds respectively to each of the plurality of input signals;

determining, by the processor, whether at least one of the plurality of input signals contains speech;

calculating, by the processor, a plurality of current smoothed power values respectively for the plurality of input signals at a current time interval based on the plurality of previous smoothed power values associated to the previous time interval; and

mixing the plurality of input signals based on the plurality of current smoothed power values, wherein the calculating the plurality of current smoothed power values comprises calculating a current smoothed power value for each input signal of the plurality of input signals based on whether at least one of the plurality of input signals contains speech; and

wherein the mixing the plurality of input signals comprises calculating a plurality of mixing gains for the plurality of input signals and a mixing gain among the plurality of mixing gains for an input signal among the plurality of input signals is determined based on a current smoothed power value among the plurality of current smoothed power values corresponding to the input signal.

15 . The apparatus according to claim 14 , wherein the calculating the plurality of current smoothed power values comprises:

when it is determined that at least one of the plurality of input signals contains speech and echo dominates over speech, calculating the current smoothed power value for each input signal based on a current power value among a plurality of current power values and a previous smoothed power value among the plurality of previous smoothed power values, wherein each of the plurality of previous smoothed power values corresponds respectively to each of the plurality of input signals, the current power value and the previous smoothed power value correspond to each input signal; and

when it is determined that none of the plurality of input signals contains speech or it is determined that at least one of the plurality of input signals contains speech but the echo dominates over speech, calculating for each input signal current smoothed power value based on a determined value and the previous smoothed power value corresponding to each input signal.

16 . The apparatus according to claim 15 , wherein the calculating for each input signal current smoothed power value based on the determined value and the previous smoothed power value corresponding to each input signal comprises that the current smoothed power value is determined by smoothing between the determined value and the previous smoothed power value corresponding to each input signal and wherein the determined value is an average of the plurality of previous smoothed power values.

17 . The apparatus according to claim 14 , wherein one of the plurality of mixing gains are further determined based on a power ratio between the current smoothed power value of a input signal and an average of the plurality of current smoothed power values of the plurality of input signals.

18 . The apparatus according to claim 17 , wherein one of the plurality of mixing gains is further determined by a first updated power ratio, and the first updated power ratio equals to a square root of the power ratio divided by a number of the plurality of input signals.

19 . The apparatus according to claim 18 , wherein one of the plurality of mixing gains are further determined by a second updated power ratio, wherein if the first updated power ratio is greater than a high threshold the second updated power ratio is determined to be the high threshold, if the first updated power ratio is not greater than a low threshold, the second updated power ratio is determined to be the low threshold, and if the second updated power ratio is determined to be not greater than the high threshold and greater than the low threshold, the second updated power ratio is determined to be the first updated power ratio, wherein the high threshold is greater than the low threshold.

20 . A non-transitory computer program product comprising instructions executable to perform a method for mixing the plurality of input signals associated respectively with a plurality of microphones wherein each of the plurality of input signals comprises sound events generated by a one or more sound sources, the method comprising:

receiving, by a processor, a plurality of previous smoothed power values associated to a previous time interval, wherein each of a plurality of current power values corresponds respectively to each of the plurality of input signals;

determining, by the processor, whether at least one of the plurality of input signals contains speech;

calculating, by the processor, a plurality of current smoothed power values respectively for the plurality of input signals at a current time interval based on the plurality of previous smoothed power values associated to the previous time interval; and

mixing the plurality of input signals based on the plurality of current smoothed power values,

wherein the calculating the plurality of current smoothed power values comprises calculating a current smoothed power value for each input signal of the plurality of input signals based on whether at least one of the plurality of input signals contains speech; and

wherein the mixing the plurality of input signals comprises calculating a plurality of mixing gains for the plurality of input signals and a mixing gain among the plurality of mixing gains for an input signal among the plurality of input signals is determined based on a current smoothed power value among the plurality of current smoothed power values corresponding to the input signal.