IP Library Granted Patent US 12,444,423
Granted Patent B2
US 12,444,423 · App. 17/974,674 · Granted Oct 14, 2025

System and method for single channel distant speech processing

Inventor: Dushyant Sharma (Mountain House, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L17/20G10L17/02G10L21/028H04S3/008H04S2400/01H04S2400/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,423
App. No.
17/974,674
Granted
Oct 14, 2025
Kind
B2
Abstract

A method, computer program product, and computing system for receiving a signal from a single microphone. A plurality of modified signals may be generated from the single microphone signal, where the plurality of modified signals include at least one of: a speaker-specific signal, an acoustic parameter-specific signal, and a speech enhanced signal. Speech processing may be performed on the plurality of modified signals.

Claims (39)

1. A computer-implemented method, executed on a computing device, comprising:

receiving a signal from a single microphone;

generating a plurality of modified signals from the signal received from the single microphone, wherein the plurality of modified signals include at least a first acoustic parameter-specific signal and a second acoustic parameter-specific signal, the first acoustic parameter-specific signal having a higher C50 value than the second acoustic parameter-specific signal;

weighting each of the plurality of modified signals, wherein the first acoustic parameter-specific signal is assigned a higher weight than the second acoustic parameter-specific signal based on the first acoustic parameter-specific signal having a higher C50 value than the second acoustic parameter-specific signal;

generating a combined modified signal by combining the plurality of modified signals based upon, at least in part, the weighting of each of the plurality of modified signals; and

performing speech recognition on the combined modified signal.

2. The computer-implemented method of claim 1 , wherein the plurality of modified signals includes a plurality of speaker-specific signals associated with a plurality of speakers.

3. The computer-implemented method of claim 2 , wherein generating the plurality of speaker-specific signals associated with the plurality of speakers includes constructing a mask for each speaker in the plurality of speakers to attenuate speech associated with other speakers in the plurality of speakers, thus defining a plurality of speaker masks.

4. The computer-implemented method of claim 3 , wherein constructing the mask for each speaker in the plurality of speakers includes constructing a quantized mask with an attenuation factor based upon, at least in part, a speaker identification confidence value.

5. The computer-implemented method of claim 1 , wherein generating the plurality of modified signals includes performing a plurality of speech enhancements on the signal received from the single microphone.

6. The computer-implemented method of claim 1 , wherein combining the plurality of modified signals includes:

extracting feature information from each modified signal in the plurality of modified signals; and

combining the extracted feature information.

7. The computer-implemented method of claim 1 , wherein performing speech recognition on the combined modified signal includes:

performing speech recognition on the combined modified signal using a multi-channel speech processing system.

8. A computing system comprising:

a memory storing programming instructions; and

a processor configured to execute the programming instructions stored by the memory, wherein the programming instructions, upon execution by the processor, cause the computing system to:

receive a signal from a single microphone;

generate a plurality of modified signals from the signal received from the single microphone, wherein the plurality of modifies signals include at least a first acoustic parameter-specific signal and a second acoustic parameter-specific signal, the first acoustic parameter-specific signal having a higher C50 value than the second acoustic parameter-specific signal;

weight each of the plurality of modified signals, wherein the first acoustic parameter-specific signal is assigned a higher weight than the second acoustic parameter-specific signal based on the first acoustic parameter-specific signal having a higher C50 value than the second acoustic parameter-specific signal;

generate a combined modified signal by combining the plurality of modified signals based upon, at least in part, the weighting of each of the plurality of modified signals; and

perform speech recognition on the combined modified signal.

9. The computing system of claim 8 , wherein the plurality of modified signals includes a plurality of speaker-specific signals associated with a plurality of speakers.

10. The computing system of claim 9 , wherein generating the plurality of speaker-specific signals associated with the plurality of speakers includes constructing a mask for each speaker in the plurality of speakers to attenuate speech associated with other speakers in the plurality of speakers, thus defining a plurality of speaker masks.

11. The computing system of claim 10 , wherein constructing the mask for each speaker in the plurality of speakers includes constructing a quantized mask with an attenuation factor based upon, at least in part, a speaker identification confidence value.

12. A computer program product comprising a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

receiving a signal from a single microphone;

generating a plurality of modified signals from the signal received from the single microphone, wherein the plurality of modifies signals include at least a first acoustic parameter-specific signal and a second acoustic parameter-specific signal, the first acoustic parameter-specific signal having a higher C50 value than the second acoustic parameter-specific signal;

weighting each of the plurality of modified signals, wherein the first acoustic parameter-specific signal is assigned a higher weight than the second acoustic parameter-specific signal based on the first acoustic parameter-specific signal having a higher C50 value than the second acoustic parameter-specific signal; and

generating a combined modifies signal by combining the plurality of modified signals based upon, at least in part, the weighting of each of the plurality of modified signals and

performing speech recognition on the combined modifies signal.

13. The computer program product of claim 12 , wherein the plurality of modified signals includes a plurality of speaker-specific signals associated with a plurality of speakers, and wherein generating the plurality of speaker-specific signals associated with the plurality of speakers includes constructing a mask for each speaker to attenuate speech associated with other speakers, thus defining a plurality of speaker masks.

14. The method of claim 1 , wherein a C50 value of the first acoustic parameter-specific signal is a ratio between sound energy received within the first fifty (50) milliseconds (ms) of a room impulse response (RIR) corresponding to the first acoustic parameter-specific signal and sound energy received after the first 50 ms of the RIR corresponding to the first acoustic parameter-specific signal.

15. The method of claim 1 , wherein the plurality of modified signals include at least two acoustic parameter-specific signals associated with different segmental signal to noise ratios (SNRs), and wherein the at least two acoustic parameter-specific signals are assigned different weights based on the different segmental SNRs associated with the at least two acoustic parameter-specific signals.

16. The method of claim 1 , wherein the plurality of modified signals include at least two acoustic parameter-specific signals having different reverberation parameters, and wherein the at least two acoustic parameter-specific signals are assigned different weights based on the different reverberation parameters associated with the at least two acoustic parameter-specific signals.

17. The method of claim 16 , wherein the reverberation parameters associated with the at least two acoustic parameter-specific signals include reverberation times of the at least two acoustic parameter-specific signals.

18. The method of claim 17 , wherein the reverberation times of the at least two acoustic parameter-specific signals correspond to a length of time required for a sound intensity of the at least two acoustic parameter-specific signals to fall below a threshold.

19. The method of claim 1 , wherein speech recognition is performed on the combined modified signal using a multi-channel speech processing system.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2025
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070762/0862 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2022
From: SHARMA, DUSHYANT
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 061558/0841 →
Continuity (1)
Related Publication 20240144936A1 · May 2, 2024
References Cited (28)
US 10885902B1 · Papania-Davis · 2021 [cited by applicant]
US 12087319B1 · Looney · 2024 [cited by examiner]
US 20140278397A1 · Chen · 2014 [cited by examiner]
US 20160098987A1 · Stolcke · 2016 [cited by applicant]
US 20170278527A1 · Sharma · 2017 [cited by examiner]
US 20210035551A1 · Stanton · 2021 [cited by applicant]
US 20210335337A1 · Gkoulalas-Divanis · 2021 [cited by applicant]
US 20210350815A1 · Sharma · 2021 [cited by examiner]
US 20220399026A1 · Gong et al. · 2022 [cited by applicant]
US 20230267944A1 · Dushyant · 2023 [cited by examiner]
US 20230410789A1 · Sharma · 2023 [cited by examiner]
US 20230410814A1 · Yin · 2023 [cited by applicant]
US 20240296826A1 · Sharma · 2024 [cited by applicant]
CN 107071636A · 2017 [cited by examiner]
WO WO2021119102A1 · 2021 [cited by examiner]
WO WO2021183657A1 · 2021 [cited by examiner]
WO WO2023177095A1 · 2023 [cited by examiner]
Gong, et al., “Self-Attention Channel Combinator Frontend for End-to-End Multichannel Far-field Speech Recognition”, In Repository of arXiv:2109.04783v1, Sep. 10, 2021, 5 Pages. [cited by applicant]
Pascual, et al., “Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks”, In Repository of arXiv:1904.03416v1, Apr. 6, 2019, 5 Pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US23/034600, Mar. 14, 2024, 21 pages. [cited by applicant]
U.S. Appl. No. 18/318,185, filed May 16, 2023. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/017370, (MS#412809-PCT01) May 6, 2024, 12 pages. [cited by applicant]
Shahin-Shamsabadi, et al., “Differentially private speaker anonymization.” arXiv:2202.11823, Feb. 23, 2022, pp. 1-17. [cited by applicant]
Grais et al., “Deep neural networks for single channel source separation,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3734-3738, May 4, 2014. [cited by applicant]
Grais et al., “Single-Channel Audio Source Separation Using Deep Neural Network Ensembles,” AES Convention 140, 7 Pages, May 26, 2016. [cited by applicant]
Invitation to Pay Additional Fees received for PCT Application No. PCT/US2023/034600, mailed on Jan. 22, 2024, 14 pages. [cited by applicant]
International Preliminary Report on Patentability (Chapter 1) received for PCT Application No. PCT/US2023/034600, May 8, 2025, 14 pages. [cited by applicant]
Non-Final Office Action mailed on Jul. 16, 2025, in U.S. Appl. No. 18/318,185, 45 pages. [cited by applicant]