IP Library Granted Patent US 12,308,035
Granted Patent B2
US 12,308,035 · App. 17/539,451 · Granted May 20, 2025

System and method for self-attention-based combining of multichannel signals for speech processing

Inventors: Rong Gong (Vienna, AT); Carl Benjamin Quillen (Brookline, MA); Dushyant Sharma (Mountain House, CA); Ljubomir Milanovic (Vienna, AT)
Assignee: Microsoft Technology Licensing, LLC
G10L19/008G10L25/30H04R3/005H04R2203/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,308,035
App. No.
17/539,451
Granted
May 20, 2025
Kind
B2
Abstract

A method, computer program product, and computing system for receiving a plurality of signals from a plurality of microphones, thus defining a plurality of channels. A weighted multichannel representation of the plurality of channels may be generated. A plurality of weights for each channel of the plurality of channels may be generated based upon, at least in part, the weighted multichannel representation of the plurality of channels. A single channel representation of the plurality of channels may be generated based upon, at least in part, the weighted multichannel representation of the plurality of channels and the plurality of weights generated for each channel of the plurality of channels.

Claims (23)

1. A computer-implemented method, executed on a computing device, comprising:

receiving a plurality of signals from a plurality of microphones, thus defining a plurality of channels;

generating a weighted multichannel representation of the plurality of channels via one or more fixed beamformers, wherein generating the weighted multichannel representation includes defining a plurality of attention weights that correspond to a direction of one or more sound sources using a first self-attention machine learning model and multiplying the plurality of attention weights by the plurality of channels;

generating a plurality of weights for each channel of the plurality of channels based upon, at least in part, the weighted multichannel representation of the plurality of channels via a second self-attention machine learning model includes:

defining a frequency dimension contracted representation of the weighted multichannel representation, and

generating the plurality of weights as a weighted sum of the frequency dimension contracted representation of the weighted multichannel representation by performing a softmax of a product of the plurality of attention weights and the weighted multichannel representation; and

generating a single channel representation of the plurality of channels for processing by a backend speech processing system based upon, at least in part, the weighted multichannel representation of the plurality of channels and the plurality of weights generated for each channel of the plurality of channels.

2. The computer-implemented method of claim 1 , wherein generating the weighted multichannel representation of the plurality of channels includes defining each channel of the weighted multichannel representation of the plurality of channels as a linear combination of the plurality of channels.

3. The computer-implemented method of claim 1 , further comprising: utilizing the plurality of attention weights corresponding to the direction of the one or more sound sources for speech processing.

4. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

receiving a plurality of signals from a plurality of microphones, thus defining a plurality of channels;

generating a weighted multichannel representation of the plurality of channels via one or more fixed beamformers, wherein generating the weighted multichannel representation includes defining a plurality of attention weights that correspond to a direction of one or more sound sources using a first self-attention machine learning model and multiplying the plurality of attention weights by the plurality of channels;

generating a plurality of weights for each channel of the plurality of channels based upon, at least in part, the weighted multichannel representation of the plurality of channels via a second self-attention machine learning model includes:

defining a frequency dimension contracted representation of the weighted multichannel representation, and

generating the plurality of weights as a weighted sum of the frequency dimension contracted representation of the weighted multichannel representation by performing a softmax of a product of the plurality of attention weights and the weighted multichannel representation; and

generating a single channel representation of the plurality of channels for processing by a backend speech processing system based upon, at least in part, the weighted multichannel representation of the plurality of channels and the plurality of weights generated for each channel of the plurality of channels.

5. The computer program product of claim 4 , wherein generating the weighted multichannel representation of the plurality of channels includes defining each channel of the weighted multichannel representation of the plurality of channels as a linear combination of the plurality of channels.

6. The computer program product of claim 4 , wherein the operations further comprise:

utilizing the plurality of attention weights corresponding to the direction of the one or more sound sources for speech processing.

7. A computing system comprising:

a memory; and

a processor configured to receive a plurality of signals from a plurality of microphones, thus defining a plurality of channels, wherein the processor is further configured to generate a weighted multichannel representation of the plurality of channels via one or more fixed beamformers, wherein generating the weighted multichannel representation includes defining a plurality of attention weights that correspond to a direction of one or more sound sources using a first self-attention machine learning model and multiplying the plurality of attention weights by the plurality of channels, wherein the processor is further configured to generate a plurality of weights for each channel of the plurality of channels based upon, at least in part, the weighted multichannel representation of the plurality of channels via a second self-attention machine learning model includes: defining a frequency dimension contracted representation of the weighted multichannel representation, and generating the plurality of weights as a weighted sum of the frequency dimension contracted representation of the weighted multichannel representation by performing a softmax of a product of the plurality of attention weights and the weighted multichannel representation, and wherein the processor is further configured to generate a single channel representation of the plurality of channels for processing by a backend speech processing system based upon, at least in part, the weighted multichannel representation of the plurality of channels and the plurality of weights generated for each channel of the plurality of channels.

8. The computing system of claim 7 , wherein generating the weighted multichannel representation of the plurality of channels includes defining each channel of the weighted multichannel representation of the plurality of channels as a linear combination of the plurality of channels.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2025
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070746/0933 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2021
From: GONG, RONG; QUILLEN, CARL BENJAMIN; MILANOVIC, LJUBOMIR; SHARMA, DUSHYANT
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 058255/0480 →