IP Library Granted Patent US 12,542,130
Granted Patent B2
US 12,542,130 · App. 18/298,488 · Granted Feb 3, 2026

System and method for dynamically adjusting speech processing neural networks

Inventors: Felix Weninger (Wellesley, MA); Marco Gaudesi (Turin, IT); Puming Zhan (Acton, MA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/16G10L15/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,542,130
App. No.
18/298,488
Granted
Feb 3, 2026
Kind
B2
Abstract

A method, computer program product, and computing system for dividing a speech signal into a plurality of chunks. A context window is defined for processing a chunk of the plurality of chunks using a neural network of a speech processing system. A processing load associated with the speech processing system is determined. The context window is dynamically adjusted based upon, at least in part, the processing load associated with the speech processing system.

Claims (36)

1 . A computer-implemented method, executed on a computing device, comprising:

dividing a speech signal into a plurality of chunks;

defining a context window for processing a chunk of the plurality of chunks using a neural network of a speech processing system;

at a first time instance, determining a first processing load associated with the speech processing system;

performing a first dynamic adjustment to the context window based on the first processing load associated with the speech processing system, wherein the first dynamic adjustment includes an increase to the context window;

at a second time instance, determining a second processing load associated with the speech processing system, wherein the second processing load is lower than the first processing load;

performing a second dynamic adjustment to the context window based on the second processing load associated with the speech processing system, wherein the second dynamic adjustment includes, relative to the first dynamic adjustment, a decrease to the context window, and wherein decreasing the context window includes adjusting a period of future context to a first selected number of chunks that is greater than or equal to one and adjusting a period of past context to a second selected number of chunks; and

processing the speech signal using the neural network of the speech processing system based on the adjusted context window, which is adjusted by the second dynamic adjustment.

2 . The computer-implemented method of claim 1 , wherein defining the context window includes defining a chunk size in terms of frames of the speech signal for the context window.

3 . The computer-implemented method of claim 1 , wherein the first or second dynamic adjustment to the context window includes selecting from a predefined combination of the chunk size, the period of future context, and the period of past context from a plurality of predefined combinations based on the first or second processing load associated with the speech processing system.

4 . The computer-implemented method of claim 1 , further comprising:

processing the speech signal using the neural network of the speech processing system based on the adjusted context window, which is adjusted by the first dynamic adjustment.

5 . A computing system comprising:

one or more processors; and

one or more hardware storage devices that store instructions that are executable by the one or more processors to cause the computing system to:

divide a speech signal into a plurality of chunks;

define a context window for processing a chunk of the plurality of chunks using a neural network of an automated speech recognition (ASR) system;

at a first time instance, determine a first processing load associated with the speech processing system;

perform a first dynamic adjustment to the context window based on the first processing load associated with the speech processing system, wherein the first dynamic adjustment includes an increase or a decrease to the context window;

at a second time instance, determine a second processing load associated with the speech processing system, wherein the second processing load is lower than the first processing load;

perform a second dynamic adjustment to the context window based on the second processing load associated with the speech processing system, wherein the second dynamic adjustment includes, relative to the first dynamic adjustment, a decrease to the context window, and wherein decreasing the context window includes adjusting a period of future context to a first selected number of chunks that is greater than or equal to one and adjusting a period of past context to a second selected number of chunks; and

process the speech signal using the neural network of the ASR system based on the adjusted context window, which is adjusted by the second dynamic adjustment.

6 . The computing system of claim 5 , wherein defining the context window includes defining a chunk size in terms of frames of the speech signal for the context window.

7 . The computing system of claim 5 , wherein the first or second dynamic adjustment to the context window includes selecting from a predefined combination of the chunk size, the period of future context, and the period of past context from a plurality of predefined combinations based on the first or second processing load associated with the speech processing system.

8 . One or more hardware storage devices that store instructions that are executable by one or more processors to cause the one or more processors to:

divide a speech signal into a plurality of chunks;

define a context window for processing a chunk of the plurality of chunks using a neural network of a speech processing system;

at a first time instance, determine a first processing load associated with the speech processing system;

perform a first dynamic adjustment to the context window based on the first processing load associated with the speech processing system, wherein the first dynamic adjustment includes one of an increase or a decrease to the context window;

process the speech signal using the neural network of the speech processing system based on the adjusted context window, which is adjusted by the first dynamic adjustment;

at a second time instance, determine a second processing load associated with the speech processing system, wherein the second processing load is lower than the first processing load;

perform a second dynamic adjustment to the context window based on the second processing load associated with the speech processing system, wherein the second dynamic adjustment includes, relative to the first dynamic adjustment, a decrease to the context window, and wherein decreasing the context window includes adjusting a period of future context to a first selected number of chunks that is greater than or equal to one and adjusting a period of past context to a second selected number of chunks;

process the speech signal using the neural network of the speech processing system based on the adjusted context window, which is adjusted by the second dynamic adjustment; and

process the speech signal using the neural network of the speech processing system based on the adjusted context window, which is adjusted by the second dynamic adjustment.

9 . The one or more hardware storage devices of claim 8 , wherein defining the context window includes defining a chunk size in terms of frames of the speech signal for the context window.

10 . The computer program product of claim 8 , wherein the first or second dynamic adjustment to the context window includes selecting from a predefined combination of the chunk size, the period of future context, and the period of past context from a plurality of predefined combinations based on the first or second processing load associated with the speech processing system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2023
From: WENINGER, FELIX; GAUDESI, MARCO; ZHAN, PUMING
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 063285/0523 →