IP Library › Granted Patent US 11,532,312
Granted Patent B2
US 11,532,312 · App. 17/123,087 · Granted Dec 20, 2022

User-perceived latency while maintaining accuracy

Inventors: Hosam Adel Khalil (Issaquah, WA); Emilian Stoimenov (Bellevue, WA); Christopher Hakan Basoglu (Everett, WA); Kshitiz Kumar (Redmond, WA); Jian Wu (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/30G10L15/16G10L19/167G10L25/51G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,532,312
App. No.
17/123,087
Granted
Dec 20, 2022
Kind
B2
Abstract

Disclosed speech recognition techniques improve user-perceived latency while maintaining accuracy by: receiving an audio stream, in parallel, by a primary (e.g., accurate) speech recognition engine (SRE) and a secondary (e.g., fast) SRE; generating, with the primary SRE, a primary result; generating, with the secondary SRE, a secondary result; appending the secondary result to a word list; and merging the primary result into the secondary result in the word list. Combining output from the primary and secondary SREs into a single decoder as described herein improves user-perceived latency while maintaining or improving accuracy, among other advantages.

Claims (71)

1. A method of speech recognition, the method comprising:

receiving an audio stream, in parallel, by a primary speech recognition engine (SRE) and a secondary SRE;

generating a concatenated output including an enclosed output of the secondary SRE and an early-stage encoded output of the primary SRE;

processing the concatenated output by the primary SRE;

processing encoded output of the secondary SRE by the secondary SRE;

generating, with the processed output of the primary SRE, a primary result;

generating, with the processed output of the secondary SRE, a secondary result;

appending the secondary result to a word list; and

merging the primary result into the secondary result in the word list, wherein the merging comprises:

synchronizing the primary result with the secondary result;

determining, within the primary result or the secondary result, that at least some words belong to a class model;

based on at least the synchronizing, determining a word in the primary result that corresponds with a corresponding word in the secondary result; and

based on at least determining that the corresponding word in the secondary result does not belong to the class model, replacing the corresponding word in the secondary result with the word in the primary result.

2. The method of claim 1 , further comprising:

displaying the word list.

3. The method of claim 1 , further comprising:

based at least on determining that the corresponding word in the secondary result belongs to a same grammar model as the word in the primary result, replacing the corresponding word in the secondary result with the word in the primary result.

4. The method of claim 1 , wherein generating the concatenated output includes determining that the encoded output of the secondary SRE is ahead of the early-stage encoded output in time, and performing the concatenation based on the determining.

5. The method of claim 1 , wherein synchronizing the primary result with the secondary result comprises comparing a sync marker of the primary result with a sync marker of the secondary result.

6. The method of claim 1 , further comprising:

determining whether the word in the primary result differs from the corresponding word in the secondary result, wherein replacing the corresponding word in the secondary result with the word in the primary result comprises:

based on at least determining that the word in the primary result differs from the corresponding word in the secondary result and determining that the corresponding word in the secondary result does not belong to a class model, replacing the corresponding word in the secondary result with the word in the primary result.

7. The method of claim 1 , wherein the class model is selected from a list comprising:

a contact name, a date, a time, an application name, a filename, a location, a commonly-recognized name.

8. A system for speech recognition, the system comprising:

a processor; and

a computer-readable medium storing instructions that are operative upon execution by the processor to:

receive an audio stream, in parallel, by a primary speech recognition engine (SRE) and a secondary SRE;

generate a concatenated output including an encoded output of the secondary SRE and an early-stage encoded output of the primary SRE;

process the concatenated output by the primary SRE;

process the encoded output of the secondary SRE by the secondary SRE;

generate, with the processed output of the primary SRE, a primary result;

generate, with the processed output of the secondary SRE, a secondary result;

append the secondary result to a word list; and

merge the primary result into the secondary result in the word list, wherein the merging comprises:

synchronizing the primary result with the secondary result;

determining, within the primary result or the secondary result, that at least some words belong to a class model;

based on at least the synchronizing, determining a word in the primary result that corresponds with a corresponding word in the secondary result; and

based on at least determining that the corresponding word in the secondary result does not belong to the class model, replacing the corresponding word in the secondary result with the word in the primary result.

9. The system of claim 8 , wherein the secondary SRE comprises a recurrent neural network transducer (RNN-T).

10. The system of claim 8 , wherein the instructions are further operative to:

based at least on determining that the corresponding word in the secondary result belongs to a same grammar model as the word in the primary result, replacing the corresponding word in the secondary result with the word in the primary result.

11. The system of claim 8 , wherein generating the concatenated output includes determining that the encoded output of the secondary SRE is ahead of the early-stage encoded output in time and performing the concatenating based on the determination.

12. The system of claim 8 , wherein synchronizing the primary result with the secondary result comprises comparing a sync marker of the primary result with a sync marker of the secondary result.

13. The system of claim 8 , wherein the instructions are further operative to:

determine whether the word in the primary result differs from the corresponding word in the secondary result, wherein replacing the corresponding word in the secondary result with the word in the primary result comprises:

based on at least determining that the word in the primary result differs from the corresponding word in the secondary result and determining that the corresponding word in the secondary result does not belong to a class model, replacing the corresponding word in the secondary result with the word in the primary result.

14. The system of claim 8 , wherein the class model is selected from a list comprising:

a contact name, a date, a time, an application name, a filename, a location, a commonly-recognized name.

15. A computing device having computer-executable instructions stored thereon, which, on execution by a computer, cause the computer to perform operations comprising:

receiving an audio stream by a secondary speech recognition engine (SRE) on the computing device;

transmitting the audio stream to a remote node for processing by a primary SRE;

transmitting an encoded output of the secondary SRE to the remote node, wherein the remote node generates a concatenated output including the encoded output of the secondary SRE and an early-stage encoded output of the primary SRE and further processes the concatenated output;

based on the further processing of the concatenated output, receiving, from the remote node, the primary result;

further processing the encoded output of the secondary SRE by the secondary SRE;

generating, with the further processed output of the secondary SRE, a secondary result;

appending the secondary result to a word list; and

merging the primary result into the secondary result in the word list, wherein the merging comprises:

synchronizing the primary result with the secondary result;

determining, within the primary result or the secondary result, that at least some words belong to a class model;

based on at least the synchronizing, determining a word in the primary result that corresponds with a corresponding word in the secondary result; and

based on at least determining that the corresponding word in the secondary result does not belong to the class model, replacing the corresponding word in the secondary result with the word in the primary result.

16. The computing device of claim 15 , wherein the operations further comprise:

displaying the word list.

17. The computing device of claim 15 , wherein generating the concatenated output includes determining that the encoded output of the secondary SRE is ahead of the early-stage encoded output in time and performing the concatenating based on the determination.

18. The computing device of claim 15 , wherein synchronizing the primary result with the secondary result comprises comparing a sync marker of the primary result with a sync marker of the secondary result.

19. The computing device of claim 15 , wherein the operations further comprise:

determining whether the word in the primary result differs from the corresponding word in the secondary result, wherein replacing the corresponding word in the secondary result with the word in the primary result comprises:

based on at least determining that the word in the primary result differs from the corresponding word in the secondary result and determining that the corresponding word in the secondary result does not belong to a class model, replacing the corresponding word in the secondary result with the word in the primary result.

20. The computing device of claim 15 , wherein the class model is selected from a list comprising:

a contact name, a date, a time, an application name, a filename, a location, a commonly-recognized name.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2021
From: KHALIL, HOSAM ADEL; STOIMENOV, EMILIAN; BASOGLU, CHRISTOPHER HAKAN; KUMAR, KSHITIZ; WU, JIAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 054805/0535 →
Continuity (1)
Related Publication 20220189467A1 · Jun 16, 2022
Cited By (2)
US 12,390,724 US 12,488,581