IP Library Granted Patent US 12,223,953
Granted Patent B2
US 12,223,953 · App. 17/737,587 · Granted Feb 11, 2025

End-to-end automatic speech recognition system for both conversational and command-and-control speech

Inventors: Alejandro Coucheiro Limeres (Aachen, DE); Junho Park (Bedford, MA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/197G10L15/16G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,953
App. No.
17/737,587
Granted
Feb 11, 2025
Kind
B2
Abstract

A contextual end-to-end automatic speech recognition (ASR) system includes: an audio encoder configured to process input audio signal to produce as output encoded audio signal; a bias encoder configured to produce as output at least one bias entry corresponding to a word to bias for recognition by the ASR system; a transcription token probability prediction network configured to produce as output a probability of a selected transcription token, based at least in part on the output of the bias encoder and the output of the audio encoder; a first attention mechanism configured to receive the at least one bias entry and determine whether the at least one bias entry is suitable to be transcribed at a specific moment of an ongoing transcription; and a second attention mechanism configured to produce prefix penalties for restricting the first attention mechanism to only entries fitting a current transcription context.

Claims (44)

1. A contextual end-to-end automatic speech recognition (ASR) system, comprising:

an audio encoder processing input audio signal to produce as output encoded audio signal;

a bias encoder producing as output at least one bias entry corresponding to a word to bias for recognition by the ASR system; and

a transcription token probability prediction network producing as output a probability of a selected transcription token, based at least in part on the output of the bias encoder and the output of the audio encoder.

2. The system according to claim 1 , wherein the transcription token probability prediction network comprises a joiner module and a Softmax function layer.

3. The system according to claim 1 , wherein the transcription token probability prediction network comprises a decoder module and a Softmax function layer.

4. The system according to claim 2 , further comprising:

a text predictor producing as output a prediction of a next transcription token, based on a previous transcription token, wherein the output of the text predictor is supplied to the joiner module.

5. The system according to claim 4 , further comprising:

a first attention mechanism receiving the at least one bias entry and determine whether the at least one bias entry is suitable to be transcribed at a specific moment of an ongoing transcription.

6. The system according to claim 5 , further comprising:

a second attention mechanism producing prefix penalties for restricting the first attention mechanism to only entries fitting a current transcription context.

7. The system according to claim 6 , further comprising:

a label encoder encoding a current state of a transcription;

wherein the second attention mechanism produces the prefix penalties at least in part based on the encoded current state of the transcription.

8. The system according to claim 3 , further comprising:

a first attention mechanism receiving the at least one bias entry and determine whether the at least one bias entry is suitable to be transcribed at a specific moment of an ongoing transcription.

9. The system according to claim 8 , further comprising:

a second attention mechanism producing prefix penalties for restricting the first attention mechanism to only entries fitting a current transcription context.

10. The system according to claim 9 , further comprising:

a label encoder encoding a current state of a transcription configured to encode the current state of the transcription;

wherein the second attention mechanism produces the prefix penalties at least in part based on the encoded current state of the transcription.

11. A method of operating a contextual end-to-end automatic speech recognition (ASR) system, comprising:

processing, by an audio encoder, an input audio signal to produce as output encoded audio signal;

producing, by a bias encoder, as output at least one bias entry corresponding to a word to bias for recognition by the ASR system; and

producing, by a transcription token probability prediction network, as output a probability of a selected transcription token, based at least in part on the output of the bias encoder and the output of the audio encoder.

12. The method according to claim 11 , wherein the transcription token probability prediction network comprises a joiner module and a Softmax function layer.

13. The method according to claim 11 , wherein the transcription token probability prediction network comprises a decoder module and a Softmax function layer.

14. The method according to claim 12 , further comprising:

producing, by a text predictor, as output a prediction of a next transcription token, based on a previous transcription token, wherein the output of the text predictor is supplied to the joiner module.

15. The method according to claim 14 , further comprising:

determining, by a first attention mechanism, whether the at least one bias entry output by the bias encoder is suitable to be transcribed at a specific moment of an ongoing transcription.

16. The method according to claim 15 , further comprising:

producing, by a second attention mechanism, prefix penalties for restricting the first attention mechanism to only entries fitting a current transcription context.

17. The method according to claim 16 , further comprising:

encoding, by a label encoder, a current state of a transcription;

wherein the prefix penalties are produced by the second attention mechanism at least in part based on the encoded current state of the transcription.

18. The method according to claim 13 , further comprising:

determining, by a first attention mechanism, whether the at least one bias entry output by the bias encoder is suitable to be transcribed at a specific moment of an ongoing transcription.

19. The method according to claim 18 , further comprising:

producing, by a second attention mechanism, prefix penalties for restricting the first attention mechanism to only entries fitting a current transcription context.

20. The method according to claim 19 , further comprising:

encoding, by a label encoder, a current state of a transcription;

wherein the prefix penalties are produced by the second attention mechanism at least in part based on the encoded current state of the transcription.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065191/0383 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2023
From: LIMERES, ALEJANDRO COUCHEIRO; PARK, JUNHO
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 063182/0500 →