IP Library Patent Application 18826655
Patent Application
App. No. 18/826,655

Transducer-Based Streaming Deliberation for Cascaded Encoders

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/826,655
Abstract

A method includes receiving a sequence of acoustic frames and generating, by a first encoder, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The method also includes generating, by a first pass transducer decoder, a first pass speech recognition hypothesis for a corresponding first higher order feature representation and generating, by a text encoder, a text encoding for a corresponding first pass speech recognition hypothesis. The method also includes generating, by a second encoder, a second higher order feature representation for a corresponding first higher order feature representation. The method also includes generating, by a second pass transducer decoder, a second pass speech recognition hypothesis using a corresponding second higher order feature representation and a corresponding text encoding.

Claims (43)

1 . A transducer-based deliberation model comprising:

an encoder configured to:

receive, as input, a sequence of acoustic frames; and

generate, at each of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames;

a first pass transducer decoder configured to:

receive, as input, the higher order feature representation generated by the encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a first pass speech recognition hypothesis for a corresponding higher order feature representation;

a text encoder configured to:

receive, as input, the first pass speech recognition hypothesis generated at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a text encoding for a corresponding first pass speech recognition hypothesis; and

a second pass transducer decoder configured to:

receive, as input, the text encoding generated by the text encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a second pass speech recognition hypothesis.

2 . The transducer-based deliberation model of claim 1 , further comprising a prediction network shared by the first pass transducer decoder and the second pass transducer decoder, the prediction network configured to:

receive, as input, a sequence of non-blank symbols output by a final softmax layer; and

generate, at each of the plurality of output steps, a dense representation.

3 . The transducer-based deliberation model of claim 2 , wherein the second pass transducer decoder further comprises a joint network configured to:

receive, as input, the dense representation generated by the prediction network at each of the plurality of output steps, the text encoding generated by the text encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, the second pass speech recognition hypothesis.

4 . The transducer-based deliberation model of claim 1 , wherein the encoder comprises a stack of multi-headed attention layers.

5 . The transducer-based deliberation model of claim 5 , wherein the stack of multi-headed attention layers comprises a stack of Conformer layers.

6 . The transducer-based deliberation model of claim 5 , wherein the stack of multi-head attention layers comprises a stack of Transformer layers.

7 . The transducer-based deliberation model of claim 1 , wherein the encoder comprises a causal encoder.

8 . The transducer-based deliberation model of claim 1 , wherein the second pass transducer decoder trains without using any text-only data.

9 . The transducer-based deliberation model of claim 1 , wherein receiving the text encoding generated by the text encoder at each of the plurality of output steps comprises receiving a partial sequence of the text encoding in a streaming fashion.

10 . The transducer-based deliberation model of claim 1 , wherein the first and second pass speech recognition hypotheses each correspond to a partial speech recognition result.

11 . A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

receiving a sequence of acoustic frames;

generating, by an encoder, at each of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames;

generating, by a first pass transducer decoder, at each of the plurality of output steps, a first pass speech recognition hypothesis for a corresponding higher order feature representation;

generating, by a text encoder, at each of the plurality of output steps, a text encoding for a corresponding first pass speech recognition hypothesis; and

generating, by a second pass transducer decoder, at each of the plurality of output steps, a second pass speech recognition hypothesis using a corresponding text encoding.

12 . The computer-implemented method of claim 11 , wherein the operations further comprise:

generating, by a prediction network based on a sequence of non-blank symbols output by a final softmax layer, at each of the plurality of output steps, a dense representation,

wherein the first pass transducer decoder and the second pass transducer decoder share the prediction network.

13 . The computer-implemented method of claim 12 , wherein the operations further comprise, at each of the plurality of output steps, generating, by a joint network, the second pass speech recognition hypothesis based on the dense representation generated by the prediction network at each of the plurality of output steps and the text encoding generated by the text encoder at each of the plurality of output steps.

14 . The computer-implemented method of claim 11 , wherein the encoder comprises a stack of multi-headed attention layers.

15 . The computer-implemented method of claim 14 , wherein the stack of multi-headed attention layers comprises a stack of Conformer layers.

16 . The computer-implemented method of claim 14 , wherein the stack of multi-headed attention layers comprises a stack of Transformer layers.

17 . The computer-implemented method of claim 11 , wherein the encoder comprises a causal encoder.

18 . The computer-implemented method of claim 11 , wherein the second pass transducer decoder trains without using any text-only data.

19 . The computer-implemented method of claim 11 , wherein receiving the text encoding comprises receiving a partial sequence of the text encoding in a streaming fashion.

20 . The computer-implemented method of claim 11 , wherein the first and second pass speech recognition hypotheses each correspond to a partial speech recognition result.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2024
From: HU, KE; SAINATH, TARA N.; NARAYANAN, ARUN; PANG, RUOMING; STROHMAN, TREVOR
To: GOOGLE LLC
Reel/Frame 068509/0564 →