IP Library › Granted Patent US 11,594,212
Granted Patent B2
US 11,594,212 · App. 17/155,010 · Granted Feb 28, 2023

Attention-based joint acoustic and text on-device end-to-end model

Inventors: Tara N. Sainath (Jersey City, NJ); Ruoming Pang (New York, NY); Ron Weiss (New York, NY); Yanzhang He (Mountain View, CA); Chung-Cheng Chiu (Sunnyvale, CA); Trevor Strohman (Mountain View, CA)
Assignee: Google LLC
G10L15/063G06N3/08G10L15/16G10L15/197G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,594,212
App. No.
17/155,010
Granted
Feb 28, 2023
Kind
B2
Abstract

A method includes receiving a training example for a listen-attend-spell (LAS) decoder of a two-pass streaming neural network model and determining whether the training example corresponds to a supervised audio-text pair or an unpaired text sequence. When the training example corresponds to an unpaired text sequence, the method also includes determining a cross entropy loss based on a log probability associated with a context vector of the training example. The method also includes updating the LAS decoder and the context vector based on the determined cross entropy loss.

Claims (79)

1. A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving a training example for a listen-attend-spell (LAS) decoder of a two-pass streaming neural network model, the two-pass streaming neural network model comprising:

a recurrent neural network-transducer (RNN-T) decoder configured to generate hypotheses during a first pass of the two-pass streaming neural network model; and

the LAS decoder configured to rescore the generated hypotheses during a second pass of the two-pass streaming neural network model;

determining whether the training example corresponds to a supervised audio-text pair or an unpaired text sequence; and

when the training example corresponds to an unpaired text sequence:

generating, based on the unpaired text sequence, a linguistic context vector;

generating a first log probability from the linguistic context vector;

synthesizing, using a text-to-speech (TTS) system, audio data from the unpaired text sequence;

generating, based on the synthesized audio data, an acoustic context vector;

generating a second log probability from the acoustic context vector;

determining an interpolation of the first log probability and the second log probability; and

updating the LAS decoder based on the determined interpolation.

2. The computer-implemented method of claim 1 , wherein the operations further comprise:

receiving a second training example for the LAS decoder of the two-pass streaming neural network model;

determining that the second training example corresponds to the supervised audio-text pair;

generating, based on the supervised audio-text pair, a second acoustic context vector; and

updating the LAS decoder and acoustic context vector parameters associated with the second acoustic context vector based on a log probability for the second acoustic context vector.

3. The computer-implemented method of claim 1 , wherein determining whether the training example corresponds to the supervised audio-text pair or the unpaired text sequence comprises identifying a domain identifier that indicates whether the training example corresponds to the supervised audio-text pair or the unpaired text sequence.

4. The computer-implemented method of claim 1 , wherein updating the LAS decoder reduces a word error rate (WER) of the two-pass streaming neural network model with respect to long tail entities.

5. The computer-implemented method of claim 1 , wherein the LAS decoder operates in a beam search mode based on a hypothesis generated by the RNN-T decoder during the first pass of the two-pass streaming neural network model.

6. The computer-implemented method of claim 1 , wherein the operations further comprise generating the acoustic context vector with an attention mechanism configured to summarize encoder features from an encoded acoustic frame.

7. A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving a training example for a listen-attend-spell (LAS) decoder of a two-pass streaming neural network model, the two-pass streaming neural network model comprising:

a recurrent neural network-transducer (RNN-T) decoder configured to generate hypotheses during a first pass of the two-pass streaming neural network model; and

the LAS decoder configured to rescore the generated hypotheses during a second pass of the two-pass streaming neural network model;

determining whether the training example corresponds to a supervised audio-text pair or unpaired training data; and

when the training example corresponds to the unpaired training data:

generating a missing portion of the unpaired training data to form a generated audio-text pair;

generating, based on the generated audio-text pair, an acoustic context vector and a linguistic context vector;

generating a first log probability from the linguistic context vector;

generating a second log probability from the acoustic context vector;

determining an interpolation of the first log probability and the second log probability; and

updating the LAS decoder based on the determined interpolation.

8. The computer-implemented method of claim 7 , wherein generating the missing portion of the unpaired training data comprises synthesizing, using a text-to-speech (TTS) system, audio data from the present portion of the unpaired training data.

9. The computer-implemented method of claim 7 , wherein determining whether the training example corresponds to the supervised audio-text pair or the unpaired training data comprises identifying a domain identifier that indicates whether the training example corresponds to the supervised audio-text pair or the unpaired training data.

10. The computer-implemented method of claim 7 , wherein updating the LAS decoder reduces a word error rate (WER) of the two-pass streaming neural network model with respect to long tail entities.

11. The computer-implemented method of claim 7 , wherein the operations further comprise generating the acoustic context vector using an attention mechanism configured to summarize encoder features from an encoded acoustic frame.

12. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a training example for a listen-attend-spell (LAS) decoder of a two-pass streaming neural network model, the two-pass streaming neural network model comprising:

a recurrent neural network-transducer (RNN-T) decoder configured to generate hypotheses during a first pass of the two-pass streaming neural network model; and

the LAS decoder configured to rescore the generated hypotheses during a second pass of the two-pass streaming neural network model;

determining whether the training example corresponds to a supervised audio-text pair or an unpaired text sequence; and

when the training example corresponds to an unpaired text sequence:

generating, based on the unpaired text sequence, a linguistic context vector;

generating a first log probability from the linguistic context vector;

synthesizing, using a text-to-speech (TTS) system, audio data from the unpaired text sequence;

generating, based on the synthesized audio data, an acoustic context vector;

generating a second log probability from the acoustic context vector;

determining an interpolation of the first log probability and the second log probability; and

updating the LAS decoder based on the determined interpolation.

13. The system of claim 12 , wherein the operations further comprise:

receiving a second training example for the LAS decoder of the two-pass streaming neural network model;

determining that the second training example corresponds to the supervised audio-text pair;

generating, based on the supervised audio-text pair, a second acoustic context vector; and

updating the LAS decoder and acoustic context vector parameters associated with the second acoustic context vector based on a log probability for the second acoustic context vector.

14. The system of claim 12 , wherein determining whether the training example corresponds to the supervised audio-text pair or the unpaired text sequence comprises identifying a domain identifier that indicates whether the training example corresponds to the supervised audio-text pair or the unpaired text sequence.

15. The system of claim 12 , wherein updating the LAS decoder reduces a word error rate (WER) of the two-pass streaming neural network model with respect to long tail entities.

16. The system of claim 12 , wherein the LAS decoder operates in a beam search mode based on a hypothesis generated by the RNN-T decoder during the first pass of the two-pass streaming neural network model.

17. The system of claim 12 , wherein the operations further comprise generating the acoustic context vector with an attention mechanism configured to summarize encoder features from an encoded acoustic frame.

18. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a training example for a listen-attend-spell (LAS) decoder of a two-pass streaming neural network model, the two-pass streaming neural network model comprising:

a recurrent neural network-transducer (RNN-T) decoder configured to generate hypotheses during a first pass of the two-pass streaming neural network model; and

the LAS decoder configured to rescore the generated hypotheses during a second pass of the two-pass streaming neural network model;

determining whether the training example corresponds to a supervised audio-text pair or unpaired training data; and

when the training example corresponds to unpaired training data:

generating a missing portion of the unpaired training data to form a generated audio-text pair;

generating, based on the generated audio-text pair, an acoustic context vector and a linguistic context vector;

generating a first log probability from the acoustic context vector;

generating a second log probability from the linguistic context vector;

determining an interpolation of the first log probability and the second log probability; and

updating the LAS decoder based on the determined interpolation.

19. The system of claim 18 , wherein determining whether the training example corresponds to the supervised audio-text pair or the unpaired training data comprises identifying a domain identifier that indicates whether the training example corresponds to the supervised audio-text pair or the unpaired training data.

20. The system of claim 18 , wherein updating the LAS decoder reduces a word error rate (WER) of the two-pass streaming neural network model with respect to long tail entities.

21. The system of claim 18 , wherein the operations further comprise generating the acoustic context vector using an attention mechanism configured to summarize encoder features from an encoded acoustic frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2021
From: SAINATH, TARA N.; PANG, RUOMING; WEISS, RON; HE, YANZHANG; CHIU, CHUNG-CHENG; STROHMAN, TREVOR
To: GOOGLE LLC
Reel/Frame 054997/0053 →
Continuity (2)
Provisional Application 62964567 · Jan 22, 2020
Related Publication 20210225362A1 · Jul 22, 2021
Cited By (1)
US 12,361,215