IP Library › Granted Patent US 11,335,333
Granted Patent B2
US 11,335,333 · App. 16/717,746 · Granted May 17, 2022

Speech recognition with sequence-to-sequence models

Inventors: Wei Han (Mountain View, CA); Chung-Cheng Chiu (Sunnyvale, CA); Yu Zhang (Mountain View, CA); Yonghui Wu (Fremont, CA); Patrick Nguyen (Mountain View, CA); Sergey Kishchenko (Mountain View, CA)
Assignee: Google LLC
G10L15/16G10L15/04G10L15/063G10L15/22G10L15/02G10L15/187G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,335,333
App. No.
16/717,746
Granted
May 17, 2022
Kind
B2
Abstract

A method includes obtaining audio data for a long-form utterance and segmenting the audio data for the long-form utterance into a plurality of overlapping segments. The method also includes, for each overlapping segment of the plurality of overlapping segments: providing features indicative of acoustic characteristics of the long-form utterance represented by the corresponding overlapping segment as input to an encoder neural network; processing an output of the encoder neural network using an attender neural network to generate a context vector; and generating word elements using the context vector and a decoder neural network. The method also includes generating a transcription for the long-form utterance by merging the word elements from the plurality of overlapping segments and providing the transcription as an output of the automated speech recognition system.

Claims (50)

1. A method for transcribing a long-form utterance using an automatic speech recognition system, the method comprising:

obtaining, at data processing hardware, audio data for the long-form utterance;

segmenting, by the data processing hardware, the audio data for the long-form utterance into a plurality of overlapping segments;

for each overlapping segment of the plurality of overlapping segments:

providing, by the data processing hardware, features indicative of acoustic characteristics of the long-form utterance represented by the corresponding overlapping segment as input to an encoder neural network;

processing, by the data processing hardware, an output of the encoder neural network using an attender neural network to generate a context vector; and

generating, by the data processing hardware, word elements using the context vector and a decoder neural network;

generating, by the data processing hardware, a transcription for the long-form utterance by merging the word elements from the plurality of overlapping segments; and

providing, by the data processing hardware, the transcription as an output of the automated speech recognition system.

2. The method of claim 1 , wherein segmenting the audio data for the long-form utterance into the plurality of overlapping segments comprises applying a 50-percent overlap between overlapping segments.

3. The method of claim 1 , wherein generating the transcription for the long-form utterance by merging the word elements from the plurality of overlapping segments comprises:

for each overlapping pair of segments of the plurality of overlapping segments, identifying one or more matching word elements; and

generating the transcription for the long-form utterance based on the one or more matching word elements identified from each overlapping pair of segments.

4. The method of claim 1 , further comprising, for each overlapping segment of the plurality of overlapping segments, assigning, by the data processing hardware, a confidence score to each generated word element based on a relative location of the corresponding generated word element in the corresponding overlapping segment.

5. The method of claim 4 , wherein assigning the confidence score to each generated word element comprises assigning higher confidence scores to generated word elements located further from starting and ending boundaries of the corresponding overlapping segment.

6. The method of claim 4 , wherein generating the transcription for the long-form utterance by merging the word elements from the plurality of overlapping segments comprises:

identifying non-matching word elements between a first overlapping segment of the plurality of overlapping segments and a subsequent second overlapping segment of the plurality of overlapping segments, the first overlapping segment associated with one of an odd number or an even number and the subsequent second overlapping segment associated with the other one of the odd number or even number; and

selecting the non-matching word element from for use in the transcription that is associated with a highest assigned confidence score.

7. The method of claim 1 , wherein the encoder neural network, the attender neural network, and the decoder neural network are jointly trained on a plurality of training utterances, each training utterances of the plurality of training utterances comprising a duration that is shorter than a duration of the long-form utterance.

8. The method of claim 1 , wherein the encoder neural network comprises a recurrent neural network including long short-term memory (LSTM) elements.

9. The method of claim 1 , further comprising applying, by the data processing hardware, a monotonicity constraint to the attender neural network.

10. The method of claim 1 , wherein:

providing features indicative of acoustic characteristics of the long-form utterance represented by the corresponding overlapping segment as input to the encoder neural network comprises providing a series of features vectors that represent a corresponding portion of the long-form utterance represented by the overlapping segment; and

generating word elements using the context vector and the decoder neural network comprises beginning decoding of word elements representing the utterance after the encoder neural network has completed generating output encodings for each of the feature vectors in the series of features vectors that represent the corresponding portion of the long-form utterance represented by the overlapping segment.

11. An automated speech recognition (ASR) system for transcribing a long-form utterance, the ASR system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining audio data for the long-form utterance;

segmenting the audio data for the long-form utterance into a plurality of overlapping segments;

for each overlapping segment of the plurality of overlapping segments:

providing features indicative of acoustic characteristics of the long-form utterance represented by the corresponding overlapping segment as input to an encoder neural network;

processing an output of the encoder neural network using an attender neural network to generate a context vector; and

generating word elements using the context vector and a decoder neural network;

generating a transcription for the long-form utterance by merging the word elements from the plurality of overlapping segments; and

providing the transcription as an output of the automated speech recognition system.

12. The ASR system of claim 11 , wherein segmenting the audio data for the long-form utterance into the plurality of overlapping segments comprises applying a 50-percent overlap between overlapping segments.

13. The ASR system of claim 11 , wherein generating the transcription for the long-form utterance by merging the word elements from the plurality of overlapping segments comprises:

for each overlapping pair of segments of the plurality of overlapping segments, identifying one or more matching word elements; and

generating the transcription for the long-form utterance based on the one or more matching word elements identified from each overlapping pair of segments.

14. The ASR system of claim 11 , wherein the operations further comprise, for each overlapping segment of the plurality of overlapping segments, assigning a confidence score to each generated word element based on a relative location of the corresponding generated word element in the corresponding overlapping segment.

15. The ASR system of claim 14 , wherein assigning the confidence score to each generated word element comprises assigning higher confidence scores to generated word elements located further from starting and ending boundaries of the corresponding overlapping segment.

16. The ASR system of claim 14 , wherein generating the transcription for the long-form utterance by merging the word elements from the plurality of overlapping segments comprises:

identifying non-matching word elements between a first overlapping segment of the plurality of overlapping segments and a subsequent second overlapping segment of the plurality of overlapping segments, the first overlapping segment associated with one of an odd number or an even number and the subsequent second overlapping segment associated with the other one of the odd number or even number; and

selecting the non-matching word element for use in the transcription that is associated with a highest assigned confidence score.

17. The ASR system of claim 11 , wherein the encoder neural network, the attender neural network, and the decoder neural network are jointly trained on a plurality of training utterances, each training utterances of the plurality of training utterances comprising a duration that is shorter than a duration of the long-form utterance.

18. The ASR system of claim 11 , wherein the encoder neural network comprises a recurrent neural network including long short-term memory (LSTM) elements.

19. The ASR system of claim 11 , wherein the operations further comprise applying a monotonicity constraint to the attender neural network.

20. The ASR system of claim 1 , wherein:

providing features indicative of acoustic characteristics of the long-form utterance represented by the corresponding overlapping segment as input to the encoder neural network comprises providing a series of features vectors that represent a corresponding portion of the long-form utterance represented by the overlapping segment; and

generating word elements using the context vector and the decoder neural network comprises beginning decoding of word elements representing the utterance after the encoder neural network has completed generating output encodings for each of the feature vectors in the series of features vectors that represent the corresponding portion of the long-form utterance represented by the overlapping segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2020
From: HAN, WEI; CHIU, CHUNG-CHENG; ZHANG, YU; WU, YONGHUI; NGUYEN, PATRICK; KISHCHENKO, SERGEY
To: GOOGLE LLC
Reel/Frame 051940/0781 →
Continuity (3)
Continuation In Part 16516390 · Jul 19, 2019
Provisional Application 62701237 · Jul 20, 2018
Related Publication 20200126538A1 · Apr 23, 2020
Cited By (1)
US 12,315,497