IP Library › Granted Patent US 10,529,320
Granted Patent B2
US 10,529,320 · App. 16/251,430 · Granted Jan 7, 2020

Complex evolution recurrent neural networks

Inventors: Izhak Shafran (Menlo Park, CA); Thomas E. Bagby (San Francisco, CA); Russell John Wyatt Skerry-Ryan (Mountain View, CA)
Assignee: Google LLC
G10L15/16G06N3/02G10H1/00G10L15/02G10L19/0212G10H2210/036G10H2210/046G10H2250/235G10H2250/311G10L17/18G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,529,320
App. No.
16/251,430
Granted
Jan 7, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech recognition using complex evolution recurrent neural networks. In some implementations, audio data indicating acoustic characteristics of an utterance is received. A first vector sequence comprising audio features determined from the audio data is generated. A second vector sequence is generated, as output of a first recurrent neural network in response to receiving the first vector sequence as input, where the first recurrent neural network has a transition matrix that implements a cascade of linear operators comprising (i) first linear operators that are complex-valued and unitary, and (ii) one or more second linear operators that are non-unitary. An output vector sequence of a second recurrent neural network is generated. A transcription for the utterance is generated based on the output vector sequence generated by the second recurrent neural network. The transcription for the utterance is provided.

Claims (44)

1. A method performed by one or more computers, wherein the method comprises:

receiving, by the one or more computers, audio data indicating acoustic characteristics of an utterance;

generating, by the one or more computers, a first vector sequence comprising audio features determined from the audio data;

generating, by the one or more computers, a second vector sequence that a first recurrent neural network outputs in response to receiving the first vector sequence as input, wherein the first recurrent neural network has a transition matrix that implements a cascade of linear operators comprising (i) first linear operators that are complex-valued and unitary, and (ii) one or more second linear operators that are non-unitary;

generating, by the one or more computers, an output vector sequence that a second recurrent neural network outputs in response to receiving the second vector sequence as input, wherein the second recurrent neural network comprises one or more layers including long short-term memory cells;

determining, by the one or more computers, a transcription for the utterance based on the output vector sequence generated by the second recurrent neural network; and

providing, by the one or more computers, the transcription for the utterance.

2. The method of claim 1 , wherein the one or more second linear operators that are non-unitary are diagonal matrix multiplication operators, and wherein all other linear operators in the cascade are unitary operators.

3. The method of claim 1 , wherein one or more second linear operators that are non-unitary introduce decay of retained data in memory of the first recurrent neural network.

4. The method of claim 1 , wherein the cascade of linear operators includes at least one of each of the operators in a set comprising a Fourier transformation, an inverse Fourier transformation, a diagonal matrix multiplication, a column permutation, and a Householder reflection.

5. The method of claim 1 , wherein the cascade of linear operators comprises a sequence of operators comprising a first diagonal matrix multiplication, a Fourier transform, a first Householder reflection, a column permutation, a second diagonal matrix multiplication, an inverse Fourier transformation, a second Householder reflection, and a third diagonal matrix multiplication.

6. The method of claim 1 , wherein the cascade of linear operators is limited to operators selected from a set consisting of Fourier transformations, inverse Fourier transformations, diagonal matrix multiplications, column permutations, and Householder reflections; and

wherein the one or more second linear operators that are non-unitary are limited to diagonal matrix multiplications.

7. The method of claim 1 , wherein the audio data comprises audio data for the utterance acquired using two or more microphones, wherein the first vector sequence comprises audio features determined using audio data from the two or more microphones, and wherein the first recurrent neural network is configured to perform beamforming processing.

8. The method of claim 1 , wherein the first recurrent neural network is configured to perform de-reverberation processing.

9. The method of claim 1 , wherein the first recurrent neural network is configured to perform noise reduction processing.

10. The method of claim 1 , wherein receiving the audio data comprises receiving the audio data from a client device over a network; and

wherein providing the transcription comprises providing the transcription to a client device over a network.

11. The method of claim 1 , wherein providing the transcription comprises providing the transcription to a computer system implementing a digital conversational assistant.

12. The method of claim 1 , further comprising:

determining an action specified by the transcription; and

performing the determined action by the one or more computers or instructing a client device or server system to perform the determined action.

13. A system comprising:

one or more computers; and

one or more computer-readable media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

receiving, by the one or more computers, audio data indicating acoustic characteristics of an utterance;

generating, by the one or more computers, a first vector sequence comprising audio features determined from the audio data;

generating, by the one or more computers, a second vector sequence that a first recurrent neural network outputs in response to receiving the first vector sequence as input, wherein the first recurrent neural network has a transition matrix that implements a cascade of linear operators comprising (i) first linear operators that are complex-valued and unitary, and (ii) one or more second linear operators that are non-unitary;

generating, by the one or more computers, an output vector sequence that a second recurrent neural network outputs in response to receiving the second vector sequence as input, wherein the second recurrent neural network comprises one or more layers including long short-term memory cells;

determining, by the one or more computers, a transcription for the utterance based on the output vector sequence generated by the second recurrent neural network; and

providing, by the one or more computers, the transcription for the utterance.

14. The system of claim 13 , wherein the one or more second linear operators that are non-unitary are diagonal matrix multiplication operators, and wherein all other linear operators in the cascade are unitary operators.

15. The system of claim 13 , wherein one or more second linear operators that are non-unitary introduce decay of retained data in memory of the first recurrent neural network.

16. The system of claim 13 , wherein the cascade of linear operators includes at least one of each of the operators in the set comprising a Fourier transformation, an inverse Fourier transformation, a diagonal matrix multiplication, a column permutation, and a Householder reflection.

17. The system of claim 13 , wherein the cascade of linear operators comprises a sequence of operators comprising a first diagonal matrix multiplication, a Fourier transform, a first Householder reflection, a column permutation, a second diagonal matrix multiplication, an inverse Fourier transformation, a second Householder reflection, and a third diagonal matrix multiplication.

18. One or more non-transitory computer-readable media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

receiving, by the one or more computers, audio data indicating acoustic characteristics of an utterance;

generating, by the one or more computers, a first vector sequence comprising audio features determined from the audio data;

generating, by the one or more computers, a second vector sequence that a first recurrent neural network outputs in response to receiving the first vector sequence as input, wherein the first recurrent neural network has a transition matrix that implements a cascade of linear operators comprising (i) first linear operators that are complex-valued and unitary, and (ii) one or more second linear operators that are non-unitary;

generating, by the one or more computers, an output vector sequence that a second recurrent neural network outputs in response to receiving the second vector sequence as input, wherein the second recurrent neural network comprises one or more layers including long short-term memory cells;

determining, by the one or more computers, a transcription for the utterance based on the output vector sequence generated by the second recurrent neural network; and

providing, by the one or more computers, the transcription for the utterance.

19. The one or more non-transitory computer-readable media of claim 18 , wherein the one or more second linear operators that are non-unitary are diagonal matrix multiplication operators, and wherein all other linear operators in the cascade are unitary operators.

20. The one or more non-transitory computer-readable media of claim 18 , wherein one or more second linear operators that are non-unitary introduce decay of retained data in memory of the first recurrent neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2019
From: SHAFRAN, IZHAK; BAGBY, THOMAS E.; SKERRY-RYAN, RUSSELL JOHN
To: GOOGLE LLC
Reel/Frame 048438/0190 →
Continuity (3)
Continuation In Part 16171629 · Oct 26, 2018
Continuation 15386979 · Dec 21, 2016
Related Publication 20190156819A1 · May 23, 2019
Cited By (1)
US 12,566,244