IP Library Granted Patent US 11,062,725
Granted Patent B2
US 11,062,725 · App. 16/278,830 · Granted Jul 13, 2021

Multichannel speech recognition using neural networks

Inventors: Ehsan Variani (Mountain View, CA); Kevin William Wilson (Cambridge, MA); Ron J. Weiss (New York, NY); Tara N. Sainath (Jersey City, NJ); Arun Narayanan (Santa Clara, CA)
Assignee: Google LLC
G10L25/30G10L15/16G10L15/20G10L19/008G10L21/028G10L21/0388G10L2021/02087G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,062,725
App. No.
16/278,830
Granted
Jul 13, 2021
Kind
B2
Abstract

This specification describes computer-implemented methods and systems. One method includes receiving, by a neural network of a speech recognition system, first data representing a first raw audio signal and second data representing a second raw audio signal. The first raw audio signal and the second raw audio signal describe audio occurring at a same period of time. The method further includes generating, by a spatial filtering layer of the neural network, a spatial filtered output using the first data and the second data, and generating, by a spectral filtering layer of the neural network, a spectral filtered output using the spatial filtered output. Generating the spectral filtered output comprises processing frequency-domain data representing the spatial filtered output. The method still further includes processing, by one or more additional layers of the neural network, the spectral filtered output to predict sub-word units encoded in both the first raw audio signal and the second raw audio signal.

Claims (44)

1. A method comprising:

receiving, at data processing hardware, a multi-channel audio input comprising a first audio signal and a second audio signal occurring during a same period of time;

obtaining, by the data processing hardware, a first time-domain representation of the first audio signal and a second time-domain representation of the second audio signal;

generating, by the data processing hardware, using a spatial filtering convolutional layer of a neural network configured to perform spatial filtering, a corresponding spatial filtered output for each of multiple spatial directions by processing the first time-domain representation of the first audio signal and the second time-domain representation of the second audio signal, the spatial filtering convolutional layer using a stride parameter being set to an integer greater than one;

converting, by the data processing hardware, the corresponding spatial filtered output generated for each of the multiple spatial directions into corresponding frequency-domain data; and

processing, by the data processing hardware, using one or more additional neural network layers of the neural network, the corresponding frequency-domain data converted from the spatial filtered output generated for each of the multiple spatial directions to predict speech content encoded in the first audio signal and the second audio signal.

2. The method of claim 1 , wherein converting the corresponding spatial filtered output generated for each of the multiple spatial directions into corresponding frequency-domain data comprises computing a discrete Fourier transform for the corresponding spatial filtered output generated for each of the multiple spatial directions.

3. The method of claim 2 , wherein computing the discrete Fourier transform for the corresponding spatial filtered output generated for each of the multiple spatial directions comprises computing a fast Fourier transform for the corresponding spatial filtered output generated for each of the multiple spatial directions.

4. The method of claim 1 , wherein the neural network is part of a speech recognition model.

5. The method of claim 1 , wherein the neural network is part of an acoustic model configured to indicate probabilities of sub-word units.

6. The method of claim 1 , wherein the one or more additional neural network layers comprise one or more deep neural network layers that provide output to one or more long short-term memory layers.

7. The method of claim 1 , wherein the corresponding spatial filtered output generated for each of the multiple spatial directions comprises a single channel of time-domain data.

8. The method of claim 1 , wherein at least one additional neural network layer of the one or more additional neural network layers is configured to perform feature extraction.

9. The method of claim 8 , wherein the at least one additional neural network layer of the one or more additional neural network layers that is configured to perform feature extraction is also configured to apply a transformation to the corresponding frequency-domain data converted from the spatial filtered output generated for each of the multiple spatial directions.

10. The method of claim 9 , wherein the transformation is a linear transformation.

11. The method of claim 9 , wherein the transformation is a projection.

12. The method of claim 9 , wherein the transformation is a complex linear projection.

13. The method of claim 9 , wherein the transformation is a linear projection of energy.

14. The method of claim 1 , wherein the neural network comprises:

the spatial filtering convolutional layer;

at least one feature extraction neural network layer configured to determine frequency-based characteristics of the corresponding frequency-domain data converted from the spatially filtered output generated for each of the multiple spatial directions; and

one or more neural network layers configured to receive output of the at least one feature extraction neural network layer and determine speech content using one or more recurrent neural network layers and one or more deep neural network layers.

15. The method of claim 1 , further comprising:

detecting, by the data processing hardware, the first audio signal and the second audio signal using multiple microphones of a computing device,

wherein the data processing hardware resides on the computing device.

16. The method of claim 1 , further comprising:

detecting, by the data processing hardware, the first audio signal and the second audio signal using multiple microphones of a computing device;

wherein the neural network is stored or implemented on the computing device.

17. The method of claim 1 , wherein processing the corresponding frequency-domain data converted from the spatial filtered output generated for each of the multiple spatial directions comprises identifying a voice command indicated by the first audio signal and the second audio signal.

18. The method of claim 1 , wherein the spatial filtering convolutional layer and the one or more additional layers have been jointly trained during training of the neural network.

19. A system comprising:

one or more computing devices; and

one or more computer-readable media storing instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform operations comprising:

receiving a multi-channel audio input comprising a first audio signal and a second audio signal occurring during a same period of time;

obtaining a first time-domain representation of the first audio signal and a second time-domain representation of the second audio signal;

generating, using a spatial filtering convolutional layer of a neural network configured to perform spatial filtering, a corresponding spatial filtered output for each of multiple spatial directions by processing the first time-domain representation of the first audio signal and the second time-domain representation of the second audio signal, the spatial filtering convolutional layer using a stride parameter being set to an integer greater than one;

converting the corresponding spatial filtered output generated for each of the multiple spatial directions into corresponding frequency-domain data; and

processing, using one or more additional neural network layers of the neural network, the corresponding frequency-domain data converted from the spatial filtered output generated for each of the multiple spatial directions to predict speech content encoded in the first audio signal and the second audio signal.

20. One or more non-transitory computer-readable media storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising:

receiving a multi-channel audio input comprising a first audio signal and a second audio signal occurring during a same period of time;

obtaining a first time-domain representation of the first audio signal and a second time-domain representation of the second audio signal;

generating, using a spatial filtering convolutional layer of a neural network configured to perform spatial filtering, a corresponding spatial filtered output for each of multiple spatial directions by processing the first time-domain representation of the first audio signal and the second time-domain representation of the second audio signal, the spatial filtering convolutional layer using a stride parameter being set to an integer greater than one;

converting the corresponding spatial filtered output generated for each of the multiple spatial directions into corresponding frequency-domain data; and

processing, using one or more additional neural network layers of the neural network, the corresponding frequency-domain data converted from the spatial filtered output generated for each of the multiple spatial directions to predict speech content encoded in the first audio signal and the second audio signal.

Assignments (2)
ENTITY CONVERSION Recorded Feb 19, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 049951/0325 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2019
From: VARIANI, EHSAN; WILSON, KEVIN WILLIAM; WEISS, RON J.; SAINATH, TARA N.; NARAYANAN, ARUN
To: GOOGLE INC.
Reel/Frame 048371/0305 →
Continuity (3)
Continuation 15350293 · Nov 14, 2016
Provisional Application 62384461 · Sep 7, 2016
Related Publication 20190259409A1 · Aug 22, 2019
Cited By (3)
US 12,340,303 US 12,511,312 US 12,645,940