IP Library Granted Patent US 11,200,889
Granted Patent B2
US 11,200,889 · App. 16/685,135 · Granted Dec 14, 2021

Dilated convolutions and gating for efficient keyword spotting

Inventors: Alice Coucke (Paris, FR); Mohammed Chlieh (Paris, FR); Thibault Gisselbrecht (Paris, FR); David Leroy (Paris, FR); Mathieu Poumeyrol (Paris, FR); Thibaut Lavril (Paris, FR)
Assignee: Sonos, Inc.
G10L15/16G10L15/063G10L15/22G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,200,889
App. No.
16/685,135
Granted
Dec 14, 2021
Kind
B2
Abstract

A method for detection of a keyword in a continuous stream of audio signal, by using a dilated convolutional neural network (DCNN), implemented by one or more computers embedded on a device, the dilated convolutional network (DCNN) comprising a plurality of dilation layers (DL), including an input layer (IL) and an output layer (OL), each layer of the plurality of dilation layers (DL) comprising gated activation units, and skip-connections to the output layer (OL), the dilated convolutional network (DCNN) being configured to generate an output detection signal when a predetermined keyword is present in the continuous stream of audio signal, the generation of the output detection signal being based on a sequence (SSM) of successive measurements (SM) provided to the input layer (IL), each successive measurement (SM) of the sequence (SSM) being measured on a corresponding frame from a sequence of successive frames extracted from the continuous stream of audio signal, at a plurality of successive time steps.

Claims (10)

1. A method for detection of a keyword in a continuous stream of audio signal, by using a dilated convolutional neural network (DCNN), implemented by one or more computers embedded on a device, the dilated convolutional network (DCNN) comprising a plurality of dilation layers (DL), including an input layer (IL) and an output layer (OL), each layer of the plurality of dilation layers (DL) comprising gated activation units, and skip-connections to the output layer (OL), the dilated convolutional network (DCNN) being configured to generate an output detection signal when a predetermined keyword is present in the continuous stream of audio signal, the generation of the output detection signal being based on a sequence (SSM) of successive measurements (SM) provided to the input layer (IL), each successive measurement (SM) of the sequence (SSM) being measured on a corresponding frame from a sequence of successive frames extracted from the continuous stream of audio signal, at a plurality of successive time steps.

2. The method according to claim 1 , wherein the dilated convolutional neural network (DCNN) is configured to compute, at a time step, a dilated convolution based on a convolution kernel for each dilation layer, and to put in a cache memory the result of the computation at the time step, so that, at a next time step, the result of the computation is used to compute a new dilated convolution based on a shifted convolution kernel for each dilation layer.

3. A computer implemented method for training a dilated convolutional neural network (DCNN), the dilated convolutional neural network (DCNN) being implemented by one or more computers embedded on a device, for keyword detection in a continuous stream of audio signal, the method comprising a data set preparation phase followed by a training phase based on the result of the data set preparation phase, the data set preparation phase comprising a labelling step comprising a step of associating a first label to successive frames which occur inside a predetermined time period centred on a time step at which an end (EK) of the keyword occurs, and in associating a second label to frames occurring outside the predetermined time period and inside a positive audio sample containing a formulation of the keyword, the positive audio samples comprising a first sequence of frames, the frames of the first sequence of frames occurring at successive time steps in between the beginning of the positive audio sample and the end of the positive audio sample.

4. A computer implemented method according to claim 3 , wherein the labelling step further comprises a step of associating the second label to frames inside a negative audio sample not containing a formulation of the keyword, the negative audio sample comprising a second sequence of frames, the frames of the second sequence of frames occurring at successive time steps in between a beginning time step of the positive audio sample and an ending time step of the positive audio sample.

5. A computer implemented method according to claim 3 , wherein the width of the predetermined time period is optimised during a further step of validation based on a set of validation data.

6. A computer implemented method according to claim 3 , wherein, during the training phase, the training of the dilated convolutional neural network (DCNN) is configured to learn only from the frames included in the second sequence of frames and from the frames which are associated to the first label and which are included in the first sequence of frames, and not to learn from the frames which are included in the first sequence of frames and which are associated to the second label.

7. A computer implemented method according to claim 4 , wherein the width of the predetermined time period is optimised during a further step of validation based on a set of validation data.

8. A computer implemented method according to claim 7 , wherein, during the training phase, the training of the dilated convolutional neural network (DCNN) is configured to learn only from the frames included in the second sequence of frames and from the frames which are associated to the first label and which are included in the first sequence of frames, and not to learn from the frames which are included in the first sequence of frames and which are associated to the second label.

9. A computer implemented method according to claim 4 , wherein, during the training phase, the training of the dilated convolutional neural network (DCNN) is configured to learn only from the frames included in the second sequence of frames and from the frames which are associated to the first label and which are included in the first sequence of frames, and not to learn from the frames which are included in the first sequence of frames and which are associated to the second label.

10. A computer implemented method according to claim 5 , wherein, during the training phase, the training of the dilated convolutional neural network (DCNN) is configured to learn only from the frames included in the second sequence of frames and from the frames which are associated to the first label and which are included in the first sequence of frames, and not to learn from the frames which are included in the first sequence of frames and which are associated to the second label.

Assignments (2)
CHANGE OF NAME Recorded Mar 5, 2020
From: SNIPS
To: SONOS VOX FRANCE SAS
Reel/Frame 052097/0755 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2020
From: COUCKE, ALICE; CHLIEH, MOHAMMED; GISSELBRECHT, THIBAULT; LEROY, DAVID; POUMEYROL, MATHIEU; LAVRIL, THIBAUT
To: SNIPS
Reel/Frame 051707/0795 →
Priority Claims (1)
EP 18306501 · Nov 15, 2018 · regional
Continuity (1)
Related Publication 20200160847A1 · May 21, 2020
Cited By (2)
US 12,524,852 US 12,592,224