IP Library › Granted Patent US 10,157,309
Granted Patent B2
US 10,157,309 · App. 15/402,128 · Granted Dec 18, 2018

Online detection and classification of dynamic gestures with recurrent convolutional neural networks

Inventors: Pavlo Molchanov (San Jose, CA); Xiaodong Yang (San Jose, CA); Shalini De Mello (San Francisco, CA); Kihwan Kim (Sunnyvale, CA); Stephen Walter Tyree (St. Louis, MO); Jan Kautz (Lexington, MA)
Assignee: NVIDIA CORPORATION
G06K9/00355G06K9/6251G06K9/6256G06K9/6277G06N3/0445G06N3/0454G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,157,309
App. No.
15/402,128
Granted
Dec 18, 2018
Kind
B2
Abstract

A method, computer readable medium, and system are disclosed for detecting and classifying hand gestures. The method includes the steps of receiving an unsegmented stream of data associated with a hand gesture, extracting spatio-temporal features from the unsegmented stream by a three-dimensional convolutional neural network (3DCNN), and producing a class label for the hand gesture based on the spatio-temporal features.

Claims (36)

1. A computer-implemented method for detecting and classifying gestures, comprising:

receiving, by a processor, an unsegmented stream of data associated with a hand gesture, wherein the processor is configured as a three-dimensional convolutional neural network (3D-CNN) for unified detection and classification of gestures;

extracting spatio-temporal features from the unsegmented stream by the 3D-CNN; and

producing a class label for the hand gesture based on the spatio-temporal features.

2. The method of claim 1 , wherein the spatio-temporal features are local spatio-temporal features extracted from a frame within the unsegmented stream of data and the processor is further configured as a recurrent neural network layer, and producing comprises processing the local spatio-temporal features by the recurrent neural network layer to generate global spatio-temporal features based on a set of frames within the unsegmented stream of data.

3. The method of claim 2 , wherein the processor is further configured as a softmax layer and the global spatio-temporal features are processed by the softmax layer to generate the class label.

4. The method of claim 2 , wherein the producing further comprises processing global spatio-temporal features generated by the recurrent neural network layer for a previous set of frames within the unsegmented stream of data.

5. The method of claim 1 , wherein the unsegmented stream of data comprises a sequence of frames partitioned into clips of m frames, and further comprising producing a class label for each one of the clips.

6. The method of claim 5 , wherein the spatio-temporal features comprise local spatio-temporal features corresponding to each frame within a first clip of the clips and global spatio-temporal features corresponding to two or more of the clips.

7. The method of claim 1 , wherein the unsegmented stream of data comprises color values corresponding to the hand gesture.

8. The method of claim 1 , wherein the unsegmented stream of data comprises depth values corresponding to the hand gesture.

9. The method of claim 1 , wherein the unsegmented stream of data comprises optical flow values corresponding to the hand gesture.

10. The method of claim 1 , wherein the unsegment stream of data comprises stereo-infrared pairs and/or disparity values corresponding to the hand gesture.

11. The method of claim 1 , wherein the class label is produced before the hand gesture ends.

12. The method of claim 1 , wherein the unsegmented stream of data is one of color data and depth data, and further comprising:

receiving a second unsegmented stream of stereo-infrared data;

generating a first class-conditional probability vector for the unsegmented stream of data; and

generating a second class-conditional probability vector for the second unsegmented stream.

13. The method of claim 12 , wherein producing the class label comprises combining the first class-conditional probability vector with the second class-conditional probability vector.

14. The method of claim 1 , wherein the 3D-CNN is trained using weakly-segmented streams of data captured using a first sensor and the unsegmented stream of data associated with the hand gesture is obtained using a second sensor that is different than the first sensor.

15. The method of claim 1 , wherein a CNN function is used during training of the 3D-CNN.

16. The method of claim 1 , wherein during training of the 3D-CNN, the 3D-CNN generates first spatio-temporal features and a portion of the first spatio-temporal features are removed before the first spatio-temporal features are processed by a recurrent neural network.

17. A system for detecting and classifying gestures, comprising:

a memory configured to store an unsegmented data stream associated with a hand gesture; and

a processor that is coupled to the memory and configured as a three-dimensional convolutional neural network (3D-CNN) for unified detection and classification of gestures to:

receive the unsegmented stream of data;

extract spatio-temporal features from the unsegmented stream using the 3D-CNN; and

produce a class label for the hand gesture based on the spatio-temporal features.

18. The system of claim 17 , wherein the spatio-temporal features are local spatio-temporal features extracted from a frame within the unsegmented stream of data and the processor is further configured as a recurrent neural network layer, and

the recurrent neural network layer processes the local spatio-temporal features to generate global spatio-temporal features based on a set of frames within the unsegmented stream of data.

19. A non-transitory computer-readable media storing computer instructions for detecting and classifying gestures that, when executed by one or more processors, cause the one or more processors to perform the steps of:

receiving an unsegmented stream of data associated with a hand gesture, wherein the one or more processors are each configured as a three-dimensional convolutional neural network (3D-CNN) for unified detection and classification of gestures;

extracting spatio-temporal features from the unsegmented stream by the 3D-CNN; and

producing a class label for the hand gesture based on the spatio-temporal features.

20. The non-transitory computer-readable media of claim 19 , wherein the spatio-temporal features are local spatio-temporal features extracted from a frame within the unsegmented stream of data and the processor is further configured as a recurrent neural network layer, and

the recurrent neural network layer processes the local spatio-temporal features to generate global spatio-temporal features based on a set of frames within the unsegmented stream of data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2017
From: MOLCHANOV, PAVLO; YANG, XIAODONG; DE MELLO, SHALINI; KIM, KIHWAN; TYREE, STEPHEN WALTER; KAUTZ, JAN
To: NVIDIA CORPORATION
Reel/Frame 041355/0576 →
Continuity (2)
Provisional Application 62278924 · Jan 14, 2016
Related Publication 20170206405A1 · Jul 20, 2017
Cited By (2)
US 12,405,866 US 12,638,936