IP Library Granted Patent US 11,087,488
Granted Patent B2
US 11,087,488 · App. 16/421,158 · Granted Aug 10, 2021

Automated gesture identification using neural networks

Inventors: Trevor Chandler (Thornton, CO); Dallas Nash (Frisco, TX); Michael Menefee (Richardson, TX)
Assignee: AVODAH, INC.
G06T7/73G06F3/017G06F3/0304G06K9/00248G06K9/00342G06K9/00355G06N3/0454G06T7/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,087,488
App. No.
16/421,158
Granted
Aug 10, 2021
Kind
B2
Abstract

Disclosed are methods, apparatus and systems for gesture recognition based on neural network processing. One exemplary method for identifying a gesture communicated by a subject includes receiving a plurality of images associated with the gesture, providing the plurality of images to a first 3-dimensional convolutional neural network (3D CNN) and a second 3D CNN, where the first 3D CNN is operable to produce motion information, where the second 3D CNN is operable to produce pose and color information, and where the first 3D CNN is operable to implement an optical flow algorithm to detect the gesture, fusing the motion information and the pose and color information to produce an identification of the gesture, and determining whether the identification corresponds to a singular gesture across the plurality of images using a recurrent neural network that comprises one or more long short-term memory units.

Claims (33)

1. An artificial intelligence system adapted for processing images associated with a gesture performed by a subject, comprising:

a plurality of pipeline structures, each of the pipeline structures configured to include three components:

an associated pre-rule component configured to process an input to the pipeline structure,

an associated pipeline component, and

an associated post-rule component configured to process an output of the pipeline component,

wherein a first pipeline structure comprises:

a first pre-rule component configured to determine whether an input stream includes a plurality of input images comprising pixels,

a first pipeline component configured to generate recognition information comprising at least one characteristic in each of the plurality of input images, the at least one characteristic comprising a pose, a color or a gesture type, and

a first post-rule component configured to determine whether the at least one characteristic is associated with the gesture,

wherein a second pipeline structure comprises:

a second pre-rule component configured to determine whether the plurality of input images is associated with the gesture, and

a second pipeline component configured to perform a facial recognition or an emotional recognition operation on the plurality of input images and generate a first result,

wherein a third pipeline structure comprises:

a third pre-rule component configured to determine whether the first result from the second pipeline structure is compatible with the recognition information generated by the first pipeline structure, and

a third pipeline component configured to determine, using a feedback connection, whether the recognition information corresponds to a singular gesture across the plurality of input images.

2. The system of claim 1 , wherein the first pipeline structure comprises a three-dimensional convolutional neural network (3D CNN) configured to receive the input stream and output the recognition information, wherein the second pipeline structure comprises a facial or emotional recognition (FER) module, and wherein the third pipeline structure comprises a recurrent neural network (RNN) comprising an output that is coupled to an input of the RNN to provide the feedback connection.

3. The system of claim 1 , wherein a fourth pipeline structure of the plurality of pipeline structures comprises a pre-processing module configured to receive the input stream and output the plurality of input images, and wherein the fourth pipeline structure comprises:

a fourth pre-rule component configured to determine whether the input stream includes a plurality of captured images comprising pixels,

a fourth pipeline component configured to perform pose estimation on each of the plurality of captured images, and

a fourth post-rule component configured to overlay pose estimation pixels onto the plurality of captured images to generate the plurality of input images.

4. The system of claim 3 , wherein the pose estimation identifies the subject's body, fingers and face.

5. The system of claim 3 , wherein the fourth pipeline component is further configured, prior to performing the pose estimation, to:

process the plurality of captured images to extract pixels corresponding to a subject, a foreground and a background in each of the plurality of captured images; and

perform spatial filtering, upon determining that the background or the foreground includes no information that pertains to the gesture being identified, to remove the pixels corresponding to the background or the foreground.

6. The system of claim 1 , wherein the first pipeline structure is operable to use an optical flow algorithm to detect the gesture.

7. The system of claim 6 , wherein the optical flow algorithm comprises sharpening, line, edge, corner and shape enhancements.

8. The system of claim 1 , wherein the second pipeline component is operable to use 32 reference points on a face of a subject performing the gesture to generate the first result.

9. The system of claim 1 , wherein the third pipeline component comprises one or more long short-term memory (LTSM) units.

10. The system of claim 1 , wherein a fifth pipeline structure of the plurality of pipeline structures comprises an audio recognition module, and wherein the fifth pipeline structure comprises:

a fifth pre-rule component configured to determine whether an input includes audio samples,

a fifth pipeline component configured to perform audio recognition on the audio samples, and

a fifth post-rule component configured to determine whether the audio samples are associated with the gesture and generate a second result, and

wherein the third pre-rule of the third pipeline structure is further configured to determine whether the second result is compatible with the recognition information generated by the first pipeline structure.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2019
From: AVODAH LABS, INC.
To: AVODAH PARTNERS, LLC
Reel/Frame 051057/0353 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2019
From: AVODAH PARTNERS, LLC
To: AVODAH, INC.
Reel/Frame 051057/0359 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2019
From: CHANDLER, TREVOR; MENEFEE, MICHAEL; NASH, DALLAS; EVALTEC GLOBAL LLC
To: AVODAH LABS, INC.
Reel/Frame 051060/0568 →
Continuity (4)
Continuation 16258514 · Jan 25, 2019
Provisional Application 62693821 · Jul 3, 2018
Provisional Application 62629398 · Feb 12, 2018
Related Publication 20200126250A1 · Apr 23, 2020
Cited By (1)
US 12,694,664