IP Library Granted Patent US 12,694,664
Granted Patent B2
US 12,694,664 · App. 18/629,724 · Granted Jul 28, 2026

Real-time gesture recognition method and apparatus

Inventors: Trevor Chandler (Thornton, CO); Dallas Nash (Frisco, TX); Michael Menefee (Richardson, TX)
Assignee: AVODAH, INC.
G06V10/82G06F3/013G06F3/017G06F3/167G06F40/40G06F40/58G06N3/045G06N3/08G06T7/20G06T7/73G06V10/764G06V40/165G06V40/176G06V40/20G06V40/28G09B21/00G09B21/009G10L15/22G10L15/24G10L15/26H04N23/90G06N3/0442G06N20/00G06T3/4046G06T17/00G06T2207/20084G10L13/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,664
App. No.
18/629,724
Filed
Apr 8, 2024
Granted
Jul 28, 2026
Kind
B2
Art Unit
2663
USPC
345/419
Abstract

Disclosed are methods, apparatus and systems for real-time gesture recognition. One exemplary method for the real-time identification of a gesture communicated by a subject includes receiving, by a first thread of the one or more multi-threaded processors, a first set of image frames associated with the gesture, the first set of image frames captured during a first time interval, performing, by the first thread, pose estimation on each frame of the first set of image frames including eliminating background information from each frame to obtain one or more areas of interest, storing information representative of the one or more areas of interest in a shared memory accessible to the one or more multi-threaded processors, and performing, by a second thread of the one or more multi-threaded processors, a gesture recognition operation on a second set of image frames associated with the gesture.

Claims (54)

1 . A method for real-time recognition, using one or more multi-threaded processors, of a gesture communicated by a subject, the method comprising:

receiving, by a first thread of the one or more multi-threaded processors, a first set of image frames associated with the gesture, the first set of image frames captured during a first time interval;

performing, by the first thread, pose estimation on each frame of the first set of image frames including eliminating background information from each frame to obtain one or more areas of interest;

storing information representative of the one or more areas of interest in a shared memory accessible to the one or more multi-threaded processors; and

performing, by a second thread of the one or more multi-threaded processors, a gesture recognition operation on a second set of image frames associated with the gesture, the second set of image frames captured during a second time interval that is different from the first time interval,

wherein performing the gesture recognition operation comprises:

using a first processor of the one or more multi-threaded processors that implements a first three-dimensional convolutional neural network (3D CNN) to perform an optical flow operation on the information representative of the one or more areas of interest that is accessed from the shared memory, wherein the optical flow operation is enabled to recognize a motion associated with the gesture;

using a second processor of the one or more multi-threaded processors that implements a second 3D CNN to perform spatial and color processing operations on the information representative of the one or more areas of interest that is accessed from the shared memory;

fusing results of the optical flow operation and results of the spatial and color processing operations to produce an identification of the gesture; and

using a recurrent neural network (RNN) to determine that the identification corresponds to a singular gesture across at least the first and second sets of image frames,

wherein the first thread and the second thread execute concurrently on the one or more multi-threaded processors, wherein the first thread continuously receives and processes new image frames from the first set of image frames while the second thread performs the gesture recognition operation on the second set of image frames, and wherein the pose estimation and the gesture recognition operation are performed in parallel on different sets of image frames captured at different time intervals, and wherein the first time interval and the second time interval are non-overlapping time intervals.

2 . The method of claim 1 , wherein the first set of image frames are captured using a set of visual sensing devices that include multiple apertures oriented with respect to the subject to receive optical signals corresponding to the gesture from multiple angles.

3 . The method of claim 2 , further comprising:

collecting depth information corresponding to the gesture in one or more planes perpendicular to an image plane captured by the set of visual sensing devices, wherein eliminating the background information is further based on the depth information.

4 . The method of claim 1 , wherein the first 3D CNN has been trained on a limited set of training data, and wherein generating the limited set of training data comprises:

generating a 3D scene that includes a 3D model;

using a value indicative of a total number of images in the limited set of training data to determine a plurality of variations of the 3D scene;

applying each of plurality of variations to the 3D scene to produce a plurality of modified 3D scenes; and

capturing an image of each of the plurality of modified 3D scenes to generate the limited set of training data.

5 . The method of claim 4 , further comprising:

generating, for each image of the limited set of training data, a label that corresponds to a feature of interest, wherein the label comprises one or more bounding lines that delineates a precise boundary of the feature of interest.

6 . The method of claim 5 , wherein the precise boundary of the feature of interest is generated based on a group of polygons that collectively form the feature of interest in the 3D model.

7 . The method of claim 4 , wherein determining the plurality of variations of the 3D scene is based on a set of parameters that specify at least one of: a position of the 3D model, an angle of 3D model, a position of a camera, an orientation of a camera, a lighting attribute, a texture of a subsection of the 3D model, or a background of the 3D scene.

8 . The method of claim 4 , further comprising:

obtaining, after generating the limited set of training data, an evaluation of the gesture recognition operation; and

re-generating another limited set of training data upon a determination that the gesture recognition operation fails to meet one or more predetermined criteria.

9 . The method of claim 1 , wherein the optical flow operation comprises sharpening, line, edge, corner and shape enhancements.

10 . The method of claim 1 , wherein performing the pose estimation produces overlay pixels corresponding to a body, fingers and face of the subject.

11 . The method of claim 10 , wherein the overlay pixels comprise pixels with different colors for each finger of the subject.

12 . The method of claim 1 , wherein the spatial and color processing operations comprise recognizing one or more characteristics of the gesture in data corresponding to a single image frame of the second set of image frames.

13 . The method of claim 1 , wherein the information representative of the one or more areas of interest are accessed by the first 3D CNN and the second 3D CNN from the shared memory without copying data corresponding to the information representative of the one or more areas of interest to any other memory location.

14 . The method of claim 1 , wherein each of the first set of image frames and the second set of image frames comprises a frame number or a Society of Motion Picture and Television Engineers (SMPTE) timecode.

15 . The method of claim 1 , wherein the RNN comprises one or more long short-term memory (LSTM) units.

16 . An apparatus for real-time recognition of a gesture communicated by a subject, the apparatus comprising:

one or more multi-threaded processors; and

a non-transitory memory with instructions stored thereon, the instructions upon execution by the one or more multi-threaded processors, causing the one or more multi-threaded processors to:

receive, by a first thread of the one or more multi-threaded processors, a first set of image frames associated with the gesture, the first set of image frames captured during a first time interval;

perform, by the first thread, pose estimation on each frame of the first set of image frames including eliminating background information from each frame to obtain one or more areas of interest;

store information representative of the one or more areas of interest in a shared memory accessible to the one or more multi-threaded processors; and

perform, by a second thread of the one or more multi-threaded processors, a gesture recognition operation on a second set of image frames associated with the gesture, the second set of image frames captured during a second time interval that is different from the first time interval,

wherein the instructions upon execution by the one or more multi-threaded processors cause the one or more multi-threaded processors, as part of performing the gesture recognition operation, to:

use a first processor of the one or more multi-threaded processors that implements a first three-dimensional convolutional neural network (3D CNN) to perform an optical flow operation on the information representative of the one or more areas of interest that is accessed from the shared memory, wherein the optical flow operation is enabled to recognize a motion associated with the gesture;

use a second processor of the one or more multi-threaded processors that implements a second 3D CNN to perform spatial and color processing operations on the information representative of the one or more areas of interest that is accessed from the shared memory;

fuse results of the optical flow operation and results of the spatial and color processing operations to produce an identification of the gesture; and

use a recurrent neural network (RNN) to determine that the identification corresponds to a singular gesture across at least the first and second sets of image frames,

wherein the first thread and the second thread execute concurrently on the one or more multi-threaded processors, wherein the first thread continuously receives and processes new image frames from the first set of image frames while the second thread performs the gesture recognition operation on the second set of image frames, and wherein the pose estimation and the gesture recognition operation are performed in parallel on different sets of image frames captured at different time intervals, and wherein the first time interval and the second time interval are non-overlapping time intervals.

17 . The apparatus of claim 16 , wherein the first set of image frames are captured using a set of visual sensing devices that include multiple apertures oriented with respect to the subject to receive optical signals corresponding to the gesture from multiple angles.

18 . The apparatus of claim 16 , wherein the first 3D CNN has been trained on a limited set of training data, and wherein the instructions upon execution by the one or more multi-threaded processors cause the one or more multi-threaded processors, as part of generating the limited set of training data, to:

generate a 3D scene that includes a 3D model;

use a value indicative of a total number of images in the limited set of training data to determine a plurality of variations of the 3D scene;

apply each of plurality of variations to the 3D scene to produce a plurality of modified 3D scenes; and

capture an image of each of the plurality of modified 3D scenes to generate the limited set of training data.

19 . The apparatus of claim 18 , wherein the instructions upon execution by the one or more multi-threaded processors cause the one or more multi-threaded processors to:

generate, for each image of the limited set of training data, a label that corresponds to a feature of interest, wherein the label comprises one or more bounding lines that delineates a precise boundary of the feature of interest.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2026
From: EVALTEC GLOBAL LLC
To: AVODAH LABS, INC.
Reel/Frame 073854/0629 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2026
From: AVODAH LABS, INC.
To: AVODAH PARTNERS, LLC
Reel/Frame 073854/0770 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2026
From: AVODAH PARTNERS, LLC
To: AVODAH, INC.
Reel/Frame 073854/0864 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2026
From: CHANDLER, TREVOR; NASH, DALLAS; MENEFEE, MICHAEL
To: AVODAH LABS, INC.
Reel/Frame 074949/0377 →
Continuity (14)
Continuation 17367974 · Jul 6, 2021
Continuation 16730587 · Dec 30, 2019
Continuation 16270532 · Feb 7, 2019
Continuation In Part 16258524 · Jan 25, 2019
Continuation In Part 16258514 · Jan 25, 2019
Continuation In Part 16258509 · Jan 25, 2019
Continuation In Part 16258531 · Jan 25, 2019
Provisional Application 62693841 · Jul 3, 2018
Provisional Application 62693821 · Jul 3, 2018
Provisional Application 62664883 · Apr 30, 2018
Provisional Application 62660739 · Apr 20, 2018
Provisional Application 62654174 · Apr 6, 2018
Provisional Application 62629398 · Feb 12, 2018
Related Publication 20250022265A1 · Jan 16, 2025
References Cited (131)
US 5481454A · Inoue et al. · 1996 [cited by applicant]
US 5544050A · Abe et al. · 1996 [cited by applicant]
US 5659764A · Sakiyama et al. · 1997 [cited by applicant]
US 5704012A · Bigus · 1997 [cited by applicant]
US 5887069A · Sakou et al. · 1999 [cited by applicant]
US 6477239B1 · Ohki et al. · 2002 [cited by applicant]
US 6628244B1 · Hirosawa et al. · 2003 [cited by applicant]
US 7027054B1 · Cheiky et al. · 2006 [cited by applicant]
US 7702506B2 · Yoshimine · 2010 [cited by applicant]
US 8488023B2 · Bacivarov et al. · 2013 [cited by applicant]
US 8553037B2 · Smith et al. · 2013 [cited by applicant]
US 8751215B2 · Tardif · 2014 [cited by applicant]
US D719472S · Sakaue et al. · 2014 [cited by applicant]
US D721290S · Varacca · 2015 [cited by applicant]
US D722315S · Liang et al. · 2015 [cited by applicant]
US D752460S · Gnauck · 2016 [cited by applicant]
US 9305229B2 · Delean et al. · 2016 [cited by applicant]
US 9418458B2 · Chertok et al. · 2016 [cited by applicant]
US 9507432B2 · Hildreth · 2016 [cited by examiner]
US 9715252B2 · Reeves et al. · 2017 [cited by applicant]
US 10037458B1 · Mahmoud et al. · 2018 [cited by applicant]
US 10289903B1 · Chandler et al. · 2019 [cited by applicant]
US 10304208B1 · Chandler et al. · 2019 [cited by applicant]
US 10346198B1 · Chandler et al. · 2019 [cited by applicant]
US 10489639B2 · Menefee et al. · 2019 [cited by applicant]
US 10521264B2 · Chandler et al. · 2019 [cited by applicant]
US 10521928B2 · Chandler et al. · 2019 [cited by applicant]
US 10580213B2 · Browry et al. · 2020 [cited by applicant]
US 10599921B2 · Menefee et al. · 2020 [cited by applicant]
US 10956725B2 · Menefee et al. · 2021 [cited by applicant]
US 11036973B2 · Chandler et al. · 2021 [cited by applicant]
US 11055521B2 · Chandler et al. · 2021 [cited by applicant]
US 11087488B2 · Chandler et al. · 2021 [cited by applicant]
US 11954904B2 · Chandler et al. · 2024 [cited by applicant]
US 20020069067A1 · Klinefelter et al. · 2002 [cited by applicant]
US 20030191779A1 · Sagawa et al. · 2003 [cited by applicant]
US 20040210603A1 · Roston · 2004 [cited by applicant]
US 20050258319A1 · Jeong · 2005 [cited by applicant]
US 20060134585A1 · Adamo-Vilani · 2006 [cited by applicant]
US 20060139348A1 · Harada et al. · 2006 [cited by applicant]
US 20060204033A1 · Yoshimine · 2006 [cited by applicant]
US 20080013793A1 · Hillis et al. · 2008 [cited by applicant]
US 20080013826A1 · Hillis et al. · 2008 [cited by applicant]
US 20080024388A1 · Bruce · 2008 [cited by applicant]
US 20080201144A1 · Song et al. · 2008 [cited by applicant]
US 20090022343A1 · Van Schaack et al. · 2009 [cited by applicant]
US 20100044121A1 · Simon · 2010 [cited by examiner]
US 20100194679A1 · Wu et al. · 2010 [cited by applicant]
US 20100296706A1 · Kaneda et al. · 2010 [cited by applicant]
US 20100310157A1 · Kim et al. · 2010 [cited by applicant]
US 20110221974A1 · Stern et al. · 2011 [cited by applicant]
US 20110228463A1 · Matagne · 2011 [cited by applicant]
US 20110274311A1 · Lee et al. · 2011 [cited by applicant]
US 20110301934A1 · Tardif · 2011 [cited by applicant]
US 20120068917A1 · Huang · 2012 [cited by examiner]
US 20120206456A1 · Crocker · 2012 [cited by applicant]
US 20120206457A1 · Crocker · 2012 [cited by applicant]
US 20130100130A1 · Crocker · 2013 [cited by applicant]
US 20130124149A1 · Carr et al. · 2013 [cited by applicant]
US 20130142417A1 · Kutliroff · 2013 [cited by examiner]
US 20130318525A1 · Palanisamy et al. · 2013 [cited by applicant]
US 20140101578A1 · Kwak et al. · 2014 [cited by applicant]
US 20140225890A1 · Ronot et al. · 2014 [cited by applicant]
US 20140253429A1 · Dai et al. · 2014 [cited by applicant]
US 20140309870A1 · Ricci et al. · 2014 [cited by applicant]
US 20150092008A1 · Manley et al. · 2015 [cited by applicant]
US 20150187135A1 · Magder et al. · 2015 [cited by applicant]
US 20150244940A1 · Lombardi et al. · 2015 [cited by applicant]
US 20150317304A1 · An et al. · 2015 [cited by applicant]
US 20150324002A1 · Quiet et al. · 2015 [cited by applicant]
US 20160042228A1 · Opalka et al. · 2016 [cited by applicant]
US 20160196672A1 · Chertok et al. · 2016 [cited by applicant]
US 20160267349A1 · Shoaib et al. · 2016 [cited by applicant]
US 20160320852A1 · Poupyrev · 2016 [cited by applicant]
US 20160379082A1 · Rodriguez et al. · 2016 [cited by applicant]
US 20170090995A1 · Jubinski et al. · 2017 [cited by applicant]
US 20170153711A1 · Dai et al. · 2017 [cited by applicant]
US 20170192665A1 · Karmon · 2017 [cited by examiner]
US 20170206405A1 · Molchanov · 2017 [cited by examiner]
US 20170220836A1 · Phillips et al. · 2017 [cited by applicant]
US 20170236450A1 · Jung et al. · 2017 [cited by applicant]
US 20170255832A1 · Jones et al. · 2017 [cited by applicant]
US 20170351910A1 · Elwazer et al. · 2017 [cited by applicant]
US 20180018529A1 · Hiramatsu · 2018 [cited by applicant]
US 20180032846A1 · Yang et al. · 2018 [cited by applicant]
US 20180047208A1 · Marin et al. · 2018 [cited by applicant]
US 20180101520A1 · Fuchizaki · 2018 [cited by applicant]
US 20180107901A1 · Nakamura et al. · 2018 [cited by applicant]
US 20180137644A1 · Rad et al. · 2018 [cited by applicant]
US 20180144214A1 · Hsieh et al. · 2018 [cited by applicant]
US 20180181809A1 · Ranjan et al. · 2018 [cited by applicant]
US 20180189974A1 · Clark et al. · 2018 [cited by applicant]
US 20180268601A1 · Rad et al. · 2018 [cited by applicant]
US 20180373985A1 · Yang et al. · 2018 [cited by applicant]
US 20180374236A1 · Ogata et al. · 2018 [cited by applicant]
US 20190026956A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190043472A1 · Garcia · 2019 [cited by applicant]
US 20190064851A1 · Tran et al. · 2019 [cited by applicant]
US 20190066733A1 · Somanath et al. · 2019 [cited by applicant]
US 20190116322A1 · Holzer · 2019 [cited by examiner]
US 20190251343A1 · Menefee et al. · 2019 [cited by applicant]
US 20190251344A1 · Menefee et al. · 2019 [cited by applicant]
US 20190251702A1 · Chandler et al. · 2019 [cited by applicant]
US 20190332419A1 · Chandler et al. · 2019 [cited by applicant]
US 20200005028A1 · Gu · 2020 [cited by applicant]
US 20200034609A1 · Chandler et al. · 2020 [cited by applicant]
US 20200043086A1 · Sorensen · 2020 [cited by examiner]
US 20200104582A1 · Menefee et al. · 2020 [cited by applicant]
US 20200126250A1 · Chandler et al. · 2020 [cited by applicant]
US 20210374393A1 · Chandler et al. · 2021 [cited by applicant]
US 20220026992A1 · Chandler et al. · 2022 [cited by applicant]
JP 2017111660A · 2017 [cited by applicant]
WO 2015191468A1 · 2015 [cited by applicant]
Ng et al., “Beyond Short Snippets: Deep Networks for Video Classification” (Year: 2015). [cited by examiner]
Yang et al., “Online Detection and Classification of Dynamic Hand Gestures with Recurrent 3D Convolutional Neural Networks” (Year: 2016). [cited by examiner]
Hagar et al., “One-Shot-Learning Gesture Recognition using HOG-HOF Features” (Year: 2014). [cited by examiner]
International Application No. PCT/US2019/017299, International Search Report and Written Opinion mailed May 31, 2019 (12 pages). [cited by applicant]
Menefee, M. et al. U.S. Appl. No. 16/270,540, Non-Final Office Action mailed Jul. 29, 2019, (9 pages). [cited by applicant]
Chandler, T. et al. U.S. Appl. No. 16/270,532, Notice of Allowance mailed Aug. 12, 2019, (11 pages). [cited by applicant]
Chandler, T. et al. U.S. Appl. No. 16/505,484, Notice of Allowance mailed Aug. 21, 2019, (16 pages). [cited by applicant]
Menefee, M. et al. U.S. Appl. No. 16/694,965, Notice of Allowance mailed Nov. 10, 2020, (pp. 1-6). [cited by applicant]
Menefee, M. et al. U.S. Appl. No. 16/694,965 Non-Final Office Action mailed Mar. 5, 2020 (pp. 1-6). [cited by applicant]
Menefee, M. et al., U.S. Appl. No. 29/678,367, Non-Final Office Action mailed Apr. 6, 2020, (pp. 1-28). [cited by applicant]
Menefee, M. et al., U.S. Appl. No. 29/678,367, Notice of Allowance, mailed Oct. 22, 2020, (pp. 1-7). [cited by applicant]
International Application No. PCT/US2019/01 7299, International Preliminary Report on Patentability mailed Aug. 27, 2020 (pp. 1-9). [cited by applicant]
International Application No. PCT/US2020/016271, International Search Report and Written Opinion, mailed Jun. 30, 2020 (pp. 1-12). [cited by applicant]
Chandler, Trevor, et al. U.S. Appl. No. 16/410,147 Ex Parte Quayle Action mailed Dec. 15, 2020, pp. 1-6. [cited by applicant]
Chandler, Trevor, et al. U.S. Appl. No. 16/410,147 Notice of Allowance, mailed Feb. 11, 2021, pp. 1-8. [cited by applicant]
Chandler, T. et al. U.S. Appl. No. 16/258,514 Notice of Allowance mailed Mar. 27, 2019, (pp. 1-8). [cited by applicant]
Chandler, T. et al. U.S. Appl. No. 16/258,524 Notice of Allowance mailed Apr. 23, 2019, (pp. 1-16). [cited by applicant]
Chandler, T. et al. U.S. Appl. No. 16/258,531 Notice of Allowance mailed Mar. 25, 2019, (pp. 1-8). [cited by applicant]