IP Library › Granted Patent US 11,423,701
Granted Patent B2
US 11,423,701 · App. 17/118,578 · Granted Aug 23, 2022

Gesture recognition method and terminal device and computer readable storage medium using the same

Inventors: Miaochen Guo (Shenzhen, CN); Jingtao Zhang (Shenzhen, CN); Shuping Hu (Shenzhen, CN); Dong Wang (Shenzhen, CN); Zaiwang Gu (Shenzhen, CN); Jianxin Pang (Shenzhen, CN); Youjun Xiong (Shenzhen, CN)
Assignee: UBTECH ROBOTICS CORP LTD
G06V40/28G06K9/6201G06K9/6217G06T5/002G06T7/64G06T7/73G06T7/90G06V10/22G06V10/56G06V20/41G06V20/46H04N1/6075G06T2207/10016G06T2207/10024G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,701
App. No.
17/118,578
Granted
Aug 23, 2022
Kind
B2
Abstract

The present disclosure provides a gesture recognition method as well as a terminal device and a computer-readable storage medium using the same. The method includes: obtaining a video stream collected by an image recording device in real time; performing a hand recognition on the video stream to determine static gesture information of a recognized hand in each video frame of the video stream; encoding the static gesture information in the video frames of the video stream in sequence to obtain an encoded information sequence of the recognized hands; and performing a slide detection on the encoded information sequence using a preset sliding window to determine a dynamic gesture category of each recognized hand. In this manner, static gesture recognition and dynamic gesture recognition are effectively integrated in the same process. The dynamic gesture recognition is realized through the slide detection of the sliding window without complex network calculations.

Claims (88)

1. A computer-implemented gesture recognition method, comprising steps of:

obtaining a video stream collected by an image recording device;

performing, a hand recognition on the video stream to determine static gesture information of at least a recognized hand in each video frame of the video stream, wherein the static gesture information comprises a static gesture category and a hand centroid position:

encoding the static gesture information in the video frames of the video stream in sequence to obtain an encoded information sequence of the recognized hands, wherein the static gesture information is encoded in a form having the hand centroid position and a category label of the recognized hand; and

performing a slide detection on the encoded information sequence using a preset sliding window to determine a dynamic gesture category of each recognized hand.

2. The method of claim 1 , wherein the step of performing the hand recognition on the video stream to determine the static gesture information of the recognized hand in each video frame of the video stream comprises:

performing the hand recognition on a target video frame through a neural network classifier to obtain the static gesture category of the recognized hand in the target video frame, in response to the hand recognition being in a short-distance recognition mode, wherein the target video frame is any one of the video frames of the video stream; and

calculating the hand centroid position in the target vidw frame by performing an image processing on the target video frame.

3. The method of claim 2 , wherein the step of calculating the hand centroid position in the target video frame by performing the image processing on the target video frame comprises:

performing a preprocessing on the target video frame to obtain a hand mask image in the target video frame;

binarizing the hand mask image to obtain a binarized hand image;

extracting a contour from the binarized hand image, and selecting the contour with the largest area from the extracted contour as a hand contour; and

performing a distance transformation on the hand contour to calculate the hand centroid position in the target video frame.

4. The method of claim 3 , wherein the step of performing a preprocessing on the target video frame to obtain the hand mask image in the target video frame comprises:

smoothing the target video frame using Gaussian filtering to obtain a smooth image;

converting the smooth image from the RGB color space to the HSV color space to obtain a space-converted image;

performing a feature extraction on the space-converted image using a preset elliptical skin color model to obtain a feature image; and

filtering out an impurity area in the feature image through a morphological opening and closing operation to obtain the hand mask image.

5. The method of claim 1 , wherein the step of performing the hand recognition on the video stream to determine the static gesture information of the recognized hand in each video frame of the video stream comprises:

performing a hand recognition on each video frame of the video stream through a neural network target detector to determine the static gesture category, and the hand centroid position of each recognized hand in each video frame of the video stream, in response to the hand recognition being in a long-distance recognition mode; and

determining a matchingness of each recognized hand in each video frame, and building a tracker corresponding to each recognized hand.

6. The method of claim 5 , wherein the step of determining the matchingness of each recognized hand in each video frame, and building the tracker corresponding to each recognized hand comprises:

calculating a predicted detection frame in the next video frame according to a hand bounding box in the current video frame using a Kalman filter, and determining a hand bounding box in the next video frame according to the predicted detection frame;

performing a Hungarian matching between the hand bounding box in the current video frame and the hand bounding box in the next video frame to determine the matchingness of each recognized hand in the current video frame and the next video frame; and

building the tracker corresponding to the matched recognized hand, in response to the matched recognized hand being successfully matched in a plurality of the video frames consecutive in the video stream.

7. The method of claim 1 , wherein the step of performing the slide detection on the encoded information sequence using the preset sliding window to determine the dynamic gesture category of each recognized hand comprises:

detecting a key frame in the encoded information sequence within the current sliding window, and determining the dynamic gesture category corresponding to the encoded information sequence in response to the detected key frame meeting a preset pattern characteristic; and

sliding the sliding window with one frame backward in the encoded information sequence, and returning to the step of detecting the key frame in the encoded information sequence within the current sliding window until a gesture recognition process is terminated.

8. A terminal device, comprising:

an image recording device;

a memory;

a processor; and

one or more computer programs stored in the memory, and executable on the processor, wherein the one or more computer programs comprise:

instructions for obtaining a video stream collected by the image recording device,

instructions for performing a hand recognition on the video stream to determine static gesture information of at least a recognized hand in each video frame of the video stream, wherein the static gesture information comprises a static gesture category and a hand centroid position;

instructions for encoding the static gesture information in the video frames of the video stream in sequence to obtain an encoded information sequence of the recognized hands, wherein the static gesture information is encoded in a form having the hand centroid position and a category label of the recognized hand; and

instructions for performing a slide detection on the encoded information sequence using a preset sliding window to determine a dynamic gesture category of each recognized hand.

9. The terminal device of claim 8 , wherein the instructions for performing the hand recognition on the video stream to determine the static gesture information of the recognized hand in each video frame of the video stream comprise:

instructions for performing the hand recognition on a target video frame through a neural network classifier to obtain the static gesture category of the recognized hand in the target video frame, in response to the hand recognition being in a short-distance recognition mode, wherein the target video frame is any one of the video frames of the video stream; and

instructions for calculating the hand centroid position in the target video frame by performing an image processing on the target video frame.

10. The terminal device of claim 9 , wherein the instructions for calculating the hand centroid position in the target video frame by performing the image processing on the target video frame comprise:

instructions for performing a preprocessing oar the target video frame to obtain a hand mask image in the target video frame;

instructions for binarizing the hand mask image to obtain a binarized hand image;

instructions for extracting a contour from the binarized hand image, and selecting the contour with the largest area from the extracted contour as a hand contour; and

instructions for performing a distance transformation on the hand contour to calculate the hand centroid position in the target video frame.

11. The terminal device of claim 10 , wherein the instructions for performing a preprocessing on the target video frame to obtain the hand mask image in the target video frame comprise:

instructions for smoothing the target video frame using Gaussian filtering to obtain a smooth image;

instructions for converting the smooth image from the RGB color space to the color space to obtain a space-converted image;

instructions for performing a feature extraction on the space-converted image using preset elliptical skin color model to obtain a feature image; and

instructions for filtering out an impurity area in the feature image through a morphological opening and closing operation to obtain the hand mask image.

12. The terminal device of claim 8 , wherein the instructions for performing the hand recognition on the video stream to determine the static gesture information of the recognized hand in each video frame of the video stream comprise:

instructions for performing a hand recognition on each video frame of the video stream through a neural network target detector to determine the static gesture category and the hand centroid position of each recognized hand in each video frame of the video stream, in response to the hand recognition being in a long-distance recognition mode; and

instructions for determining a matchingness of each recognized hand in each video frame, and building a tracker corresponding to each recognized hand.

13. The terminal device of claim 12 , wherein the instructions for determining the matchingness of each recognized hand in each video frame, and building the tracker corresponding to each recognized hand comprise:

instructions for calculating a predicted detection frame in the next video frame according to a hand bounding box in the current video frame using a Kalman filter, and determining a hand bounding box in the next video frame according to the predicted detection frame;

instructions for performing a Hungarian matching between the hand bounding box in the current video frame and the hand bounding box in the next video frame to determine the matchingness of each recognized hand in the current video frame and the next video frame; and

instructions for building the tracker corresponding to the matched recognized hand, in response to the matched recognized hand being successfully matched in a plurality of the video fames consecutive in the video stream.

14. The terminal device of claim 8 , wherein the instructions for performing the slide detection on the encoded information sequence using the preset sliding window to determine the dynamic gesture category of each recognized hand comprise:

instructions for detecting a key frame in the encoded information sequence within the current sliding window, and determining the dynamic gesture category corresponding to the encoded information sequence in response to the detected key frame meeting a preset pattern characteristic; and

instructions for sliding the sliding window with one frame backward the encoded information sequence, and returning to detect the key frame in the encoded information sequence within the current sliding, window until a gesture recognition process is terminated.

15. A non-transitory computer readable storage medium for storing one or more computer programs, wherein the one or more computer programs comprise:

instructions for obtaining a video stream collected by an image recording device;

instructions for performing a hand recognition on the video stream to determine static gesture information of at least a recognized hand in each video frame of the video stream;

instructions for encoding the static gesture information in the video frames of the video stream in sequence to obtain an encoded information sequence of the recognized hands; and

instructions for performing a slide detection on the encoded information sequence using a preset sliding window to determine a dynamic gesture category of each recognized hand;

wherein the instructions for performing the slide detection on the encoded information sequence using, the preset sliding window to determine the dynamic gesture category of each recognized hand comprise:

instructions for detecting a key frame in the encoded information sequence within the current sliding window, and determining the dynamic gesture category correspondina to the encoded information sequence in response to the detected key frame meeting a preset pattern characteristic; and

instructions for slid the sliding window with one frame backward in the encoded information sequence, and returning to detect the key frame in the encoded information sequence within the current sliding window until a gesture recognition process is terminated.

16. The storage medium of claim 15 , wherein the static gesture information comprises a static gesture category and a hand centroid position, and the instructions for performing the hand recognition on the video stream to determine the static gesture information of the recognized hand in each video frame of the video stream comprise:

instructions for performing the hand recognition on a target video frame through a neural network classifier to obtain the static gesture category of the recognized hand in the target video frame, in response to the hand recognition being in a short-distance recognition mode, wherein the target video frame is any one of the video frames of the video stream; and

instructions for calculating the hand centroid position in the target video frame by performing an image processing on the target video frame.

17. The storage medium of claim 16 , wherein the instructions for calculating the hand centroid position in the target video frame by performing the image processing on the target video frame comprise:

instructions for performing a preprocessing on the target video frame to obtain a hand mask image in the target video frame:

instructions for binarizing the hand mask image to obtain a binarized hand image;

instructions for extracting a contour from the binarized hand image, and selecting the contour with the largest area from the extracted contour as a hand contour; and

instructions for performing a distance transformation on the hand contour to calculate the hand centroid position in the target video frame.

18. The storage medium of claim 17 , wherein the instructions for performing a preprocessing on the target video frame to obtain the hand mask image in the target video frame comprise:

instructions for smoothing the target video frame using Gaussian filtering to obtain a smooth image;

instructions for converting the smooth image from the RGB color space to the HSV color space to obtain a space-converted image;

instructions for performing a feature extraction on the space-converted image using a preset elliptical skin color model to obtain a feature image; and

instructions for filtering out an impurity area in the feature image through a morphological opening and closing operation to obtain the hand mask image.

19. The storage medium of claim 15 , wherein the static gesture information comprises a static gesture category and a hand centroid position, and the instructions for performing the hand recognition on the video stream to determine the static gesture information of the recognized hand in each video frame of the video stream comprise:

instructions for performing a hand recognition on each video frame of the video stream through a neural network target detector to determine the static gesture category and the hand centroid position of each recognized hand in each video frame of the video stream, in response to the hand recognition being in a long-distance recognition mode; and

instructions for determining a matchingness of each recognized hand in each video frame, and building a tracker corresponding to each recognized hand.

20. The storage medium of claim 19 , wherein the instructions for determining the matchingness of each recognized hand in each video frame, and building the tracker corresponding to each recognized hand comprise:

instructions for calculating a predicted detection frame in the next video frame according to a hand bounding box in the current video frame using a Kalman filter, and determining a hand bounding box in the next video frame according to the predicted detection frame;

instructions for performing a Hungarian matching between the hand bounding box in the current video frame and the hand bounding box in the next video frame to determine the matchingness of each recognized hand in the current video frame and the next video frame; and

instructions for building the tracker corresponding to the matched recognized hand, in response to the matched recognized hand being successfully matched in a plurality of the video frames consecutive in the video stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2020
From: GUO, MIAOCHEN; ZHANG, JINGTAO; HU, SHUPING; WANG, DONG; GU, ZAIWANG; PANG, JIANXIN; XIONG, YOUJUN
To: UBTECH ROBOTICS CORP LTD
Reel/Frame 054612/0139 →
Priority Claims (1)
CN 202010320878.9 · Apr 22, 2020 · national
Continuity (1)
Related Publication 20210334524A1 · Oct 28, 2021
Cited By (1)
US 12,192,612