IP Library Granted Patent US 11,914,788
Granted Patent B2
US 11,914,788 · App. 17/827,939 · Granted Feb 27, 2024

Methods and systems for hand gesture-based control of a device

Inventors: Juwei Lu (North York, CA); Sayem Mohammad Siam (North York, CA); Wei Zhou (Richmond Hill, CA); Peng Dai (Markham, CA); Xiaofei Wu (Shenzhen, CN); Songcen Xu (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06F3/017G06V10/25G06V10/82G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,914,788
App. No.
17/827,939
Granted
Feb 27, 2024
Kind
B2
Abstract

Methods and systems for gesture-based control of a device are described. A virtual gesture-space is determined in a received input frame. The virtual gesture-space is associated with a primary user from a ranked user list of users. The received input frame is processed in only the virtual gesture-space, to detect and track a hand. Using a hand bounding box generated by detecting and tracking the hand, gesture classification is performed to determine a gesture input associated with the hand. A command input associated with the determined gesture input is processed. The device may be a smart television, a smart phone, a tablet, etc.

Claims (66)

1. A method for processing gesture input, the method comprising:

determining a virtual gesture-space defined in a received input frame of a sequence of frames, the virtual gesture-space being associated with a primary user from a ranked user list of one or more users and the virtual gesture-space being smaller than a total field of view captured in the received input frame;

processing only the virtual gesture-space in one or more frames in the sequence of frames to detect and track a hand;

using a hand bounding box generated by detecting and tracking the hand, performing gesture classification to determine a gesture input associated with the hand; and

outputting the determined gesture input to cause processing of a command input associated with the determined gesture input.

2. The method of claim 1 , wherein determining the virtual gesture space comprises:

processing the input frame to detect the one or more users;

generating the ranked user list based on the detected one or more users, the primary user being identified as a highest ranked user on the ranked user list; and

generating the virtual gesture-space based on a detected anatomical feature of the primary user.

3. The method of claim 2 , wherein processing the input frame comprises:

selecting a region of interest (ROI) for processing the input frame, the ROI defining an area smaller than a total area of the input frame;

wherein the ROI is selected from a defined ROI sequence, the ROI sequence defining a plurality of ROIs for processing a respective plurality of sequentially received input frames.

4. The method of claim 1 , wherein determining the gesture input comprises:

determining, in at least one subsequent frame, an invalid gesture input associated with the hand detected in the virtual gesture-space associated with the primary user;

selecting a next highest ranked user in the ranked user list as a new primary user; and

repeating the method using the new primary user.

5. The method of claim 1 , wherein processing the input frame one or more frames in only the virtual gesture-space comprises:

determining a low-light condition associated with at least one frame of the one or more frames; and

automatically performing image adjustment to adjust pixel values of the at least one frame in response to the low-light condition.

6. The method of claim 1 , wherein the one or more frames are processed using a trained joint neural network to detect and track the hand, and wherein the trained joint neural network includes a trained gesture classification convolutional neural network.

7. The method of claim 1 , wherein performing gesture classification comprises:

identifying a gesture class associated with the detected hand;

determining a state transition from a previous gesture state to a current gesture state, the state transition being determined based on the identified gesture class; and

determining the gesture input associated with the current gesture state.

8. The method of claim 1 , wherein the one or more frames are received and processed at a frequency lower than a frame capture frequency of an image capturing device used to capture the sequence of frames.

9. An apparatus comprising:

a processing device coupled to a memory storing machine-executable instructions thereon, wherein the instructions, when executed by the processing device, cause the apparatus to:

determining a virtual gesture-space defined in a received input frame of a sequence of frames, the virtual gesture-space being associated with a primary user from a ranked user list of one or more users and the virtual gesture-space being smaller than a total field of view captured in the received input frame;

processing only the virtual gesture-space in one or more frames in the sequence of frames to detect and track a hand;

using a hand bounding box generated by detecting and tracking the hand, perform gesture classification to determine a gesture input associated with the hand;

wherein the determined gesture input causes processing of a command input associated with the determined gesture input.

10. The apparatus of claim 9 , wherein the instructions further cause the apparatus to determine the virtual gesture space by:

processing the input frame to detect the one or more users;

generating the ranked user list based on the detected one or more users, the primary user being identified as a highest ranked user on the ranked user list; and

generating the virtual gesture-space based on a detected anatomical feature of the primary user.

11. The apparatus of claim 10 , wherein the instructions further cause the apparatus to process the input frame by:

selecting a region of interest (ROI) for processing the input frame, the ROI defining an area smaller than a total area of the input frame;

wherein the ROI is selected from a defined ROI sequence, the ROI sequence defining a plurality of ROIs for processing a respective plurality of sequentially received input frames.

12. The apparatus of claim 9 , wherein the instructions further cause the apparatus to determine the gesture input by:

determining, in at least one subsequent frame, an invalid gesture input associated with the hand detected in the virtual gesture-space associated with the primary user;

selecting a next highest ranked user in the ranked user list as a new primary user; and

repeating the method using the new primary user.

13. The apparatus of claim 9 , wherein the instructions further cause the apparatus to process the one or more frames in only the virtual gesture-space by:

determining a low-light condition associated with at least one frame of the one or more frames; and

automatically performing image adjustment to adjust pixel values of the at least one frame in response to the low-light condition.

14. The apparatus of claim 9 , where the one or more frames are processed using a trained joint neural network to detect and track the hand, and wherein the trained joint neural network includes a trained gesture classification convolutional neural network.

15. The apparatus of claim 9 , wherein the instructions further cause the apparatus to perform gesture classification by:

identifying a gesture class associated with the detected hand;

determining a state transition from a previous gesture state to a current gesture state, the state transition being determined based on the identified gesture class; and

determining the gesture input associated with the current gesture state.

16. The apparatus of claim 9 , wherein the one or more frames are received and processed at a frequency lower than a frame capture frequency of an image capturing device used to capture the sequence of frames.

17. The apparatus of claim 9 , wherein the apparatus is a gesture-controlled device.

18. The apparatus of claim 17 , further comprising a camera for capturing the sequence of frames.

19. The apparatus of claim 17 , wherein the gesture-controlled device is one of:

a television;

a smartphone;

a tablet;

a vehicle-coupled device;

an internet of things device;

an artificial reality device; or

a virtual reality device.

20. A non-transitory computer-readable medium having machine-executable instructions stored thereon, the instructions, when executed by a processing device of an apparatus, cause the apparatus to:

determine a virtual gesture-space defined in a received input frame of a sequence of frames, the virtual gesture-space being associated with a primary user from a ranked user list of one or more users and the virtual gesture-space being smaller than a total field of view captured in the received input frame;

process only the virtual gesture-space in one or more frames in the sequence of frames to detect and track a hand;

using a hand bounding box generated by detecting and tracking the hand, perform gesture classification to determine a gesture input associated with the hand; and

output the determined gesture input to cause processing of a command input associated with the determined gesture input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2022
From: LU, JUWEI; SIAM, SAYEM MOHAMMAD; ZHOU, WEI; DAI, PENG; WU, XIAOFEI; XU, SONGCEN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 061589/0426 →
Continuity (2)
Continuation PCTCN2020080416 · Mar 20, 2020
Related Publication 20220291755A1 · Sep 15, 2022