IP Library › Granted Patent US 12,524,078
Granted Patent B2
US 12,524,078 · App. 18/345,794 · Granted Jan 13, 2026

Devices and methods for gesture-based selection

Inventors: Juwei Lu (North York, CA); Sayem Mohammad Siam (North York, CA); Deepak Sridhar (San Diego, CA); Sidharth Singla (Mississauga, CA); Yannick Verdie (Toronto, CA); Xiaofei Wu (Shenzhen, CN); Srikanth Muralidharan (Thornhill, CA); Zihao Yang (Richmond Hill, CA); Peng Dai (Markham, CA); Songcen Xu (Shenzhen, GD)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06F3/017G06T7/70G06V20/46G06V40/10G06V40/28G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,078
App. No.
18/345,794
Granted
Jan 13, 2026
Kind
B2
Abstract

Methods and devices for machine vision-based selection of content are described. One or more hands are detected in a current frame of video data. A respective fingertip location is determined for each of up to two of the detected hands. A content selection gesture is determined corresponding to the up to two detected hands. Selected content is extracted, as indicated by the content selection gesture and based on the up to two fingertip locations. The device may be a smartphone, a tablet, a laptop, a smart light device, a reader device, etc.

Claims (56)

1 . A method for content selection, the method comprising:

in at least one prior frame of video data, determining there is a content change, obtaining content data recognized from non-digital visual content captured in the at least one prior frame of video data and storing the content data;

detecting one or more hands in an obtained current frame of video data;

determining a respective fingertip location associated with each of up to two detected hands of the detected one or more hands;

identifying a content selection gesture corresponding to the up to two detected hands; and

extracting selected content data corresponding to selected non-digital visual content in the current frame of video data indicated by the content selection gesture, wherein extracting the selected content data comprises extracting the selected content data from the stored content data, wherein indication of the selected non-digital visual content is further based on the respective up to two fingertip locations.

2 . The method of claim 1 , wherein detecting the respective fingertip location associated with each of the up to two detected hands comprises:

using hand bounding boxes corresponding to the up to two detected hands, performing hand classification and hand pose detection to determine the respective fingertip location associated with each of the up to two detected hands.

3 . The method of claim 2 , wherein hand classification and hand pose detection is performed to also determine a respective gesture label with each of the up to two detected hands, and wherein the content selection gesture is identified based on the respective up to two gesture labels.

4 . The method of claim 3 , wherein each of the up to two gesture labels represents a gesture class selected from: a point gesture, or an open hand gesture.

5 . The method of claim 1 , wherein extracting the selected content data corresponding to the selected non-digital visual content further comprises extracting a portion of the current frame of video data indicated by the content selection gesture and performing visual content recognition in the portion of the current frame of video data.

6 . The method of claim 1 , wherein determining the content change comprises:

determining a difference between a statistical characteristic of the at least one prior frame of video data and another frame of video data captured prior to the at least one prior frame, wherein the determined difference is greater than a preset threshold.

7 . The method of claim 1 , wherein determining the content change comprises:

detecting a hand in the at least one prior frame of video data; and

identifying a content change gesture, corresponding to the detected hand in the at least one prior frame of video data, indicating the content change.

8 . The method of claim 1 , further comprising:

determining, from an obtained current frame of depth data, for each fingertip location associated with each of the up to two detected hands, whether the respective fingertip location is associated with a first touch state; and

determining the content selection gesture when the respective fingertip location is considered to have the first touch state.

9 . The method of claim 8 , wherein the respective fingertip location is determined to be associated with the first touch state when a respective fingertip depth associated with the fingertip location is within a predetermined depth margin of a background depth map.

10 . The method of claim 1 , further comprising:

providing an output based on the selected content data;

wherein the output comprises:

a translation of text included in the selected content data;

an audio reading of text included in the selected content data; or

a virtual overlay indicating the selected content data.

11 . A device comprising:

a processing unit coupled to a memory storing machine-executable instructions thereon, wherein the instructions, when executed by the processing device, cause the device to:

in at least one prior frame of video data, determine there is a content change, obtain content data recognized from non-digital visual content captured in the at least one prior frame of video data and store the content data;

detect one or more hands in an obtained current frame of video data;

determine a respective fingertip location associated with each of up to two detected hands of the detected one or more hands;

identify a content selection gesture corresponding to the up to two detected hands; and

extract selected content data corresponding to selected non-digital visual content in the current frame of video data indicated by the content selection gesture, wherein extracting the selected content data comprises extracting the selected content data from the stored content data, wherein indication of the selected non-digital visual content is further based on the respective up to two fingertip locations.

12 . The device of claim 11 , wherein the instructions cause the device to detect the respective fingertip location associated with each of the up to two detected hands by:

using hand bounding boxes corresponding to the up to two detected hands, performing hand classification and hand pose detection to determine the respective fingertip location associated with each of the up to two detected hands.

13 . The device of claim 12 , wherein hand classification and hand pose detection is performed to also determine a respective gesture label with each of the up to two detected hands, and wherein the content selection gesture is identified based on the respective up to two gesture labels.

14 . The device of claim 13 , wherein each of the up to two gesture labels represents a gesture class selected from: a point gesture, or an open hand gesture.

15 . The device of claim 11 , wherein the instructions cause the device to determine the content change by:

determining a difference between a statistical characteristic of the at least one prior frame of video data and another frame of video data captured prior to the at least one prior frame, wherein the determined difference is greater than a preset threshold.

16 . The device of claim 11 , wherein the instructions cause the device to determine the content change by:

detecting a hand in the at least one prior frame of video data; and

identifying a content change gesture, corresponding to the detected hand in the at least one prior frame of video data, indicating the content change.

17 . The device of claim 11 , wherein the instructions further cause the device to:

determine, from an obtained current frame of depth data, for each fingertip location associated with each of the up to two detected hands, whether the respective fingertip location is associated with a first touch state; and

determine the content selection gesture when the respective fingertip location is considered to have the first touch state.

18 . A non-transitory computer-readable medium having machine-executable instructions stored thereon, the instructions, when executed by a processing unit of a device, cause the device to:

in at least one prior frame of video data, determine there is a content change, obtain content data recognized from non-digital visual content captured in the at least one prior frame of video data and store the content data;

detect one or more hands in an obtained current frame of video data;

determine a respective fingertip location associated with each of up to two detected hands of the detected one or more hands;

identify a content selection gesture corresponding to the up to two detected hands; and

extract selected content data corresponding to selected non-digital visual content indicated in the current frame of video data by the content selection gesture, wherein extracting the selected content data comprises extracting the selected content data from the stored content data, wherein indication of the selected non-digital visual content is further based on the respective up to two fingertip locations.

19 . The non-transitory computer-readable medium of claim 18 , wherein the instructions cause the device to determine the content change by:

determining a difference between a statistical characteristic of the at least one prior frame of video data and another frame of video data captured prior to the at least one prior frame, wherein the determined difference is greater than a preset threshold.

20 . The non-transitory computer-readable medium of claim 18 , wherein the instructions cause the device to determine the content change by:

detecting a hand in the at least one prior frame of video data; and

identifying a content change gesture, corresponding to the detected hand in the at least one prior frame of video data, indicating the content change.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2025
From: HUAWEI TECHNOLOGIES CANADA CO., LTD.
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 073153/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2025
From: LU, JUWEI; SIAM, SAYEM MOHAMMAD; SRIDHAR, DEEPAK; VERDIE, YANNICK; WU, XIAOFEI; MURALIDHARAN, SRIKANTH; YANG, ZIHAO; DAI, PENG; XU, SONGCEN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 073153/0812 →
Continuity (2)
Continuation PCTCN2021106778 · Jul 16, 2021
Related Publication 20230350499A1 · Nov 2, 2023
References Cited (28)
US 7519223B2 · Dehlin · 2009 [cited by examiner]
US 7665041B2 · Wilson · 2010 [cited by examiner]
US 8830312B2 · Hummel · 2014 [cited by examiner]
US 8854433B1 · Rafii · 2014 [cited by examiner]
US 10401967B2 · Katz · 2019 [cited by examiner]
US 20050271279A1 · Fujimura · 2005 [cited by examiner]
US 20100259493A1 · Chang · 2010 [cited by examiner]
US 20120242800A1 · Ionescu · 2012 [cited by examiner]
US 20130283213A1 · Guendelman · 2013 [cited by examiner]
US 20130343605A1 · Dal Mutto · 2013 [cited by examiner]
US 20140145929A1 · Minnen · 2014 [cited by examiner]
US 20140168084A1 · Burr · 2014 [cited by examiner]
US 20150242096A1 · Carro · 2015 [cited by examiner]
US 20160091964A1 · Iyer · 2016 [cited by examiner]
US 20160239080A1 · Marcolina · 2016 [cited by examiner]
US 20170045952A1 · Zhang · 2017 [cited by examiner]
US 20170371403A1 · Wetzler · 2017 [cited by examiner]
US 20180047182A1 · Gelb · 2018 [cited by examiner]
US 20180088674A1 · Chen · 2018 [cited by examiner]
US 20180120950A1 · Karmon · 2018 [cited by examiner]
US 20200117336A1 · Mani · 2020 [cited by examiner]
US 20200409469A1 · Kannan · 2020 [cited by examiner]
US 20210056408A1 · Brandt · 2021 [cited by examiner]
US 20210096652A1 · Wang · 2021 [cited by examiner]
US 20210264140A1 · Hill · 2021 [cited by examiner]
US 20210271892A1 · Luo · 2021 [cited by examiner]
US 20220083197A1 · Rockel · 2022 [cited by examiner]
US 20220405502A1 · Wang · 2022 [cited by examiner]