IP Library Granted Patent US 12,554,333
Granted Patent B2
US 12,554,333 · App. 18/211,507 · Granted Feb 17, 2026

Hand-gesture activation of actionable items

Inventors: Brett D. Miller (San Carlos, CA); Daniel K. Boothe (San Francisco, CA); Martin E. Johnson (Los Gatos, CA)
Assignee: APPLE INC.
G06F3/017G06F3/011G06F3/167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,333
App. No.
18/211,507
Granted
Feb 17, 2026
Kind
B2
Abstract

In one implementation, a method of performing an action is performed at a device including an image sensor, one or more processors, and non-transitory memory. The method includes receiving, from the image sensor, one or more images of a physical environment. The method includes detecting, in the one or more images of the physical environment, one or more actionable items respectively associated with one or more actions. The method includes detecting, in the one or more images of the physical environment, a hand gesture indicating a particular actionable item. The method includes in response to detecting the hand gesture, performing an action associated with the particular actionable item.

Claims (42)

1 . A method comprising:

at a device including an image sensor, one or more processors, and non-transitory memory:

receiving, from the image sensor, one or more images of a physical environment;

detecting, in the one or more images of the physical environment, one or more actionable items respectively associated with one or more actions, wherein detecting the one or more actionable items includes classifying the one or more actionable items to determine corresponding actions for the one or more actionable items, and wherein classifying the one or more actionable items includes performing computer vision on the one or more images of the physical environment;

detecting, in the one or more images of the physical environment, a first hand gesture selecting a particular actionable item of the one or more actionable items, the particular actionable item associated with a plurality of actions including a first action and a second action different from the first action;

in response to detecting the first hand gesture selecting the particular actionable item, without displaying a user interface element to perform the first action associated with the particular actionable item, performing the first action associated with the particular actionable item;

detecting, in the one or more images of the physical environment, a second hand gesture selecting the particular actionable item; and

in response to detecting the second hand gesture selecting the particular actionable item, without displaying a user interface element to perform the second action associated with the particular actionable item, performing the second action associated with the particular actionable item.

2 . The method of claim 1 , wherein detecting the one or more actionable items includes detecting machine-readable content.

3 . The method of claim 1 , wherein detecting the one or more actionable items includes detecting an object.

4 . The method of claim 3 , wherein performing the first action includes changing a state of the object.

5 . The method of claim 1 , wherein performing the first action includes playing audio based on the particular actionable item.

6 . The method of claim 5 , wherein the audio includes at least one of: a reading of the particular actionable item, a definition of the particular actionable item, or a translation of the particular actionable item.

7 . The method of claim 1 , wherein performing the first action includes initiating a phone call based on the particular actionable item.

8 . The method of claim 1 , wherein performing the first action includes storing, in the non-transitory memory, information based on the particular actionable item.

9 . The method of claim 1 , wherein performing the first action is further performed in response to a vocal command.

10 . The method of claim 9 , wherein performing the first action includes selecting, based on the vocal command, the first action from a plurality of actions associated with the particular actionable item.

11 . The method of claim 1 , wherein the device includes a communication interface, wherein the particular actionable item corresponds to another device, and wherein performing the first action includes transmitting data to the other device via the communication interface.

12 . The method of claim 11 , wherein the data indicates a request to change a state of the other device.

13 . The method of claim 1 , wherein the first hand gesture is a swipe hand gesture.

14 . The method of claim 1 , wherein the first hand gesture is a circling hand gesture.

15 . The method of claim 1 , wherein the first hand gesture is a tap hand gesture and the second hand gesture is a double-tap hand gesture.

16 . The method of claim 1 , wherein the first hand gesture is a circle hand gesture in which a finger contacts a thumb to form a circle.

17 . A device comprising:

an image sensor;

a non-transitory memory; and

one or more processors to:

receive, from the image sensor, one or more images of a physical environment;

detect, in the one or more images of the physical environment, one or more actionable items respectively associated with one or more actions, wherein detecting the one or more actionable items includes classifying the one or more actionable items to determine corresponding actions for the one or more actionable items, and wherein classifying the one or more actionable items includes performing computer vision on the one or more images of the physical environment;

detect, in the one or more images of the physical environment, a first hand gesture selecting a particular actionable item of the one or more actionable items, the particular actionable item associated with a plurality of actions including a first action and a second action different from the first action;

in response to detecting the first hand gesture selecting the particular actionable item, without displaying a user interface element to perform a the first action associated with the particular actionable item, perform the first action associated with the particular actionable item;

detect, in the one or more images of the physical environment, a second hand gesture selecting the particular actionable item; and

in response to detecting the second hand gesture selecting the particular actionable item, without displaying a user interface element to perform the second action associated with the particular actionable item, perform the second action associated with the particular actionable item.

18 . The device of claim 17 , wherein the one or more processors are to detect the one or more actionable items by detecting machine-readable content or an object.

19 . The device of claim 17 , wherein the device does not include a display.

20 . A non-transitory memory storing one or more programs, which, when executed by one or more processors of a device including an image sensor cause the device to:

receive, from the image sensor, one or more images of a physical environment;

detect, in the one or more images of the physical environment, one or more actionable items respectively associated with one or more actions, wherein detecting the one or more actionable items includes classifying the one or more actionable items to determine corresponding actions for the one or more actionable items, and wherein classifying the one or more actionable items includes performing computer vision on the one or more images of the physical environment;

detect, in the one or more images of the physical environment, a first hand gesture selecting a particular actionable item of the one or more actionable items, the particular actionable item associated with a plurality of actions including a first action and a second action different from the first action;

in response to detecting the first hand gesture selecting the particular actionable item, without displaying a user interface element to perform a the first action associated with the particular actionable item, perform the first action associated with the particular actionable item;

detect, in the one or more images of the physical environment, a second hand gesture selecting the particular actionable item; and

in response to detecting the second hand gesture selecting the particular actionable item, without displaying a user interface element to perform the second action associated with the particular actionable item, perform the second action associated with the particular actionable item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2023
From: MILLER, BRETT D.; JOHNSON, MARTIN E.; BOOTHE, DANIEL K.
To: APPLE INC.
Reel/Frame 064318/0794 →
Continuity (2)
Provisional Application 63354007 · Jun 21, 2022
Related Publication 20230409122A1 · Dec 21, 2023
References Cited (34)
US 9317124B2 · Kongqiao et al. · 2016 [cited by applicant]
US 10362299B1 · Niemeyer · 2019 [cited by examiner]
US 10466794B2 · Maeda et al. · 2019 [cited by applicant]
US 11783548B2 · Richter · 2023 [cited by examiner]
US 11842729B1 · Richter · 2023 [cited by examiner]
US 20120249741A1 · Maciocci · 2012 [cited by examiner]
US 20130194259A1 · Bennett · 2013 [cited by examiner]
US 20140062962A1 · Jang · 2014 [cited by examiner]
US 20140152557A1 · Yamamoto et al. · 2014 [cited by applicant]
US 20150062165A1 · Saito · 2015 [cited by examiner]
US 20150103003A1 · Kerr · 2015 [cited by examiner]
US 20160313902A1 · Hill · 2016 [cited by examiner]
US 20170287218A1 · Nuernberger · 2017 [cited by examiner]
US 20190019515A1 · Kim et al. · 2019 [cited by applicant]
US 20190087015A1 · Lam · 2019 [cited by examiner]
US 20190250716A1 · Kim · 2019 [cited by applicant]
US 20190333278A1 · Palangie · 2019 [cited by examiner]
US 20190369714A1 · Pla I. Conesa · 2019 [cited by examiner]
US 20190391666A1 · Kim · 2019 [cited by applicant]
US 20200097083A1 · Mao · 2020 [cited by examiner]
US 20200117336A1 · Mani · 2020 [cited by examiner]
US 20210097766A1 · Palangie · 2021 [cited by examiner]
US 20210097768A1 · Malia · 2021 [cited by examiner]
US 20210097776A1 · Faulkner · 2021 [cited by examiner]
US 20210149496A1 · Sen · 2021 [cited by examiner]
US 20210192802A1 · Nepveu · 2021 [cited by examiner]
US 20210232232A1 · Wang et al. · 2021 [cited by applicant]
US 20230297607A1 · Salter · 2023 [cited by examiner]
US 20230409122A1 · Miller · 2023 [cited by examiner]
US 20240070931A1 · Lozada · 2024 [cited by examiner]
US 20240201787A1 · Medarametla Lakshmi · 2024 [cited by examiner]
Ajune et al., Augmented Reality Real-time Drawing Application with a Hand Gesture on a Handheld Interface, 2021, IEEE,6 pages. (Year: 2021). [cited by examiner]
Rani et al., Hand Gesture Control of Virtual Object in Augmented Reality, 2017, IEEE, 6 pages. (Year: 2017). [cited by examiner]
Extended European Search Report dated Oct. 24, 2023, EP Application No. 23180382.6, pp. 1-11. [cited by applicant]