IP Library Granted Patent US 12,293,023
Granted Patent B2
US 12,293,023 · App. 18/061,226 · Granted May 6, 2025

Systems, apparatus, and methods for gesture-based augmented reality, extended reality

Inventors: Te-Won Lee (San Diego, CA); Edwin Chongwoo Park (San Diego, CA)
Assignee: SoftEye, Inc.
G06F3/017G06F3/013G06F3/167G06T11/00G06V10/25G06V10/26G06V10/28G06V10/70G06V10/82G06V40/18G06V40/20G06V40/28H04N23/651
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,023
App. No.
18/061,226
Filed
Dec 2, 2022
Granted
May 6, 2025
Kind
B2
Art Unit
2627
USPC
345/156
Abstract

Systems, apparatus, and methods for a gesture-based augmented reality and/or extended reality (AR/XR) user interface. Conventional image processing scales quadratically based on image resolution. Processing complexity directly corresponds to memory size, power consumption, and heat dissipation. As a result, existing smart glasses solutions have short run-times (<1 hr) and may have battery weight and heat dissipation issues that are uncomfortable for continuous wear. The disclosed solution provides a system and method for low-power image processing via the use of scalable processing. In one specific implementation, gesture detection is divided into multiple stages. Each stage conditionally enables subsequent stages for more complex processing. By scaling processing complexity at each stage, high complexity processing can be performed on an “as-needed” basis.

Claims (52)

1. An apparatus for multi-stage hands-free gesture-based user interface processing, comprising:

an eye-tracking camera sensor;

an outward-facing camera sensor;

a power source;

a first processor;

a second processor;

a first non-transitory computer-readable medium comprising a first set of instructions that when executed by the first processor, causes the first processor to:

detect a gesture based on a first image captured by the outward-facing camera sensor;

identify the gesture is a valid gesture associated with one or more gesture-specific tasks;

determine a target based on a second image captured by the eye-tracking camera sensor; and

wake the second processor in response to identification of the valid gesture; and

a second non-transitory computer-readable medium different from the first non-transitory computer-readable medium comprising a second set of instructions that when executed by the second processor, causes the second processor to:

perform the one or more gesture-specific tasks of the valid gesture; and

enter a first sleep state in response to a completion of the one or more gesture-specific tasks.

2. The apparatus of claim 1 , where the first image is captured at a binned resolution and the one or more gesture-specific tasks comprises retrieving a previously captured full resolution image by the outward-facing camera sensor, the previously captured full resolution image captured prior to waking the second processor.

3. The apparatus of claim 1 , where the first image is captured at a binned resolution and the one or more gesture-specific tasks performed by the second processor comprises capturing a full resolution image by the outward-facing camera sensor.

4. The apparatus of claim 1 , further comprising a network interface and a display,

where the one or more gesture-specific tasks comprises downloading media content associated with the target via the network interface rendering the media content via the display.

5. The apparatus of claim 1 , where the second set of instructions further causes the second processor to put the first processor into a second sleep state after the one or more gesture-specific tasks are completed.

6. The apparatus of claim 1 , where the one or more gesture-specific tasks comprises a chained sequence of gestures.

7. A method for multi-stage hands-free gesture-based user interface processing, comprising:

detecting a gesture based on a first image captured by an outward-facing camera sensor of smart glasses, the gesture associated with one or more gesture-specific tasks;

determining a target based on a second image captured by an eye-tracking camera sensor of the smart glasses;

waking a processor of the smart glasses to perform the one or more gesture-specific tasks based on the gesture; and

putting the processor to sleep after the one or more gesture-specific tasks are completed.

8. The method of claim 7 , further comprising:

determining a gaze point captured in at least the second image; and

determining the target based on the gaze point, where the gesture comprises a two-handed gesture.

9. The method of claim 8 , where the gaze point is distinct from the two-handed gesture.

10. The method of claim 8 , where:

the two-handed gesture comprises a frame formed by hands of a user, and

determining the target comprises determining the gaze point is not on or within the frame created by the two-handed gesture.

11. The method of claim 8 , where the two-handed gesture comprises a shorthand finger positioning.

12. The method of claim 7 , where the target is determined based on a gaze point captured in at least the second image and where the gesture comprises a one-handed gesture.

13. The method of claim 12 , where the one-handed gesture comprises a shorthand finger positioning.

14. The method of claim 7 , further comprising determining an indistinct target based on determining there is an indistinct gaze point or a gaze point fixated on one or both hands of a user.

15. An apparatus for multi-stage hands-free gesture-based user interface processing, comprising:

an eye-tracking camera sensor comprising a first neural network trained to determine a target based on a first image captured by the eye-tracking camera sensor;

an outward-facing camera sensor comprising a second neural network trained to detect a gesture based on a second image captured by the outward-facing camera sensor;

a power source;

a processor; and

a non-transitory computer-readable medium comprising a set of instructions that when executed by the processor, causes the processor to:

perform one or more gesture-specific tasks based on the target determined by the eye-tracking camera sensor and the gesture determined by the outward-facing camera sensor; and

enter a first sleep state after the one or more gesture-specific tasks are completed.

16. The apparatus of claim 15 , further comprising a network interface and a display,

where the one or more gesture-specific tasks comprises downloading media content associated with the target via the network interface and rendering the media content via the display.

17. The apparatus of claim 15 , where the set of instructions further causes the processor to put the outward-facing camera sensor into a second sleep state after the one or more gesture-specific tasks are completed.

18. The apparatus of claim 15 , where:

the second neural network of the outward-facing camera sensor is trained to detect an initial gesture associated with a chained sequence of gestures, and

the one or more gesture-specific tasks performed by the processor comprises detecting the chained sequence of gestures in response to detection of the initial gesture by the outward-facing camera sensor, each gesture of the chained sequence of gestures comprising separate actions and associated with separate gesture specific tasks based on previously performed gestures of the chained sequence of gestures.

19. The apparatus of claim 15 , where the target is determined based on a gaze point, the gesture comprises a two-handed gesture, and the gaze point is distinct from the two-handed gesture.

20. The apparatus of claim 15 , where the target is determined based on a gaze point, the gesture comprises a one-handed gesture, and the gaze point is indistinct from the one-handed gesture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2023
From: LEE, TE-WON; PARK, EDWIN CHONGWOO
To: SOFTEYE, INC.
Reel/Frame 064328/0935 →
Continuity (3)
Provisional Application 63340470 · May 11, 2022
Provisional Application 63285453 · Dec 2, 2021
Related Publication 20230176658A1 · Jun 8, 2023
References Cited (54)
US 9471153B1 · Ivanchenko · 2016 [cited by applicant]
US 9619105B1 · Dal Mutto · 2017 [cited by examiner]
US 9741169B1 · Holz · 2017 [cited by examiner]
US 11262885B1 · Burckel · 2022 [cited by applicant]
US 11847266B2 · Lee et al. · 2023 [cited by applicant]
US 20130106681A1 · Eskilsson · 2013 [cited by examiner]
US 20140118257A1 · Baldwin · 2014 [cited by examiner]
US 20150049017A1 · Weber · 2015 [cited by examiner]
US 20150338915A1 · Publicover · 2015 [cited by examiner]
US 20160274762A1 · Lopez et al. · 2016 [cited by applicant]
US 20170249009A1 · Parshionikar · 2017 [cited by examiner]
US 20180368074A1 · Gong · 2018 [cited by examiner]
US 20190043415A1 · Sinha · 2019 [cited by examiner]
US 20190324529A1 · Stellmach et al. · 2019 [cited by applicant]
US 20190362557A1 · Lacey · 2019 [cited by examiner]
US 20190370529A1 · Kumar et al. · 2019 [cited by applicant]
US 20200051260A1 · Shen et al. · 2020 [cited by applicant]
US 20200092453A1 · Gordon et al. · 2020 [cited by applicant]
US 20200341546A1 · Yuan et al. · 2020 [cited by applicant]
US 20200382717A1 · Chiu et al. · 2020 [cited by applicant]
US 20210072831A1 · Edwards · 2021 [cited by applicant]
US 20210201661A1 · Jazaery et al. · 2021 [cited by applicant]
US 20210248427A1 · Guo et al. · 2021 [cited by applicant]
US 20210373657A1 · Connor et al. · 2021 [cited by applicant]
US 20220101593A1 · Rockel et al. · 2022 [cited by applicant]
US 20220326367A1 · Matuszak et al. · 2022 [cited by applicant]
US 20220391697A1 · Mousavi et al. · 2022 [cited by applicant]
US 20230038159A1 · Holland et al. · 2023 [cited by applicant]
US 20230049339A1 · Ganguly et al. · 2023 [cited by applicant]
US 20230118074A1 · Park et al. · 2023 [cited by applicant]
US 20230305632A1 · Lee et al. · 2023 [cited by applicant]
US 20230333384A1 · Zhu et al. · 2023 [cited by applicant]
US 20230368326A1 · Lee et al. · 2023 [cited by applicant]
US 20230368328A1 · Lee et al. · 2023 [cited by applicant]
US 20230370752A1 · Lee et al. · 2023 [cited by applicant]
U.S. Appl. No. 18/061,203, filed Dec. 2, 2022, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 18/061,226, filed Dec. 2, 2022, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 10/061,257, filed Dec. 2, 2022, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 18/185,362, filed Mar. 16, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 18/185,364, filed Mar. 16, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 18/185,366, filed Mar. 16, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 18/316,181, filed May 11, 2023, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 18/316,203, filed May 11, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 18/316,206, filed May 11, 2023, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 18/316,214, filed May 11, 2023, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 18/316,218, filed May 11, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 18/316,221, filed May 11, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 18/315,225, filed May 11, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 63/491,733, filed Mar. 22, 2023, Edwin Chongwoo Park. [cited by applicant]
U.S. Appl. No. 18/745,027, filed Jun. 17, 2024, SoftEye, Inc., Entire Document [cited by applicant]
U.S. Appl. No. 18/745,233, filed Jun. 17, 2024, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/745,353, filed Jun. 17, 2024, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/745,462, filed Jun. 17, 2024, SoftEye, Inc., Entire Document. [cited by applicant]
Vaswani et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017). [cited by applicant]