IP Library Granted Patent US 12,307,019
Granted Patent B2
US 12,307,019 · App. 18/366,314 · Granted May 20, 2025

Systems, apparatus, and methods for gesture-based augmented reality, extended reality

Inventors: Te-Won Lee (San Diego, CA); Edwin Chongwoo Park (San Diego, CA)
Assignee: SoftEye, Inc.
G06F3/017G06F3/013G06F3/167G06T11/00G06V10/25G06V10/26G06V10/28G06V10/70G06V10/82G06V40/18G06V40/20G06V40/28H04N23/651
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,019
App. No.
18/366,314
Filed
Aug 7, 2023
Granted
May 20, 2025
Kind
B2
Examiner
SITTA, GRANT
Art Unit
2622
USPC
345/156
Abstract

Systems, apparatus, and methods for a gesture-based augmented reality and/or extended reality (AR/XR) user interface. Conventional image processing scales quadratically based on image resolution. Processing complexity directly corresponds to memory size, power consumption, and heat dissipation. As a result, existing smart glasses solutions have short run-times (<1 hr) and may have battery weight and heat dissipation issues that are uncomfortable for continuous wear. The disclosed solution provides a system and method for low-power image processing via the use of scalable processing. In one specific implementation, gesture detection is divided into multiple stages. Each stage conditionally enables subsequent stages for more complex processing. By scaling processing complexity at each stage, high complexity processing can be performed on an “as-needed” basis.

Claims (40)

1. A gesture-driven scalable processing apparatus, comprising:

a scalable processing subsystem comprising a machine learning processor, and a power source configured to provide power in a standby state, a capture state, and a context state; and

a non-transitory computer-readable medium comprising a first set of instructions that when executed by the scalable processing subsystem, causes the scalable processing subsystem to:

train the machine learning processor to recognize a gaze fixation event and a hand motion event;

obtain eye-tracking images at a first frame rate in the standby state;

cause the power source to change to the capture state when the gaze fixation event is detected within the eye-tracking images by the machine learning processor;

obtain front-facing images at a second frame rate in the capture state; and

cause the power source to change to the context state when the hand motion event is detected within the front-facing images by the machine learning processor.

2. The gesture-driven scalable processing apparatus of claim 1 , where the machine learning processor is trained during an offline set-up sequence that prompts a user to perform a plurality of user interactions.

3. The gesture-driven scalable processing apparatus of claim 1 , where the machine learning processor is trained to recognize at least one of a physical speed and a range of motion of the hand motion event.

4. The gesture-driven scalable processing apparatus of claim 1 , where the machine learning processor is trained to recognize the gaze fixation event based on at least one of a threshold duration or a threshold amplitude.

5. The gesture-driven scalable processing apparatus of claim 1 , where the first set of instructions, when executed by the scalable processing subsystem, further causes the scalable processing subsystem to determine a gesture based on at least one of a gaze point, a region of interest, or a hand movement in the context state.

6. The gesture-driven scalable processing apparatus of claim 5 , where the scalable processing subsystem is trained to recognize a set of gestures that can be combined into a larger set of gestures.

7. The gesture-driven scalable processing apparatus of claim 5 , where the first set of instructions, when executed by the scalable processing subsystem, further causes the scalable processing subsystem to cause the power source to return to the standby state when the hand motion event is not detected within the front-facing images by the machine learning processor.

8. A method for gesture-driven scalable processing, comprising:

training a machine learning processor to recognize a first user interaction at a first exposure setting and a first frame rate and a second user interaction at a second exposure setting and a second frame rate;

capturing a first set of images in a standby power state at the first exposure setting and the first frame rate;

changing a scalable processing subsystem to a capture power state when the first user interaction is detected by the machine learning processor;

capturing a second set of images in the capture power state at the second exposure setting and the second frame rate; and

changing the scalable processing subsystem to a context power state when the second user interaction is detected by the machine learning processor.

9. The method of claim 8 , where the first set of images are captured by an inward-facing camera sensor and the first user interaction comprises a gaze fixation event.

10. The method of claim 9 , where the machine learning processor is trained to detect the gaze fixation event based on at least one of a threshold duration or a threshold amplitude.

11. The method of claim 8 , where the first set of images are captured by an outward-facing camera sensor and the first user interaction comprises a hand motion event.

12. The method of claim 11 , where the machine learning processor is trained to detect the hand motion event based on at least one of a physical speed and a range of motion of a user's hand.

13. The method of claim 8 , further comprising determining a gesture based on at least one of a gaze point, a region of interest, or a hand movement in the context power state.

14. The method of claim 8 , further comprising capturing a sequential chain of additional gestures during the context power state.

15. A scalable processing subsystem configured to:

train a machine learning processor to recognize a set of user-specific user interactions captured at a plurality of resolutions and image qualities, the set of user-specific user interactions associated with a corresponding set of power states;

cause the scalable processing subsystem to change to a second power state when a first user-specific user interaction is recognized during a first power state;

cause the scalable processing subsystem to change to a third power state when a second user-specific user interaction is recognized during the second power state; and

cause the scalable processing subsystem to change to a fourth power state when a third user-specific user interaction is recognized during the third power state.

16. The scalable processing subsystem of claim 15 , further configured to:

cause the scalable processing subsystem to return to the first power state when the second user-specific user interaction is not recognized during the second power state; and

cause the scalable processing subsystem to return to the second power state when the third user-specific user interaction is not recognized during the third power state.

17. The scalable processing subsystem of claim 15 , further configured to:

cause the scalable processing subsystem to return to the first power state when the second user-specific user interaction is not recognized during the second power state; and

cause the scalable processing subsystem to return to the first power state when the third user-specific user interaction is not recognized during the third power state.

18. The scalable processing subsystem of claim 15 , where the machine learning processor is trained during an offline set-up sequence that prompts a user to perform a plurality of user interactions.

19. The scalable processing subsystem of claim 18 , where the plurality of user interactions is a subset of, but can be combined into, the set of user-specific user interactions associated with the corresponding set of power states.

20. The scalable processing subsystem of claim 15 , where the machine learning processor is trained to recognize the set of user-specific user interactions at a resolution, frame rate, exposure setting or image quality associated with the corresponding set of power states.

Continuity (4)
Continuation 18061257 · Dec 2, 2022
Provisional Application 63340470 · May 11, 2022
Provisional Application 63285453 · Dec 2, 2021
Related Publication 20240019939A1 · Jan 18, 2024
References Cited (55)
US 9471153B1 · Ivanchenko · 2016 [cited by examiner]
US 9619105B1 · Dal Mutto · 2017 [cited by applicant]
US 9741169B1 · Holz · 2017 [cited by applicant]
US 11262885B1 · Burckel · 2022 [cited by applicant]
US 11847266B2 · Lee et al. · 2023 [cited by applicant]
US 20130106681A1 · Eskilsson et al. · 2013 [cited by applicant]
US 20140118257A1 · Baldwin · 2014 [cited by applicant]
US 20150049017A1 · Weber et al. · 2015 [cited by applicant]
US 20150309663A1 · Seo · 2015 [cited by examiner]
US 20150338915A1 · Publicover et al. · 2015 [cited by applicant]
US 20160274762A1 · Lopez et al. · 2016 [cited by applicant]
US 20170249009A1 · Parshionikar · 2017 [cited by applicant]
US 20180368074A1 · Gong et al. · 2018 [cited by applicant]
US 20190043415A1 · Sinha et al. · 2019 [cited by applicant]
US 20190324529A1 · Stellmach et al. · 2019 [cited by applicant]
US 20190362557A1 · Lacey et al. · 2019 [cited by applicant]
US 20190370529A1 · Kumar et al. · 2019 [cited by applicant]
US 20200051260A1 · Shen et al. · 2020 [cited by applicant]
US 20200092453A1 · Gordon et al. · 2020 [cited by applicant]
US 20200341546A1 · Yuan et al. · 2020 [cited by applicant]
US 20200382717A1 · Chiu · 2020 [cited by examiner]
US 20210072831A1 · Edwards · 2021 [cited by applicant]
US 20210201661A1 · Jazaery et al. · 2021 [cited by applicant]
US 20210248427A1 · Guo et al. · 2021 [cited by applicant]
US 20210373657A1 · Connor et al. · 2021 [cited by applicant]
US 20220101593A1 · Rockel et al. · 2022 [cited by applicant]
US 20220326367A1 · Matuszak · 2022 [cited by examiner]
US 20220391697A1 · Mousavi et al. · 2022 [cited by applicant]
US 20230038159A1 · Holland · 2023 [cited by examiner]
US 20230049339A1 · Ganguly et al. · 2023 [cited by applicant]
US 20230118074A1 · Park et al. · 2023 [cited by applicant]
US 20230305632A1 · Lee et al. · 2023 [cited by applicant]
US 20230333384A1 · Zhu et al. · 2023 [cited by applicant]
US 20230368326A1 · Lee et al. · 2023 [cited by applicant]
US 20230368328A1 · Lee et al. · 2023 [cited by applicant]
US 20230370752A1 · Lee et al. · 2023 [cited by applicant]
U.S. Appl. No. 18/061,203, filed Dec. 2, 2022, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 18/061,226, filed Dec. 2, 2022, Te-Won Lee. [cited by applicant]
U.S. Appl. No. 18/061,257, filed Dec. 2, 2022, Te-Won Lee, Entire Document. [cited by applicant]
U.S. Appl. No. 18/185,362, filed Mar. 16, 2023, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/185,364, filed Mar. 16, 2023, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/185,366, filed Mar. 16, 2023, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/316m181, filed May 11, 2023, Te-Won Lee, Entire Document. [cited by applicant]
U.S. Appl. No. 18/316,203, filed May 11, 2023, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/316,206, filed May 11, 2023, Te-Won Lee, Entire Docuement. [cited by applicant]
U.S. Appl. No. 18/316,214, filed May 11, 2023, Te-Won Lee, Entire Document. [cited by applicant]
U.S. Appl. No. 18/316,218, filed May 11, 2023, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/316,221, filed May 11, 2023, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/316,225, filed May 11, 2023, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/745,027, filed Jun. 17, 2024, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/745,233, filed Jun. 17, 2024, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/745,353, filed Jun. 17, 2024, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 18/745,462, filed Jun. 17, 2024, SoftEye, Inc., Entire Document. [cited by applicant]
U.S. Appl. No. 63/491,733, filed Mar. 22, 2023, Edwin Chongwoo Park, Entire Document. [cited by applicant]
Vaswani et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017). [cited by applicant]