IP Library › Granted Patent US 12,498,836
Granted Patent B1
US 12,498,836 · App. 19/271,494 · Granted Dec 16, 2025

Systems and methods for device control using natural voice commands, motion-based gesture recognition, and tokenized processing to provide universal application control

Inventors: Manoj Kumar Jhawar (Encinitas, CA); Aman Manoj Jhawar (San Marcos, CA)
Assignee: AUGMENTALIS INC.
G06F3/0481G06F3/011G06F3/013G06F3/017G06F3/167G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,498,836
App. No.
19/271,494
Filed
Jul 16, 2025
Granted
Dec 16, 2025
Kind
B1
Art Unit
2699
USPC
345/156
Abstract

The disclosure describes techniques for integrating natural voice command and motion-based gesture recognition, a tokenizer, and a LLM to enable hands-free control of computing devices without requiring specialized programming. In some implementations, a method includes: obtaining data associated with input by a user at an application interface indicating an intent to control the application interface; tokenizing, using a tokenizer, the data associated with the input by the user to create tokenized user input data; obtaining tokenized UI element data corresponding to a tokenized record of actionable UI elements associated with the application interface; generating, using a LLM, based at least on the tokenized UI element data and the tokenized user input data, events to inject into the application interface to control one or more of the actionable UI elements; injecting the events into the application interface; and providing feedback to the user in accordance with injecting the events.

Claims (66)

1 . One or more non-transitory computer readable mediums storing instructions that, when executed by one or more processors of a system, cause the system to perform operations comprising:

obtaining data associated with an input by a user at an application interface indicating an intent to control the application interface;

tokenizing, using a tokenizer, the data associated with the input by the user to create tokenized user input data;

obtaining tokenized user interface (UI) element data corresponding to a tokenized record of actionable UI elements associated with the application interface;

generating, using a trained large language model (LLM), based at least on the tokenized UI element data and the tokenized user input data, one or more events to inject into the application interface to control one or more of the actionable UI elements;

injecting the one or more events into the application interface to control the one or more actionable UI elements in accordance with the input by the user; and

providing feedback to the user in accordance with injecting the one or more events into the application interface,

wherein obtaining the tokenized UI element data corresponding to the tokenized record of actionable UI elements associated with the application interface comprises:

scraping the application interface to create a record of the actionable UI elements associated with the application interface; and

tokenizing, using the tokenizer, the record of the actionable UI elements associated with the application interface to create the tokenized UI element data.

2 . The one or more non-transitory computer readable mediums of claim 1 , wherein scraping the application interface to create the record of the actionable UI elements associated with the application interface comprises:

identifying a first actionable UI element that does not have associated metadata in the application interface;

identifying an action associated with the first actionable UI element;

assigning a unique identifier to the first actionable UI element; and

storing, in a datastore, the unique identifier and the action associated with the first actionable UI element.

3 . The one or more non-transitory computer readable mediums of claim 2 , wherein the first actionable UI element is a UI element that can be invoked via an accessibility service of an operating system on which the application interface runs.

4 . The one or more non-transitory computer readable mediums of claim 1 , wherein:

the operations further comprise:

receiving, using a microphone of the system, a vocal command from the user to control the application interface; and

converting the vocal command to vocal command text; and

obtaining the data associated with input by the user comprises obtaining the vocal command text.

5 . The one or more non-transitory computer readable mediums of claim 4 , wherein:

the data associated with the input by the user is fused input data; and

obtaining the data associated with input by the user further comprises:

obtaining user motion data in response to an inertial measurement unit (IMU) of the system detecting head motion of the user, or obtaining user gaze data in response to an eye tracking unit of the system detecting eye motion of the user; and

fusing the vocal command text with the user motion data or the user gaze data to obtain the fused input data.

6 . The one or more non-transitory computer readable mediums of claim 1 , wherein

the one or more events to inject into the application interface to control one or more of the actionable UI elements comprise a plurality of events to inject into the application interface to control a plurality of the actionable UI elements; and

injecting the one or more events into the application interface comprises: injecting the plurality of events into the application interface to control the plurality of actionable UI elements in accordance with the input by the user.

7 . The one or more non-transitory computer readable mediums of claim 6 , wherein:

the plurality of events comprise a first event that is initiated by an operating system on which the application interface runs, and is not interpretable by the application interface; and

injecting the plurality of events comprises injecting, via the operating system, the first event into the application interface.

8 . The one or more non-transitory computer readable mediums of claim 7 , wherein:

the plurality of events further comprise a second event that is interpretable by the application interface; and

injecting the plurality of events further comprises injecting, via the application interface, the second event into the application interface.

9 . The one or more non-transitory computer readable mediums of claim 5 , wherein the system is a head mounted display system.

10 . The one or more non-transitory computer readable mediums of claim 1 , wherein the tokenizer and the trained LLM are configured to run locally on the system.

11 . The one or more non-transitory computer readable mediums of claim 1 , wherein providing feedback to the user in accordance with injecting the one or more events into the application interface comprises: presenting, on a display of the system, an updated rendering of the application interface in response to controlling the one or more actionable UI elements.

12 . The one or more non-transitory computer readable mediums of claim 11 , wherein providing the feedback to the user in accordance with injecting the one or more events into the application interface further comprises: playing sound on a speaker of the system, or delivering haptic feedback to the user using a haptic feedback device of the system.

13 . A method, comprising:

obtaining data associated with an input by a user at an application interface indicating an intent to control the application interface;

tokenizing, using a tokenizer, the data associated with the input by the user to create tokenized user input data;

obtaining tokenized user interface (UI) element data corresponding to a tokenized record of actionable UI elements associated with the application interface;

generating, using a trained large language model (LLM) model, based at least on the tokenized UI element data and the tokenized user input data, one or more events to inject into the application interface to control one or more of the actionable UI elements;

injecting the one or more events into the application interface to control the one or more actionable UI elements in accordance with the input by the user; and

providing feedback to the user in accordance with injecting the one or more events into the application interface,

wherein obtaining the tokenized UI element data corresponding to the tokenized record of actionable UI elements associated with the application interface comprises:

scraping the application interface to create a record of the actionable UI elements associated with the application interface; and

tokenizing, using the tokenizer, the record of the actionable UI elements associated with the application interface to create the tokenized UI element data.

14 . One or more non-transitory computer readable mediums storing instructions that, when executed by one or more processors of a system, cause the system to perform operations comprising:

obtaining audio data corresponding to a vocal command by a user indicating an intent to perform a task with an object present in a field of view of a camera of the system;

capturing, using the camera, one or more images of the object;

recognizing, based on the one or more images, using a trained image classification model, the object;

tokenizing, using a tokenizer, the audio data to obtain tokenized audio data;

obtaining tokenized object document data corresponding to a tokenized record of object document data obtained from one or more documents describing the object;

generating, using a trained large language model (LLM), based at least on the tokenized object document data and the tokenized audio data, a workflow to perform to complete the task, the workflow comprising a series of steps; and

displaying, via an application interface, a first graphical representation of the workflow and a second graphical representation of the object.

15 . The one or more non-transitory computer readable mediums of claim 14 , wherein the operations further comprise: instantiating one or more actionable UI elements for manipulating the first graphical representation and the second graphical representation in the application interface.

16 . The one or more non-transitory computer readable mediums of claim 15 , wherein the operations further comprise:

obtaining multimodal user input data associated with multimodal input by the user at the application interface indicating an intent to control the one or more actionable UI elements; and

processing the multimodal user input data to provide feedback to the user, the feedback comprising: presenting, on a display, an updated rendering of the application interface.

17 . The one or more non-transitory computer readable mediums of claim 16 , wherein the obtaining the multimodal user input data comprises processing a natural language voice command by the user.

18 . The one or more non-transitory computer readable mediums of claim 15 , wherein obtaining the tokenized object document data comprises:

extracting the object document data from one or more documents describing the object; and

tokenizing the object document data that was extracted to obtain the tokenized object document data.

19 . The one or more non-transitory computer readable mediums of claim 18 , wherein the tokenized object document data comprises spatial tokens relating one or more equipment components to one or more respective physical locations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2025
From: JHAWAR, MANOJ KUMAR; JHAWAR, AMAN MANOJ
To: AUGMENTALIS INC.
Reel/Frame 071748/0790 →
Continuity (1)
Provisional Application 63672935 · Jul 18, 2024
References Cited (51)
US 8442820B2 · Kim et al. · 2013 [cited by applicant]
US 9031847B2 · Sarin et al. · 2015 [cited by applicant]
US 9285589B2 · Osterhout et al. · 2016 [cited by applicant]
US 9733895B2 · Kim et al. · 2017 [cited by applicant]
US 9842589B2 · Inutsuka · 2017 [cited by applicant]
US 9911418B2 · Chi · 2018 [cited by applicant]
US 10474961B2 · Brigham et al. · 2019 [cited by applicant]
US 10656806B2 · Jhawar · 2020 [cited by examiner]
US 10732721B1 · Clements · 2020 [cited by applicant]
US 10788902B2 · Kawano et al. · 2020 [cited by applicant]
US 10908419B2 · Gross et al. · 2021 [cited by applicant]
US 11068932B2 · Shen · 2021 [cited by examiner]
US 11200900B2 · Smith · 2021 [cited by applicant]
US 11393472B2 · Chakladar et al. · 2022 [cited by applicant]
US 11972095B2 · Klein et al. · 2024 [cited by applicant]
US 12008026B1 · Sanz et al. · 2024 [cited by applicant]
US 12183323B2 · Fu et al. · 2024 [cited by applicant]
US 12411699B2 · Jacob · 2025 [cited by examiner]
US 20110071830A1 · Kim et al. · 2011 [cited by applicant]
US 20120016678A1 · Gruber et al. · 2012 [cited by applicant]
US 20150002676A1 · Yoo · 2015 [cited by applicant]
US 20190324613A1 · Jhawar · 2019 [cited by examiner]
US 20200192493A1 · Wu et al. · 2020 [cited by applicant]
US 20200204613A1 · Hatambeiki et al. · 2020 [cited by applicant]
US 20210240783A1 · Ricci · 2021 [cited by applicant]
US 20230128422A1 · Li et al. · 2023 [cited by applicant]
US 20230168859A1 · Bao · 2023 [cited by applicant]
US 20230245654A1 · Shrivastava et al. · 2023 [cited by applicant]
US 20240029729A1 · Shetty · 2024 [cited by applicant]
US 20240312345A1 · Ogawa · 2024 [cited by applicant]
US 20240320444A1 · Maschmeyer · 2024 [cited by examiner]
US 20240362036A1 · Jacob · 2024 [cited by examiner]
US 20240412720A1 · Vasylyev · 2024 [cited by applicant]
US 20250029170A1 · Chachek et al. · 2025 [cited by applicant]
US 20250060618A1 · Lee et al. · 2025 [cited by applicant]
US 20250130636A1 · Zurauskas · 2025 [cited by applicant]
US 20250148811A1 · Santoro et al. · 2025 [cited by applicant]
CN 111209861B · 2022 [cited by applicant]
WO WO2015116972A1 · 2015 [cited by applicant]
Android Accessibility Help, “Get started with Voice Access spoken commands,” Google Support (2025), https://support.google.com/accessibility/android/answer/6151848?hl=en, 2 pages. [cited by applicant]
Delgado, Carlos, “Artyom.js Library—A speech recognition, voice commands and speech synthesis javascript library,” https://sdkcarlos.github.io/sites/artyom.html (2016), 17 pages. [cited by applicant]
“Dragon Professional v16: Premier speech-recognition for Windows 11 and 10,” Nuance® Dragon® Professional v16 Data Sheet, Nuance Communications, Inc. (Dec. 2, 2022), https://www.nuance.com/asset/en_us/collateral/dragon/… [cited by applicant]
Hume, Tom, “Use Voice Access to control your Android device with your voice,” Google Blog, The Keyword, Outreach and Initiative, Accessibility (Dec. 3, 2020), https://blog.google/outreach-initiatives/accessibility/voice… [cited by applicant]
IPhone User Guide, “Find out what Siri can do on iPhone” Apple Support (2025), https://support.apple.com/guide/iphone/find-out-what-siri-can-do-ipha48873ed6/ios. [cited by applicant]
Mortensen, Ditte Hvas, “How to Design Voice User Interfaces,” Interaction Design Foundation (2020), https://www.interaction-design.org/literature/article/how-to-design-voice-user-interfaces?srsltid=AfmBOopBReuFaY-ORMVhi… [cited by applicant]
“RealWear Navigator™ 500 Series User Guide,” RealWear, Inc., Version 1.2, (Jun. 1, 2022), https://support.realwear.com/hubfs/RealWear%20Navigator%20500%20User%20Guide%20v1.2%20English%20(2).pdf, 148 pages. [cited by applicant]
“RealWear Navigator® Z1: User Guide,” RealWear, Inc., Rev03 (2024), https://support.realwear.com/hubfs/Z1%20User%20Guides/20240403%20RealWear%20NavZ1_UserGuide%20(English).pdf, 41 pages. [cited by applicant]
RealWear Inc. Website (2024), https://www.realwear.com/, 6 pages. [cited by applicant]
Vuzix Corporation Website (2025), https://www.vuzix.com/, 9 pages. [cited by applicant]
“Web Speech API,” Mozilla, MDN Web Docs, Apr. 4, 2025, https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API, 7 pages. [cited by applicant]
Zaheer, Sarah, “Advancing AR Interfaces Integrating Gesture, Voice, and Eye-Tracking for Enhanced Interaction,” International Journal Research of Leading Publication (IJLRP), vol. 5, Issue 6, Jun. 2024, 12 pages. [cited by applicant]