IP Library Granted Patent US 12693732
Granted Patent B1
US 12693732 · App. 19/280,193 · Granted Jul 28, 2026

System and method for providing real-time insights to user

Inventors: Thiyam Akuvan (Imphal, IN); Gyanendro Khomdram (Imphal West, IN); Thoudam Kheljeet Singh (Imphal West, IN); Oinam Roshan (Imphal, IN); Phijam Nongthanganba (Imphal West, IN)
Assignee: Akumen Artificial Intelligence Private Limited
G06F3/011G06F1/163G06V10/764G06V30/153G08B21/0446G08B25/016G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12693732
App. No.
19/280,193
Granted
Jul 28, 2026
Kind
B1
Abstract

The present invention discloses a system and a method for providing real-time insights to a user. The system obtains real-time sensor data. The system preprocesses the obtained real-time sensor data. The system analyses the real-time sensor data through machine learning models and artificial intelligence models. The system recognises obstacles in an environmental space. The system generates contextual responses for the user. The system provides processed feedback data including real-time haptic feedback and audio feedback. The system detects abrupt motion changes indicative of a fall event. The system provides an optimized processing performance on edge devices and servers. The system transmits alert notifications to preregistered emergency contacts upon detection of the abrupt motion changes. The system transfers the processed feedback data for providing the real-time insights to the user.

Claims (61)

1 . A system for providing one or more real-time insights to a user, comprising:

one or more wearable devices, each wearable device of the one or more wearable devices, comprise at least one of:

a plurality of sensors configured to generate real-time sensor data comprises at least one of: visual data, audio data, depth perception data, and motion tracking data; and

a first hardware processor operatively connected to the plurality of sensors, configured to:

control optimized electric power distribution from a power supply unit through a power management unit (PMU) to at least one of: the plurality of sensors, one or more haptic feedback units, one or more audio output units, one or more input control units, and a radio frequency (RF) communication unit;

transmit the real-time sensor data via the radio frequency (RF) communication unit to at least one of: one or more edge devices and one or more servers; and

receive processed feedback data from at least one of: the one or more edge devices and the one or more servers through the radio frequency (RF) communication unit;

one or more second hardware processors associated with the one or more edge devices;

a memory unit operatively connected to the one or more second hardware processors, wherein the memory unit comprises a set of computer-readable instructions in form of a plurality of subsystems, configured to be executed by the one or more second hardware processors, wherein the plurality of subsystems comprises:

a data obtaining subsystem configured to obtain the real-time sensor data from each wearable device of the one or more wearable devices for on-device processing on the one or more edge devices;

a data pre-processing subsystem configured to preprocess the obtained real-time sensor data, including at least one of: denoising the visual data, altering contrast in the visual data, filtering irrelevant audio data, and normalising the depth perception data; and

a data processing subsystem configured to analyse the real-time sensor data through at least one of: one or more machine learning models and one or more artificial intelligence models, wherein the data processing subsystem comprises at least one of:

an object recognition module configured to recognise and classify one or more obstacles in the environmental space by processing at least one of: the visual data, the depth perception data, and the motion tracking data through a convolutional neural network (CNN), wherein the convolutional neural network comprises at least one of: an EfficientDet model, a MobileNet model, and a You Only Look Once (YOLO) model, configured for execution on the one or more edge devices to perform real-time object detection;

an interaction response module configured to generate one or more contextual responses for the user by processing the visual data, the audio data through at least one of: one or more multi-modal artificial intelligence language models and one or more text recognition models;

a navigation and obstacle detection module configured to provide processed feedback data including at least one of: real-time haptic feedback and audio feedback, based on processing the depth perception data and the visual data through a predictive motion model, wherein the real-time haptic feedback comprises discrete vibration patterns produced corresponding to different types of processed feedback data, including one of: proximity of the one or more obstacles and navigation instructions, allowing the user to distinguish between the discrete vibration patterns; and

a fall detection module configured to detect abrupt motion changes indicative of a fall event based on the motion tracking data and vision-based gait analysis, wherein the fall detection module comprising a multi-sensor fusion procedure configured to combine the motion tracking data, the visual data, and the depth perception data to alleviate false positives in detection of the abrupt motion changes;

one or more artificial intelligence (AI) data processing units (DPUs) associated with the one or more servers;

a high bandwidth memory unit operatively connected to the one or more artificial intelligence (AI) data processing units (DPUs), wherein the high bandwidth memory unit comprises the set of computer-readable instructions in form of the plurality of subsystems, configured to be executed by the high bandwidth memory unit, wherein the plurality of subsystems comprises:

a data acquisition subsystem configured to obtain the pre-processed sensor data for the cloud-based processing in the one or more servers;

a cloud data processing subsystem configured with at least one of: the one or more machine learning models, one or more large language models (LLMs), one or more conversational artificial intelligence models, one or more speech recognition models, and one or more text to speech models, for complex processing of the sensor data to generate the processed feedback data; and

an artificial intelligence (AI) model optimization subsystem configured to optimise at least one of: the one or more machine learning models, the one or more artificial intelligence models, the one or more large language models (LLMs), the one or more conversational artificial intelligence models, the one or more speech recognition models, and the one or more text to speech models, for providing an optimized processing performance on at least one of: the one or more edge devices and the one or more servers;

a notification subsystem associated with the one or more edge devices configured to transmit one or more alert notifications to one or more preregistered emergency contacts with location coordinates of the user upon detection of the abrupt motion changes; and

a feedback subsystem associated with the one or more edge devices configured transfer the processed feedback data to the radio frequency (RF) communication unit associated with each wearable device of the one or more wearable devices for providing the one or more real-time insights to the user.

2 . The system as claimed in claim 1 , wherein the one or more real-time insights comprise at least one of: assistive navigation, scene understanding, real-time environmental interaction.

3 . The system as claimed in claim 1 , wherein the one or more wearable devices are selected from a group comprises at least one of: a pair of smart glasses, a headgear, a headband, a chest-mounted device, a wearable pendant, a smart watch, an wrist worn device, a cane, shoes, and earbuds.

4 . The system as claimed in claim 1 , wherein the plurality of sensors comprise at least one of:

a visual data capturing unit configured to capture the visual data in the environmental space;

an audio input unit ( 110 a ) configured to record ambient audio signals to generate the audio data for at least one of: speech recognition, sound classification, and environmental awareness;

an obstacle detection unit configured to generate the depth perception data by computing a distance between the user and the one or more obstacles for detecting the one or more obstacles in the environmental space; and

an inertial measurement unit (IMU) configured to determine acceleration and orientation of the user for generating the motion tracking data.

5 . The system as claimed in claim 1 , wherein the convolutional neural network (CNN) is optimized for the one or more edge devices as an on-device processing artificial intelligence model,

the on-device processing artificial intelligence model configured to provide the processed feedback data by combining at least one of: the visual data, the audio data, the depth perception data, and the motion tracking data, with natural language processing (NLP) procedures.

6 . The system as claimed in claim 1 , wherein

the one or more multi-modal artificial intelligence language models comprise at least one of: vision-language models, multi-modal transformer models, and artificial intelligence (AI)-driven fusion networks;

the one or more text recognition models comprise at least one of: optical character recognition (OCR), handwriting recognition models, and the one or more text to speech models;

the predictive motion model comprises at least one of: recurrent neural network (RNN)-based motion prediction models, Kalman filtering-based motion estimation models, and light detection and ranging (LiDAR)-based dynamic movement tracking;

the one or more conversational artificial intelligence models comprise at least one of:

the one or more large language models (LLMs), artificial intelligence (AI)-powered natural language understanding (NLU) modules, task-oriented dialogue models, and emotion-aware artificial intelligence (AI) models; and

the one or more speech recognition models comprise at least one of: end-to-end deep learning-based automatic speech recognition (ASR) models, noise-resistant speech-to-text models, multi-language voice recognition engines, and wake-word detection models.

7 . The system as claimed in claim 1 , wherein the fall detection module comprises the vision-based gait analysis configured to detect abnormal walking patterns of the user to predict the fall event.

8 . The system as claimed in claim 1 , wherein the artificial intelligence (AI) model optimization subsystem comprises:

a federated learning framework configured to train and update at least one of: the one or more machine learning models, the one or more artificial intelligence models, the one or more large language models (LLMs), the one or more conversational artificial intelligence models, the one or more speech recognition models, and the one or more text to speech models locally on the one or more edge devices while periodically synchronizing updates with the one or more servers; and

a model compression engine configured to dynamically condense size of at least one of: the one or more machine learning model and the one or more artificial intelligence models through at least one of: quantization and pruning to provide the optimized processing performance on the one or more edge devices.

9 . A method for providing one or more real-time insights to a user, comprising:

generating, by a plurality of sensors associated with each wearable device of the one or more wearable devices, real-time sensor data comprises at least one of: visual data, audio data, depth perception data, and motion tracking data;

controlling, by a first hardware processor associated with each wearable device of the one or more wearable devices, optimized electric power distribution from a power supply unit through a power management unit (PMU) to at least one of: the plurality of sensors, one or more haptic feedback units, one or more audio output units ( 110 b ) associated with the plurality of sensors, one or more input control units, and a radio frequency (RF) communication unit;

transmitting, by the first hardware processor, the real-time sensor data via the radio frequency (RF) communication unit to at least one of: one or more edge devices and one or more servers;

obtaining, by one or more second hardware processors associated with the one or more edge devices through a data obtaining subsystem, the real-time sensor data from each wearable device of the one or more wearable devices for on-device processing on the one or more edge devices;

preprocessing, by the one or more second hardware processors through a data pre-processing subsystem, the obtained real-time sensor data, including at least one of: denoising the visual data, altering contrast in the visual data, filtering irrelevant audio data, normalising the depth perception data;

analysing, by the one or more second hardware processors through a data processing subsystem, the real-time sensor data through at least one of: one or more machine learning models and one or more artificial intelligence models,

the analysing comprises:

recognising, by an object recognition module, one or more obstacles in the environmental space by processing at least one of: the visual data, the depth perception data, and the motion tracking data through a convolutional neural network (CNN), wherein the convolutional neural network comprises at least one of: an EfficientDet model, a MobileNet model, and a You Only Look Once (YOLO) model, for execution on the one or more edge devices to perform real-time object detection;

generating, by an interaction response module, one or more contextual responses for the user by processing the visual data, the audio data through at least one of: one or more multi-modal artificial intelligence language models and one or more text recognition models;

providing, by a navigation and obstacle detection module, the processed feedback data including at least one of: real-time haptic feedback and audio feedback, based on processing the depth perception data, the visual data through a predictive motion model, wherein the real-time haptic feedback comprises discrete vibration patterns produced corresponding to different types of processed feedback data, including one of: proximity of the one or more obstacles and navigation instructions, allowing the user to distinguish between the discrete vibration patterns; and

detecting, by a fall detection module, abrupt motion changes indicative of a fall event based on the motion tracking data and vision-based gait analysis, wherein the detecting comprises combining the motion tracking data, the visual data, and the depth perception data through a multi-sensor fusion procedure to alleviate false positives in detection of the abrupt motion changes;

obtaining, by one or more artificial intelligence (AI) data processing units (DPUs) associated with the one or more servers through a data acquisition subsystem, the pre-processed sensor data for the cloud-based processing in the one or more servers;

complex processing, by the one or more artificial intelligence (AI) data processing units (DPUs) through a cloud data processing subsystem, the sensor data to generate the processed feedback data through at least one of: the one or more machine learning models, one or more large language models (LLMs), one or more conversational artificial intelligence models, one or more speech recognition models, and one or more text to speech models;

optimising, by the one or more artificial intelligence (AI) data processing units (DPUs) through an artificial intelligence (AI) model optimization subsystem, at least one of: the one or more machine learning models, the one or more artificial intelligence models, the one or more large language models (LLMs), the one or more conversational artificial intelligence models, the one or more speech recognition models, and the one or more text to speech models, for providing an optimized processing performance on at least one of: the one or more edge devices and the one or more servers;

transmitting, by a notification subsystem, one or more alert notifications to one or more preregistered emergency contacts with location coordinates of the user upon detection of the abrupt motion changes;

transferring, by a feedback subsystem, the processed feedback data to the radio frequency (RF) communication unit associated with each wearable device of the one or more wearable devices; and

receiving, by the first hardware processor associated with each wearable device of the one or more wearable devices, the processed feedback data from at least one of: the one or more edge devices and the one or more servers through the radio frequency (RF) communication unit to provide the one or more real-time insights to the user.