IP Library › Granted Patent US 12,472,988
Granted Patent B2
US 12,472,988 · App. 18/426,901 · Granted Nov 18, 2025

Method and system for predicting gesture of subjects surrounding an autonomous vehicle

Inventors: Suresh Sundaram (Bangalore, IN); Nishant Bhattacharya (Boulder, CO); Yuvika Dev Sharma (Bangalore, IN)
Assignees: Wipro Limited; Indian Institute of Science
B60W60/00272B60W50/0097G06V20/58G06V40/28B60W2554/4041
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,472,988
App. No.
18/426,901
Granted
Nov 18, 2025
Kind
B2
Abstract

Embodiments of present disclosure relates to method and gesture prediction system for predicting gesture of subjects for controlling AV during navigation. The gesture prediction system obtains data from sensors and generates context information of environment of AV. The gesture prediction system generates parameters for sampling frames for determining gesture of subjects. The gesture prediction system estimates current pose, and subsequent poses by extrapolating current pose and subsequent poses using deep learning techniques. Further, the gesture prediction system infers current behaviour of subjects by classifying current pose and subsequent poses into one of predefined gesture categories. The gesture prediction system predicts gesture of subjects based on current behaviour, context information and location information of subjects for controlling AV during navigation. Thus, the present disclosure forecast gestures of subjects at a faster rate and controls the AV during the navigation.

Claims (42)

1 . A method of predicting a gesture of subjects for controlling an Autonomous Vehicle (AV) during navigation, the method comprising:

obtaining, by a gesture prediction system, data associated with a current environment of the AV from one or more sensors associated with the AV, wherein the data comprises a plurality of frames and location information of the AV;

generating, by the gesture prediction system, context information associated with the current environment of the AV based on the obtained data, wherein the context information comprises information associated with the current environment and information associated with one or more subjects present in the current environment;

generating, by the gesture prediction system, a plurality of parameters based on the location information and the context information for sampling the plurality of frames;

estimating, by the gesture prediction system, a current pose and one or more subsequent poses of each of one or more subjects from each sampled frames using a deep learning technique, wherein estimating the one or more subsequent poses comprising:

predicting the one or more subsequent poses to be performed by each of the one or more subjects based on one or more factors associated with corresponding previous poses of the one or more subjects using the deep learning technique;

inferring, by the gesture prediction system, a gesture category from a plurality of predefined gesture categories, for the one or more subsequent poses of each of the one or more subjects, based on corresponding current pose, the one or more subsequent poses and pre-existing gesture data; and

predicting, by the gesture prediction system, a gesture to be performed by each of the one or more subjects based on respective gesture category, the information associated with the one or more subjects, and the context information, for controlling the AV during navigation.

2 . The method as claimed in claim 1 , wherein the context information further comprises, number of objects present around the AV and information related to the objects, and wherein the information associated with the one or more subjects comprises a category of the one or more subjects, location of the one or more subjects, and movement information of the one or more subjects.

3 . The method as claimed in claim 1 , wherein the plurality of parameters comprises relational parameters and temporal parameters configured based on the context information and location information for sampling the plurality of frames by controlling number of the plurality of frames and a time interval between each of the plurality of frames, respectively.

4 . The method as claimed in claim 1 , wherein predicting the one or more subsequent poses of the one or more subjects comprises:

estimating, by the gesture prediction system, motion associated with the current pose and the one or more subsequent poses estimated corresponding to the one or more subjects; and

predicting, by the gesture prediction system, a subsequent pose of the one or more subjects by extrapolating the respective motion using the deep learning technique.

5 . The method as claimed in claim 1 , wherein the one or more factors for the one or more subjects comprise a number of one or more estimated poses and time interval between each of the one or more estimated poses.

6 . The method as claimed in claim 1 , wherein controlling the AV during navigation comprises:

determining, by the gesture prediction system, a navigational path for the AV dynamically based on a priority of the gesture associated with the one or more subjects surrounding the AV, and the location information.

7 . A gesture prediction system for predicting a gesture of subjects for controlling an Autonomous Vehicle (AV) during navigation, comprising:

a processor; and

a memory communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution, cause the processor to:

obtain data associated with a current environment of the AV from one or more sensors associated with the AV, wherein the data comprises a plurality of frames and location information of the AV;

generate context information associated with the current environment of the AV based on the obtained data, wherein the context information comprises information associated with the current environment and information associated with one or more subjects present in the current environment;

generate a plurality of parameters based on the location information and the context information for sampling the plurality of frames;

estimate a current pose and one or more subsequent poses of each of one or more subjects from each sampled frames using a deep learning technique, wherein the processor is configured to estimate the one or more subsequent poses by:

predicting the one or more subsequent poses to be performed by each of the one or more subjects based on one or more factors associated with corresponding previous poses of the one or more subjects using the deep learning technique;

infer a gesture category from a plurality of predefined gesture categories, for the one or more subsequent poses of each of the one or more subjects, based on corresponding current pose, the one or more subsequent poses and pre-existing gesture data; and

predict a gesture to be performed by each of the one or more subjects based on respective gesture category, the information associated with the one or more subjects, and the context information, for controlling the AV during navigation.

8 . The gesture prediction system as claimed in claim 7 , wherein the context information further comprises number of objects present around the AV and information related to the objects, and, wherein the information associated with the one or more subjects comprises a category of the one or more subjects, location of the one or more subjects, and movement information of the one or more subjects.

9 . The gesture prediction system as claimed in claim 7 , wherein the plurality of parameters comprises relational parameters and temporal parameters configured based on the context information and location information for sampling the plurality of frames by controlling number of the plurality of frames and a time interval between each of the plurality of frames, respectively.

10 . The gesture prediction system as claimed in claim 7 , wherein the processor is configured to predict the one or more subsequent poses of the one or more subjects by:

estimating motion associated with the current pose and the one or more subsequent poses estimated corresponding to each of the one or more subjects; and

predicting a subsequent pose of each of the one or more subjects by extrapolating the respective motion using the deep learning technique.

11 . The gesture prediction system as claimed in claim 7 , wherein the one or more factors for the one or more subjects comprise a number of one or more estimated poses and time interval between each of the one or more estimated poses.

12 . The gesture prediction system as claimed in claim 7 , wherein the processor controls the AV during navigation by:

determining a navigational path for the AV dynamically based on a priority of the gesture associated with the one or more subjects surrounding the AV, and the location information.

13 . A non-transitory computer readable medium including instruction stored thereon that when processed by at least one processor cause a gesture prediction system to perform operation comprising:

obtaining data associated with a current environment of the AV from one or more sensors associated with the AV, wherein the data comprises a plurality of frames and location information of the AV;

generating context information associated with the current environment of the AV based on the obtained data, wherein the context information comprises information associated with the current environment and information associated with one or more subjects present in the current environment;

generating a plurality of parameters based on the location information and the context information for sampling the plurality of frames;

estimating a current pose and one or more subsequent poses of each of one or more subjects from each sampled frames using a deep learning technique, wherein estimating the one or more subsequent poses comprising:

predicting the one or more subsequent poses to be performed by each of the one or more subjects based on one or more factors associated with corresponding previous poses of the one or more subjects using the deep learning technique;

inferring a gesture category from a plurality of predefined gesture categories, for the one or more subsequent poses of each of the one or more subjects, based on corresponding current pose, the one or more subsequent poses and pre-existing gesture data; and

predicting a gesture to be performed by each of the one or more subjects based on respective gesture category, the information associated with the one or more subjects, and the context information, for controlling the AV during navigation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2024
From: SUNDARAM, SURESH; BHATTACHARYA, NISHANT; SHARMA, YUVIKA DEV
To: WIPRO LIMITED
Reel/Frame 066347/0771 →
Priority Claims (1)
IN 202341057673 · Aug 28, 2023 · national
Continuity (1)
Related Publication 20250074475A1 · Mar 6, 2025
References Cited (14)
US 11577722B1 · Packer · 2023 [cited by examiner]
US 12240497B1 · Lin · 2025 [cited by examiner]
US 20170327112A1 · Yokoyama · 2017 [cited by examiner]
US 20180143644A1 · Li · 2018 [cited by examiner]
US 20190354194A1 · Wang et al. · 2019 [cited by applicant]
US 20200086879A1 · Lakshmi Narayanan · 2020 [cited by examiner]
US 20220148319A1 · Chan · 2022 [cited by examiner]
US 20220318560A1 · Kishon · 2022 [cited by examiner]
US 20230067485A1 · Willoughby · 2023 [cited by examiner]
US 20230150550A1 · Shi · 2023 [cited by examiner]
US 20230219597A1 · Cohen · 2023 [cited by examiner]
WO WO2024261695A1 · 2024 [cited by examiner]
Bhattacharya et al., “CGAP2: Context and Gap Aware Predictive Pose Framework for Early Detection of Gestures”, Indian Institute of Science, arxiv.org/pdf/2011.09216.pdf, Nov. 2020, 9 pages. [cited by applicant]
Kim et al., “Real-Time Human Pose Estimation and Gesture Recognition from Depth Images Using Superpixels and SVM Classifier”, Sensors, vol. 15, No. 6, https://doi.org/10.3390/s150612410, May 2015, 18 pages. [cited by applicant]