IP Library › Granted Patent US 12,725,304
Granted Patent B2
US 12,725,304 · App. 18/812,869 · Granted Sep 1, 2026

Multi-modal full body pose tracking

Inventors: Jinka Sai Sagar (Hyderabad, IN); Vishnu Menon (Thrissur, IN); Kinal Mehta (Pune, IN); Chiranjib Choudhuri (Bangalore, IN); Srenivas Varadarajan (Bangalore, IN); Ajit Deepak Gupte (Bangalore, IN)
Assignee: QUALCOMM Incorporated
G06T7/74G06F3/0346G06T7/246G06T2207/20221G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,304
App. No.
18/812,869
Granted
Sep 1, 2026
Kind
B2
Abstract

Techniques and systems are provided for pose prediction. For instance, a process can include combining image features detected from an obtained image with estimated image features to generate combined features; generating temporally encoded features by temporally encoding the combined features; combining detected motion tracking information with estimated motion tracking information to generate combined motion tracking information; generating temporally encoded motion tracking information by temporally encoding the combined motion tracking information; generating spatially encoded multi-modal information by spatially encoding the temporally encoded features and the temporally encoded motion tracking information; and predicting a body pose by regressing the spatially encoded multi-modal information.

Claims (46)

1 . An apparatus for pose prediction, comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

combine image features detected from an obtained image with estimated image features to generate combined features;

generate temporally encoded features by temporally encoding the combined features;

combine detected motion tracking information with estimated motion tracking information to generate combined motion tracking information;

generate temporally encoded motion tracking information by temporally encoding the combined motion tracking information;

generate spatially encoded multi-modal information by spatially encoding the temporally encoded features and the temporally encoded motion tracking information; and

predict a body pose by regressing the spatially encoded multi-modal information.

2 . The apparatus of claim 1 , wherein the at least one processor is configured to:

generate an estimated image feature based on the spatially encoded multi-modal information; and

generate an estimated motion tracking information based on the spatially encoded multi-modal information.

3 . The apparatus of claim 2 , wherein, to combine the image features and the estimated image features, the at least one processor is configured to:

determine an image feature is missing from the image features detected from the obtained image; and

blend an estimated image feature corresponding with the missing image feature with a previous image feature.

4 . The apparatus of claim 3 , wherein the estimated image feature and the previous image feature are blended using an exponential moving average combiner.

5 . The apparatus of claim 2 , wherein, to combine the detected motion tracking information with the estimated motion tracking information, the at least one processor is configured to:

determine that motion tracking information is missing from the detected motion tracking information; and

blend the estimated motion tracking information corresponding with the missing motion tracking information with the detected motion tracking information.

6 . The apparatus of claim 5 , wherein the estimated motion tracking information are blended with the detected motion tracking information using an exponential moving average combiner.

7 . The apparatus of claim 1 , wherein, to regress the spatially encoded multi-modal information to predict a body pose, the at least one processor is configured to regress the spatially encoded multi-modal information to a skeletal pose.

8 . The apparatus of claim 1 , wherein the detected motion tracking information is received from at least one of a head-mounted display or a handheld controller.

9 . The apparatus of claim 1 , wherein the image features are encoded into a first multi-dimensional matrix, and wherein the detected motion tracking information are encoded into a second multi-dimensional matrix.

10 . The apparatus of claim 1 , wherein the detected motion tracking information comprises 6 degrees of freedom (6 DoF) information.

11 . The apparatus of claim 1 , wherein the estimated image features are estimated based on a previous image, and wherein the at least one processor is configured to output the body pose.

12 . The apparatus of claim 1 , wherein the detected motion tracking information comprises global information, wherein the image features provide local information, and wherein the spatially encoded multi-modal information fuses the global information and local information.

13 . A method for pose prediction, comprising:

combining image features detected from an obtained image with estimated image features to generate combined features;

generating temporally encoded features by temporally encoding the combined features;

combining detected motion tracking information with estimated motion tracking information to generate combined motion tracking information;

generating temporally encoded motion tracking information by temporally encoding the combined motion tracking information;

generating spatially encoded multi-modal information by spatially encoding the temporally encoded features and the temporally encoded motion tracking information; and

predicting a body pose by processing the spatially encoded multi-modal information.

14 . The method of claim 13 , further comprising:

generating an estimated image feature based on the spatially encoded multi-modal information; and

generating an estimated motion tracking information based on the spatially encoded multi-modal information.

15 . The method of claim 14 , wherein combining the image features and the estimated image features comprises:

determining an image feature is missing from the image features detected from the obtained image; and

blending an estimated image feature corresponding with the missing image feature with a previous image feature.

16 . The method of claim 14 , wherein combining the detected motion tracking information with the estimated motion tracking information comprises:

determining that motion tracking information is missing from the detected motion tracking information; and

blending the estimated motion tracking information corresponding with the missing motion tracking information with the detected motion tracking information.

17 . The method of claim 16 , wherein the estimated motion tracking information are blended with the detected motion tracking information using an exponential moving average combiner.

18 . The method of claim 13 , wherein regressing the spatially encoded multi-modal information to predict a body pose comprises regressing the spatially encoded multi-modal information to a skeletal pose.

19 . The method of claim 13 , wherein the detected motion tracking information is received from at least one of a head-mounted display or a handheld controller.

20 . The method of claim 13 , wherein the image features are encoded into a first multi-dimensional matrix, and wherein the detected motion tracking information are encoded into a second multi-dimensional matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2024
From: SAI SAGAR, JINKA; MENON, VISHNU; MEHTA, KINAL; CHOUDHURI, CHIRANJIB; VARADARAJAN, SRENIVAS; GUPTE, AJIT DEEPAK
To: QUALCOMM INCORPORATED
Reel/Frame 068619/0427 →
Continuity (1)
Related Publication 20260057546A1 · Feb 26, 2026
References Cited (18)
US 11507203B1 · Bosworth · 2022 [cited by examiner]
US 12118733B2 · Raghoebardajal · 2024 [cited by examiner]
US 12189867B2 · Lau · 2025 [cited by examiner]
US 12303255B1 · Schnan Mastronardi · 2025 [cited by examiner]
US 12367670B2 · Dasgupta · 2025 [cited by examiner]
US 20140176436A1 · Raffa · 2014 [cited by examiner]
US 20180053056A1 · Rabinovich · 2018 [cited by examiner]
US 20190012806A1 · Lehmann · 2019 [cited by examiner]
US 20190026904A1 · Chen · 2019 [cited by examiner]
US 20210035326A1 · Koike · 2021 [cited by examiner]
US 20210089116A1 · Erivantcev · 2021 [cited by examiner]
US 20210090284A1 · Ning · 2021 [cited by examiner]
US 20220157016A1 · Sharma · 2022 [cited by examiner]
US 20230051704A1 · Luan · 2023 [cited by examiner]
US 20240029358A1 · Sharma · 2024 [cited by examiner]
US 20240281996A1 · Nandipati · 2024 [cited by examiner]
US 20240296585A1 · Lee · 2024 [cited by examiner]
US 20250069259A1 · Sun · 2025 [cited by examiner]