IP Library › Granted Patent US 12,456,204
Granted Patent B1
US 12,456,204 · App. 19/097,799 · Granted Oct 28, 2025

Computer vision-driven interactive full-body motion tracking

Inventors: Luís António Correia de Oliveira (Oporto, PT); Pedro Henrique Oliveira Santos (Oporto, PT); Gustavo Seixas De Sá Burmester (Oporto, PT); Ricardo Miguel Pontes Leonardo (Oporto, PT); Pedro Fillipe da Silva Rodrigues (Oporto, PT)
Assignee: SWORD HEALTH, S.A.
G06T7/20G06T5/90G06T7/50G06T7/70G06T2207/20208G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,204
App. No.
19/097,799
Granted
Oct 28, 2025
Kind
B1
Abstract

Systems and methods described herein relate to computer vision-driven full-body motion tracking using a portable computing device. The portable computing device has a first image sensor and a second image sensor to capture images of a full body of the user while the user is performing exercises. The images are captured while the portable computing device is positioned in a substantially vertical orientation. The images are processed to generate high dynamic range (HDR) image data and to determine depth information associated with the user. Motion tracking data is generated in real time, using the HRD image data and the depth information, while the user is performing an exercise. Real-time interactive feedback may be provided during exercises.

Claims (45)

1. One or more non-transitory machine-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

accessing images of a full body of a user captured by a first image sensor and a second image sensor of a portable computing device, the images captured while the portable computing device is positioned in a substantially vertical orientation and while the user is performing one or more exercises;

processing the images of the first image sensor and the second image sensor to:

generate high dynamic range (HDR) image data, and

determine depth information associated with the user; and

generating, in real time and while the user is performing the one or more exercises, motion tracking data based on the HDR image data and the depth information,

wherein generating the motion tracking data comprises generating pose estimation data by analyzing body landmarks of the user while the user performs the one or more exercises, the pose estimation data obtained by processing first input comprising the HDR image data generated from the images of the first image sensor and the second image sensor and second input comprising the depth information determined from the images of the first image sensor and the second image sensor.

2. The one or more non-transitory machine-readable storage media of claim 1 , wherein the first image sensor and the second image sensor simultaneously capture pairs of images at different exposure levels, and the HDR image data is generated for each pair of images by a process comprising:

aligning a first image from the first image sensor and a second image from the second image sensor to obtain aligned image data, and

generating a final image using the aligned image data.

3. The one or more non-transitory machine-readable storage media of claim 2 , wherein generating the final image comprises performing at least one of exposure fusion or tone mapping.

4. The one or more non-transitory machine-readable storage media of claim 2 , wherein at least one of aligning the first image with the second image or generating the final image comprises executing one or more machine learning models.

5. The one or more non-transitory machine-readable storage media of claim 1 , wherein determining the depth information comprises generating a depth map to estimate a distance between one or more body parts of the user and the portable computing device.

6. The one or more non-transitory machine-readable storage media of claim 1 , wherein determining the depth information comprises executing one or more machine learning models.

7. The one or more non-transitory machine-readable storage media of claim 1 , wherein processing of the images comprises performing stereoscopic correction using a first image from the first image sensor and a second image from the second image sensor.

8. The one or more non-transitory machine-readable storage media of claim 1 , wherein processing the first input comprising the HDR image data and the second input comprising the depth information comprises:

generating the pose estimation data based on the HDR image data; and

adjusting the pose estimation data through post-processing using the depth information.

9. The one or more non-transitory machine-readable storage media of claim 1 , wherein the images comprise a plurality of images captured over time at a fixed, substantially vertical axis.

10. The one or more non-transitory machine-readable storage media of claim 1 , wherein the one or more exercises comprise at least two of:

a first exercise performed in a substantially upright position in front of the portable computing device,

a second exercise performed in a supported position in front of the portable computing device, and

a third exercise performed in a floor-based position in front of the portable computing device,

wherein the portable computing device captures full-body motion of the user without any angular adjustment of the portable computing device between the at least two exercises.

11. The one or more non-transitory machine-readable storage media of claim 1 , wherein each of the first image sensor and the second image sensor has a focal length and sensor size selected to provide:

a diagonal field of view (DFOV) of greater than 110 degrees;

a horizontal field of view (HFOV) of greater than 90 degrees; and

a vertical field of view (VFOV) of greater than 70 degrees.

12. The one or more non-transitory machine-readable storage media of claim 11 , wherein the DFOV is greater than 130 degrees.

13. The one or more non-transitory machine-readable storage media of claim 11 , wherein the HFOV is greater than 100 degrees.

14. The one or more non-transitory machine-readable storage media of claim 11 , wherein the VFOV is greater than 100 degrees.

15. The one or more non-transitory machine-readable storage media of claim 1 , the operations further comprising:

displaying, on a display screen of the portable computing device and while the user is performing the one or more exercises, a user interface comprising a digital video feed representing the images together with at least one of motion tracking output or instructions for performing the one or more exercises.

16. The one or more non-transitory machine-readable storage media of claim 1 , wherein the first image sensor and the second image sensor are front-facing cameras that are spaced apart, along a longitudinal axis of the portable computing device, at a baseline of between 15 mm and 65 mm.

17. The one or more non-transitory machine-readable storage media of claim 1 , wherein the portable computing device is a substantially rectangular tablet computing device, and the first image sensor and the second image sensor are front-facing cameras positioned at or adjacent a long edge of the portable computing device.

18. The one or more non-transitory machine-readable storage media of claim 1 , wherein the images are captured while a support stand holds a front face of the portable computing device in the substantially vertical orientation.

19. A computer-implemented method comprising:

accessing images of a full body of a user captured by a first image sensor and a second image sensor of a portable computing device, the images captured while the portable computing device is positioned in a substantially vertical orientation and while the user is performing one or more exercises;

processing the images of the first image sensor and the second image sensor to:

generate high dynamic range (HDR) image data, and

determine depth information associated with the user; and

generating, in real time and while the user is performing the one or more exercises, motion tracking data based on the HDR image data and the depth information,

wherein generating the motion tracking data comprises generating pose estimation data by analyzing body landmarks of the user while the user performs the one or more exercises, the pose estimation data obtained by processing first input comprising the HDR image data generated from the images of the first image sensor and the second image sensor and second input comprising the depth information determined from the images of the first image sensor and the second image sensor.

20. The one or more non-transitory machine-readable storage media of claim 1 , wherein the body landmarks comprise three-dimensional (3D) landmark positions, and generating the motion tracking data comprises:

while the user performs the one or more exercises, executing a machine learning model that processes the HDR image data as the first input and the depth information as the second input to generate the 3D landmark positions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2025
From: DE OLIVEIRA, LUÍS ANTÓNIO CORREIA; SANTOS, PEDRO HENRIQUE OLIVEIRA; BURMESTER, GUSTAVO SEIXAS DE SÁ; LEONARDO, RICARDO MIGUEL PONTES; RODRIGUES, PEDRO FILLIPE DA SILVA
To: SWORD HEALTH, S.A.
Reel/Frame 070702/0567 →
References Cited (13)
US 20110299761A1 · Myokan · 2011 [cited by examiner]
US 20210001172A1 · Namboodiri · 2021 [cited by examiner]
US 20210346761A1 · Sterling · 2021 [cited by examiner]
US 20220030148A1 · Gruhlke · 2022 [cited by examiner]
US 20240112427A1 · Berliner · 2024 [cited by examiner]
US 20240257309A1 · Holland · 2024 [cited by examiner]
Wu, Po-Jung, Kuang-Tsu Shih, and Homer Chen. “Dual-camera HDR synthesis guided by long-exposure image.” 2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA). IEEE, 2016. … [cited by examiner]
Chari, Pradyumna, Anil Kumar Vadathya, and Kaushik Mitra. “Optimal HDR and depth from dual cameras.” arXiv preprint arXiv:2003.05907 (2020). (Year: 2020). [cited by examiner]
“Depth Map from Stereo Images”, OpenCV.org, [Online]. Accessed online Mar. 28, 2025, Retrieved from the Internet: <https://docs.opencv.org/4.x/dd/d53/tutorial_py depthmap.html>, 2 pgs. [cited by applicant]
“MediaPipe BlazePose GHUM 3D”, Model Card, Proceedings of FAT Conference ACM, New York, NY, USA, [Online]. Retrieved from the Internet: <https://developers.google.com/static/ml-kit/images/vision/pose-detection/pose_mode… [cited by applicant]
“MoveNet.SinglePose”, Tensorflow, [Online]. Accessed online Mar. 28, 2025, Retrieved from the Internet: <https://storage.googleapis.com/movenet/MoveNet.SinglePose%20Model%20Card.pdf>. [cited by applicant]
D'Eusanio, Andrea, et al., “RefiNet: 3D Human Pose Refinement with Depth Maps”, 2020 25th International Conference on Pattern Recognition (ICPR), (Jan. 2021), 9 pages. [cited by applicant]
Yao, Yao, et al., “MVSNet Depth Inference for Unstructured Multi-view Stereo”, Proceedings of the European Conference on Computer Vision (ECCV), (Apr. 2018), 17 pages. [cited by applicant]