IP Library › Granted Patent US 12,736,676
Granted Patent B2
US 12,736,676 · App. 17/884,406 · Granted Sep 15, 2026

System and method of capturing large scale scenes using wearable inertial measurement devices and light detection and ranging sensors

Inventors: Chenglu Wen (Xiamen, CN); Yudi Dai (Xiamen, CN); Lan Xu (Shanghai, CN); Cheng Wang (Xiamen, CN); Jingyi Yu (Shanghai, CN)
Assignees: Xiamen University; Shanghaitech University
G01S17/58G01S17/89G06T19/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,736,676
App. No.
17/884,406
Granted
Sep 15, 2026
Kind
B2
Abstract

Described herein are systems and methods of capturing motions of humans in a scene. A plurality of IMU devices and a LiDAR sensor are mounted on a human. IMU data is captured by the IMU devices and LiDAR data is captured by the LiDAR sensor. Motions of the human are estimated based on the IMU data and the LiDAR data. A three-dimensional scene map is built based on the LiDAR data. An optimization is performed to obtain optimized motions of the human and optimized scene map.

Claims (88)

1 . A method for capturing motions of humans in a scene, the method comprising:

mounting a plurality of IMU devices and a LiDAR sensor on a body part of a human;

obtaining IMU data captured by the IMU devices and LiDAR data captured by the LiDAR sensor;

estimating motions of the human based on the IMU data and the LiDAR data by aligning a LiDAR coordinate system of the LiDAR data and an IMU coordinate system of the IMU data, wherein the estimating the motions of the human comprises:

calibrating rigid offsets between the LiDAR coordinate system and the IMU coordinate system;

estimating motion of the LiDAR sensor from the LiDAR data to obtain a LiDAR trajectory;

aligning the LiDAR trajectory to the IMU coordinate system using the calibrated rigid offsets; and

estimating the motions of the human based on the IMU data and the LiDAR trajectory aligned to the IMU coordinate system;

building a three-dimensional scene map based on the LiDAR data; and

performing a joint optimization to obtain optimized motions of the human and optimized scene map.

2 . The method of claim 1 , further comprising:

jumping by the human during the step of obtaining IMU data captured by the IMU devices and the LiDAR data captured by the LiDAR sensor; and

synchronizing the LiDAR data and the IMU data based on a peak derived from the LiDAR data and a peak derived from the IMU data.

3 . The method of claim 1 , further comprising:

performing a graph-based optimization to fuse an LiDAR trajectory and an IMU trajectory, wherein the LiDAR trajectory comprises a movement of a center of the human derived from the LiDAR data, and the IMU trajectory comprises a movement of a center of the human derived from the IMU data.

4 . The method of claim 1 , wherein the joint optimization is based on a contact constraint and a sliding constraint.

5 . The method of claim 4 , wherein the contact constraint defines a contact loss as a distance from the body part derived from the motions of the human to a nearest surface derived from the three-dimensional scene map, and the sliding constraint defines a sliding loss as a distance between two successive body parts derived from the motions of the human, and the method further comprises:

minimizing a sum of the contact loss and the sliding loss with a gradient descent algorithm to iteratively optimize the motions of the human.

6 . The method of claim 5 , wherein the body part is a foot of the human, and the surface is a ground.

7 . The method of claim 1 , further comprising:

mounting a plurality of second IMU devices and a second LiDAR sensor on a second human;

obtaining second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor;

estimating motions of the second human based on the second IMU data;

building a three-dimensional second scene map based on the second LiDAR data;

fusing the three-dimensional scene map and the three-dimensional second scene map to obtain a combined scene map; and

performing an optimization to obtain optimized motions of the human and the second human in an optimized combined scene map.

8 . The method of claim 1 , further comprising:

creating a metaverse based on the optimized motions of the human and the optimized scene map.

9 . A system for capturing motions comprising:

a plurality of IMU devices to be worn by a human;

a LiDAR sensor to be mounted on a body part of the human;

a processor; and

a memory storing instructions that, when executed by the processor, cause the system to perform a method for capturing motions of humans in a scene comprising:

obtaining IMU data captured by the IMU devices and LiDAR data captured by the LiDAR sensor; and

estimating motions of the human based on the IMU data and the LiDAR data by aligning a LiDAR coordinate system of the LiDAR data and an IMU coordinate system of the IMU data, wherein the estimating the motions of the human comprises:

calibrating rigid offsets between the LiDAR coordinate system and the IMU coordinate system;

estimating motion of the LiDAR sensor from the LiDAR data to obtain a LiDAR trajectory;

aligning the LiDAR trajectory to the IMU coordinate system using the calibrated rigid offsets; and

estimating the motions of the human based on the IMU data and the LiDAR trajectory aligned to the IMU coordinate system.

10 . The system of claim 9 , further comprising: a L-shaped bracket configured to mount the LiDAR sensor on a hip of the human, wherein the LiDAR sensor and the IMU devices are configured to have a rigid transformation.

11 . The system of claim 9 , further comprising:

a wireless receiver coupled to the system, wherein the wireless receiver is configured to receive the IMU data captured by the IMU devices.

12 . The system of claim 9 , wherein the instructions, when executed, further cause the system to perform:

building a three-dimensional scene map based on the LiDAR data; and

performing an optimization to obtain optimized motions of the human and an optimized scene map.

13 . The system of claim 12 , wherein the instructions, when executed, further causes the system to perform:

perform a graph-based optimization to fuse an LiDAR trajectory and an IMU trajectory, wherein the LiDAR trajectory comprises a movement of a center of the human derived from the LiDAR data, and the IMU trajectory comprises a movement of a center of the human derived from the IMU data.

14 . The system of claim 12 , wherein the optimization is based on a contact constraint and a sliding constraint.

15 . The system of claim 14 , wherein the contact constraint defines a contact loss as a distance from the body part derived from the motions of the human to a nearest surface derived from the three-dimensional scene map, and the sliding constraint defines a sliding loss as a distance between two successive body parts derived from the motions of the human, and the method further comprises:

minimizing a sum of the contact loss and the sliding loss with a gradient descent algorithm to iteratively optimize the motions of the human.

16 . The system of claim 15 , wherein the body part is a foot of the human, and the surface is a ground.

17 . The system of claim 12 , wherein the method further comprises:

creating a metaverse based on the optimized motions of the human and the optimized scene map.

18 . The system of claim 9 , further comprising:

a plurality of second IMU devices configured to be worn by a second human;

a second LiDAR sensor configured to be mounted on the second human;

wherein the method further comprises:

obtaining second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor;

estimating motions of the second human based on the second IMU data;

building a three-dimensional second scene map based on the second LiDAR data;

fusing the three-dimensional scene map and the three-dimensional second scene map to obtain a combined scene map; and

performing an optimization to obtain optimized motions of the human and the second human in an optimized combined scene map.

19 . A method of optimizing motions of humans in a scene, the method comprising:

obtaining a three-dimensional scene map and motions of a human, wherein the three-dimensional scene map is obtained based on LiDAR data captured by a LiDAR sensor mounted on a body part of the human, and the motions of the human are obtained based on IMU data and the LiDAR data by aligning a LiDAR coordinate system of the LiDAR data and an IMU coordinate system of the IMU data, wherein obtaining the motions of the humans comprises:

calibrating rigid offsets between the LiDAR coordinate system and the IMU coordinate system;

estimating motion of the LiDAR sensor from the LiDAR data to obtain a LiDAR trajectory;

aligning the LiDAR trajectory to the IMU coordinate system using the calibrated rigid offsets; and

estimating the motions of the human based on the IMU data and the LiDAR trajectory aligned to the IMU coordinate system;

performing a graph-based optimization to fuse an LiDAR trajectory and an IMU trajectory; and

performing a joint optimization based on a plurality of physical constraints to obtain optimized motions of the human and an optimized scene map.

20 . The method of claim 19 , further comprising:

calibrating the three-dimensional scene map and the motions of the human.

21 . The method of claim 19 , further comprising:

synchronizing the three-dimensional scene map and the motion of the human.

22 . The method of claim 19 , wherein the LiDAR trajectory comprises a movement of a center of the human derived from the LiDAR data, and the IMU trajectory comprises a movement of a center of the human derived from the IMU data.

23 . The method of claim 19 , wherein the joint optimization is based on a contact constraint and a sliding constraint.

24 . The method of claim 23 , wherein the contact constraint defines a contact loss as a distance from the body part derived from the motions of the human to a nearest surface derived from the three-dimensional scene map, and the sliding constraint defines a sliding loss as a distance between two successive body parts derived from the motions of the human, and the method further comprises:

minimizing a sum of the contact loss and the sliding loss with a gradient descent algorithm to iteratively optimize the motions of the human.

25 . The method of claim 24 , wherein the body part is a foot of the human, and the surface is a ground.

26 . The method of claim 19 , further comprising:

mounting a plurality of second IMU devices and a second LiDAR sensor on a second human;

obtaining second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor;

estimating motions of the second human based on the second IMU data;

building a three-dimensional second scene map based on the second LiDAR data;

fusing the three-dimensional scene map and the three-dimensional second scene map to obtain a combined scene map; and

performing an optimization to obtain optimized motions of the human and the second human in an optimized combined scene map.

27 . The method of claim 19 , further comprising:

creating a metaverse based on the optimized motions of the human and the optimized scene map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2022
From: WEN, CHENGLU; DAI, YUDI; XU, LAN; WANG, CHENG; YU, JINGYI
To: XIAMEN UNIVERSITY; SHANGHAITECH UNIVERSITY
Reel/Frame 061130/0558 →
Priority Claims (1)
WO PCT/CN2022/078083 · Feb 25, 2022 · international
Continuity (2)
Continuation PCTCN2022079151 · Mar 3, 2022
Related Publication 20230273315A1 · Aug 31, 2023
References Cited (31)
US 10937178B1 · Srinivasan · 2021 [cited by applicant]
US 11537819B1 · Das · 2022 [cited by examiner]
US 11628855B1 · Pradhan et al. · 2023 [cited by applicant]
US 20020025053A1 · Lydecker · 2002 [cited by examiner]
US 20180072313A1 · Stenneth · 2018 [cited by applicant]
US 20180204338A1 · Narang · 2018 [cited by examiner]
US 20180217663A1 · Chandrasekhar et al. · 2018 [cited by applicant]
US 20190163968A1 · Hua et al. · 2019 [cited by applicant]
US 20200160559A1 · Urtasun et al. · 2020 [cited by applicant]
US 20200183007A1 · Nagashima · 2020 [cited by applicant]
US 20200226785A1 · Woods · 2020 [cited by examiner]
US 20210299873A1 · Badiozamani et al. · 2021 [cited by applicant]
US 20220066459A1 · Jain et al. · 2022 [cited by applicant]
US 20220143467A1 · Green · 2022 [cited by applicant]
US 20220321343A1 · Bahrami et al. · 2022 [cited by applicant]
CN 110140099A · 2019 [cited by applicant]
CN 110596683A · 2019 [cited by applicant]
CN 110873879A · 2020 [cited by applicant]
CN 111665512A · 2020 [cited by applicant]
CN 113466890A · 2020 [cited by applicant]
CN 112923934A · 2021 [cited by applicant]
WO 2020087041A1 · 2020 [cited by applicant]
Brubaker, Marcus Anthony. Physical Models of Human Motion for Estimation and Scene Analysis. Diss. 2012. (Year: 2012). [cited by examiner]
Patil, A.K.; Balasubramanyam, A.; Ryu, J.Y.; B N, P.K.; Chakravarthi, B.; Chai, Y.H. Fusion of Multiple Lidars and Inertial Sensors for the Real-Time Pose Tracking of Human Motion. Sensors 2020, 20, 5342. https://doi.or… [cited by examiner]
PCT International Search Report and the Written Opinion mailed Nov. 10, 2022, issued in related International Application No. PCT/CN2022/078083 (9 pages). [cited by applicant]
Non-Final Office Action dated Sep. 10, 2024, issued in related U.S. Appl. No. 17/884,273 (21 pages). [cited by applicant]
Liuyuan Deng et al. (Fusing Geometrical and Visual Information via Superpoints for the Semantic Segmentation of 3D Road Scenes, 2020) ( Year: 2020). [cited by applicant]
David Griffiths et al. (A Review on Deep Learning Techniques for 3D Sensed Data Classification, 2019) (Year: 2019). [cited by applicant]
PCT International Search Report and the Written Opinion mailed Nov. 29, 2022, issued in related International Application No. PCT/CN2022/079151 (10 pages). [cited by applicant]
First Office Action and Search Report dated May 14, 2026, issued in Chinese Patent Application No. 202280006556.4, with English machine translation (28 pages). [cited by applicant]
Yudi Dai et al., “HSC4D: Human-centered 4D Scene Capture in Large-scale Indoor-outdoor Space Using Wearable IMUs and LiDAR,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun.… [cited by applicant]