System and method of capturing large scale scenes using wearable inertial measurement devices and light detection and ranging sensors
Described herein are systems and methods of capturing motions of humans in a scene. A plurality of IMU devices and a LiDAR sensor are mounted on a human. IMU data is captured by the IMU devices and LiDAR data is captured by the LiDAR sensor. Motions of the human are estimated based on the IMU data and the LiDAR data. A three-dimensional scene map is built based on the LiDAR data. An optimization is performed to obtain optimized motions of the human and optimized scene map.
1 . A method for capturing motions of humans in a scene, the method comprising:
mounting a plurality of IMU devices and a LiDAR sensor on a body part of a human;
obtaining IMU data captured by the IMU devices and LiDAR data captured by the LiDAR sensor;
estimating motions of the human based on the IMU data and the LiDAR data by aligning a LiDAR coordinate system of the LiDAR data and an IMU coordinate system of the IMU data, wherein the estimating the motions of the human comprises:
calibrating rigid offsets between the LiDAR coordinate system and the IMU coordinate system;
estimating motion of the LiDAR sensor from the LiDAR data to obtain a LiDAR trajectory;
aligning the LiDAR trajectory to the IMU coordinate system using the calibrated rigid offsets; and
estimating the motions of the human based on the IMU data and the LiDAR trajectory aligned to the IMU coordinate system;
building a three-dimensional scene map based on the LiDAR data; and
performing a joint optimization to obtain optimized motions of the human and optimized scene map.
2 . The method of claim 1 , further comprising:
jumping by the human during the step of obtaining IMU data captured by the IMU devices and the LiDAR data captured by the LiDAR sensor; and
synchronizing the LiDAR data and the IMU data based on a peak derived from the LiDAR data and a peak derived from the IMU data.
3 . The method of claim 1 , further comprising:
performing a graph-based optimization to fuse an LiDAR trajectory and an IMU trajectory, wherein the LiDAR trajectory comprises a movement of a center of the human derived from the LiDAR data, and the IMU trajectory comprises a movement of a center of the human derived from the IMU data.
4 . The method of claim 1 , wherein the joint optimization is based on a contact constraint and a sliding constraint.
5 . The method of claim 4 , wherein the contact constraint defines a contact loss as a distance from the body part derived from the motions of the human to a nearest surface derived from the three-dimensional scene map, and the sliding constraint defines a sliding loss as a distance between two successive body parts derived from the motions of the human, and the method further comprises:
minimizing a sum of the contact loss and the sliding loss with a gradient descent algorithm to iteratively optimize the motions of the human.
6 . The method of claim 5 , wherein the body part is a foot of the human, and the surface is a ground.
7 . The method of claim 1 , further comprising:
mounting a plurality of second IMU devices and a second LiDAR sensor on a second human;
obtaining second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor;
estimating motions of the second human based on the second IMU data;
building a three-dimensional second scene map based on the second LiDAR data;
fusing the three-dimensional scene map and the three-dimensional second scene map to obtain a combined scene map; and
performing an optimization to obtain optimized motions of the human and the second human in an optimized combined scene map.
8 . The method of claim 1 , further comprising:
creating a metaverse based on the optimized motions of the human and the optimized scene map.
9 . A system for capturing motions comprising:
a plurality of IMU devices to be worn by a human;
a LiDAR sensor to be mounted on a body part of the human;
a processor; and
a memory storing instructions that, when executed by the processor, cause the system to perform a method for capturing motions of humans in a scene comprising:
obtaining IMU data captured by the IMU devices and LiDAR data captured by the LiDAR sensor; and
estimating motions of the human based on the IMU data and the LiDAR data by aligning a LiDAR coordinate system of the LiDAR data and an IMU coordinate system of the IMU data, wherein the estimating the motions of the human comprises:
calibrating rigid offsets between the LiDAR coordinate system and the IMU coordinate system;
estimating motion of the LiDAR sensor from the LiDAR data to obtain a LiDAR trajectory;
aligning the LiDAR trajectory to the IMU coordinate system using the calibrated rigid offsets; and
estimating the motions of the human based on the IMU data and the LiDAR trajectory aligned to the IMU coordinate system.
10 . The system of claim 9 , further comprising: a L-shaped bracket configured to mount the LiDAR sensor on a hip of the human, wherein the LiDAR sensor and the IMU devices are configured to have a rigid transformation.
11 . The system of claim 9 , further comprising:
a wireless receiver coupled to the system, wherein the wireless receiver is configured to receive the IMU data captured by the IMU devices.
12 . The system of claim 9 , wherein the instructions, when executed, further cause the system to perform:
building a three-dimensional scene map based on the LiDAR data; and
performing an optimization to obtain optimized motions of the human and an optimized scene map.
13 . The system of claim 12 , wherein the instructions, when executed, further causes the system to perform:
perform a graph-based optimization to fuse an LiDAR trajectory and an IMU trajectory, wherein the LiDAR trajectory comprises a movement of a center of the human derived from the LiDAR data, and the IMU trajectory comprises a movement of a center of the human derived from the IMU data.
14 . The system of claim 12 , wherein the optimization is based on a contact constraint and a sliding constraint.
15 . The system of claim 14 , wherein the contact constraint defines a contact loss as a distance from the body part derived from the motions of the human to a nearest surface derived from the three-dimensional scene map, and the sliding constraint defines a sliding loss as a distance between two successive body parts derived from the motions of the human, and the method further comprises:
minimizing a sum of the contact loss and the sliding loss with a gradient descent algorithm to iteratively optimize the motions of the human.
16 . The system of claim 15 , wherein the body part is a foot of the human, and the surface is a ground.
17 . The system of claim 12 , wherein the method further comprises:
creating a metaverse based on the optimized motions of the human and the optimized scene map.
18 . The system of claim 9 , further comprising:
a plurality of second IMU devices configured to be worn by a second human;
a second LiDAR sensor configured to be mounted on the second human;
wherein the method further comprises:
obtaining second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor;
estimating motions of the second human based on the second IMU data;
building a three-dimensional second scene map based on the second LiDAR data;
fusing the three-dimensional scene map and the three-dimensional second scene map to obtain a combined scene map; and
performing an optimization to obtain optimized motions of the human and the second human in an optimized combined scene map.
19 . A method of optimizing motions of humans in a scene, the method comprising:
obtaining a three-dimensional scene map and motions of a human, wherein the three-dimensional scene map is obtained based on LiDAR data captured by a LiDAR sensor mounted on a body part of the human, and the motions of the human are obtained based on IMU data and the LiDAR data by aligning a LiDAR coordinate system of the LiDAR data and an IMU coordinate system of the IMU data, wherein obtaining the motions of the humans comprises:
calibrating rigid offsets between the LiDAR coordinate system and the IMU coordinate system;
estimating motion of the LiDAR sensor from the LiDAR data to obtain a LiDAR trajectory;
aligning the LiDAR trajectory to the IMU coordinate system using the calibrated rigid offsets; and
estimating the motions of the human based on the IMU data and the LiDAR trajectory aligned to the IMU coordinate system;
performing a graph-based optimization to fuse an LiDAR trajectory and an IMU trajectory; and
performing a joint optimization based on a plurality of physical constraints to obtain optimized motions of the human and an optimized scene map.
20 . The method of claim 19 , further comprising:
calibrating the three-dimensional scene map and the motions of the human.
21 . The method of claim 19 , further comprising:
synchronizing the three-dimensional scene map and the motion of the human.
22 . The method of claim 19 , wherein the LiDAR trajectory comprises a movement of a center of the human derived from the LiDAR data, and the IMU trajectory comprises a movement of a center of the human derived from the IMU data.
23 . The method of claim 19 , wherein the joint optimization is based on a contact constraint and a sliding constraint.
24 . The method of claim 23 , wherein the contact constraint defines a contact loss as a distance from the body part derived from the motions of the human to a nearest surface derived from the three-dimensional scene map, and the sliding constraint defines a sliding loss as a distance between two successive body parts derived from the motions of the human, and the method further comprises:
minimizing a sum of the contact loss and the sliding loss with a gradient descent algorithm to iteratively optimize the motions of the human.
25 . The method of claim 24 , wherein the body part is a foot of the human, and the surface is a ground.
26 . The method of claim 19 , further comprising:
mounting a plurality of second IMU devices and a second LiDAR sensor on a second human;
obtaining second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor;
estimating motions of the second human based on the second IMU data;
building a three-dimensional second scene map based on the second LiDAR data;
fusing the three-dimensional scene map and the three-dimensional second scene map to obtain a combined scene map; and
performing an optimization to obtain optimized motions of the human and the second human in an optimized combined scene map.
27 . The method of claim 19 , further comprising:
creating a metaverse based on the optimized motions of the human and the optimized scene map.