IP Library Granted Patent US 12681564
Granted Patent B2
US 12681564 · App. 19/047,844 · Granted Jul 14, 2026

Safety boundary generation

Inventors: Senbo Wang (Shenzhen, CN); Xiaqiang Dai (Shenzhen, CN); Ziqi Li (Shenzhen, CN); Pan Ji (Shenzhen, CN); Hongdong Li (Shenzhen, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G06F3/011G06T19/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12681564
App. No.
19/047,844
Granted
Jul 14, 2026
Kind
B2
Abstract

In a safety boundary generation method, an image of a surrounding environment captured by a camera in a head-mounted display device is received. Three-dimensional reconstruction is performed based on the image of the surrounding environment to obtain a three-dimensional environment model. The three-dimensional environment model indicates a three-dimensional representation of the surrounding environment of a user that wears the head-mounted display device. Ground detection on the three-dimensional environment model is performed to obtain a plurality of ground boundary points. Each of the plurality of ground boundary points is an intersection point between a ground plane and an obstacle in the surrounding environment. A safety boundary is determined based on the plurality of ground boundary points. A region enclosed by the safety boundary is a safety region for the user to interact with the head-mounted display device. Apparatus and non-transitory computer-readable storage medium counterpart embodiments are also contemplated.

Claims (83)

1 . A safety boundary generation method, comprising:

receiving an image of a surrounding environment captured by a camera in a head-mounted display device;

performing three-dimensional reconstruction based on the image of the surrounding environment to obtain a three-dimensional environment model, the three-dimensional environment model indicating a three-dimensional representation of the surrounding environment of a user that wears the head-mounted display device;

performing, by processing circuitry, ground detection on the three-dimensional environment model to determine a ground plane from the three-dimensional environment model;

determining a plurality of ground boundary points, each of the plurality of ground boundary points being an intersection point between the ground plane and an obstacle in the surrounding environment; and

determining a safety boundary based on the plurality of ground boundary points, a region enclosed by the safety boundary being a safety region for the user to interact with the head-mounted display device.

2 . The method according to claim 1 , wherein the performing the three-dimensional reconstruction comprises:

initializing a stereoscopic space based on the image of the surrounding environment;

determining a signed distance field (SDF) value for each voxel of a plurality of voxels in the stereoscopic space based on a distance from the respective voxel to the camera and a depth value of a voxel projection point on a surface of the obstacle, the voxel projection point being an intersection point between a three-dimensional line from the voxel to an optical center of the camera and the surface of the obstacle, the SDF value representing a positional relationship between the respective voxel and the surface of the obstacle; and

performing the three-dimensional reconstruction based on the SDF values of the plurality of voxels to obtain the three-dimensional environment model.

3 . The method according to claim 2 , wherein the performing the three-dimensional reconstruction comprises:

performing a three-dimensional reconstruction calculation based on a marching cubes (MC) algorithm, to obtain an initial three-dimensional environment model;

performing repair processing on the initial three-dimensional environment model to obtain a repaired three-dimensional environment model; and

performing simplification processing on the repaired three-dimensional environment model in which a plurality of redundant patches is removed to obtain the three-dimensional environment model.

4 . The method according to claim 3 , wherein the performing the repair processing on the initial three-dimensional environment model comprises:

dividing the initial three-dimensional environment model into a plurality of patches, each of the plurality of patches being obtained by connecting a plurality of connection points;

removing duplicate patches and independent patches from the plurality of patches to obtain remaining patches, each independent patch having no connection relationship with any of the other patches;

removing duplicate points and independent points from the plurality of connection points to obtain remaining connection points, each independent point being located outside the plurality of patches; and

reconstructing connection relationships between the remaining patches and the remaining connection points to obtain the repaired three-dimensional environment model.

5 . The method according to claim 4 , wherein the performing the simplification processing on the repaired three-dimensional environment model comprises:

removing a first reconstructed patch in the repaired three-dimensional environment model when a total quantity of connections of the first reconstructed patch is less than a connection threshold, the first reconstructed patch being obtained via the repair processing;

removing a second reconstructed patch when a distance between the second reconstructed patch and the camera is greater than a distance threshold;

removing a plurality of reconstructed patches belonging to an independent region when a region volume of the independent region is less than a volume threshold; and

determining the repaired three-dimensional environment model in which the plurality of reconstructed patches is removed as the three-dimensional environment model.

6 . The method according to claim 2 , wherein the determining the SDF value for each voxel of the plurality of voxels comprises:

determining an i th frame initial SDF value for each voxel based on a difference between an i th distance from the respective voxel to the camera and an i th depth value of the voxel projection point indicated by an i th frame of a depth map, i being a positive integer; and

updating the i th frame initial SDF value based on an i th frame of update weight corresponding to the respective voxel and an (i−1) th frame of SDF value to obtain an i th frame of SDF value, the i th frame of update weight being determined based on an angle between the three-dimensional line from the respective voxel to the optical center of the camera and an i th line-of-sight direction of the camera.

7 . The method according to claim 2 , wherein the initializing the stereoscopic space comprises:

generating a depth map corresponding to the surrounding environment based on the image of the surrounding environment; and

initializing the stereoscopic space based on a maximum depth value in the depth map.

8 . The method according to claim 1 , wherein the performing the ground detection on the three-dimensional environment model comprises:

iteratively determining a ground plane based on a direction of gravity and a plurality of three-dimensional points in the three-dimensional environment model;

identifying a plurality of ground points in the three-dimensional environment model based on the ground plane; and

establishing a connection relationship between the plurality of ground points to obtain the plurality of ground boundary points on the ground plane.

9 . The method according to claim 8 , wherein the iteratively determining the ground plane comprises:

establishing an i th plane based on the direction of gravity as a normal vector and an i th three-dimensional point in the three-dimensional environment model, i being a positive integer;

verifying the i th plane based on a plurality of remaining three-dimensional points in the three-dimensional environment model to obtain a quantity of three-dimensional points conforming to the i th plane, the remaining three-dimensional points being three-dimensional points in the three-dimensional environment model other than the i th three-dimensional point; and

determining a plane corresponding to the largest quantity of conforming three-dimensional points as the ground plane.

10 . The method according to claim 1 , wherein the determining the safety boundary based on the plurality of ground boundary points comprises:

establishing an occupancy grid map including a plurality of grids corresponding to the ground plane;

performing boundary filling in the occupancy grid map based on coordinates of the plurality of ground boundary points to obtain a plurality of occupied grids; and

performing curve fitting based on centers of gravity of the occupied grids to obtain the safety boundary.

11 . The method according to claim 10 , further comprising:

determining a prompt region based on a plurality of unoccupied grids adjacent to the plurality of occupied grids, a quantity of grids between each unoccupied grid and any occupied grid being less than a quantity threshold; and

providing a safety prompt to the user when the user enters the prompt region, the safety prompt indicating the user is approaching the safety boundary.

12 . The method according to claim 1 , wherein the camera is a multi-lens camera or a depth camera.

13 . An apparatus, comprising:

processing circuitry configured to:

receive an image of a surrounding environment captured by a camera in a head-mounted display device;

perform three-dimensional reconstruction based on the image of the surrounding environment to obtain a three-dimensional environment model, the three-dimensional environment model indicating a three-dimensional representation of the surrounding environment of a user that wears the head-mounted display device;

perform ground detection on the three-dimensional environment model to determine a ground plane from the three-dimensional environment model;

determine a plurality of ground boundary points, each of the plurality of ground boundary points being an intersection point between the ground plane and an obstacle in the surrounding environment; and

determine a safety boundary based on the plurality of ground boundary points, a region enclosed by the safety boundary being a safety region for the user to interact with the head-mounted display device.

14 . The apparatus according to claim 13 , wherein the processing circuitry is configured to:

initialize a stereoscopic space based on the image of the surrounding environment;

determine a signed distance field (SDF) value for each voxel of a plurality of voxels in the stereoscopic space based on a distance from the respective voxel to the camera and a depth value of a voxel projection point on a surface of the obstacle, the voxel projection point being an intersection point between a three-dimensional line from the voxel to an optical center of the camera and the surface of the obstacle, the SDF value representing a positional relationship between the respective voxel and the surface of the obstacle; and

perform the three-dimensional reconstruction based on the SDF values of the plurality of voxels to obtain the three-dimensional environment model.

15 . The apparatus according to claim 14 , wherein the processing circuitry is configured to:

perform a three-dimensional reconstruction calculation based on a marching cubes (MC) algorithm, to obtain an initial three-dimensional environment model;

perform repair processing on the initial three-dimensional environment model to obtain a repaired three-dimensional environment model; and

perform simplification processing on the repaired three-dimensional environment model in which a plurality of redundant patches is removed to obtain the three-dimensional environment model.

16 . The apparatus according to claim 15 , wherein the processing circuitry is configured to:

divide the initial three-dimensional environment model into a plurality of patches, each of the plurality of patches being obtained by connecting a plurality of connection points;

remove duplicate patches and independent patches from the plurality of patches to obtain remaining patches, each independent patch having no connection relationship with any of the other patches;

remove duplicate points and independent points from the plurality of connection points to obtain remaining connection points, each independent point being located outside the plurality of patches; and

reconstruct connection relationships between the remaining patches and the remaining connection points to obtain the repaired three-dimensional environment model.

17 . The apparatus according to claim 16 , wherein the processing circuitry is configured to:

remove a first reconstructed patch in the repaired three-dimensional environment model when a total quantity of connections of the first reconstructed patch is less than a connection threshold, the first reconstructed patch being obtained via the repair processing;

remove a second reconstructed patch when a distance between the second reconstructed patch and the camera is greater than a distance threshold;

remove a plurality of reconstructed patches belonging to an independent region when a region volume of the independent region is less than a volume threshold; and

determine the repaired three-dimensional environment model in which the plurality of reconstructed patches is removed as the three-dimensional environment model.

18 . The apparatus according to claim 14 , wherein the processing circuitry is configured to:

determine an i th frame initial SDF value for each voxel based on a difference between an i th distance from the respective voxel to the camera and an i th depth value of the voxel projection point indicated by an i th frame of a depth map, i being a positive integer; and

update the i th frame initial SDF value based on an i th frame of update weight corresponding to the respective voxel and an (i−1) th frame of SDF value to obtain an i th frame of SDF value, the i th frame of update weight being determined based on an angle between the three-dimensional line from the respective voxel to the optical center of the camera and an i th line-of-sight direction of the camera.

19 . The apparatus according to claim 14 , wherein the processing circuitry is configured to:

generate a depth map corresponding to the surrounding environment based on the image of the surrounding environment; and

initialize the stereoscopic space based on a maximum depth value in the depth map.

20 . A non-transitory computer-readable storage medium storing instructions which, when executed by a processor, cause the processor to perform:

receiving an image of a surrounding environment captured by a camera in a head-mounted display device;

performing three-dimensional reconstruction based on the image of the surrounding environment to obtain a three-dimensional environment model, the three-dimensional environment model indicating a three-dimensional representation of the surrounding environment of a user that wears the head-mounted display device;

performing ground detection on the three-dimensional environment model to determine a ground plane from the three-dimensional environment model;

determining a plurality of ground boundary points, each of the plurality of ground boundary points being an intersection point between the ground plane and an obstacle in the surrounding environment; and

determining a safety boundary based on the plurality of ground boundary points, a region enclosed by the safety boundary being a safety region for the user to interact with the head-mounted display device.