IP Library › Granted Patent US 12,456,210
Granted Patent B2
US 12,456,210 · App. 17/893,007 · Granted Oct 28, 2025

Scene contour recognition in video based on depth information

Inventors: Shenghao Zhang (Shenzhen, CN); Yonggen Ling (Shenzhen, CN); Wanchao Chi (Shenzhen, CN); Yu Zheng (Shenzhen, CN); Xinyang Jiang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06T7/55G06V10/44G06V20/10G06T2207/10024G06T2207/10028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,210
App. No.
17/893,007
Filed
Aug 22, 2022
Granted
Oct 28, 2025
Kind
B2
Art Unit
2674
USPC
382/103
Abstract

A scene contour recognition method is provided. In the method, a plurality of scene images of an environment is obtained. Three-dimensional information of a target plane in the plurality of scene images is determined based on depth information for each of the plurality of scene images, The target plane corresponds to a target object in the plurality of scene images. A three-dimensional contour corresponding to the target object is generated by fusing the target plane in each of the plurality of scene images based on the three-dimensional information of the target plane in each of the plurality of scene images. A contour diagram of the target object is generated by projecting the three-dimensional contour onto a two-dimensional plane.

Claims (82)

1. A scene contour recognition method, the method comprising:

obtaining a plurality of scene images of an environment;

determining three-dimensional (3D) information of a target plane in the plurality of scene images based on depth information for each of the plurality of scene images, the target plane corresponding to a target object in the plurality of scene images, the 3D information of the target plane including a plane equation of the target plane and 3D coordinates corresponding to points on the target plane;

generating a 3D contour corresponding to the target object by fusing the target plane in each of the plurality of scene images based on the 3D information of the target plane in each of the plurality of scene images; and

generating a contour diagram of the target object by projecting the 3D contour onto a two-dimensional (2D) plane, wherein

the determining the 3D information of the target plane includes:

detecting the target plane in the plurality of scene images with plane fitting based on the depth information of the plurality of scene images; and

determining the 3D coordinates corresponding to the points on the target plane, and the plane equation of the target plane in a world coordinate system for the target plane in the plurality of scene images.

2. The method according to claim 1 , wherein the generating the 3D contour comprises:

determining 3D point information of edge feature points on the target plane based on the 3D information of the target plane for each of the plurality of scene images; and

generating the 3D contour corresponding to the target object by fusing the target plane in each of the plurality of scene images based on the 3D point information of the edge feature points on the target plane in each of the plurality of scene images.

3. The method according to claim 2 , wherein the generating the 3D contour corresponding to the target object comprises:

determining a target coordinate system based on a position of an image capturing apparatus at a current time, the image capturing apparatus being configured to capture the plurality of scene images;

determining target coordinates of the edge feature points in the target coordinate system according to the 3D point information of the edge feature points on the target plane, for the target plane in each of the plurality of scene images based on the target coordinate system; and

generating the 3D contour corresponding to the target object by fusing the target plane in each of the plurality of scene images based on the target coordinates of the edge feature points on the target plane in each of the plurality of scene images.

4. The method according to claim 1 , further comprising:

determining an equation corresponding to the target plane based on the 3D information of the target plane;

determining a center point of the target plane based on the equation corresponding to the target plane; and

obtaining a 3D contour of an adjusted target plane by adjusting the target plane based on a distance between an optical center of an image capturing apparatus configured to capture the plurality of scene images and the center point.

5. The method according to claim 4 , wherein the obtaining the 3D contour of the adjusted target plane comprises:

determining an optimization weight of the 3D contour based on the distance between the optical center of the image capturing apparatus and the center point, the optimization weight being inversely proportional to the distance;

determining an optimization parameter corresponding to the target plane based on the optimization weight;

adjusting the target plane based on the optimization parameter; and

obtaining the 3D contour of the adjusted target plane.

6. The method according to claim 1 , wherein the generating the contour diagram of the target object comprises:

detecting discrete points distributed on each plane in the 3D contour, and 3D coordinates corresponding to the discrete points;

generating 2D coordinates corresponding to the discrete points by performing dimension reduction on the 3D coordinates corresponding to the discrete points; and

generating the contour diagram of the target object by combining the 2D coordinates corresponding to the discrete points.

7. The method according to claim 1 , wherein after the generating the contour diagram of the target object, the method further comprises:

detecting a contour edge corresponding to the contour diagram based on discrete points in the contour diagram;

determining a contour range corresponding to the contour diagram based on the contour edge; and

obtaining an optimized contour diagram by eliminating discrete points outside the contour range based on the contour range.

8. The method according to claim 1 , wherein the obtaining the plurality of scene images comprises:

obtaining photographing parameters including a photographing period; and

capturing the plurality of scene images corresponding to each traveling position based on the photographing period in a traveling process, the plurality of scene images including a depth map and a color map.

9. The method according to claim 1 , wherein after the generating the contour diagram of the target object, the method further comprises:

determining an obtaining manner of a robotic device to obtain the target object based on a position of the target object corresponding to the contour diagram in a current traveling scenario; and

obtaining the target object based on the obtaining manner.

10. The method according to claim 9 , further comprising:

detecting a contour diagram corresponding to a target placement position of the target object;

determining a placement manner of the robotic device to place the target object based on the contour diagram corresponding to the target placement position; and

placing the target object at the target placement position based on the placement manner.

11. The method according to claim 1 , wherein after the generating the contour diagram of the target object, the method further comprises:

determining a traveling manner in a current traveling scenario based on the contour diagram, the traveling manner including at least one of a traveling direction, a traveling height, and a traveling distance; and

performing a traveling task based on the traveling manner.

12. A scene contour recognition apparatus, comprising:

processing circuitry configured to:

obtain a plurality of scene images of an environment;

determine three-dimensional (3D) information of a target plane in the plurality of scene images based on depth information for each of the plurality of scene images, the target plane corresponding to a target object in the plurality of scene images, the 3D information of the target plane including a plane equation of the target plane and 3D coordinates corresponding to points on the target plane;

generate a 3D contour corresponding to the target object by fusing the target plane in each of the plurality of scene images based on the 3D information of the target plane in each of the plurality of scene images; and

generate a contour diagram of the target object by projecting the 3D contour onto a two-dimensional (2D) plane, wherein

the determination of the 3D information of the target plane includes:

detecting the target plane in the plurality of scene images with plane fitting based on the depth information of the plurality of scene images; and

determining the 3D coordinates corresponding to the points on the target plane, and the plane equation of the target plane in a world coordinate system for the target plane in the plurality of scene images.

13. The scene contour recognition apparatus according to claim 12 , wherein the processing circuitry is configured to:

determine 3D point information of edge feature points on the target plane based on the 3D information of the target plane for each of the plurality of scene images; and

generate the 3D contour corresponding to the target object by fusing the target plane in each of the plurality of scene images based on the 3D point information of the edge feature points on the target plane in each of the plurality of scene images.

14. The scene contour recognition apparatus according to claim 13 , wherein the processing circuitry is configured to:

determine a target coordinate system based on a position of an image capturing apparatus at a current time, the image capturing apparatus being configured to capture the plurality of scene images;

determine target coordinates of the edge feature points in the target coordinate system according to the 3D point information of the edge feature points on the target plane, for the target plane in each of the plurality of scene images based on the target coordinate system; and

generate the 3D contour corresponding to the target object by fusing the target plane in each of the plurality of scene images based on the target coordinates of the edge feature points on the target plane in each of the plurality of scene images.

15. The scene contour recognition apparatus according to claim 12 , wherein the processing circuitry is configured to:

determine an equation corresponding to the target plane based on the 3D information of the target plane;

determine a center point of the target plane based on the equation corresponding to the target plane; and

obtain a 3D contour of an adjusted target plane by adjusting the target plane based on a distance between an optical center of an image capturing apparatus configured to capture the plurality of scene images and the center point.

16. The scene contour recognition apparatus according to claim 15 , wherein the processing circuitry is configured to:

determine an optimization weight of the 3D contour based on the distance between the optical center of the image capturing apparatus and the center point, the optimization weight being inversely proportional to the distance;

determine an optimization parameter corresponding to the target plane based on the optimization weight;

adjust the target plane based on the optimization parameter; and

obtain the 3D contour of the adjusted target plane.

17. The scene contour recognition apparatus according to claim 12 , wherein the processing circuitry is configured to:

detect discrete points distributed on each plane in the 3D contour, and 3D coordinates corresponding to the discrete points;

generate 2D coordinates corresponding to the discrete points by performing dimension reduction on the 3D coordinates corresponding to the discrete points; and

generate the contour diagram of the target object by combining the 2D coordinates corresponding to the discrete points.

18. A non-transitory computer-readable storage medium, storing instructions which when executed by a processor cause the processor to perform:

obtaining a plurality of scene images of an environment;

determining three-dimensional (3D) information of a target plane in the plurality of scene images based on depth information for each of the plurality of scene images, the target plane corresponding to a target object in the plurality of scene images, the 3D information of the target plane including a plane equation of the target plane and 3D coordinates corresponding to points on the target plane;

generating a 3D contour corresponding to the target object by fusing the target plane in each of the plurality of scene images based on the 3D information of the target plane in each of the plurality of scene images; and

generating a contour diagram of the target object by projecting the 3D contour onto a two-dimensional (2D) plane, wherein

the determining the 3D information of the target plane includes:

detecting the target plane in the plurality of scene images with plane fitting based on the depth information of the plurality of scene images; and

determining the 3D coordinates corresponding to the points on the target plane, and the plane equation of the target plane in a world coordinate system for the target plane in the plurality of scene images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2022
From: ZHANG, SHENGHAO; LING, YONGGEN; CHI, WANCHAO; ZHENG, YU; JIANG, XINYANG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 060862/0165 →
Priority Claims (1)
CN 202010899315.X · Aug 31, 2020 · national
Continuity (2)
Continuation PCTCN2021112058 · Aug 11, 2021
Related Publication 20220414910A1 · Dec 29, 2022
References Cited (20)
US 11367271B2 · Hemani · 2022 [cited by examiner]
US 11715377B2 · Chen · 2023 [cited by examiner]
US 11984028B2 · Li · 2024 [cited by examiner]
US 12086998B2 · Lang · 2024 [cited by examiner]
US 20200394918A1 · Chen · 2020 [cited by examiner]
US 20220414910A1 · Zhang · 2022 [cited by examiner]
CN 106441275A · 2017 [cited by applicant]
CN 108898661A · 2018 [cited by applicant]
CN 109215111A · 2019 [cited by applicant]
CN 109961501A · 2019 [cited by applicant]
CN 110455189A · 2019 [cited by applicant]
CN 112070782A · 2020 [cited by applicant]
WO 2014147863A · 2014 [cited by applicant]
Chinese Office Action issued Sep. 14, 2023 in Application No. 202010899315.X with English Translation (24 pages). [cited by applicant]
Bodo Rosenhahn et al: “Three-Dimensional Shape Knowledge for Joint Image Segmentation and Pose Tracking”, International Journal of Computer Vision, Kluwer Academic Publishers, BO, vol. 73, No. 3, Sep. 25, 2006, pp. 243-… [cited by applicant]
Supplementary European Search Report issued Jul. 25, 2023 in Application No. 21860143.3 (8 pages). [cited by applicant]
International Search Report issued Nov. 11, 2021 in International Application No. PCT/CN2021/112058 with English Translation (6 pages). [cited by applicant]
Nakashika T, Hori T, Takiguchi T, et al. 3D-object recognition based on LLC using depth spatial pyramid[C], 2014 22nd International Conference on Pattern Recognition. IEEE, 2014: 4224-4228. [cited by applicant]
Carbonara S, Guaragnella C. Efficient stairs detection algorithm Assisted navigation for vision impaired people[C], 2014 IEEE International Symposium on Innovations in Intelligent Systems and Applications (INISTA) Proce… [cited by applicant]
He K, Gkioxari G, Dollar P, et al. Mask r-cnn[C], Proceedings of the IEEE international conference on computer vision. 2017: 2961-2969. [cited by applicant]