IP Library Granted Patent US 11,715,224
Granted Patent B2
US 11,715,224 · App. 17/242,415 · Granted Aug 1, 2023

Three-dimensional object reconstruction method and apparatus

Inventors: Yuan Gao (Shenzhen, CN); Xiang Kai Lin (Shenzhen, CN); Lin Chao Bao (Shenzhen, CN); Wei Liu (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06T7/55G06T7/337G06V10/25G06V10/76G06V10/803G06V10/82G06V20/653G06V40/162G06V40/164G06V40/169G06V40/171G06T2207/10028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,715,224
App. No.
17/242,415
Granted
Aug 1, 2023
Kind
B2
Abstract

A three-dimensional object reconstruction method, applied to a terminal device or a server, is provided. The method includes obtaining a plurality of video frames of an object; determining three-dimensional location information of key points of the object in the plurality of video frames and physical meaning information of the key points, the physical meaning information indicating respective positions of the object; determining a correspondence between the key points having the same physical meaning information in the plurality of video frames; and generating a three-dimensional object according to the correspondence and the three-dimensional location information of the key points.

Claims (68)

1. A three-dimensional object reconstruction method, applied to a terminal device or a server, the method comprising:

obtaining a plurality of video frames of an object;

determining three-dimensional location information of key points of the object in the plurality of video frames and physical meaning information of the key points, the physical meaning information indicating respective positions of the object;

determining a correspondence between the key points having the same physical meaning information in the plurality of video frames; and

generating a three-dimensional object according to the correspondence and the three-dimensional location information of the key points,

wherein the plurality of video frames comprises a reference frame and non-reference frames, and the method further comprises determining key frames among the non-reference frames based on a quantity of key points, included in each key frame, that are matched as inliers with key points in the reference frame, and further based on weights of the inliers,

wherein the weights of the inliers are applied such that a greater weight is applied to a key point whose physical meaning information reflects a deformable feature and a lower weight is applied to a key point whose physical meaning information reflects a non-deformable feature, and

wherein the generating the three-dimensional object comprises generating the three-dimensional object according to the correspondence and three-dimensional location information of inliers in the reference frame and the key frames.

2. The method according to claim 1 , wherein each of the plurality of video frames comprises a color video subframe and a depth video subframe,

the method further comprises determining key point information of the object in the plurality of video frames according to color video subframes of the plurality of video frames, the key point information comprising two-dimensional location information of key points of the object and the physical meaning information of the key points, and

the determining the three-dimensional location information of the key points comprises determining the three-dimensional location information of the key points of the object in the plurality of video frames from depth video subframes of the plurality of video frames according to the two-dimensional location information of the key points.

3. The method according to claim 2 , wherein the determining the key point information comprises:

performing object detection on the color video subframe of a video frame by using a first network model, and determining, in the color video subframe, a target region in which the object is located; and

extracting video frame data of the target region, and determining the key point information of the object in the video frame by using a second network model.

4. The method according to claim 2 , wherein the generating the three-dimensional object comprises performing registration of point cloud data in the plurality of video frames according to the correspondence and the three-dimensional location information of the key points, and generating the three-dimensional object based on the registration of the point cloud data.

5. The method according to claim 4 , wherein the determining the key frames comprises:

obtaining relative attitudes of the object in the non-reference frames relative to the reference frame, and obtaining a quantity of key points matched as inliers in each of the non-reference frames; and

determining, for each attitude range of a plurality of attitude ranges, at least one non-reference frame as a key frame according to quantities of inliers in the non-reference frames, the plurality of attitude ranges being obtained according to the relative attitudes of the object in the non-reference frames; and

the performing the registration of the point cloud data comprises:

performing the registration of the point cloud data in the plurality of video frames according to the correspondence and the three-dimensional location information of inliers in the reference frame and the key frames.

6. The method according to claim 5 , wherein the determining the key frames further comprises:

determining inlier scores of the non-reference frames based on the weights of the inliers and the quantities of inliers in the non-reference frames; and

determining the at least one non-reference frame as the key frame in each attitude range according to the inlier scores of the non-reference frames.

7. The method according to claim 5 , wherein the performing the registration of the point cloud data in the plurality of video frames according to the correspondence and the three-dimensional location information of inliers in the reference frame and the key frames comprises:

rotating the inliers in the key frames according to the relative attitudes of the object in the key frames relative to the reference frame, to perform pre-registration with the inliers in the reference frame; and

performing the registration of the point cloud data in the plurality of video frames according to a result of the pre-registration.

8. The method according to claim 5 , wherein the plurality of attitude ranges are obtained by:

determining an angle range in a horizontal or vertical direction covering the relative attitudes of the object in the non-reference frames; and

dividing the angle range into the plurality of attitude ranges by using an angle threshold.

9. A three-dimensional object reconstruction apparatus, comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code including:

video frame obtaining code configured to cause the at least one processor to obtain a plurality of video frames of an object;

first determining code configured to cause the at least one processor to determine three-dimensional location information of key points of the object in the plurality of video frames and physical meaning information of the key points, the physical meaning information indicating respective positions of the object;

second determining code configured to cause the at least one processor to determine a correspondence between the key points having the same physical meaning information in the plurality of video frames; and

generation code configured to cause the at least one processor to generate a three-dimensional object according to the correspondence and the three-dimensional location information of the key points,

wherein the plurality of video frames comprises a reference frame and non-reference frames, and the program code comprises third determining code configured to cause the at least one processor to determine key frames among the non-reference frames based on a quantity of key points, included in each key frame, that are matched as inliers with key points in the reference frame, and further based on weights of the inliers,

wherein the weights of the inliers are applied such that a greater weight is applied to a key point whose physical meaning information reflects a deformable feature and a lower weight is applied to a key point whose physical meaning information reflects a non-deformable feature, and

wherein the generation code is configured to cause the at least one processor to generate the three-dimensional object comprises generating the three-dimensional object according to the correspondence and three-dimensional location information of inliers in the reference frame and the key frames.

10. The apparatus according to claim 9 , wherein each of the plurality of video frames comprises a color video subframe and a depth video subframe,

the program code further comprises:

fourth determining code configured to cause the at least one processor to determine key point information of the object in the plurality of video frames according to color video subframes of the plurality of video frames, the key point information comprising two-dimensional location information of key points of the object and the physical meaning information of the key points,

wherein the first determining code is further configured to cause the at least one processor to determine three-dimensional location information of the key points of the object in the plurality of video frames from depth video subframes of the plurality of video frames according to the two-dimensional location information of the key points.

11. The apparatus according to claim 10 , wherein the fourth determining code comprises:

object detection subcode configured to cause the at least one processor to perform object detection on the color video subframe of a video frame by using a first network model, and determine, in the color video subframe, a target region in which the object is located; and

extraction subcode configured to cause the at least one processor to extract video frame data of the target region, and determine the key point information of the object in the video frame by using a second network model.

12. The apparatus according to claim 10 , wherein the generation code is further configured to cause the at least one processor to perform registration of point cloud data in the plurality of video frames according to the correspondence and the three-dimensional location information of the key points, and generate the three-dimensional object based on the registration of the point cloud data.

13. The apparatus according to claim 12 , wherein the program code further comprises,

obtaining code configured to cause the at least one processor to obtain relative attitudes of the object in the non-reference frames relative to the reference frame, and obtain a quantity of key points matched as inliers in each of the non-reference frames, and

fifth determining code configured to cause the at least one processor to determine, for each attitude range of a plurality of attitude ranges, at least one non-reference frame as a key frame according to quantities of inliers in the non-reference frames, the plurality of attitude ranges being obtained according to the relative attitudes of the object in the non-reference frames; and

registration code configured to cause the at least one processor to perform the registration of the point cloud data in the plurality of video frames according to the correspondence and the three-dimensional location information of inliers in the reference frame and the key frames.

14. The apparatus according to claim 13 , wherein the fifth determining code comprises:

inlier score determining subcode configured to cause the at least one processor to determine inlier scores of the non-reference frames based on the weights of the inliers and the quantities of inliers in the non-reference frames; and

key frame determining subcode configured to cause the at least one processor to determine the at least one non-reference frame as the key frame in each attitude range according to the inlier scores of the non-reference frames.

15. The apparatus according to claim 13 , wherein the registration code is further configured to cause the at least one processor to:

rotate the inliers in the key frames according to the relative attitudes of the object in the key frames relative to the reference frame, to perform pre-registration with the inliers in the reference frame; and

perform the registration of the point cloud data in the plurality of video frames according to a result of the pre-registration.

16. The apparatus according to claim 13 , wherein the plurality of attitude ranges are obtained by:

determining an angle range in a horizontal or vertical direction covering the relative attitudes of the object in the non-reference frames; and

dividing the angle range into the plurality of attitude ranges by using an angle threshold.

17. A non-transitory computer-readable storage medium, configured to store program code executable by at least one processor to cause the at least one processor to perform:

obtaining a plurality of video frames of an object;

determining three-dimensional location information of key points of the object in the plurality of video frames and physical meaning information of the key points, the physical meaning information indicating respective positions of the object;

determining a correspondence between the key points having the same physical meaning information in the plurality of video frames; and

generating a three-dimensional object according to the correspondence and the three-dimensional location information of the key points,

wherein the plurality of video frames comprises a reference frame and non-reference frames, and the program code further causes the at least one processor to determine key frames among the non-reference frames based on a quantity of key points, included in each key frame, that are matched as inliers with key points in the reference frame, and further based on weights of the inliers,

wherein the weights of the inliers are applied such that a greater weight is applied to a key point whose physical meaning information reflects a deformable feature and a lower weight is applied to a key point whose physical meaning information reflects a non-deformable feature, and

wherein the generating the three-dimensional object comprises generating the three-dimensional object according to the correspondence and three-dimensional location information of inliers in the reference frame and the key frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2021
From: GAO, YUAN; LIN, XIANG KAI; BAO, LIN CHAO; LIU, WEI
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 056067/0095 →
Priority Claims (1)
CN 201910233202.3 · Mar 26, 2019 · national
Continuity (2)
Continuation PCTCN2020079439 · Mar 16, 2020
Related Publication 20210248763A1 · Aug 12, 2021