IP Library Granted Patent US 11,640,694
Granted Patent B2
US 11,640,694 · App. 17/208,943 · Granted May 2, 2023

3D model reconstruction and scale estimation

Inventors: Sean M. Adkinson (North Plains, OR); Flora Ponjou Tasse (London, GB); Pavan K. Kamaraju (London, GB); Ghislain Fouodji Tasse (London, GB); Ryan R. Fink (Vancouver, WA)
Assignee: STREEM, LLC
G06T17/20G06N3/04G06T7/50G06T7/70G06T15/04G06T2207/10028G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,640,694
App. No.
17/208,943
Granted
May 2, 2023
Kind
B2
Abstract

Embodiments include systems and methods for creation of a 3D mesh from a video stream or a sequence of frames. A sparse point cloud is first created from the video stream, which is then densified per frame by comparison with spatially proximate frames. A 3D mesh is then created from the densified depth maps, and the mesh is textured by projecting the images from the video stream or sequence of frames onto the mesh. Metric scale of the depth maps may be estimated where direct measurements are not able to be measured or calculated using a machine learning depth estimation network.

Claims (48)

1. A method comprising:

receiving, at a computing device, a sequence of frames of a scene captured by a camera;

estimating, by the computing device, a camera pose for the camera from the sequence of frames;

generating, by the computing device, a sparse depth map for each frame from the sequence of frames and camera pose, each sparse depth map comprised of at least one 3D point;

densifying, by the computing device for each frame in the sequence of frames, each sparse depth map by comparison of each frame from the sequence of frames with at least two neighboring frames from the sequence of frames, to create a dense depth map for each frame from the sequence of frames;

generating, by the computing device, a 3D mesh from the dense depth maps;

texturing, by the computing device, the 3D mesh by projecting one or more frames from the sequence of frames onto the 3D mesh; and

mapping a coordinate space of the 3D mesh to a coordinate space of the sequence of frames.

2. The method of claim 1 , further comprising determining, by the computing device, the at least two neighboring frames for each frame by whether each of the at least two neighboring frames shares at least a predetermined number of points in the sparse depth map with the frame.

3. The method of claim 1 , where generating each sparse depth map comprises:

detecting, by the computing device, features within a first frame and within a second frame of the sequence of frames, the second frame being temporally adjacent to the first frame; and

calculating, by the computing device, a depth value for one or more common points on a detected feature within the first frame that matches a detected feature within the second frame.

4. The method of claim 1 , further comprising receiving, at the computing device, additional camera pose data with the sequence of frames.

5. The method of claim 4 , wherein the camera pose data comprises directly measured depth data.

6. The method of claim 1 , further comprising:

passing, by the computing device, each frame of the sequence of frames through a depth estimation network to obtain an estimated depth map;

rendering, by the computing device from the sparse depth map, a depth map representing a camera view; and

fitting, by the computing device, the camera view depth map to the estimated depth map to obtain a depth map with an estimated metric scale.

7. A non-transitory computer readable medium (CRM) comprising instructions that, when executed by an apparatus, cause the apparatus to:

receive, at the apparatus, a video stream comprised of a sequence of frames, each of the sequence of frames comprising an image;

estimate a camera pose from the sequence of frames;

generate a sparse depth map from the sequence of frames and estimated camera pose;

densify, for each frame in the sequence of frames, the sparse depth map by comparison of each frame from the sequence of frames with at least two neighboring frames from the sequence of frames, to create a dense depth map;

generate a 3D mesh from the dense depth map;

texture the 3D mesh by projecting one or more frames from the sequence of frames onto the 3D mesh; and

map a coordinate space of the 3D mesh to a coordinate space of the sequence of frames.

8. The CRM of claim 7 , wherein the instructions are to further cause the apparatus to determine the at least two neighboring frames for each frame by whether each of the at least two neighboring frames shares at least a predetermined number of points in the sparse depth map with the frame.

9. The CRM of claim 7 , wherein the instructions are to further cause the apparatus to:

detect features within a first frame and within a second frame of the sequence of frames, the second frame being temporally adjacent to the first frame; and

calculate a depth value for one or more common points on a detected feature within the first frame that matches a detected feature within the second frame.

10. The CRM of claim 7 , wherein the instructions are to further cause the apparatus to receive, at the apparatus, additional camera pose data with the sequence of frames.

11. The CRM of claim 10 , wherein the camera pose data comprises directly measured depth data.

12. The CRM of claim 7 , wherein the apparatus comprises a mobile device or a server.

13. The CRM of claim 7 , wherein the instructions are to further cause the apparatus to:

pass each frame of the sequence of frames through a depth estimation network to obtain an estimated depth map;

render, from the sparse depth map, a depth map representing a camera view; and

fit the camera view depth map to the estimated depth map to obtain a depth map with an estimated metric scale.

14. The CRM of claim 13 , wherein the instructions are to further cause the apparatus to transmit the textured 3D mesh and estimated metric scale to a remote device.

15. A non-transitory computer readable medium (CRM) comprising instructions that, when executed by a computing apparatus, cause the apparatus to:

pass each frame of a sequence of frames through a depth estimation network to obtain an estimated depth map;

estimate a camera pose from the sequence of frames;

generate a depth map from the sequence of frames;

render a camera view depth map representing a camera view from the generated depth map and camera pose;

fit the camera view depth map to the estimated depth map to obtain an estimated metric scale for each point within the generated depth map; and

map a coordinate space of the camera view depth map to a coordinate space of the sequence of frames.

16. The CRM of claim 15 , wherein the generated depth map is a sparse depth map, and wherein the instructions are to further cause the apparatus to compute the sparse depth map by a comparison of features between each frame of the sequence of frames and frames in the sequence of frames that are temporally adjacent to each frame.

17. The CRM of claim 15 , wherein the depth estimation network is a deep learning network.

18. The CRM of claim 15 , wherein the apparatus is a cloud computing platform.

Continuity (2)
Provisional Application 62992324 · Mar 20, 2020
Related Publication 20210295599A1 · Sep 23, 2021
Cited By (2)
US 12,260,498 US 12,387,431