IP Library › Granted Patent US 11,880,935
Granted Patent B2
US 11,880,935 · App. 17/951,405 · Granted Jan 23, 2024

Multi-view neural human rendering

Inventors: Minye Wu (Shanghai, CN); Jingyi Yu (Shanghai, CN)
Assignee: SHANGHAITECH UNIVERSITY
G06T17/00G06F18/213G06F18/214G06T7/593G06T15/04G06T15/205G06T17/20G06V10/82G06V20/653H04N13/207G06T2200/08G06T2207/10016G06T2207/10024G06T2207/10028G06T2207/20081G06T2207/20084G06T2207/30196G06T2210/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,880,935
App. No.
17/951,405
Granted
Jan 23, 2024
Kind
B2
Abstract

An image-based method of modeling and rendering a three-dimensional model of an object is provided. The method comprises: obtaining a three-dimensional point cloud at each frame of a synchronized, multi-view video of an object, wherein the video comprises a plurality of frames; extracting a feature descriptor for each point in the point cloud for the plurality of frames without storing the feature descriptor for each frame; producing a two-dimensional feature map for a target camera; and using an anti-aliased convolutional neural network to decode the feature map into an image and a foreground mask.

Claims (51)

1. An image-based method of modeling and rendering a three-dimensional model of an object, the method comprising:

obtaining a three-dimensional point cloud at each frame of a synchronized, multi-view video of an object, wherein the video comprises a plurality of frames;

extracting, using a feature extracting (FE) network, a feature descriptor for each point in the point cloud for the plurality of frames;

producing a two-dimensional feature map for a target camera based on the feature descriptors and parameters of the target camera, wherein the producing comprises:

rasterizing points in the point clouds into pixel squares covering a subset of pixels in the two-dimensional feature map based on the feature descriptors, and

assigning a default feature vector to each of the rest of the pixels in the two-dimensional feature map;

generating a depth map corresponding to the two-dimensional feature map; and

inputting the two-dimensional feature map and the depth map into a rendering convolutional neural (RE) network to decode the feature map into an image and a foreground mask,

the method further comprises:

training the FE network and the RE network based on training data captured by a multi-camera dome system, wherein the training comprises a shared training followed by an individual training, wherein:

the shared training trains on a plurality of objects to obtain a shared RE network and different FE networks for the plurality of objects, and

the individual training trains on a specific object for fine tuning the corresponding FE network.

2. The method of claim 1 , wherein the point cloud is not triangulated to form a mesh.

3. The method of claim 1 , further comprising:

obtaining the three-dimensional point cloud at each frame of the synchronized, multi-view video of the object using a three dimension scanning sensor comprising a depth camera.

4. The method of claim 1 , further comprising:

capturing the synchronized, multi-view video of the object by a plurality of cameras from a plurality of view directions; and

reconstructing the three-dimensional point cloud at each frame of the video.

5. The method of claim 4 , further comprising:

taking into consideration the view directions in extracting the feature descriptor without storing the feature descriptor for each frame.

6. The method of claim 1 , wherein each point in the point cloud comprises color, and the method further comprises:

imposing recovered color of the points in the feature descriptor.

7. The method of claim 1 , further comprising:

extracting the feature descriptor using PointNet++.

8. The method of claim 7 , further comprising:

removing a classifier branch in PointNet++.

9. The method of claim 1 , wherein the feature descriptor comprises a feature vector comprising at least 20 dimensions.

10. The method of claim 9 , wherein the feature vector comprises 24 dimensions.

11. The method of claim 1 , wherein the producing the two-dimensional feature map further comprises

mapping the feature descriptor to a target view for the target camera, wherein back-propagation of gradient on the two-dimensional feature map is directly conducted on the three-dimensional point cloud.

12. The method of claim 1 , further comprising:

using U-Net to decode the feature map into the image and the foreground mask.

13. The method of claim 12 , wherein the rendering convolutional neural network comprises U-Net to produce the image and the foreground mask.

14. The method of claim 13 , further comprising:

replacing a downsampling operation in U-Net with MaxBlurPool and ConvBlurPool.

15. The method of claim 13 , further comprising:

replacing a convolutional layer in U-Net with a gated convolution layer.

16. The method of claim 13 , further comprising:

maintaining translation invariance for the feature map.

17. The method of claim 1 , wherein the rendering convolutional neural network comprises an anti-aliased neural network, and the method further comprises:

training the anti-aliased convolutional neural network using a ground truth foreground mask.

18. The method of claim 17 , further comprising:

performing two-dimensional image transformation to the training data.

19. The method of claim 1 , further comprising:

generating a plurality of masks from a plurality of viewpoints; and

using the masks as silhouettes to conduct visual hull reconstruction of a mesh.

20. The method of claim 19 , further comprising:

partitioning a space of interest into a plurality of discrete voxels;

classifying the voxels into two categories; and

computing a signed distance field to recover the mesh.

21. The method of claim 1 , wherein the object is a performer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: WU, MINYE; YU, JINGYI
To: SHANGHAITECH UNIVERSITY
Reel/Frame 064065/0289 →
Continuity (2)
Continuation PCTCN2020082120 · Mar 30, 2020
Related Publication 20230027234A1 · Jan 26, 2023