IP Library › Granted Patent US 12,307,599
Granted Patent B2
US 12,307,599 · App. 17/659,449 · Granted May 20, 2025

Weak multi-view supervision for surface mapping estimation

Inventors: Aidas Liaudanskas (San Francisco, CA); Nishant Rai (San Francisco, CA); Srinivas Rao (San Francisco, CA); Rodrigo Ortiz-Cayon (San Francisco, CA); Matteo Munaro (San Francisco, CA); Stefan Johannes Josef Holzer (San Mateo, CA)
Assignee: Fyusion, Inc.
G06T17/20G06T7/70G06V20/647G06T2207/20081G06T2207/20084G06T2207/30248
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,599
App. No.
17/659,449
Granted
May 20, 2025
Kind
B2
Abstract

One or more two-dimensional images of a three-dimensional object may be analyzed to estimate a three-dimensional mesh representing the object and a mapping of the two-dimensional images to the three-dimensional mesh. Initially, a correspondence may be determined between the images and a UV representation of a three-dimensional template mesh by training a neural network. Then, the three-dimensional template mesh may be deformed to determine the representation of the object. The process may involve a reprojection loss cycle in which points from the images are mapped onto the UV representation, then onto the three-dimensional template mesh, and then back onto the two-dimensional images.

Claims (37)

1. A method comprising:

determining via a processor a correspondence between one or more two-dimensional images of a three-dimensional object and a UV representation of a three-dimensional template mesh of the three-dimensional object by training a neural network, the three-dimensional template mesh including a plurality of points in three-dimensional space and a plurality of edges between the plurality of points;

determining via the processor a deformation of the three-dimensional template mesh, the deformation displacing one or more of the plurality of points, wherein the deformation is determined so as to reduce reprojection consistency loss when mapping points from the two-dimensional images back onto the two-dimensional images through both the UV representation and the three-dimensional template mesh, wherein the one or more two-dimensional images include a proximate two-dimensional image, the proximate two-dimensional image being captured from a proximate virtual camera pose, wherein a reprojection consistency loss value depends in part on a proximate reprojection consistency loss value computed for a corresponding pixel in the proximate two-dimensional image; and

storing on a storage device a deformed three-dimensional template mesh.

2. The method recited in claim 1 , wherein training the neural network comprises predicting, for a first location in a designated one of the two-dimensional images, a corresponding second location in the UV representation.

3. The method recited in claim 2 , wherein training the neural network further comprises determining a third location in the three-dimensional template mesh by mapping the second location to the third location via UV parameterization.

4. The method recited in claim 3 , wherein training the neural network further comprises determining a fourth location in the designated two-dimensional image by projecting the third location onto the virtual camera pose associated with the designated two-dimensional image.

5. The method recited in claim 4 , wherein training the neural network further comprises determining a reprojection consistency loss value representing a displacement in two-dimensional space between the first location and the fourth location.

6. The method recited in claim 5 , wherein training the neural network further comprises updating the neural network based on the reprojection consistency loss value.

7. The method recited in claim 4 , the method further comprising:

determining the virtual camera pose by analyzing the two-dimensional image to identify a virtual camera position and virtual camera orientation for the two-dimensional image relative to the three-dimensional template mesh.

8. The method recited in claim 4 , wherein the one or more two-dimensional images include at least the designated two-dimensional image and the proximate two-dimensional image.

9. The method recited in claim 1 , wherein training the neural network comprises determining a visibility loss value representing occlusion of a designated portion of the three-dimensional object within a designated one of the two-dimensional images and update the neural network based on the visibility loss value.

10. The method recited in claim 1 , the method further comprising:

determining an object type corresponding to the three-dimensional object by analyzing one or more of the one or more two-dimensional images; and

selecting the three-dimensional template mesh from a plurality of available three-dimensional template meshes, the three-dimensional template mesh corresponding with the object type.

11. The method recited in claim 10 ,

wherein the object type is a vehicle, and wherein the three-dimensional template mesh provides a generic representation of vehicles.

12. The method recited in claim 10 ,

wherein the object type is a vehicle sub-type, and wherein the three-dimensional template mesh provides a generic representation of the vehicle sub-type.

13. A computing system comprising a processor and a storage device, the computing system configured to perform a method comprising:

determining via the processor a correspondence between one or more two-dimensional images of a three-dimensional object and a UV representation of a three-dimensional template mesh of the three-dimensional object by training a neural network, the three-dimensional template mesh including a plurality of points in three-dimensional space and a plurality of edges between the plurality of points;

determining via the processor a deformation of the three-dimensional template mesh, the deformation displacing one or more of the plurality of points, wherein the deformation is determined so as to reduce reprojection consistency loss when mapping points from the two-dimensional images back onto the two-dimensional images through both the UV representation and the three-dimensional template mesh, wherein the one or more two-dimensional images include a proximate two-dimensional image, the proximate two-dimensional image being captured from a proximate virtual camera pose, wherein a reprojection consistency loss value depends in part on a proximate reprojection consistency loss value computed for a corresponding pixel in the proximate two-dimensional image; and

storing on the storage device a deformed three-dimensional template mesh.

14. The computing system recited in claim 13 , wherein training the neural network comprises predicting, for a first location in a designated one of the two-dimensional images, a corresponding second location in the UV representation, wherein training the neural network further comprises determining a third location in the three-dimensional template mesh by mapping the second location to the third location via UV parameterization, wherein training the neural network further comprises determining a fourth location in the designated two-dimensional image by projecting the third location onto a virtual camera pose associated with the designated two-dimensional image, wherein training the neural network further comprises determining a reprojection consistency loss value representing a displacement in two-dimensional space between the first location and the fourth location, wherein training the neural network further comprises updating the neural network based on the reprojection consistency loss value.

15. The computing system recited in claim 14 , the method further comprising:

determining the virtual camera pose by analyzing the two-dimensional image to identify a virtual camera position and virtual camera orientation for the two-dimensional image relative to the three-dimensional template mesh.

16. The method recited in claim 14 , wherein the one or more two-dimensional images include at least the designated two-dimensional image and the proximate two-dimensional image.

17. The computing system recited in claim 13 , wherein training the neural network comprises determining a visibility loss value representing occlusion of a designated portion of the three-dimensional object within a designated one of the two-dimensional images and update the neural network based on the visibility loss value.

18. The computing system recited in claim 13 , the method further comprising:

determining an object type corresponding to the three-dimensional object by analyzing one or more of the one or more two-dimensional images; and

selecting the three-dimensional template mesh from a plurality of available three-dimensional template meshes, the three-dimensional template mesh corresponding with the object type.

19. One or more non-transitory computer readable media having instructions stored thereon for performing a method, the method comprising:

determining using a processor a correspondence between one or more two-dimensional images of a three-dimensional object and a UV representation of a three-dimensional template mesh of the three-dimensional object by training a neural network, the three-dimensional template mesh including a plurality of points in three-dimensional space and a plurality of edges between the plurality of points;

determining via the processor a deformation of the three-dimensional template mesh, the deformation displacing one or more of the plurality of points, wherein the deformation is determined so as to reduce reprojection consistency loss when mapping points from the two-dimensional images back onto the two-dimensional images through both the UV representation and the three-dimensional template mesh, wherein the one or more two-dimensional images include a proximate two-dimensional image, the proximate two-dimensional image being captured from a proximate virtual camera pose, wherein a reprojection consistency loss value depends in part on a proximate reprojection consistency loss value computed for a corresponding pixel in the proximate two-dimensional image; and

storing on the storage device a deformed three-dimensional template mesh.

20. The one or more non-transitory computer readable media recited in claim 19 , wherein training the neural network comprises predicting, for a first location in a designated one of the two-dimensional images, a corresponding second location in the UV representation, wherein training the neural network further comprises determining a third location in the three-dimensional template mesh by mapping the second location to the third location via the UV parameterization, wherein training the neural network further comprises determining a fourth location in the designated two-dimensional image by projecting the third location onto a virtual camera pose associated with the designated two-dimensional image, wherein training the neural network further comprises determining a reprojection consistency loss value representing a displacement in two-dimensional space between the first location and the fourth location, wherein training the neural network further comprises updating the neural network based on the reprojection consistency loss value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2022
From: LIAUDANSKAS, AIDAS; RAI, NISHANT; RAO, SRINIVAS; ORTIZ-CAYON, RODRIGO; MUNARO, MATTEO; HOLZER, STEFAN JOHANNES JOSEF
To: FYUSION, INC.
Reel/Frame 059615/0791 →
Continuity (2)
Provisional Application 63177580 · Apr 21, 2021
Related Publication 20220343601A1 · Oct 27, 2022
References Cited (15)
US 20060017723A1 · Baran · 2006 [cited by applicant]
US 20130307848A1 · Tena · 2013 [cited by applicant]
US 20130314412A1 · Gravois · 2013 [cited by applicant]
US 20160027200A1 · Corazza · 2016 [cited by applicant]
US 20170148179A1 · Holzer · 2017 [cited by applicant]
US 20170374341A1 · Michail · 2017 [cited by applicant]
US 20190026917A1 · Liao · 2019 [cited by examiner]
US 20210287430A1 · Li · 2021 [cited by examiner]
Tulsiani et al. “Implicit mesh reconstruction from unannotated image collections.” arXiv preprint arXiv:2007.08504 (2020). (Year: 2020). [cited by examiner]
International Preliminary Report on Patentability issued in App. No. PCT/US2022/071754, mailing date Nov. 2, 2023, 8 pages. [cited by applicant]
Kanazawa, et al., “Learning Category-Specific Mesh Reconstruction from Image Collections,” University of California, Berkley, arXiv:1803.07549v2 [cs.CV] Jul. 30, 2018, 21 pages. [cited by applicant]
Kulkarni, et al., “Articulation-aware Canonical Surface Mapping,” arXiv:2004.00614v3 [cs.CV] May 26, 2020, 17 pages. [cited by applicant]
Kulkarni, et al., “Canonical Surface Mapping via Geometric Cycle Consistency,” Carnegie Mellon University, arXiv:1907.10043v2 [cs.CV] Aug. 15, 2019, 16 pages. [cited by applicant]
International Search Report and Written Opinion issued in App. No. PCT/US2022/071754, mailing date Jul. 13, 2022, 10 pages. [cited by applicant]
Rai et al., “Weak Multi-View Supervision for Surface Mapping Estimation”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, May 4, 2021, https://arxiv.org/pdf/2105.01388.pdf, 10 pa… [cited by applicant]