IP Library Granted Patent US 12,254,570
Granted Patent B2
US 12,254,570 · App. 17/661,878 · Granted Mar 18, 2025

Generating three-dimensional representations for digital objects utilizing mesh-based thin volumes

Inventors: Sai Bi (San Jose, CA); Yang Liu (Cambridge, GB); Zexiang Xu (San Diego, CA); Fujun Luan (Santa Clara, CA); Kalyan Sunkavalli (San Jose, CA)
Assignee: Adobe Inc.
G06T17/205G06T13/20G06T2210/21
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,570
App. No.
17/661,878
Granted
Mar 18, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media that generate three-dimensional hybrid mesh-volumetric representations for digital objects. For instance, in one or more embodiments, the disclosed systems generate a mesh for a digital object from a plurality of digital images that portray the digital object using a multi-view stereo model. Additionally, the disclosed systems determine a set of sample points for a thin volume around the mesh. Using a neural network, the disclosed systems further generate a three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the mesh.

Claims (63)

1. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating, from a set of digital images portraying a digital object, a mesh for the digital object utilizing a multi-view stereo model;

building a thin volume around the mesh by:

generating one or more rays from a source point to the mesh, the one or more rays intersecting the mesh at one or more locations;

determining a plurality of sample points positioned off the mesh and positioned on the one or more rays intersecting the mesh; and

determining a set of sample points for the thin volume that includes one or more sample points from the plurality of sample points that are positioned off the mesh and within a threshold distance of the mesh and excludes at least one sample point from the plurality of sample points that is positioned off the mesh and outside the threshold distance of the mesh,

wherein the thin volume includes digital data representing features of the digital object within the threshold distance of the mesh; and

generating, utilizing a neural network, a three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the mesh.

2. The non-transitory computer-readable medium of claim 1 , wherein building the thin volume for the mesh comprises building the thin volume having a thickness that corresponds to the threshold distance.

3. The non-transitory computer-readable medium of claim 1 , wherein determining the plurality of sample points positioned off the mesh and positioned on the one or more rays intersecting the mesh comprises, for a ray of the one or more rays:

sampling one or more points along the ray intersecting the mesh.

4. The non-transitory computer-readable medium of claim 3 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

determining a direction of the ray and a distance between the one or more sample points and the mesh for the digital object; and

generating the three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the mesh by generating the three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the direction of the ray and the distance between the one or more sample points and the mesh.

5. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

generating, from the set of digital images portraying the digital object, a UV map for the digital object utilizing the multi-view stereo model, the UV map corresponding to a plurality of feature vectors representing neural textures; and

generating, utilizing the neural network, the three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the mesh by generating the three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume, the mesh, and the UV map.

6. The non-transitory computer-readable medium of claim 5 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

determining, for a sample point from the set of sample points, a feature vector utilizing the UV map; and

generating the three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume, the mesh, and the UV map by generating the three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the feature vector for the sample point and the mesh.

7. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising generating, utilizing the neural network, the three-dimensional hybrid mesh-volumetric representation for the digital object by determining, for a sample point from the set of sample points, at least one of a color value or a volume density utilizing the neural network.

8. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

modifying the mesh for the digital object by removing a portion of the mesh; and

generating a modified three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the modified mesh.

9. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising generating an animation of the three-dimensional hybrid mesh-volumetric representation using a manipulation of one or more portions of the mesh.

10. A system comprising:

one or more memory components; and

one or more processing devices coupled to the one or more memory components, the one or more processing devices to perform operations comprising:

generating, utilizing a multi-view stereo model and a set of digital images that portray a digital object, a mesh and a UV map for the digital object;

building a thin volume for the mesh by:

generating one or more rays from a source point to the mesh, the one or more rays intersecting the mesh at one or more locations;

determining a plurality of sample points positioned off the mesh and positioned on the one or more rays intersecting the mesh; and

determining a set of sample points for the thin volume that includes one or more sample points from the plurality of sample points that are positioned off the mesh and within a threshold distance of the mesh and excludes at least one sample point from the plurality of sample points that is positioned off the mesh and outside the threshold distance of the mesh;

determining, for the set of sample points and using the UV map, feature vectors including digital data representing features of the digital object within the threshold distance of the mesh; and

generating, utilizing a neural network, a three-dimensional hybrid mesh-volumetric representation of the digital object from the feature vectors, a direction of the one or more rays, and distances between the set of sample points and the mesh.

11. The system of claim 10 , wherein determining, from for the set of sample points, the feature vectors using the UV map comprises:

determining a point on the mesh that corresponds to a sample point from the set of sample points; and

identifying a feature vector that corresponds to the point on the mesh using the UV map.

12. The system of claim 11 , wherein identifying the feature vector that corresponds to the point on the mesh using the UV map comprises:

determining UV coordinates for the point on the mesh; and

determining the feature vector from a feature map that is located at the UV coordinates.

13. The system of claim 10 , wherein the operations further comprise:

determining, for the set of sample points, one or more additional feature vectors using at least one UV map for the digital object, wherein the feature vectors and the one or more additional feature vectors correspond to a pyramid of neural textures; and

generating one or more mipmaps utilizing the feature vectors and the one or more additional feature vectors for the set of sample points.

14. The system of claim 10 , wherein determining the plurality of sample points positioned off the mesh and positioned on the one or more rays intersecting the mesh comprises:

determining one or more sample points on a first side of the mesh; and

determining one or more additional sample points on a second side of the mesh.

15. The system of claim 10 , wherein generating the three-dimensional hybrid mesh-volumetric representation of the digital object from the feature vectors, the direction of the one or more rays, and the distances between the set of sample points and the mesh comprises determining, for a sample point from the set of sample points, a view-dependent color value and a volume density from a feature vector corresponding to the sample point, a direction of a ray corresponding to the sample point, and a distance between the sample point and the mesh.

16. The system of claim 10 , wherein the operations further comprise rendering the three-dimensional hybrid mesh-volumetric representation of the digital object for display on a client device.

17. The system of claim 10 , wherein building the thin volume for the mesh comprises building the thin volume having a thickness that corresponds to the threshold distance.

18. A computer-implemented method comprising:

receiving a set of digital images that portray a digital object in a plurality of views;

building a thin volume around a mesh for the digital object by:

generating one or more rays from a source point to the mesh, the one or more rays intersecting the mesh at one or more locations;

determining a plurality of sample points positioned off the mesh and positioned on the one or more rays intersecting the mesh; and

determining a set of sample points for the thin volume that includes one or more sample points from the plurality of sample points that are positioned off the mesh and within a threshold distance of the mesh and excludes at least one sample point from the plurality of sample points that is positioned off the mesh and outside the threshold distance of the mesh,

wherein the thin volume includes digital data representing features of the digital object within the threshold distance of the mesh;

generating, utilizing a neural network, a three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the mesh; and

rendering the three-dimensional hybrid mesh-volumetric representation for the digital object for display on a computing device.

19. The computer-implemented method of claim 18 , further comprising:

receiving, via the computing device, one or more user interactions for manipulating a portion of the three-dimensional hybrid mesh-volumetric representation for the digital object; and

modifying the three-dimensional hybrid mesh-volumetric representation in accordance with the one or more user interactions by modifying a portion of a mesh for the digital object that corresponds to the portion of the three-dimensional hybrid mesh-volumetric representation.

20. The computer-implemented method of claim 18 , further comprising generating a mipmap utilizing a pyramid of neural textures associated with the three-dimensional hybrid mesh-volumetric representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2022
From: BI, SAI; LIU, YANG; LUAN, FUJUN; SUNKAVALLI, KALYAN; XU, ZEXIANG
To: ADOBE INC.
Reel/Frame 059801/0380 →
Continuity (1)
Related Publication 20230360327A1 · Nov 9, 2023
References Cited (20)
US 10937237B1 · Kim · 2021 [cited by examiner]
US 11164343B1 · Batra · 2021 [cited by examiner]
US 11810243B2 · Petkov · 2023 [cited by examiner]
US 20060066608A1 · Appolloni · 2006 [cited by examiner]
US 20110282473A1 · Pavlovskaia · 2011 [cited by examiner]
US 20120192401A1 · Pavlovskaia · 2012 [cited by examiner]
US 20180225862A1 · Petkov · 2018 [cited by examiner]
US 20190043242A1 · Risser · 2019 [cited by examiner]
US 20190043255A1 · Somasundaram · 2019 [cited by examiner]
US 20190117268A1 · Pavlovskaia · 2019 [cited by examiner]
US 20190236830A1 · Akenine-Moller · 2019 [cited by examiner]
US 20200054398A1 · Kovtun · 2020 [cited by examiner]
US 20200184651A1 · Mukasa · 2020 [cited by examiner]
US 20220245910A1 · Lombardi · 2022 [cited by examiner]
US 20230076939A1 · Mammou · 2023 [cited by examiner]
Mildenhall, Ben et al. “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis.” ECCV (2020). [cited by applicant]
Baatz, H. et al.; NeRF-Tex: Neural Reflectance Field Textures, Computer Graphics forum; vol. 40 (2021), No. 4 pp. 1-13; 2021. [cited by applicant]
Lombardi, Stephen et al. “Mixture of volumetric primitives for efficient neural rendering.” ACM Transactions on Graphics (TOG) 40 (2021): 1-13. [cited by applicant]
Liu, Lingjie et al. “Neural actor: neural free-view synthesis of human actors with pose control”; ACM Transactions on Graphicsvol. 40; Issue 6; Dec. 2021 Article No. 219; pp. 1-16; https://doi.org/10.1145/3478513.348052… [cited by applicant]
J.L. Schönberger et al., Pixelwise View Selection for Unstructured Multi-view Stereo, Proceedings of the European Conference on Computer Vision, pp. 501-518, 2016. [cited by applicant]
Cited By (1)
US 12,499,567