IP Library › Granted Patent US 12,217,351
Granted Patent B2
US 12,217,351 · App. 17/984,521 · Granted Feb 4, 2025

Representing 3D shapes with probabilistic directed distance fields

Inventors: Tristan Ty Aumentado-Armstrong (Toronto, CA); Stavros Tsogkas (Toronto, CA); Sven Josef Dickinson (Toronto, CA); Allan Jepson (Oakville, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T15/06G06T15/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,351
App. No.
17/984,521
Granted
Feb 4, 2025
Kind
B2
Abstract

The present disclosure provides methods, apparatuses, and computer-readable mediums for representing shapes with probabilistic directed distance fields. In some embodiments, a method includes obtaining a camera representation and a latent shape vector representation of a scene. The camera representation indicates position information and direction information of a view of the scene. The method further includes calculating, based on the latent shape vector representation of the scene, a visibility score and a depth for each ray of a plurality of rays emanating from a corresponding plurality of positions and directions. The plurality of positions and directions are determined from the camera representation of the scene. The method further includes generating renders of geometric information of the scene using the visibility score and the depth of the plurality of rays.

Claims (55)

1. A method of representing shapes with probabilistic directed distance fields to be performed by a processor, comprising:

obtaining a camera representation and a latent shape vector representation of a scene, the camera representation indicating position information and direction information of a view of the scene;

calculating, based on the latent shape vector representation of the scene, a visibility score and a depth for each ray of a plurality of rays emanating from a corresponding plurality of positions and directions, the plurality of positions and directions being determined from the camera representation of the scene; and

generating renders of geometric information of the scene using the visibility score and the depth of the plurality of rays.

2. The method of claim 1 , further comprising:

receiving a plurality of queries requesting the visibility score and the depth for each ray of the plurality of rays emanating from the corresponding plurality of positions and directions,

wherein the calculating of the visibility score and the depth of the plurality of rays comprises calculating, in response to the receiving of a query of the plurality of queries, the visibility score and the depth of a ray of the plurality of rays corresponding to the corresponding position and direction indicated by the query.

3. The method of claim 1 , further comprising:

correcting depth information of the renders of the geometric information of the scene across at least one occlusion boundary, based on a switching mechanism over a set of estimated depth values.

4. The method of claim 1 , wherein the obtaining of the camera representation and the latent shape vector representation of the scene comprises:

encoding, using a neural encoder, an image comprising the scene.

5. The method of claim 1 , wherein the calculating of the visibility score and the depth of the plurality of rays comprises:

combining a plurality of shape representations of the scene; and

calculating the visibility score and the depth for each ray of the plurality of rays based on a combination of the plurality of shape representations of the scene.

6. The method of claim 1 , wherein the calculating of the visibility score and the depth of the plurality of rays comprises:

performing, for each ray of the plurality of rays, a single forward pass of a conditional coordinate neural network to calculate the visibility score and the depth of that ray.

7. The method of claim 1 , wherein the calculating of the visibility score and the depth of the plurality of rays comprises:

calculating a lowest distance for each ray of the plurality of rays intersecting the scene.

8. The method of claim 1 , wherein:

the visibility score indicates whether a corresponding ray intersects the scene; and

the depth indicates a distance from the corresponding position of the corresponding ray to a nearest intersection point of the corresponding ray with the scene.

9. The method of claim 1 , further comprising:

calculating, based on the latent shape vector representation of the scene, a reflectance value for each ray of the plurality of rays emanating from the corresponding plurality of positions and directions.

10. An apparatus for representing shapes with probabilistic directed distance fields, comprising:

a memory storage storing computer-executable instructions; and

a processor communicatively coupled to the memory storage, wherein the processor is configured to execute the computer-executable instructions and cause the apparatus to:

obtain a camera representation and a latent shape vector representation of a scene, the camera representation indicating position information and direction information of a view of the scene;

calculate, based on the latent shape vector representation of the scene, a visibility score and a depth for each ray of a plurality of rays emanating from a corresponding plurality of positions and directions, the plurality of positions and directions being determined from the camera representation of the scene; and

generate renders of geometric information of the scene using the visibility score and the depth of the plurality of rays.

11. The apparatus of claim 10 , wherein the processor is further configured to execute further computer-executable instructions and further cause the apparatus to:

receive a plurality of queries requesting the visibility score and the depth for each ray of the plurality of rays emanating from the corresponding plurality of positions and directions; and

calculate, in response to the receiving of a query of the plurality of queries, the visibility score and the depth of a ray of the plurality of rays corresponding to the corresponding position and direction indicated by the query.

12. The apparatus of claim 10 , wherein the processor is further configured to execute further computer-executable instructions and further cause the apparatus to:

correct depth information of the renders of the geometric information of the scene across at least one occlusion boundary, based on a switching mechanism over a set of estimated depth values.

13. The apparatus of claim 10 , wherein the processor is further configured to execute further computer-executable instructions and further cause the apparatus to:

encode, using a neural encoder, an image comprising the scene.

14. The apparatus of claim 10 , wherein the processor is further configured to execute further computer-executable instructions and further cause the apparatus to:

combine a plurality of shape representations of the scene; and

calculate the visibility score and the depth for each ray of the plurality of rays based on a combination of the plurality of shape representations of the scene.

15. The apparatus of claim 10 , wherein the processor is further configured to execute further computer-executable instructions and further cause the apparatus to:

perform, for each ray of the plurality of rays, a single forward pass of a conditional coordinate neural network to calculate the visibility score and the depth of that ray.

16. The apparatus of claim 10 , wherein the processor is further configured to execute further computer-executable instructions and further cause the apparatus to:

calculate a lowest distance for each ray of the plurality of rays intersecting the scene.

17. The apparatus of claim 10 , wherein:

the visibility score indicates whether a corresponding ray intersects the scene; and

the depth indicates a distance from the corresponding position of the corresponding ray to a nearest intersection point of the corresponding ray with the scene.

18. The apparatus of claim 10 , wherein the processor is further configured to execute further computer-executable instructions and further cause the apparatus to:

calculate, based on the latent shape vector representation of the scene, a reflectance value for each ray of the plurality of rays emanating from the corresponding plurality of positions and directions.

19. A non-transitory computer-readable storage medium storing computer-executable instructions for representing shapes with probabilistic directed distance fields by a device, the computer-executable instructions being configured, when executed by one or more processors of the device, to cause the device to:

obtain a camera representation and a latent shape vector representation of a scene, the camera representation indicating position information and direction information of a view of the scene;

calculate, based on the latent shape vector representation of the scene, a visibility score and a depth for each ray of a plurality of rays emanating from a corresponding plurality of positions and directions, the plurality of positions and directions being determined from the camera representation of the scene; and

generate renders of geometric information of the scene using the visibility score and the depth of the plurality of rays.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the computer-executable instructions further cause the device to:

receive a plurality of queries requesting the visibility score and the depth for each ray of the plurality of rays emanating from the corresponding plurality of positions and directions; and

calculate, in response to the receiving of a query of the plurality of queries, the visibility score and the depth of a ray of the plurality of rays corresponding to the corresponding position and direction indicated by the query.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: AUMENTADO-ARMSTRONG, TRISTAN TY; TSOGKAS, STAVROS; DICKINSON, SVEN JOSEF; JEPSON, ALLAN DOUGLAS
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062131/0400 →
Continuity (2)
Provisional Application 63280010 · Nov 16, 2021
Related Publication 20230154102A1 · May 18, 2023
References Cited (21)
US 8659593B2 · Furukawa et al. · 2014 [cited by applicant]
US 10902679B2 · Molyneaux et al. · 2021 [cited by applicant]
US 10950036B2 · Ha et al. · 2021 [cited by applicant]
US 11010961B2 · Stachniak et al. · 2021 [cited by applicant]
US 11315319B2 · Yu et al. · 2022 [cited by applicant]
US 20150332505A1 · Wang · 2015 [cited by examiner]
US 20190197765A1 · Molyneaux · 2019 [cited by examiner]
US 20190286231A1 · Burns et al. · 2019 [cited by applicant]
US 20210090322A1 · Hunt · 2021 [cited by applicant]
US 20220130127A1 · Zhang et al. · 2022 [cited by applicant]
KR 1020190112894A · 2019 [cited by applicant]
Yue Jiang et al., “SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape Optimization”, arXiv:1912.07109v1 [cs.CV], Dec. 15, 2019, 10 pages. [cited by applicant]
Julian Chibane et al., “Neural Unsigned Distance Fields for Implicit Function Learning”, arXiv:2010.13938v1 [cs.CV], Oct. 26, 2020, 15 pages. [cited by applicant]
Tristan Aumentado-Armstrong et al., “Cycle-Consistent Generative Rendering for 2D-3D Modality Translation”, arXiv:2011.08026v1 [cs.CV], Nov. 16, 2020, 22 pages. [cited by applicant]
Tristan Aumentado-Armstrong et al., “Representing 3D Shapes with Probabilistic Directed Distance Fields”, arXiv:2112.05300v1 [cs.CV], Dec. 10, 2021, 22 pages. [cited by applicant]
International Search Report and Written Opinion (PCT/ISA/220,PCT/ISA/210, and PCT/ISA/237) issued by the International Searching Authority on Feb. 23, 2023 in corresponding International Application No. PCT/KR2022/01796… [cited by applicant]
Communication issued Oct. 24, 2024 by the European Patent Office in European Patent Application No. 22896014.2. [cited by applicant]
Sitzmann, Vincent et al., “Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering”, arXiv:2106.02634v1 [cs.CV], Jun. 4, 2021, XP081984210. (10 pages total). [cited by applicant]
Aumentado-Armstrong, Tristan et al., “Representing 3D Shapes with Probabilistic Directed Distance Fields”, arXiv:2112.05300v1 [cs.CV], Dec. 10, 2021, XP093067478. (22 pages total). [cited by applicant]
Sitzmann, Vincent et al., “Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations”, arXiv: 1906.01618v2 [cs.CV], Jan. 28, 2020, XP081587552. (23 pages total). [cited by applicant]
Lin, Chieh Hubert et al., “COCO-GAN: Generation by Parts via Conditional Coordinating”, arXiv:1904.00284v4 [cs.LG], Jan. 5, 2020, XP081571766. (25 pages total). [cited by applicant]