IP Library › Granted Patent US 12,008,790
Granted Patent B2
US 12,008,790 · App. 16/836,028 · Granted Jun 11, 2024

Encoding three-dimensional data for processing by capsule neural networks

Inventors: Nitish Srivastava (San Francisco, CA); Ruslan Salakhutdinov (Pittsburgh, PA); Hanlin Goh (Sunnyvale, CA)
Assignee: APPLE INC.
G06T9/002G06N3/047G06N3/08H04N13/111H04N13/161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,008,790
App. No.
16/836,028
Granted
Jun 11, 2024
Kind
B2
Abstract

A method includes defining a geometric capsule that is interpretable by a capsule neural network, wherein the geometric capsule includes a feature representation and a pose. The method also includes determining multiple viewpoints relative to the geometric capsule and determining a first appearance representation of the geometric capsule for each of the multiple viewpoints. The method also includes determining a transform for each of the multiple viewpoints that moves each of the multiple viewpoints to a respective transformed viewpoint and determining second appearance representations that each correspond to one of the transformed viewpoints. The method also includes combining the second appearance representations to define an agreed appearance representation. The method also includes updating the feature representation for the geometric capsule based on the agreed appearance representation.

Claims (69)

1. A method, comprising:

defining a geometric capsule, wherein the geometric capsule is an encoded form of three-dimensional data that is interpretable by a capsule neural network, and the geometric capsule includes a three-dimensional feature representation and a pose;

determining multiple viewpoints relative to the geometric capsule;

determining a first appearance representation of the geometric capsule for each of the multiple viewpoints based on the three-dimensional feature representation for the geometric capsule;

determining a transform for each of the multiple viewpoints that moves each of the multiple viewpoints to a respective transformed viewpoint;

determining second appearance representations that each correspond to one of the transformed viewpoints based on the three-dimensional feature representation for the geometric capsule;

combining the second appearance representations to define an agreed appearance representation; and

updating the three-dimensional feature representation for the geometric capsule based on the agreed appearance representation.

2. The method of claim 1 , wherein defining the geometric capsule includes:

receiving a group of elements that represent a three-dimensional visual entity, and

determining an assignment of at least some elements from the group of elements to the geometric capsule.

3. The method of claim 2 , wherein updating the three-dimensional feature representation for the geometric capsule based on the agreed upon appearance representation includes updating the assignment of at least some elements from the group of elements to the geometric capsule.

4. The method of claim 2 , wherein the group of elements is a point cloud and each element from the group of elements is a point that is included in the point cloud.

5. The method of claim 2 , wherein the group of elements is a group of lower-level geometric capsules.

6. The method of claim 1 , wherein determining the transform for each of the multiple viewpoints is performed using a trained neural network.

7. The method of claim 6 , wherein the trained neural network is configured to determine the transform for each of the multiple viewpoints such that the second appearance representations corresponding to the transformed viewpoints are constrained to match while the transformed viewpoints do not match.

8. A non-transitory computer-readable storage device including program instructions executable by one or more processors that, when executed, cause the one or more processors to perform operations, the operations comprising:

defining a geometric capsule, wherein the geometric capsule is an encoded form of three-dimensional data that is interpretable by a capsule neural network, and the geometric capsule includes a three-dimensional feature representation and a pose;

determining multiple viewpoints relative to the geometric capsule;

determining a first appearance representation of the geometric capsule for each of the multiple viewpoints based on the three-dimensional feature representation for the geometric capsule;

determining a transform for each of the multiple viewpoints that moves each of the multiple viewpoints to a respective transformed viewpoint;

determining second appearance representations that each correspond to one of the transformed viewpoints based on the three-dimensional feature representation for the geometric capsule;

combining the second appearance representations to define an agreed appearance representation; and

updating the three-dimensional feature representation for the geometric capsule based on the agreed appearance representation.

9. The non-transitory computer-readable storage device of claim 8 , wherein defining the geometric capsule includes:

receiving a group of elements that represent a three-dimensional visual entity, and

determining an assignment of at least some elements from the group of elements to the geometric capsule.

10. The non-transitory computer-readable storage device of claim 9 , wherein updating the three-dimensional feature representation for the geometric capsule based on the agreed upon appearance representation includes updating the assignment of at least some elements from the group of elements to the geometric capsule.

11. The non-transitory computer-readable storage device of claim 9 , wherein the group of elements is a point cloud and each element from the group of elements is a point that is included in the point cloud.

12. The non-transitory computer-readable storage device of claim 9 , wherein the group of elements is a group of lower-level geometric capsules.

13. The non-transitory computer-readable storage device of claim 8 , wherein determining the transform for each of the multiple viewpoints is performed using a trained neural network.

14. The non-transitory computer-readable storage device of claim 13 , wherein the trained neural network is configured to determine the transform for each of the multiple viewpoints such that the second appearance representations corresponding to the transformed viewpoints are constrained to match while the transformed viewpoints do not match.

15. A system, comprising:

a memory that includes program instructions; and

a processor that is operable to execute the program instructions, wherein the program instructions, when executed by the processor, cause the processor to:

define a geometric capsule, wherein the geometric capsule is an encoded form of three-dimensional data that is interpretable by a capsule neural network, and the geometric capsule includes a three-dimensional feature representation and a pose;

determine multiple viewpoints relative to the geometric capsule;

determine a first appearance representation of the geometric capsule for each of the multiple viewpoints based on the three-dimensional feature representation for the geometric capsule;

determine a transform for each of the multiple viewpoints that moves each of the multiple viewpoints to a respective transformed viewpoint;

determine second appearance representations that each correspond to one of the transformed viewpoints based on the three-dimensional feature representation for the geometric capsule;

combine the second appearance representations to define an agreed appearance representation; and

update the three-dimensional feature representation for the geometric capsule based on the agreed appearance representation.

16. The system of claim 15 , wherein the program instructions to define the geometric capsule further cause the processor to:

receive a group of elements that represent a three-dimensional visual entity, and

determine an assignment of at least some elements from the group of elements to the geometric capsule.

17. The system of claim 16 , wherein updating the three-dimensional feature representation for the geometric capsule based on the agreed upon appearance representation includes updating the assignment of at least some elements from the group of elements to the geometric capsule.

18. The system of claim 16 , wherein the group of elements is a point cloud and each element from the group of elements is a point that is included in the point cloud.

19. The system of claim 16 , wherein the group of elements is a group of lower-level geometric capsules.

20. The system of claim 15 , wherein determining the transform for each of the multiple viewpoints is performed using a trained neural network.

21. The system of claim 20 , wherein the trained neural network is configured to determine the transform for each of the multiple viewpoints such that the second appearance representations corresponding to the transformed viewpoints are constrained to match while the transformed viewpoints do not match.

22. The method of claim 1 , wherein the geometric capsule represents a portion of a three-dimensional visual entity, the three-dimensional feature representation is encoded in the form of a pose-invariant feature vector, the pose describes a location and orientation of the portion of the three-dimensional visual entity, and determining the transform for each of the multiple viewpoints is based on the first appearance representations for each of the multiple viewpoints.

23. The method of claim 1 , wherein:

the first appearance representations are each described by a respective first Gaussian distribution,

the second appearance representations are each described by a respective second Gaussian distribution, and

combining the second appearance representations to define the agreed appearance representation includes setting the agreed appearance representation equal to a product of the second Gaussian distributions.

24. The non-transitory computer-readable storage device of claim 8 , wherein the geometric capsule represents a portion of a three-dimensional visual entity, the three-dimensional feature representation is encoded in the form of a pose-invariant feature vector, the pose describes a location and orientation of the portion of the three-dimensional visual entity, and determining the transform for each of the multiple viewpoints is based on the first appearance representations for each of the multiple viewpoints.

25. The non-transitory computer-readable storage device of claim 8 , wherein:

the first appearance representations are each described by a respective first Gaussian distribution,

the second appearance representations are each described by a respective second Gaussian distribution, and

combining the second appearance representations to define the agreed appearance representation includes setting the agreed appearance representation equal to a product of the second Gaussian distributions.

26. The system of claim 15 , wherein:

the geometric capsule represents a portion of a three-dimensional visual entity, the three-dimensional feature representation is encoded in the form of a pose-invariant feature vector, the pose describes a location and orientation of the portion of the three-dimensional visual entity, and determining the transform for each of the multiple viewpoints is based on the first appearance representations for each of the multiple viewpoints.

27. The system of claim 15 , wherein:

the first appearance representations are each described by a respective first Gaussian distribution,

the second appearance representations are each described by a respective second Gaussian distribution, and

combining the second appearance representations to define the agreed appearance representation includes setting the agreed appearance representation equal to a product of the second Gaussian distributions.

28. The method of claim 1 , wherein the three-dimensional feature representation is encoded in the form of a feature vector that describes a three-dimensional surface.

29. The non-transitory computer-readable storage device of claim 8 , wherein the three-dimensional feature representation is encoded in the form of a feature vector that describes a three-dimensional surface.

30. The system of claim 15 , wherein the three-dimensional feature representation is encoded in the form of a feature vector that describes a three-dimensional surface.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2020
From: SRIVASTAVA, NITISH; SALAKHUTDINOV, RUSLAN; GOH, HANLIN
To: APPLE INC.
Reel/Frame 052284/0563 →
Continuity (2)
Provisional Application 62904890 · Sep 24, 2019
Related Publication 20210090302A1 · Mar 25, 2021