IP Library › Granted Patent US 12,731,329
Granted Patent B2
US 12,731,329 · App. 18/086,475 · Granted Sep 8, 2026

Systems and methods for a shape completion model

Inventors: Hongyu Li (Chestnut Hill, MA); Nawid Jamali (San Francisco, CA); Soshi Iba (Mountain View, CA)
Assignee: Honda Motor Co., Ltd.
G06T17/00B25J9/1697G06T7/50G06T2207/10028G06T2207/20084G06T2210/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,329
App. No.
18/086,475
Filed
Dec 21, 2022
Granted
Sep 8, 2026
Kind
B2
Art Unit
2665
USPC
345/424
Abstract

Systems and methods for shape completion are provided. In one embodiment, a computer implemented method includes receiving sensor data for a visualized area of an object as at least one point cloud representation. The computer implemented method also includes transforming the at least one point cloud representation into an input voxel grid of the visualized area of the object. The input voxel grid is a volumetric representation. The computer implemented method further includes encoding the input voxel grid into a partial latent vector that lies on a partial latent space. The computer implemented method yet further includes determining a mapping between the partial latent space and a complete latent space based on the sensor data. The computer implemented method includes predicting a complete latent vector based on the complete latent space. The computer implemented method also includes estimating a complete shape of an object based on the complete latent space.

Claims (41)

1 . A system for shape completion, comprising:

a processor, and

a memory storing instructions that when executed by the processor cause the processor to:

receive sensor data for a visualized area of an object as at least one point cloud representation;

transform the at least one point cloud representation into an input voxel grid of the visualized area of the object, wherein the input voxel grid is a volumetric representation;

encode the input voxel grid into a partial latent vector that lies on a partial latent space;

extract visual features from the sensor data;

input the partial latent vector, the visual features, and a Gaussian latent code into a conditional generative model;

generate, by the conditional generative model, a complete latent vector in a complete latent space; and

decode, by a decoder of an autoencoder and using the generated complete latent vector as input to the decoder, a voxel representation of the object, wherein the reconstructed object includes the visualized area of the object and a previously occluded portion of the object.

2 . The system of claim 1 , wherein the mapping is further based on visual features extracted from the sensor data.

3 . The system of claim 2 , wherein the system of claim 1 includes an autoencoder having a generator, and wherein visual features of the sensor data are input into the generator as conditional input, the generator being part of the conditional generative model.

4 . The system of claim 1 , wherein the system of claim 1 includes an autoencoder having an encoder to encode the input voxel grid and a decoder to estimate the complete shape, wherein the encoder generates the partial latent vector and the decoder reconstructs the voxel representation of the object.

5 . The system of claim 4 , wherein the autoencoder is optimized by minimizing Jaccard index loss between reconstructed voxel grids and ground truth voxel grids.

6 . The system of claim 1 , wherein the input voxel grid is encoded into the partial latent vector using a set of three-dimensional convolutional layers.

7 . The system of claim 1 , wherein the at least one point cloud representation is normalized based on a centroid of the at least one point cloud representation and a farthest distance of the at least one point cloud representation from the centroid.

8 . A computer implemented method for shape completion, comprising:

receiving sensor data from an agent for a visualized area of an object as at least one point cloud representation;

transforming the at least one point cloud representation into an input voxel grid of the visualized area of the object, wherein the input voxel grid is a volumetric representation;

encoding the input voxel grid into a partial latent vector that lies on a partial latent space;

determining a mapping between the partial latent space and a complete latent space based on the sensor data;

generating a complete latent vector representing the object based on the mapping; and

decoding the complete latent vector to reconstruct a voxel representation of the object, wherein the reconstructed object includes the visualized area of the object and a previously occluded portion of the object.

9 . The computer implemented method of claim 8 , wherein the mapping is based on visual features extracted from the sensor data.

10 . The computer implemented method of claim 8 , further comprising extracting visual features from the sensor data, wherein the generating the complete latent vector is further based on the extracted visual features as conditional input.

11 . The computer implemented method of claim 8 , further comprising performing optimization by minimizing Jaccard index loss.

12 . The computer implemented method of claim 8 , wherein the input voxel grid is encoded into the partial latent vector using a set of three-dimensional convolutional layers.

13 . The computer implemented method of claim 8 , wherein the at least one point cloud representation is normalized based on a centroid of the at least one point cloud representation and a farthest distance of the at least one point cloud representation from the centroid.

14 . The computer implemented method of claim 8 , wherein the voxel representation is used for object shape completion in a scenario including one or more of self-occluded object shape completion, in-hand object shape completion, and cluttered object shape completion by the agent.

15 . A non-transitory computer readable storage medium storing instructions that when executed by a computer having a processor to perform a method for shape completion, the method comprising:

receiving sensor data for a visualized area of an object as at least one point cloud representation;

transforming the at least one point cloud representation into an input voxel grid of the visualized area of the object, wherein the input voxel grid is a volumetric representation;

encoding the input voxel grid into a partial latent vector that lies on a partial latent space;

determining a mapping between the partial latent space and a complete latent space based on the sensor data;

generating a complete latent vector representing the object based on the mapping; and

decoding the complete latent vector to reconstruct a voxel representation of the object, wherein the reconstructed object includes the visualized area of the object and a previously occluded portion of the object.

16 . The non-transitory computer readable storage medium of claim 15 , wherein the mapping is based on visual features extracted from the sensor data.

17 . The non-transitory computer readable storage medium of claim 15 , the method further comprising extracting visual features from the sensor data, wherein the generating the complete latent vector is further based on the visual features as conditional input.

18 . The non-transitory computer readable storage medium of claim 15 , the method further comprising performing optimization by minimizing Jaccard index loss.

19 . The non-transitory computer readable storage medium of claim 15 , wherein the input voxel grid is encoded into the partial latent vector using a set of three-dimensional convolutional layers.

20 . The non-transitory computer readable storage medium of claim 15 , wherein the at least one point cloud representation is normalized based on a centroid of the at least one point cloud representation and a farthest distance of the at least one point cloud representation from the centroid.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2022
From: LI, HONGYU; JAMALI, NAWID; IBA, SOSHI
To: HONDA MOTOR CO., LTD.
Reel/Frame 062176/0984 →
Continuity (1)
Related Publication 20260038197A1 · Feb 5, 2026
References Cited (57)
US 12322068B1 · Kim · 2025 [cited by examiner]
US 20200167956A1 · Herman et al. · 2020 [cited by applicant]
US 20200242736A1 · Liu et al. · 2020 [cited by applicant]
US 20220058827A1 · Montserrat et al. · 2022 [cited by applicant]
US 20220084241A1 · Dikhale et al. · 2022 [cited by applicant]
US 20230145048A1 · Walker et al. · 2023 [cited by applicant]
US 20240096017A1 · Gao · 2024 [cited by examiner]
US 20240135721A1 · Ambrus · 2024 [cited by examiner]
US 20250111628A1 · Han et al. · 2025 [cited by applicant]
Watkins, David. Learning Mobile Manipulation. Columbia University, 2022. (Year: 2022). [cited by examiner]
Mohamed Tahoun, Omar Tahri, Juan Antonio Corrales Ramón, and Youcef Mezouar. Visual-Tactile Fusion for 3D Objects Reconstruction from a Single Depth View and a Single Gripper Touch for Robotics Tasks. In 2021 IEEE/RSJ I… [cited by applicant]
P. J. Besl and N. D. McKay, “Method for registration of 3-D shapes,” in Sensor Fusion IV: Control Paradigms and Data Structures, vol. 1611. SPIE, Apr. 1992, pp. 586-606. [cited by applicant]
J. Bimbo, S. Luo, K. Althoefer, and H. Liu, “In-Hand Object Pose Estimation Using Covariance-Based Tactile To Geometry Matching,” IEEE Robotics and Automation Letters, vol. 1, No. 1, pp. 570-577, Jan. 2016, conference N… [cited by applicant]
Y. Cai, K.-Y. Lin, C. Zhang, Q. Wang, X. Wang, and H. Li, “Learning a Structured Latent Space for Unsupervised Point Cloud Completion,” 2022, pp. 5543-5553. [cited by applicant]
B. Calli, A. Singh, A. Walsman, S. Srinivasa, P. Abbeel, and A. M. Dollar, “The YCB object and Model set: Towards common benchmarks for manipulation research,” in 2015 International Conference on Advanced Robotics (ICAR… [cited by applicant]
Y. F. Chen, M. Everett, M. Liu, and J. P. How, “Socially aware motion planning with deep reinforcement learning,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 2017, pp. 1343-1… [cited by applicant]
H. Chen, P. Wang, F. Wang, W. Tian, L. Xiong, and H. Li, “EPro-PnP: Generalized End-to-End Probabilistic Perspective-N-Points for Monocular Object Pose Estimation,” 2022, pp. 2781-2790. [cited by applicant]
T. Chen, J. Xu, and p. Agrawal, “A System for General In-Hand Object Re-Orientation,” in Proceedings of the 5th Conference on Robot Learning. PMLR, Jan. 2022, pp. 297-307, ISSN: 2640-3498. [cited by applicant]
Epic Games, “Unreal Engine,” Apr. 2019. [cited by applicant]
M. A. Fischler and R. C. Bolles, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,” Communications of the ACM, vol. 24, No. 6, pp. 381-395, Jun. 1981. [cited by applicant]
G. Gao, M. Lauri, Y. Wang, X. Hu, J. Zhang, and S. Frintrop, “6D Object Pose Regression via Supervised Learning on Point Clouds,” in 2020 IEEE International Conference on Robotics and Automation (ICRA), May 2020, pp. 36… [cited by applicant]
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? The KITTI vision benchmark suite,” in 2012 EEE Conference on Computer Vision and Pattern Recognition, Jun. 2012, pp. 3354-3361, iSSN: 1063-6919. [cited by applicant]
S. Hinterstoisser, S. Holzer, C. Cagniart, S. llic, K. Konolige, N. Navab, and V. Lepetit, “Multimodal templates for real-time detection of textureless objects in heavily cluttered scenes,” in 2011 International Confere… [cited by applicant]
Y. Hu, J. Hugonot, p. Fua, and M. Salzmann, “Segmentation-Driven 6D Object Pose Estimation,” 2019, pp. 3385-3394. [cited by applicant]
Z. Jiang, Y. Zhu, M. Svetlik, K. Fang, and Y. Zhu, “Synergies Between Affordance and Geometry: 6-DoF Grasp Detection via Implicit Representations,” Jul. 2021, arXiv:2104.01542 [cs]. [cited by applicant]
H. Li, Z. Li, N. U. Akmandor, H. Jiang, Y. Wang, and T. Padir, “Stereovoxelnet: Real-time obstacle detection based on occupancy voxels from a stereo camera using deep neural networks,” in 2023 International Conference o… [cited by applicant]
Y. Li, G. Wang, X. Ji, Y. Xiang, and D. Fox, “DeepIM: Deep Iterative Matching for 6D Pose Estimation,” 2018, pp. 683-698. [cited by applicant]
K. Park, T. Patten, and M. Vincze, “Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation,” 2019, pp. 7668-7677. [cited by applicant]
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space,” in Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wa… [cited by applicant]
Y. Qin, H. Su, and X. Wang, “From One Hand to Multiple Hands: Imitation Learning for Dexterous Manipulation from Single-Camera Teleoperation,” Apr. 2022, arXiv:2204.12490 [cs]. [cited by applicant]
M. Tatarchenko, A. Dosovitskiy, and T. Brox, “Octree Generating Networks: Efficient Convolutional Architectures for High-Resolution 3D Outputs,” in Proceedings of the IEEE International Conference on Computer Vision (IC… [cited by applicant]
J. Varley, C. DeChant, A. Richardson, J. Ruales, and P. Allen, “Shape completion enabled robotic grasping,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 2017, pp. 2442-2447, I… [cited by applicant]
M. Bauza Villalonga, A. Rodriguez, B. Lim, E. Valls, and T. Sechopoulos, “Tactile Object Pose Estimation from the First Touch with Geometric Contact Rendering,” in Proceedings of the 2020 Conference on Robot Learning. P… [cited by applicant]
G. Wang, F. Manhardt, F. Tombari, and X. Ji, “GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose Estimation,” 2021, pp. 16 611-16 621. [cited by applicant]
H. Wang, S. Sridhar, J. Huang, J. Valentin, S. Song, and L. J. Guibas, “Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,” 2019, pp. 2642-2651. [cited by applicant]
S. Wang, J. Wu, X. Sun, W. Yuan, W. T. Freeman, J. B. Tenenbaum, and E. H. Adelson, “3D Shape Perception from Monocular Vision, Touch, and Shape Priors,” in 2018 IEEE/RSJ International Conference on Intelligent Robots a… [cited by applicant]
Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox, “PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,” May 2018, arXiv:1711.00199 [cs]. [cited by applicant]
A. Yamaguchi and C. G. Atkeson, “Recent progress in tactile sensing and sensors for robotic manipulation: can we turn tactile sensing into vision?” Advanced Robotics, vol. 33, No. 14, pp. 661-673, 2019, publisher: Taylo… [cited by applicant]
J. Zhang, X. Chen, Z. Cai, L. Pan, H. Zhao, S. Yi, C. K. Yeo, B. Dai, and C. C. Loy, “Unsupervised 3D Shape Completion Through GAN Inversion,” 2021, pp. 1768-1777. [cited by applicant]
K. Zhang, Y. Fu, S. Borse, H. Cai, F. Porikli, and X. Wang, “Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the Wild,” Oct. 2022, arXiv:2210.07199 [cs]. [cited by applicant]
Q.-Y. Zhou, J. Park, and V. Koltun, “Open3D: A Modern Library for 3D Data Processing,” Jan. 2018, arXiv:1801.09847 [cs]. [cited by applicant]
Snehal Dikhale, Karankumar Patel, Daksh Dhingra, Itoshi Naramura, Akinobu Hayashi, Soshi Iba, and Nawid Jamali. Visuo Tactile 6D Pose Estimation of an In-Hand Object Using Vision and Tactile Sensor Data. IEEE Robotics a… [cited by applicant]
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Networks, Jun. 2014. URL http://arxiv.org/abs/1406.2661. arXiv: … [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. pp. 770-778, 2016. URL https://openaccess.thecvf.com/content_cvpr_2016/html/He_Deep_Residual_Learning_CVPR_2016_paper.… [cited by applicant]
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R. Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors, Jul. 2012. URL http://arxiv.org/abs/1207. … [cited by applicant]
Michelle A. Lee, Yuke Zhu, Peter Zachares, Matthew Tan, Krishnan Srinivasan, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Jeannette Bohg. Making Sense of Vision and Touch: Learning Multimodal Representations for Conta… [cited by applicant]
Xudong Mao, Qing Li, Haoran Xie, Raymond Y. K. Lau, Zhen Wang, and Stephen Paul Smolley. Least Squares Generative Adversarial Networks, Apr. 2017. URL http://arxiv.org/abs/1611.04076. arXiv:1611.04076 [cs]. [cited by applicant]
Mehdi Mirza and Simon Osindero. Conditional Generative Adversarial Nets, Nov. 2014. URL http://arxiv.org/abs/1411.1784. arXiv: 1411.1784 [cs, stat]. [cited by applicant]
Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martin-Martin, Cewu Lu, Li Fei-Fei, and Silvio Savarese. DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion. pp. 3343-3352, 2019. URL https://openaccess.thecvf.com/c… [cited by applicant]
David Watkins-Valls, Jacob Varley, and Peter Allen. Multi-Modal Geometric Learning for Grasping and Manipulation, Feb. 2019. URL http://arxiv.org/abs/1803.07671. arXiv:1803.07671 [cs]. [cited by applicant]
Rundi Wu, Xuelin Chen, Yixin Zhuang, and Baoquan Chen. Multimodal Shape Completion via Conditional Generative Adversarial Networks, Jul. 2020. URL http://arxiv.org/abs/2003.07717. arXiv:2003.07717 [cs]. [cited by applicant]
Office Action of U.S. Appl. No. 18/311,024 dated Dec. 11, 2025, 24 pages. [cited by applicant]
Liuhao Ge, “Hang PointNet: 3D Hand Pose Estimation using Point Sets”, 2018 (Year: 2018). [cited by applicant]
Ankit Kumar, “Context-aware 6D Pose Estimation of Known Objects using RGB-D data”, Dec. 2022 (Year: 2022). [cited by applicant]
Office Action of U.S. Appl. No. 18/311,024 dated Jul. 31, 2025, 26 pages. [cited by applicant]
Ho Jin Choi, “Gaussian Process-Based Active Exploration Strategies in Vision and Touch”, Jul. 2025 (Year: 2025). [cited by applicant]
Office Action of U.S. Appl. No. 18/311,024 dated Mar. 19, 2026, 28 pages. [cited by applicant]