IP Library › Granted Patent US 12,387,435
Granted Patent B2
US 12,387,435 · App. 17/711,695 · Granted Aug 12, 2025

Digital twin sub-millimeter alignment using multimodal 3D deep learning fusion system and method

Inventors: Yiyong Tan (Mountain View, CA); Bhaskar Banerjee (Mountain View, CA); Rishi Ranjan (Mountain View, CA)
Assignee: GRIDRASTER, INC.
G06T19/006G06F18/25G06N3/045G06N5/04G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,435
App. No.
17/711,695
Filed
Apr 1, 2022
Granted
Aug 12, 2025
Kind
B2
Art Unit
2662
USPC
345/633
Abstract

A mixed reality (MR) system and method performs alignment of a digital twin and the corresponding real-world object using 3D deep neural network structures using multimodal fusion and simplified machine learning to cluster label distributions (output of 3D deep neural network trained by generic 3D benchmark dataset) that are used to reduce the training data requirements to directly train a 3D deep neural network structures. In one embodiment, multiple 3D deep neural network structures, such as PointCNN, 3D-Bonet, RandLA, etc., may be trained by different generic 3D benchmark datasets, such as ScanNet, ShapeNet, S3DIS, inadequate 3D training dataset, etc.

Claims (57)

1. A method, comprising:

retrieving a point cloud for a three dimensional scene, the three dimensional scene having an object;

retrieving a digital twin of the object for use in a mixed reality environment;

training a deep learning computer system having a plurality of learning models each trained with a dataset, the deep learning computer system being used to provide improved input with higher corresponding confidence in points pairs between the three dimensional scene and a digital world to align the digital twin in the mixed reality environment with the object;

generating, for each learning model, at least a high confidence cluster having one or more high confidence points of the point cloud that are likely part of the object and a low confidence cluster having points of the point cloud likely to be background clutter, noise or distortion;

generating a union of the high confidence cluster from each of the high confidence clusters of each of the plurality of learning models;

performing a coarse alignment of the object and the digital twin using the points in the high confidence cluster union;

summing the points in the high confidence cluster for each of the plurality of learning models to generate a set of highest accumulated confidence points in an intersection region that is associated with the one or more high confidence points of the point cloud from all of the plurality of learning models in a same region; and

refining the coarse course alignment of the object and the digital twin using the intersection region that contains the set of highest accumulated confidence points.

2. The method of claim 1 , wherein training the plurality of learning models further comprises training three learning models using three different datasets.

3. The method of claim 2 , wherein training the plurality of learning models further comprises training a PointCNN model with a Scannet dataset to generate a first trained PointCNN model, training a second PointCNN model with a S3DIS dataset to generate a second trained PointCNN model, training a RandLA model with the S3DIS dataset to generate a RandLA trained model, training a 3DBotNet model with the S3DIS dataset to generate a first trained 3DBotNet model and training a second 3DBotNet model with inadequate data of the digital twin to generate a second trained 3DBotNet model.

4. The method claim 1 , wherein training the plurality of learning models further comprises pretraining five learning models.

5. The method of claim 1 , wherein refining the coarse alignment further comprises using an affine transform matrix to refine the coarse alignment of the object and the digital twin.

6. A system, comprising:

a computer system having a processor and a memory wherein the processor executes a plurality of lines of instructions so that the processor is configured to:

retrieve a point cloud for a three dimensional scene, the three dimensional scene having an object;

retrieve a digital twin of the object for use in a mixed reality environment;

train a plurality of learning models each trained with a dataset, the plurality of learning models being used to provide improved input with higher corresponding confidence in points pairs between the three dimensional scene and a digital world to align the digital twin in the mixed reality environment with the object;

generate, for each learning model, at least a high confidence cluster having one or more high confidence points of the point cloud that are likely part of the object and a low confidence cluster having points of the point cloud likely to be background clutter, noise or distortion;

perform a coarse alignment of the object and the digital twin using the points in a union of high confidence clusters from each of the plurality of learning models;

sum the points in the high confidence cluster for each of the plurality of learning models to generate a set of highest accumulated confidence points in an intersection region that is associated with the one or more high confidence points of the point cloud from all of the plurality of learning models in a same region; and

refine the coarse alignment of the object and the digital twin using the intersection region that has the set of highest accumulated confidence points.

7. The system of claim 6 , wherein the processor is further configured to generate a union of the one or more high confidence points of the point cloud in the high confidence cluster for each of the plurality of learning models to perform the coarse alignment.

8. The system of claim 6 , wherein the processor is further configured to train at least three different learning models using at least three different datasets.

9. The system of claim 8 , wherein the processor is further configured to train a PointCNN model with a Scannet dataset to generate a first trained PointCNN model, train a second PointCNN model with a S3DIS dataset to generate a second trained PointCNN model, train a RandLA model with the S3DIS dataset to generate a RandLA trained model, train a 3DBotNet model with the S3DIS dataset to generate a first trained 3DBotNet model and train a second 3DBotNet model with inadequate data of the digital twin to generate a second trained 3DBotNet model.

10. The system claim 6 , wherein the processor is further configured to pretrain five learning models.

11. The system of claim 6 , wherein the processor is further configured to use an affine transform matrix to refine the coarse alignment of the object and the digital twin.

12. A method to display an aligned digital twin for a real world object, the method comprising:

retrieving a point cloud for a three dimensional scene, the three dimensional scene having an object;

retrieving a digital twin of the object for use in a mixed reality environment;

training a deep learning computer system having a plurality of learning models each trained with a dataset, the deep learning computer system being used to provide improved input with higher corresponding confidence in points pairs between the three dimensional scene and a digital world to align the digital twin in the mixed reality environment with the object;

generating, for each learning model, at least a high confidence cluster having one or more high confidence points of the point cloud that are likely part of the object and a low confidence cluster having points of the point cloud likely to be background clutter, noise or distortion;

performing a coarse alignment of the object and the digital twin using the points in a union of high confidence clusters for each of the plurality of learning models;

summing the points in the high confidence cluster for each of the plurality of learning models to generate a set of highest accumulated confidence points in an intersection region that is associated with the one or more high confidence points of the point cloud from all of the learning models in a same region;

refining the coarse alignment of the object and the digital twin using the intersection region that has the set of highest accumulated confidence points; and

displaying the digital twin with sub-millimeter alignment based on the refined alignment.

13. The method of claim 12 , wherein performing the coarse alignment further comprises generating a union of the one or more high confidence points of the point cloud in the high confidence cluster for each of the plurality of learning models.

14. The method of claim 12 , wherein training the plurality of learning models further comprises training three different learning models using three different datasets.

15. The method of claim 14 , wherein training the plurality of learning models further comprises training a PointCNN model with a Scannet dataset to generate a first trained PointCNN model, training a second PointCNN model with a S3DIS dataset to generate a second trained PointCNN model, training a RandLA model with the S3DIS dataset to generate a RandLA trained model, training a 3DBotNet model with the S3DIS dataset to generate a first trained 3DBotNet model and training a second 3DBotNet model with inadequate data of the digital twin to generate a second trained 3DBotNet model.

16. The method claim 12 , wherein training the plurality of learning models further comprises pretraining five learning models.

17. The method of claim 12 , wherein refining the coarse alignment further comprises using an affine transform matrix to refine the coarse alignment of the object and the digital twin.

18. A virtual reality system, comprising:

a virtual reality headset connected to a computer system;

the computer system having a processor and a memory wherein the processor executes a plurality of lines of instructions so that the processor is configured to:

retrieve a point cloud for a three dimensional scene, the three dimensional scene having an object;

retrieve a digital twin of the object for use in a mixed reality environment;

train a plurality of learning models each trained with a dataset, the plurality of learning models being used to provide improved input with higher corresponding confidence in points pairs between the three dimensional scene and a digital world to align the digital twin in the mixed reality environment with the object;

generate, for each learning model, at least a high confidence cluster having one or more high confidence points of the point cloud that are likely part of the object and a low confidence cluster having points of the point cloud likely to be background clutter, noise or distortion;

perform a coarse alignment of the object and the digital twin using a union of the points in the high confidence clusters for each of the high confidence clusters from each of the plurality of learning models;

sum the points in the high confidence cluster for each of the plurality of learning models to generate a set of highest accumulated confidence points in an intersection region that is associated with the one or more high confidence points of the point cloud from all of the plurality of learning models in a same region; and

refine the coarse alignment of the object and the digital twin using the intersection region that has the set of highest accumulated confidence points; and

the virtual reality headset able to display the digital twin with sub-millimeter alignment based on the refined alignment.

19. The system of claim 18 , wherein the processor is further configured to generate a union of the one or more high confidence points of the point cloud in the high confidence cluster for each of the plurality of learning models to perform the coarse alignment.

20. The system of claim 18 , wherein the processor is further configured to train at least three different learning models using at least three different datasets.

21. The system of claim 20 , wherein the processor is further configured to train a PointCNN model with a Scannet dataset to generate a first trained PointCNN model, train a second PointCNN model with a S3DIS dataset to generate a second trained PointCNN model, train a RandLA model with the S3DIS dataset to generate a RandLA trained model, train a 3DBotNet model with the S3DIS dataset to generate a first trained 3DBotNet model and train a second 3DBotNet model with inadequate data of the digital twin to generate a second trained 3DBotNet model.

22. The system of claim 18 , wherein the processor is further configured to pretrain five learning models.

23. The system of claim 18 , wherein the processor is further configured to use an affine transform matrix to refine the coarse alignment of the object and the digital twin.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2022
From: TAN, YIYONG; BANERJEE, BHASKAR; RANJAN, RISHI
To: GRIDRASTER, INC.
Reel/Frame 059829/0828 →
Continuity (3)
Continuation In Part 17575091 · Jan 13, 2022
Continuation 17320968 · May 14, 2021
Related Publication 20230115887A1 · Apr 13, 2023
References Cited (38)
US 6160808A · Maurya · 2000 [cited by applicant]
US 9677840B2 · Rublowsky · 2017 [cited by applicant]
US 10867217B1 · Madden et al. · 2020 [cited by applicant]
US 20030065817A1 · Benchetrit · 2003 [cited by applicant]
US 20050027870A1 · Trebes, Jr. · 2005 [cited by applicant]
US 20050047427A1 · Kashima · 2005 [cited by applicant]
US 20060050697A1 · Li · 2006 [cited by applicant]
US 20060080454A1 · Li · 2006 [cited by applicant]
US 20060098662A1 · Gupta · 2006 [cited by applicant]
US 20060159079A1 · Sachs · 2006 [cited by applicant]
US 20070130585A1 · Perret · 2007 [cited by applicant]
US 20070140171A1 · Balasubramanian · 2007 [cited by applicant]
US 20070245010A1 · Arn et al. · 2007 [cited by applicant]
US 20070247457A1 · Gustafsson · 2007 [cited by applicant]
US 20070273610A1 · Baillot · 2007 [cited by applicant]
US 20080162670A1 · Chapweske · 2008 [cited by applicant]
US 20090046140A1 · Lashmet · 2009 [cited by applicant]
US 20090182815A1 · Czechowski, III · 2009 [cited by applicant]
US 20090196338A1 · Ali · 2009 [cited by applicant]
US 20090248872A1 · Luzzatti · 2009 [cited by applicant]
US 20090254659A1 · Li et al. · 2009 [cited by applicant]
US 20100225743A1 · Florencio · 2010 [cited by applicant]
US 20100253700A1 · Bergeron · 2010 [cited by applicant]
US 20110084983A1 · Demaine · 2011 [cited by applicant]
US 20110158311A1 · Abadir · 2011 [cited by applicant]
US 20120038739A1 · Welch · 2012 [cited by applicant]
US 20120106921A1 · Sasaki · 2012 [cited by applicant]
US 20130117377A1 · Miller · 2013 [cited by applicant]
US 20130286004A1 · McCulloch · 2013 [cited by applicant]
US 20160203646A1 · Nadler · 2016 [cited by applicant]
US 20170150122A1 · Cole · 2017 [cited by applicant]
US 20180157398A1 · Kaehler · 2018 [cited by applicant]
US 20190052883A1 · Ikeda · 2019 [cited by applicant]
US 20200281539A1 · Hoernig · 2020 [cited by examiner]
US 20200372709A1 · Ponjou Tasse · 2020 [cited by applicant]
US 20210133850A1 · Ayush · 2021 [cited by applicant]
US 20220027529A1 · Zarur · 2022 [cited by examiner]
Digital Twin Demo, PTC, 2020; https://www.youtube.com/watch?v=ERa8sN837hO (Year: 2020). [cited by applicant]
Cited By (1)
US 12,711,718