IP Library Granted Patent US 12,496,721
Granted Patent B2
US 12,496,721 · App. 17/314,866 · Granted Dec 16, 2025

Virtual presence for telerobotics in a dynamic scene

Inventors: Ken Lee (Fairfax, VA); Craig Cambias (Silver Spring, MD); Xin Hou (Herndon, VA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
B25J9/1689B25J9/1653B25J9/1697B25J13/006G05D1/0038G05D1/0044
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,496,721
App. No.
17/314,866
Granted
Dec 16, 2025
Kind
B2
Abstract

Described herein are methods and systems for providing virtual presence for telerobotics in a dynamic scene. A sensor captures frames of a scene comprising one or more objects. A computing device generates a set of feature points corresponding to objects in the scene and matches the set of feature points to 3D points in a map of the scene. The computing device generates a dense mesh of the scene and the objects using the matched feature points and transmits the dense mesh the frame to a remote viewing device. The remote viewing device generates a 3D representation of the scene and the objects for display to a user and receives commands from the user corresponding to interaction with the 3D representation of the scene. The remote viewing device transmits the commands to a robot device that executes the commands to perform operations on the objects in the scene.

Claims (72)

1 . A system for providing virtual presence for telerobotics in a dynamic scene, the system comprising:

a remote viewing device and a remote controller coupled to the remote viewing device;

a sensor device that captures one or more frames of a scene comprising one or more objects, each frame comprising (i) one or more color images of the scene and the one or more objects and (ii) one or more depth maps of the scene and the one or more objects;

a computing device coupled to the sensor device, the computing device comprising a memory that stores computer-executable instructions and a processor that executes the instructions to:

receive a first set of frames from the sensor device;

for each frame in the first set of frames:

generate a set of feature points corresponding to the one or more objects in the scene,

match the set of feature points to one or more corresponding 3D points in a depth map from the frame, and

construct a dense mesh of the scene and the one or more objects using the matched feature points;

receive a second set of frames from the sensor device;

for each frame in the second set of frames:

a) calculate a geometric error between one or more segments of the dense mesh and corresponding 3D points in a depth map from the frame,

b) classify one or more segments of the dense mesh as one or more dynamic segments based on the geometric error calculated for the one or more segments,

c) non-rigidly deform only the one or more segments of the dense mesh that are classified as the one or more dynamic segments using data associated with the corresponding 3D points in the depth map from the frame,

d) calculate a second geometric error between the deformed one or more dynamic segments and corresponding 3D points in the depth map from the frame,

e) modify the deformed dense mesh using a non-linear optimization if the second geometric error does not converge,

f) repeat steps d) and e) until the second geometric error converges, and

g) update a texture of the deformed dense mesh by aligning a color image from the frame to the deformed dense mesh; and

determine net changes to the dense mesh resulting from the non-rigid deformation and transmit (i) the net changes to the dense mesh of the scene and the one or more objects and (ii) the frame to the remote viewing device,

wherein the remote viewing device configured to:

update an existing 3D representation of the scene and the one or more objects using the net changes to the dense mesh and the frame for display to a user;

receive one or more commands from the user via the remote controller, the one or more commands corresponding to interaction with the one or more objects in the updated 3D representation of the scene, and

transmit the one or more commands to a robot device, and

wherein the robot device is configured to:

execute the one or more commands received from the remote viewing device to perform one or more operations.

2 . The system of claim 1 , wherein generating the set of feature points corresponding to the one or more objects in the scene comprises detecting one or more feature points in the frame using a corner detection algorithm.

3 . The system of claim 2 , wherein matching the set of feature points to the one or more corresponding 3D points in the depth map from the frame comprises using a feature descriptor to match the set of feature points to the one or more corresponding 3D points.

4 . The system of claim 1 , wherein matching the set of feature points to the one or more corresponding 3D points in the depth map from the frame comprises minimizing a projection error between each feature point and one or more corresponding 3D points.

5 . The system of claim 4 , wherein the minimizing the projection error is performed using a nonlinear optimization algorithm.

6 . The system of claim 1 , wherein the updating the existing 3D representation of the scene and the one or more objects using the net changes to the dense mesh and the frame comprises:

detecting one or more keypoints of one or more objects in the scene using the received frame;

matching the detected one or more keypoints to one or more 3D points in a stored map to generate a point cloud;

updating the generated point cloud using the net changes to the dense mesh received from the computing device; and

mapping the frame onto a surface of the generated point cloud to update the 3D representation.

7 . The system of claim 6 , wherein updating the generated point cloud using the net changes to the dense mesh is performed using an Iterative Closest Point (ICP) algorithm.

8 . The system of claim 6 , wherein the 3D representation comprises a textured mesh of the scene and the one or more objects in the scene.

9 . The system of claim 1 , wherein the remote viewing device comprises an augmented reality (AR) viewing apparatus, a virtual reality (VR) viewing apparatus, or a mixed reality (MR) viewing apparatus.

10 . The system of claim 9 , wherein the remote viewing device is worn by the user.

11 . A computerized method for providing virtual presence for telerobotics in a dynamic scene, the method comprising:

capturing, by a sensor device, one or more frames of a scene comprising one or more objects, each frame comprising (i) one or more color images of the scene and the one or more objects and (ii) one or more depth maps of the scene and the one or more objects;

receiving, by a computing device coupled to the sensor device, a first set of frames from the sensor device;

for each frame in the first set of frames:

generating, by the computing device coupled to the sensor device, a set of feature points corresponding to the one or more objects in the scene,

matching, by the computing device, the set of feature points to one or more corresponding 3D points in a depth map from the frame, and

constructing, by the computing device, a dense mesh of the scene and the one or more objects using the matched feature points;

receiving, by the computing device, a second set of frames from the sensor device;

for each frame in the second set of frames:

a) calculating, by the computing device, a first geometric error between one or more segments of the dense mesh and corresponding 3D points in a depth map from the frame,

b) classifying, by the computing device, one or more segments of the dense mesh as one or more dynamic segments based on the first geometric error calculated for the one or more segments,

c) non-rigidly deforming, by the computing device, only the one or more segments of the dense mesh that are classified as the one or more dynamic segments using data associated with the corresponding 3D points in the depth map from the frame,

d) calculate a second geometric error between the deformed one or more dynamic segments and corresponding 3D points in the depth map from the frame,

e) modify the deformed dense mesh using a non-linear optimization if the second geometric error does not converge,

f) repeating, by the computing device, steps d) and e) until the second geometric error converges, and

g) updating, by the computing device, a texture of the deformed dense mesh by aligning a color image from the frame to the deformed dense mesh;

determining, by the computing device, net changes to the dense mesh resulting from the non-rigid deformation and transmitting (i) the net changes to the dense mesh of the scene and the one or more objects and (ii) the frame to a remote viewing device, the remote viewing device coupled to a remote controller;

updating, by the remote viewing device, an existing 3D representation of the scene and the one or more objects using the net changes to the dense mesh and the frame for display to a user;

receiving, by the remote viewing device, one or more commands from the user via the remote controller, the one or more commands corresponding to interaction with the one or more objects in the updated 3D representation of the scene;

transmitting, by the remote viewing device, the one or more commands to a robot device that interacts with the one or more objects in the scene; and

executing, by the robot device, the one or more commands received from the remote viewing device to perform one or more operations.

12 . The method of claim 11 , wherein the generating the set of feature points corresponding to the one or more objects in the scene comprises detecting one or more feature points in the frame using a corner detection algorithm.

13 . The method of claim 12 , wherein the matching the set of feature points to the one or more corresponding 3D points in the depth map from the frame comprises using a feature descriptor to match the set of feature points to the one or more corresponding 3D points.

14 . The method of claim 11 , wherein the matching the set of feature points to the one or more corresponding 3D points in the depth map from the frame comprises minimizing a projection error between each feature point and one or more corresponding 3D points.

15 . The method of claim 14 , wherein the minimizing the projection error is performed using a nonlinear optimization algorithm.

16 . The method of claim 11 , wherein the updating the existing 3D representation of the scene and the one or more objects using the net changes to the dense mesh and the frame comprises:

detecting one or more keypoints of one or more objects in the scene using the received frame;

matching the detected one or more keypoints to one or more 3D points in a stored map to generate a point cloud;

updating the generated point cloud using the net changes to the dense mesh received from the computing device; and

mapping the frame onto a surface of the generated point cloud to update the 3D representation.

17 . The method of claim 16 , wherein updating the generated point cloud using the net changes to the dense mesh is performed using an Iterative Closest Point (ICP) algorithm.

18 . The method of claim 16 , where in the 3D representation comprises a textured mesh of the scene and the one or more objects in the scene.

19 . The method of claim 11 , wherein the remote viewing device comprises an augmented reality (AR) viewing apparatus, a virtual reality (VR) viewing apparatus, or a mixed reality (MR) viewing apparatus.

20 . The method of claim 19 , wherein the remote viewing device is worn by the user.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2025
From: VANGOGH IMAGING, INC.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 070560/0391 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2023
From: LEE, KEN; CAMBIAS, CRAIG; HOU, XIN
To: VANGOGH IMAGING, INC.
Reel/Frame 062927/0615 →
Continuity (2)
Provisional Application 63022113 · May 8, 2020
Related Publication 20210347053A1 · Nov 11, 2021
References Cited (132)
US 5675326A · Juds et al. · 1997 [cited by applicant]
US 6259815B1 · Anderson et al. · 2001 [cited by applicant]
US 6275235B1 · Morgan, III · 2001 [cited by applicant]
US 6525722B1 · Deering · 2003 [cited by applicant]
US 6525725B1 · Deering · 2003 [cited by applicant]
US 7248257B2 · Elber · 2007 [cited by applicant]
US 7420555B1 · Lee · 2008 [cited by applicant]
US 7657081B2 · Blais et al. · 2010 [cited by applicant]
US 8209144B1 · Anguelov et al. · 2012 [cited by applicant]
US 8542233B2 · Brown · 2013 [cited by applicant]
US 8766979B2 · Lee et al. · 2014 [cited by applicant]
US 8942917B2 · Chrysanthakopoulos · 2015 [cited by applicant]
US 8995756B2 · Lee et al. · 2015 [cited by applicant]
US 9041711B1 · Hsu · 2015 [cited by applicant]
US 9104908B1 · Rogers et al. · 2015 [cited by applicant]
US 9171402B1 · Allen et al. · 2015 [cited by applicant]
US 9607388B2 · Lin et al. · 2017 [cited by applicant]
US 9710960B2 · Hou · 2017 [cited by applicant]
US 9886530B2 · Mehr et al. · 2018 [cited by applicant]
US 9978177B2 · Mehr et al. · 2018 [cited by applicant]
US 10467792B1 · Roche et al. · 2019 [cited by applicant]
US 20050068317A1 · Amakai · 2005 [cited by applicant]
US 20050128201A1 · Warner et al. · 2005 [cited by applicant]
US 20050253924A1 · Mashitani · 2005 [cited by applicant]
US 20060050952A1 · Blais et al. · 2006 [cited by applicant]
US 20060170695A1 · Zhou et al. · 2006 [cited by applicant]
US 20060277454A1 · Chen · 2006 [cited by applicant]
US 20070075997A1 · Rohaly et al. · 2007 [cited by applicant]
US 20080180448A1 · Anguelov et al. · 2008 [cited by applicant]
US 20080310757A1 · Wolberg · 2008 [cited by examiner]
US 20090232353A1 · Sundaresan et al. · 2009 [cited by applicant]
US 20100111370A1 · Black et al. · 2010 [cited by applicant]
US 20100198563A1 · Plewe · 2010 [cited by applicant]
US 20100209013A1 · Minear et al. · 2010 [cited by applicant]
US 20100302247A1 · Perez et al. · 2010 [cited by applicant]
US 20110052043A1 · Hyung et al. · 2011 [cited by applicant]
US 20110074929A1 · Hebert et al. · 2011 [cited by applicant]
US 20120056800A1 · Williams et al. · 2012 [cited by applicant]
US 20120063672A1 · Gordon et al. · 2012 [cited by applicant]
US 20120098937A1 · Sajadi et al. · 2012 [cited by applicant]
US 20120130762A1 · Gale et al. · 2012 [cited by applicant]
US 20120194516A1 · Newcombe et al. · 2012 [cited by applicant]
US 20120194517A1 · Izadi et al. · 2012 [cited by applicant]
US 20120306876A1 · Shotton et al. · 2012 [cited by applicant]
US 20130069940A1 · Sun et al. · 2013 [cited by applicant]
US 20130123801A1 · Umasuthan et al. · 2013 [cited by applicant]
US 20130156262A1 · Taguchi et al. · 2013 [cited by applicant]
US 20130201104A1 · Ptucha et al. · 2013 [cited by applicant]
US 20130201105A1 · Ptucha et al. · 2013 [cited by applicant]
US 20130208955A1 · Zhao et al. · 2013 [cited by applicant]
US 20130346348A1 · Buehler · 2013 [cited by examiner]
US 20140160115A1 · Keitler et al. · 2014 [cited by applicant]
US 20140176677A1 · Valkenburg et al. · 2014 [cited by applicant]
US 20140206443A1 · Sharp et al. · 2014 [cited by applicant]
US 20140240464A1 · Lee · 2014 [cited by applicant]
US 20140241617A1 · Shotton et al. · 2014 [cited by applicant]
US 20140270484A1 · Chandraker et al. · 2014 [cited by applicant]
US 20140321702A1 · Schmalstieg · 2014 [cited by applicant]
US 20150009214A1 · Lee et al. · 2015 [cited by applicant]
US 20150045923A1 · Chang et al. · 2015 [cited by applicant]
US 20150142394A1 · Mehr et al. · 2015 [cited by applicant]
US 20150213572A1 · Loss · 2015 [cited by applicant]
US 20150234477A1 · Abovitz et al. · 2015 [cited by applicant]
US 20150262405A1 · Black et al. · 2015 [cited by applicant]
US 20150269715A1 · Jeong et al. · 2015 [cited by applicant]
US 20150279118A1 · Dou et al. · 2015 [cited by applicant]
US 20150301592A1 · Miller · 2015 [cited by applicant]
US 20150325044A1 · Lebovitz · 2015 [cited by applicant]
US 20150371440A1 · Pirchheim et al. · 2015 [cited by applicant]
US 20160026253A1 · Bradski et al. · 2016 [cited by applicant]
US 20160071318A1 · Lee et al. · 2016 [cited by applicant]
US 20160171765A1 · Mehr · 2016 [cited by applicant]
US 20160173842A1 · De La Cruz et al. · 2016 [cited by applicant]
US 20160358382A1 · Lee et al. · 2016 [cited by applicant]
US 20170053447A1 · Chen et al. · 2017 [cited by applicant]
US 20170054954A1 · Keitler et al. · 2017 [cited by applicant]
US 20170054965A1 · Raab et al. · 2017 [cited by applicant]
US 20170221263A1 · Wei et al. · 2017 [cited by applicant]
US 20170243397A1 · Hou et al. · 2017 [cited by applicant]
US 20170278293A1 · Hsu · 2017 [cited by applicant]
US 20170316597A1 · Ceylan et al. · 2017 [cited by applicant]
US 20170337726A1 · Bui et al. · 2017 [cited by applicant]
US 20180005015A1 · Hou et al. · 2018 [cited by applicant]
US 20180025529A1 · Wu et al. · 2018 [cited by applicant]
US 20180114363A1 · Rosenbaum · 2018 [cited by applicant]
US 20180144535A1 · Ford et al. · 2018 [cited by applicant]
US 20180284802A1 · Tsai · 2018 [cited by examiner]
US 20180288387A1 · Somanath et al. · 2018 [cited by applicant]
US 20180300937A1 · Chien et al. · 2018 [cited by applicant]
US 20190114832A1 · Park · 2019 [cited by examiner]
US 20190208007A1 · Khalid · 2019 [cited by applicant]
US 20190244412A1 · Yago Vicente et al. · 2019 [cited by applicant]
US 20190251728A1 · Stoyles et al. · 2019 [cited by applicant]
US 20200086487A1 · Johnson et al. · 2020 [cited by applicant]
US 20200105013A1 · Chen · 2020 [cited by applicant]
CN 111383348A · 2020 [cited by examiner]
EP 1308902A2 · 2003 [cited by applicant]
KR 101054736B1 · 2011 [cited by applicant]
KR 1020110116671A · 2011 [cited by applicant]
WO 2006027339A2 · 2006 [cited by applicant]
Miroslav Trajkovic, Mark Hedley, “Fast corner detection”, Image and Vision Computing, vol. 16, Issue 2, 1998, pp. 75-87. [cited by examiner]
D. Holz, A. E. Ichim, F. Tombari, R. B. Rusu and S. Behnke, “Registration with the Point Cloud Library: A Modular Framework for Aligning in 3-D,” in IEEE Robotics & Automation Magazine, vol. 22, No. 4, pp. 110-124, Dec.… [cited by examiner]
English Machine translation of CN111383348A, Accessed Jan. 20, 2023. [cited by examiner]
S. Gaurav, Z. Al-Qurashi, A. Barapatre, G. Maratos, T. Sarma and B. D. Ziebart, “Deep Correspondence Learning for Effective Robotic Teleoperation using Virtual Reality,” 2019 IEEE-RAS 19th International Conference on Hu… [cited by examiner]
S. Lieberknecht, A. Huber, S. Ilic and S. Benhimane, “RGB-D camera-based parallel tracking and meshing,” 2011 10th IEEE International Symposium on Mixed and Augmented Reality, Basel, Switzerland, 2011, pp. 147-155, doi:… [cited by examiner]
Zollhofer, Michael, et al. “Real-time non-rigid reconstruction using an RGB-D camera.” ACM Transactions on Graphics (ToG) 33.4 (2014): 1-12. Accessed Apr. 29, 2024. [cited by examiner]
Wang, Kangkan, Guofeng Zhang, and Shihong Xia. “Templateless non-rigid reconstruction and motion tracking with a single RGB-D camera.” IEEE Transactions on Image Processing 26.12 (2017): 5966-5979. [cited by examiner]
Rossignac, J. et al., “3D Compression Made Simple: Edgebreaker on a Corner-Table,” Invited lecture at the Shape Modeling International Conference, Genoa, Italy (Jan. 30, 2001), pp. 1-6. [cited by applicant]
Melax, S., “A Simple, Fast, and Effective Polygon Reduction Algorithm,” Game Developer, Nov. 1998, pp. 44-49. [cited by applicant]
Myronenko, A. et al., “Point Set Registration: Coherent Point Drift,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 32, No. 12, Dec. 2010, pp. 2262-2275. [cited by applicant]
Bookstein, F., “Principal Warps: Thin-Plate Splines and the Decomposition of Deformations,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 11, No. 6, Jun. 1989, pp. 567-585. [cited by applicant]
Izadi, S. et al., “KinectFusion: Real-time 3D Reconstruction and Interaction Using a Moving Depth Camera,” UIST '11, Oct. 16-19, 2011, 10 pages. [cited by applicant]
Papazov, C. et al., “An Efficient RANSAC for 3D Object Recognition in Noisy and Occluded Scenes,” presented at Computer Vision—ACCV 2010—10th Asian Conference on Computer Vision, Queenstown, New Zealand, Nov. 8-12, 2010… [cited by applicant]
Biegelbauer, Georg et al., “Model-based 3D object detection—Efficient approach using superquadrics,” Machine Vision and Applications, Jun. 2010, vol. 21, Issue 4, pp. 497-516. [cited by applicant]
Kanezaki, Asako et al., “High-speed 3D Object Recognition Using Additive Features in a Linear Subspace,” 2010 EEE International Conference on Robotics and Automation, Anchorage Convention District, May 3-8, 2010, pp. 31… [cited by applicant]
International Search Report and Written Opinion from PCT patent application No. PCT/US13/062292, dated Jan. 28, 2014, 10 pages. [cited by applicant]
International Search Report and Written Opinion from PCT patent application No. PCT/US14/045591, dated Nov. 5, 2014, 9 pages. [cited by applicant]
Sumner, R. et al., “Embedded Deformation for Shape Manipulation,” Applied Geometry Group, ETH Zurich, SIGGRAPH 2007, 7 pages. [cited by applicant]
Rosten, Edward, et al., “Faster and better: a machine learning approach to corner detection,” arXiv:08102.2434v1 [cs.CV], Oct. 14, 2008, available at https://arxiv.org/pdf/0810.2434.pdf, 35 pages. [cited by applicant]
Kim, Young Min, et al., “Guided Real-Time Scanning of Indoor Objects,” Computer Graphics Forum, vol. 32, No. 7 (2013), 10 pages. [cited by applicant]
Rusinkewicz, Szymon, et al., “Real-time 3D model acquisition,” ACM Transactions on Graphics (TOG) 21.3 (2002), pp. 438-446. [cited by applicant]
European Search Report from European patent application No. EP 15839160, dated Feb. 19, 2018, 8 pages. [cited by applicant]
Liu, Song, et al., “Creating Simplified 3D Models with High Quality Textures,” arXiv:1602.06645v1 [cs.GR], Feb. 22, 2016, 9 pages. [cited by applicant]
Stoll, C., et al., “Template Deformation for Point Cloud Filtering,” Eurographics Symposium on Point-Based Graphics (2006), 9 pages. [cited by applicant]
Allen, Brett, et al., “The space of human body shapes: reconstruction and parameterization from range scans,” ACM Transactions on Graphics (TOG), vol. 22, Issue 3, Jul. 2003, pp. 587-594. [cited by applicant]
International Search Report and Written Opinion from PCT patent application No. PCT/US15/49175, dated Feb. 19, 2016, 14 pages. [cited by applicant]
Harris, Chris & Mike Stephens, “A Combined Corner and Edge Detector,” Plessey Research Roke Manor, U.K. (1988), pp. 147-151. [cited by applicant]
Bay, Herbert, et al., “Speeded-Up Robust Features (SURF),” Computer Vision and Image Understanding 110 (2008), pp. 346-359. [cited by applicant]
Rublee, Ethan, et al., “ORB: an efficient alternative to SIFT or SURF,” Willow Garage, Menlo Park, CA (2011), available from http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.370.4395&rep=rep1&type=pdf, 8 pages. [cited by applicant]
Lowe, David G., “Distinctive Image Features from Scale-Invariant Keypoints,” International Journal of Computer Vision, vol. 60, Issue 2, Nov. 2004, pp. 91-110. [cited by applicant]
Kaess, Michael, et al., “iSAM: Incremental Smoothing and Mapping,” IEEE Transactions on Robotics, Manuscript, Sep. 7, 2008, 14 pages. [cited by applicant]
Kummerle, Rainer, et al., “g20: A General Framework for Graph Optimization,” 2011 IEEE International Conference on Robotics and Automation, May 9-13, 2011, Shanghai, China, 7 pages. [cited by applicant]