IP Library Granted Patent US 12,229,333
Granted Patent B2
US 12,229,333 · App. 18/499,770 · Granted Feb 18, 2025

Technique for visualizing interactions with a technical device in an XR scene

Inventors: Daniel Rinck (Forchheim, DE); Aniol Serra Juhe (Nuremberg, DE)
Assignee: SIEMENS HEALTHINEERS AG
G06F3/011G06F3/0304
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,333
App. No.
18/499,770
Granted
Feb 18, 2025
Kind
B2
Abstract

A computer-implemented method for visualizing interactions in an extended reality (XR) scene, the computer-implemented method comprising: receiving a first dataset representing an XR scene including a technical device; displaying the XR scene on an XR headset or a head-mounted display (HMD); providing a room for a user wearing the XR headset or HMD for interacting with the XR scene, wherein the room includes a set of optical sensors including at least one optical sensor at a fixed location relative to the room; detecting optical sensor data of the user as a second dataset while the user is interacting with the XR scene in the room; and fusing the first dataset and the second dataset to generate a third dataset.

Claims (55)

1. A computer-implemented method for visualizing interactions in an extended reality (XR) scene, the computer-implemented method comprising:

receiving a first dataset, the first dataset representing the XR scene including at least a technical device;

displaying the XR scene on an XR headset or a head-mounted display (HMD);

providing a room for a user, wherein the user wears the XR headset or HMD to interact with the XR scene, wherein the XR scene is displayed on the XR headset or HMD, wherein the room includes a set of optical sensors, and wherein the set of optical sensors includes at least one optical sensor at a fixed location relative to the room;

detecting, via the set of optical sensors, optical sensor data of the user as a second dataset while the user is interacting in the room with the XR scene, wherein the XR scene is displayed on the XR headset or HMD; and

fusing the first dataset and the second dataset to generate a third dataset, wherein

the first dataset, the second dataset and the third dataset include a point cloud, and

the third dataset includes a fusion of point clouds of the first dataset and the second dataset.

2. The computer-implemented method according to claim 1 , wherein the set of optical sensors includes at least one depth camera to provide point cloud data.

3. The computer-implemented method according to claim 1 , further comprising:

generating real-time instructions based on the third dataset; and

providing the real-time instructions to the user.

4. The computer-implemented method according to claim 1 , wherein the computer-implemented method is used for at least one of product development of the technical device or for controlling the technical device.

5. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a computing device, cause the computing device to perform the computer-implemented method according to claim 1 .

6. The computer-implemented method according to claim 2 , wherein the at least one depth camera includes an RGBD camera.

7. The computer-implemented method according to claim 1 , wherein a trained neural network is used to provide output data based on input data, wherein the input data includes the third dataset and the output data represents a semantic context of the optical sensor data of the user.

8. A computer-implemented method for visualizing interactions in an extended reality (XR) scene, the computer-implemented method comprising:

receiving a first dataset, the first dataset representing the XR scene including at least a technical device;

displaying the XR scene on an XR headset or a head-mounted display (HMD);

providing a room for a user, wherein the user wears the XR headset or HMD to interact with the XR scene, wherein the XR scene is displayed on the XR headset or HMD, wherein the room includes a set of optical sensors, and wherein the set of optical sensors includes at least one optical sensor at a fixed location relative to the room;

detecting, via the set of optical sensors, optical sensor data of the user as a second dataset while the user is interacting in the room with the XR scene, wherein the XR scene is displayed on the XR headset or HMD; and

fusing the first dataset and the second dataset to generate a third dataset, wherein

a trained neural network is used to provide output data based on input data, wherein the input data includes the third dataset and the output data represents a semantic context of the optical sensor data of the user.

9. The computer-implemented method according to claim 8 , wherein the trained neural network is trained by providing the input data including the third dataset, which is labeled with content data, wherein the content data represents a semantic context of user interaction.

10. The computer-implemented method according to claim 8 , further comprising:

pre-processing the third dataset before being used as the input data for the trained neural network, wherein the pre-processing includes application of an image-to-image transfer learning algorithm to the third dataset.

11. The computer-implemented method according to claim 8 , further comprising:

pre-processing the first dataset before fusion with the second dataset, wherein the pre-processing includes application of an image-to-image transfer learning algorithm to the first dataset, wherein the generating of the third dataset includes fusing the pre-processed first dataset and the second dataset.

12. The computer-implemented method according to claim 8 , wherein the trained neural network is further configured to receive, as input data, detected optical sensor data of a user interacting with a real-world scene, wherein the real-world scene includes a real-world technical device, wherein the real-world technical device corresponds to the technical device of the XR scene.

13. The computer-implemented method according to claim 11 , wherein the fusing includes applying a calibration algorithm, which utilizes at least one registration object deployed in the room.

14. The computer-implemented method according to claim 13 , wherein the calibration algorithm uses a set of registration objects, which are provided as real, physical objects in the room and which are provided as displayed virtual objects in the XR scene, wherein for registration purposes, the real, physical objects are moved to match the displayed virtual objects in the XR scene.

15. The computer-implemented method according to claim 14 , wherein the real, physical objects include a first set of spheres, and wherein the displayed virtual objects include a second set of spheres.

16. A computing device for visualizing interactions in an extended reality (XR) scene, the computing device comprising:

a first input interface configured to receive a first dataset, the first dataset representing the XR scene including at least a technical device;

a second input interface configured to receive, from a set of optical sensors, detected optical sensor data of a user as a second dataset while the user is interacting in a room with the XR scene, wherein the XR scene is displayed to the user on a XR headset or head-mounted display (HMD), and wherein the set of optical sensors includes at least one optical sensor at a fixed location relative to the room; and

at least one processor configured to fuse the first dataset and the second dataset to generate a third dataset, wherein

the first dataset, the second dataset and the third dataset include a point cloud, and

the third dataset includes a fusion of point clouds of the first dataset and the second dataset.

17. A system for visualizing interactions in an extended reality (XR) scene, the system comprising:

at least one of at least one XR headset or at least one head-mounted display (HMD);

an XR scene generating device configured to generate a first dataset, the first data set representing the XR scene including at least a technical device, wherein the XR scene is to be displayed on the at least one of the at least one XR headset or the at least one HMD;

a set of optical sensors configured to detect optical sensor data of a user as a second dataset, while the user is interacting in a room with the XR scene, the XR scene being displayable on the at least one of the at least one XR headset or the at least one HMD, and wherein the set of optical sensors includes at least one optical sensor at a fixed location relative to the room; and

a computing device configured to execute a fusing algorithm to fuse the first dataset with the second dataset to generate a third dataset, the computing device including

a first input interface configured to receive the first dataset,

a second input interface configured to receive, from the set of optical sensors, the optical sensor data of the user as the second dataset, and

at least one processor configured to execute the fuse algorithm to fuse the first dataset with the second dataset to generate the third dataset, wherein

at least one of (i) the at least one XR headset or the at least one HMD, (ii) the computing device, or (ii) the XR scene generating device are connected via a photon unity network (PUN).

18. The system according to claim 17 ,

wherein the first dataset, the second dataset and the third dataset include a point cloud, and

wherein the third dataset includes a fusion of point clouds of the first dataset and the second dataset.

19. A computing device for visualizing interactions in an extended reality (XR) scene, the computing device comprising:

a first input interface configured to receive a first dataset, the first dataset representing the XR scene including at least a technical device;

a second input interface configured to receive, from a set of optical sensors, detected optical sensor data of a user as a second dataset while the user is interacting in a room with the XR scene, wherein the XR scene is displayed to the user on a XR headset or head-mounted display (HMD), and wherein the set of optical sensors includes at least one optical sensor at a fixed location relative to the room; and

at least one processor configured to fuse the first dataset and the second dataset to generate a third dataset, wherein

a trained neural network is used to provide output data based on input data, wherein the input data includes the third dataset and the output data represents a semantic context of the detected optical sensor data of the user.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2024
From: FRIEDRICH-ALEXANDER-UNIVERSITÄT ERLANGEN-NÜRNBERG
To: SIEMENS HEALTHINEERS AG
Reel/Frame 068709/0066 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2024
From: SERRA JUHÉ, ANIOL
To: FRIEDRICH-ALEXANDER-UNIVERSITÄT ERLANGEN-NÜRNBERG
Reel/Frame 068688/0580 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2024
From: RINCK, DANIEL; SERRA JUHÉ, ANIOL
To: SIEMENS HEALTHINEERS AG
Reel/Frame 068688/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: SIEMENS HEALTHCARE GMBH
To: SIEMENS HEALTHINEERS AG
Reel/Frame 066267/0346 →
Priority Claims (1)
EP 22205125 · Nov 2, 2022 · regional
Continuity (1)
Related Publication 20240143069A1 · May 2, 2024
References Cited (29)
US 11250637B1 · Tan et al. · 2022 [cited by applicant]
US 20160191887A1 · Casas · 2016 [cited by examiner]
US 20180232052A1 · Chizeck · 2018 [cited by examiner]
US 20190380792A1 · Poltaretskyi et al. · 2019 [cited by applicant]
US 20200275976A1 · McKinnon et al. · 2020 [cited by applicant]
US 20220265357A1 · Morvan et al. · 2022 [cited by applicant]
WO WO2020163358A1 · 2020 [cited by applicant]
WO WO2021016429A1 · 2021 [cited by applicant]
“NVIDIA Isaac ROS”, in: https://developer.nvidia.com/isaac-ros [abgerufen am Oct. 18, 2023];. [cited by applicant]
Luc, Pauline et al.: “Predicting Deeper into the Future of Semantic Segmentation.” ICCV2017-International Conference on Computer Vision, Oct. 2017, Venise, Italy, pp. 648-657, 10.1109/ICCV.2017.77. hal-01494296v2;. [cited by applicant]
Gu, Jinwei, et al. “Dynamic facial analysis: From bayesian filtering to recurrent neural network.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017.;. [cited by applicant]
Babu, Sudharshan Chandra: “A 2019 guide to Human Pose Estimation with Deep Learning”, Apr. 12, 2019, in: https://nanonets.com/blog/human-pose-estimation-2d-guide/ [abgerufen am Oct. 18, 2023];. [cited by applicant]
PUN “The Ease-of-use of Unity's Networking with the Performance & Reliability of Photon Realtime”, Download Dec. 8, 2022 https://www.photonengine.com/pun;. [cited by applicant]
Video “Holoportation” https://www.youtube.com/watch?v=7d5906cfaMO Date of Download: Aug. 26, 2022;. [cited by applicant]
Chen, Liang-Chieh, et al. “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.” IEEE transactions on pattern analysis and machine intelligence 40.4 (2017): 83… [cited by applicant]
Newcombe, R. A. et al: “KinectFusion: Real-time dense surface mapping and tracking”, Mixed and Augmented Reality (ISMAR), 2011 10th IEEE International Symposium on, IEEE, Oct. 26, 2011 (Oct. 26, 2011), pp. 127-136, XP03… [cited by applicant]
Pfister, Tomas et al. “Flowing convnets for human pose estimation in videos.” Proceedings of the IEEE international conference on computer vision. 2015.;. [cited by applicant]
Nie Y. et al.: “Rfd-net: Point scene understanding by semantic instance reconstruction”, pp. 4606-4616 (Jun. 2021), https://doi.org/10.1109/CVPR46437.2021;. [cited by applicant]
Microsoft Mesh https://www.microsoft.com/en-us/mesh;. [cited by applicant]
He, Kaiming, et al. “Mask r-cnn.” Proceedings of the IEEE international conference on computer vision. 2017.;. [cited by applicant]
Wang, Limin et al. “Action recognition with trajectory-pooled deep-convolutional descriptors.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2015.;. [cited by applicant]
Tran, Du, et al. “Learning spatiotemporal features with 3d convolutional networks.” Proceedings of the IEEE international conference on computer vision. 2015.;. [cited by applicant]
Microsoft Holoportation https://www.microsoft.com/en-us/research/project/holoportation-3/;. [cited by applicant]
Kubach, B. “Microsoft erklärt: Was ist Microsoft Mesh? Definition & Funktionen” in: https://news.microsoft.com/de-de/microsoft-erklaert-was-ist-microsoft-mesh/ [abgerufen am Oct. 18, 2023];. [cited by applicant]
Gadde, Raghudeep et al. “Semantic video cnns through representation warping.” Proceedings of the IEEE International Conference on Computer Vision. 2017.;. [cited by applicant]
“An introduction to Depthkit capture: Depthkit Core, Azure Kinect and the Refinement Algorithm”, in: https://www.depthkit.tv/tutorials/azure-kinect-microsoft-volumetric-capture-depth-workflow-depthkit [abgerufen am Oct.… [cited by applicant]
“NVIDIA Isaac Sim”, in: https://developer.nvidia.com/isaac-sim [Oct. 18, 2023];. [cited by applicant]
“NVIDIA Omniverse Replicator for DRIVE Sim Accelerates AV Development, Improves Perception Results”, in: https://blogs.nvidia.com/blog/2021/11/09/drive-sim-replicator-synthetic-data-generation/?ncid=so-yout-833884-vt03#… [cited by applicant]
Youtube: “Demo: The magic of AI neural TTS and holograms at Microsoft Inspire 2019”, in: https://www.youtube.com/watch?v=auJJrHgG9Mc&t=1s [Oct. 30, 2023]. [cited by applicant]