IP Library › Granted Patent US 12,602,866
Granted Patent B2
US 12,602,866 · App. 18/480,173 · Granted Apr 14, 2026

Digital twin authoring and editing environment for creation of AR/VR and video instructions from a single demonstration

Inventors: Karthik Ramani (West Lafayette, IN); Subramaniam Chidambaram (West Lafayette, IN); Sai Swarup Reddy (Sunnyvale, CA)
Assignee: Purdue Research Foundation
G06T17/00G06F3/014G06T7/70G06T13/40G06V40/11G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,866
App. No.
18/480,173
Granted
Apr 14, 2026
Kind
B2
Abstract

A system and method for multi-media instruction authoring is disclosed. The multi-media instruction authoring system and method provide a unified and efficient system that supports the simultaneous creation of instructional media in AR, VR, and video-based media formats by subject matter experts. The multi-media instruction authoring system and method require only a single demonstration of the task by the subject matter expert and does not require any technical expertise of the subject matter expert. The multi-media instruction authoring system and method incorporate a 3D editing interface for free hand, in-headset interaction to create 2D videos from 3D recordings, by exploring novel virtual camera interactions.

Claims (49)

1 . A method for generating multi-media instructional content, the method comprising:

recording, with at least one sensor, a demonstration by a user of a task within a real-world environment;

determining, with a processor, a first sequence of pose data for a virtual avatar based on the recorded demonstration;

generating, with the processor, based on the first sequence of pose data, at least one of augmented reality content and virtual reality content that provides instructions to perform the task, the at least one of the augmented reality content and the virtual reality content including a virtual model of the virtual avatar associated with the first sequence of pose data; and

generating, with the processor, based on the first sequence of pose data, a two-dimensional video that provides instructions to perform the task, the two-dimensional video being generated using a virtual camera that views the virtual model of the virtual avatar within a virtual environment and by animating the virtual model of the virtual avatar to move according to the first sequence of pose data.

2 . The method according to claim 1 , the generating the at least one of the augmented reality content and the virtual reality content comprising:

generating both of the augmented reality content and the virtual reality content.

3 . The method according to claim 1 , wherein the at least one sensor includes a sensor of a head-mounted augmented reality device, recorded sensor data from which is used to determine the first sequence of pose data.

4 . The method according to claim 1 , wherein the first sequence of pose data includes hand poses for the virtual avatar.

5 . The method according to claim 1 further comprising:

determining, with the processor, a second sequence of pose data for a virtual object based on the recorded demonstration, the demonstration including the user performing an operation using a real-world object corresponding to the virtual object.

6 . The method according to claim 5 , wherein the at least one sensor includes a sensor attached directly to the real-world object, recorded sensor data from which is used to determine the second sequence of pose data.

7 . The method according to claim 5 further comprising:

detecting that the user interacts with the real-world object,

wherein the recording the demonstration automatically begins in response to detecting that the user has interacted with the real-world object.

8 . The method according to claim 1 further comprising:

recording, with a camera, during the demonstration, a two-dimensional video of the demonstration,

wherein at least one of (i) the augmented reality content and (ii) the virtual reality content includes the two-dimensional video embedded therein.

9 . The method according to claim 1 further comprising:

recording, with a camera, a hand gesture performed by the user;

pairing, with the processor, based on user inputs, the hand gesture with a user-selected operation;

detecting, with the processor, after the pairing, performance of the hand gesture; and

performing, with the processor, the user-selected operation in response to detecting the performance of the hand gesture.

10 . The method according to claim 1 , wherein the augmented reality content includes a plurality of virtual models and a plurality of time sequences of pose data, the plurality of virtual models including the virtual model of the virtual avatar, each respective virtual model of the plurality of virtual models being associated with a respective sequence of pose data of the plurality of time sequences of pose data, the plurality of time sequences of pose data including the first sequence of pose data, the method further comprising:

displaying, on a display of an augmented reality device, each respective virtual model of the plurality of virtual models superimposed on the real-world environment and animated to move according to the respective sequence of pose data of the plurality of time sequences of pose data.

11 . The method according to claim 1 , wherein the virtual reality content includes a plurality of virtual models and a plurality of time sequences of pose data, the plurality of virtual models including the virtual model of the virtual avatar, each respective virtual model of the plurality of virtual models being associated with a respective sequence of pose data of the plurality of time sequences of pose data, the plurality of time sequences of pose data including the first sequence of pose data, the method further comprising:

displaying, on a display of a virtual reality device, each respective virtual model of the plurality of virtual models within a virtual environment and animated to move according to the respective sequence of pose data of the plurality of time sequences of pose data.

12 . The method according to claim 1 , the generating the two-dimensional video further comprising:

generating, with the processor, the two-dimensional video using the virtual camera, which views a plurality of virtual models within the virtual environment, the plurality of virtual models including the virtual model of the virtual avatar,

wherein, in the two-dimensional video, each respective virtual model of the plurality of virtual models is animated to move according to a respective time sequence of pose data of a plurality of time sequences of pose data, the plurality of time sequences of pose data including the first sequence of pose data.

13 . The method according to claim 12 , the generating the two-dimensional video further comprising:

displaying, on a display of at least one of an augmented reality device and a virtual reality device, a graphical user interface;

defining, based on user inputs, at least one pose of the virtual camera; and

generating, with the processor, the two-dimensional video using the virtual camera having the defined camera pose.

14 . The method according to claim 13 , the generating the two-dimensional video further comprising:

defining, based on user inputs, a timeline of the two-dimensional video including a first time segment of the two-dimensional video and a second time segment of the two-dimensional video;

defining, based on user inputs, for the first time segment of the timeline, a first pose of the virtual camera and a first time segment of the plurality of time sequences of pose data;

defining, based on user inputs, for the second time segment of the timeline, a second pose of the virtual camera and a second time segment of the plurality of time sequences of pose data; and

generating (i) the first time segment of the two-dimensional video using the virtual camera having the first camera pose and with the plurality of virtual models animated to move according to the first time segment of the plurality of time sequences of pose data and (ii) the second time segment of the two-dimensional video using the virtual camera having the second camera pose and with the plurality of virtual models animated to move according to the second time segment of the plurality of time sequences of pose data.

15 . The method according to claim 13 , wherein, for a first time segment of the two-dimensional video, the virtual camera has a first camera pose in which the virtual camera has a static position and orientation within the virtual environment.

16 . The method according to claim 13 , wherein, for a second time segment of the two-dimensional video, the virtual camera has a second camera pose in which the virtual camera moves dynamically through the virtual environment over time according to a user-defined trajectory.

17 . The method according to claim 13 , wherein, for a third time segment of the two-dimensional video, the virtual camera has a third camera pose in which the virtual camera moves dynamically through the virtual environment over time according to a trajectory that depends on a current state of at least one virtual model of the plurality of virtual models.

18 . The method according to claim 13 further comprising:

displaying, within the graphical user interface on the display, a preview of the two-dimensional video.

19 . The method according to claim 13 further comprising:

displaying, within the graphical user interface on the display, the plurality of virtual models within the virtual environment, each respective virtual model of the plurality of virtual models being animated to move according to a respective time sequence of pose data of the plurality of time sequences of pose data.

20 . The method according to claim 19 , the displaying the plurality of virtual models within the virtual environment further comprising:

selecting, based on user inputs, a selected point in time of the plurality of time sequences of pose data; and

displaying, within the graphical user interface on the display, a state of the plurality of virtual models at the selected point in time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2023
From: RAMANI, KARTHIK; CHIDAMBARAM, SUBRAMANIAM; REDDY, SAI SWARUP
To: PURDUE RESEARCH FOUNDATION
Reel/Frame 065810/0339 →
Continuity (3)
Provisional Application 63384868 · Nov 23, 2022
Provisional Application 63378101 · Oct 3, 2022
Related Publication 20240112399A1 · Apr 4, 2024
References Cited (61)
US 20160361649A1 · Hayashi · 2016 [cited by examiner]
US 20200092536A1 · Watson · 2020 [cited by examiner]
US 20200168119A1 · Ramani · 2020 [cited by examiner]
US 20210134065A1 · Ramani · 2021 [cited by examiner]
US 20210252699A1 · Ramani · 2021 [cited by examiner]
M. Whitlock, G. Fitzmaurice, T. Grossman, and J. Matejka. Authar: Concurrent authoring of tutorials for ar assembly guidance. In Graphics Interface, pp. 431-439. CHCCS/SCDHM, University of Toronto,Ontario, Canada, 2020. [cited by applicant]
L. Wright and S. Davidson. How to tell the difference between a model and a digital twin. Advanced Modeling and Simulation in Engineering Sciences, 7(1):1-13, 2020. [cited by applicant]
M. Yamaguchi, S. Mori, P. Mohr, M. Tatzgern, A. Stanescu, H. Saito, and D. Kalkofen. Video-annotated augmented reality assembly tutorials. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and T… [cited by applicant]
U. Yang and G. J. Kim. Implementation and evaluation of “just follow me”: An immersive, vr-based, motion-training system. Presence, 11(3):304-323, 2002. doi: 10.1162/105474602317473240. [cited by applicant]
YouTube. Youtube, Apr. 2021. [cited by applicant]
J. Zillner, E. Mendez, and D. Wagner. Augmented reality remote collaboration with dense reconstruction. In 2018 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct), pp. 38-39, 2018. doi: … [cited by applicant]
F. Anderson, T. Grossman, J. Matejka, and G. Fitzmaurice. Youmove: enhancing movement training with an augmented reality mirror. In Proceedings of the 26th annual ACM symposium on User interface software and technology,… [cited by applicant]
Antilatency. Antilatency, Apr. 2021. [cited by applicant]
D. Aouam, S. Benbelkacem, N. Zenati, S. Zakaria, and Z. Meftah. Voice-based augmented reality interactive system for car's components assembly. In 2018 3rd International Conference on Pattern Analysis and Intelligent Sy… [cited by applicant]
A. Bangor, P. Kortum, and J. Miller. Determining what individual sus scores mean: Adding an adjective rating scale. Journal of usability studies, 4(3):114-123, 2009. [cited by applicant]
Blackmagicdesign. Davinci resolve 17, Nov. 2020. [cited by applicant]
J. Brooke et al. Sus-a quick and dirty usability scale. Usability evaluation in industry, 189(194):4-7, 1996. [cited by applicant]
Y. Cao, X. Qian, T. Wang, R. Lee, K. Huo, and K. Ramani. An exploratory study of augmented reality presence for tutoring machine tasks. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pp… [cited by applicant]
S. Chidambaram, H. Huang, F. He, X. Qian, A. M. Villanueva, T. S. Redick, W. Stuerzlinger, and K. Ramani. ProcessAR: An Augmented Reality-Based Tool to Create in-Situ Procedural 2D/3D AR Instructions, p. 234-249. Associ… [cited by applicant]
L. L. Cone. Skycam: an aerial robotic camera system. Byte Magazine, 10:122, 1985. [cited by applicant]
A. R. Fender and C. Holz. Causality-preserving asynchronous reality. In CHI Conference on Human Factors in Computing Systems, CHI '22. Association for Computing Machinery, New York, NY, USA, 2022. doi: 10.1145/3491102.3… [cited by applicant]
M. Funk, M. Kritzler, and F. Michahelles. Holocollab: a shared virtual platform for physical assembly training using spatially-aware headmounted displays. In Proceedings of the Seventh International Conference on the In… [cited by applicant]
Q. Galvane, I.-S. Lin, F. Argelaguet, T.-Y. Li, and M. Christie. Vr as a content creation tool for movie previsualisation. In 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pp. 303-311, 2019. doi: … [cited by applicant]
R. Griffin, T. Langlotz, and S. Zollmann. 6dive: 6 degrees-of-freedom immersive video editor. Frontiers in Virtual Reality, 2:75, 2021. [cited by applicant]
T. Ha, W. Woo, Y. Lee, J. Lee, J. Ryu, H. Choi, and K. Lee. Artalet: tangible user interface based immersive augmented reality authoring tool for digilog book. In 2010 International Symposium on Ubiquitous Virtual Reali… [cited by applicant]
S. G. Hart and L. E. Staveland. Development of nasa-tlx (task load index): Results of empirical and theoretical research. In Advances in psychology, vol. 52, pp. 139-183. Elsevier, USA, 1988. [cited by applicant]
S. J. Henderson and S. K. Feiner. Augmented reality in the psychomotor phase of a procedural task. In 2011 10th IEEE International Symposium on Mixed and Augmented Reality, pp. 191-200. IEEE, IEEE Computer Society, Los … [cited by applicant]
K. Huang, J. Li, M. Sousa, and T. Grossman. Immersivepov: Filming how-to videos with a head-mounted 360° action camera. In CHI Conference on Human Factors in Computing Systems, CHI '22. Association for Computing Machine… [cited by applicant]
A. Hughes. Forging the digital twin in discrete manufacturing: A vision for unity in the virtual and real worlds, Sep. 2018. [cited by applicant]
T. Imai, A. E. Johnson, J. Leigh, D. E. Pape, and T. A. DeFanti. The virtual mail system. In Proceedings of Virtual Reality, pp. 78-78. IEEE Computer Society, Los Alamitos, CA, USA, 1999. [cited by applicant]
Instructables. instructables, Apr. 2021. [cited by applicant]
A. Ipsita, L. Erickson, Y. Dong, J. Huang, A. K. Bushinski, S. Saradhi, A. M. Villanueva, K. A. Peppler, T. S. Redick, and K. Ramani. Towards modeling of virtual reality welding simulators to promote accessible and scal… [cited by applicant]
S. Kim, H. gun Chi, and K. Ramani. Object synthesis by learning part geometry with surface and volumetric representations. Computer-Aided Design, 130:102932, 2021. doi: 10.1016/j.cad.2020.102932. [cited by applicant]
K. Lilija, H. Pohl, and K. Hornbæk. Who put that there? temporal navigation of spatial recordings by direct manipulation. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pp. 1-11. Associ… [cited by applicant]
S.-C. Lim, H.-K. Lee, and J. Park. Role of combined tactile and kinesthetic feedback in minimally invasive surgery. The International Journal of Medical Robotics and Computer Assisted Surgery, 11(3):360-374, 2015. [cited by applicant]
M. R. Marner, A. Irlitti, and B. H. Thomas. Improving procedural task performance with augmented reality annotations. In 2013 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 39-48. IEEE, IEEE Co… [cited by applicant]
A. Marwanto, Y. Wibowo, R. Djatmiko, and R. Wijaya. Kinesthetic intelligence in welding practice lectures. In Journal of Physics: Conference Series, No. 1, p. 012022. IOP Publishing, IOP Publishing, Orlando, Florida, 20… [cited by applicant]
P. Mohr, D. Mandl, M. Tatzgern, E. Veas, D. Schmalstieg, and D. Kalkofen. Retargeting video tutorials showing tools with surface contact to augmented reality. In Proceedings of the 2017 CHI Conference on Human Factors i… [cited by applicant]
M. Morozov, A. Gerasimov, and M. Fominykh. vacademia-educational virtual world with 3d recording. In 2012 International Conference on Cyberworlds, pp. 199-206. IEEE Computer Society, Los Alamitos, CA, USA, 2012. [cited by applicant]
M. Nebeling, S. Rajaram, L. Wu, Y. Cheng, and J. Herskovitz. Xrstudio: A virtual production and live streaming system for immersive instructional experiences. In Proceedings of the 2021 CHI Conference on Human Factors i… [cited by applicant]
C. Nguyen, S. DiVerdi, A. Hertzmann, and F. Liu. Vremiere: Inheadset virtual reality video editing. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, p. 5428-5438. Association for Computin… [cited by applicant]
Oculus. Oculus quest 2, 2020. Retrieved Apr. 4, 2021, from https://www.oculus.com/quest-2/. [cited by applicant]
O. Oda, C. Elvezio, M. Sukan, S. Feiner, and B. Tversky. Virtual replicas for remote assistance in virtual and augmented reality. In Proceedings of the 28th Annual ACM Symposium on User Interface Software & Technology, … [cited by applicant]
S. Ong and Z. Wang. Augmented assembly technologies based on 3d bare-hand interaction. CIRP annals, 60(1):1-4, 2011. [cited by applicant]
OptiTrack. Optitrack, Apr. 2021. [cited by applicant]
M. Perry, O. Juhlin, M. Esbjornsson, and A. Engstrom. Lean Collaboration through Video Gestures: Co-Ordinating the Production of Live Televised Sport, p. 2279-2288. Association for Computing Machinery, New York, NY, USA… [cited by applicant]
PTC. Vuforia expert capture, 2019. Retrieved May 5, 2020, from https://www.ptc.com/en/products/augmented-reality/vuforia-expert-capture. [cited by applicant]
PTC. Vuforia chalk: Remote assistance powered by augmented reality, Dec. 2020. [cited by applicant]
R. Radkowski and C. Stritzke. Interactive hand gesture-based assembly for augmented reality applications. In Proceedings of the 2012 International Conference on Advances in Computer-Human Interactions, pp. 303-308. Cite… [cited by applicant]
M. Sayed, R. Cinca, E. Costanza, and G. Brostow. Lookout! interactive camera gimbal controller for filming long takes, 2020. [cited by applicant]
M. Scheff-King. Download, edit and print your own parts from mcmaster-carr, Apr. 2014. [cited by applicant]
M. Sheets-Johnstone. Kinesthetic memory. Theoria et historia scientiarum, 7(1):69-92, 2003. [cited by applicant]
Skillshare. Skillshare, Apr. 2021. [cited by applicant]
Stereolabs. Zed mini, Apr. 2021. [cited by applicant]
B. Thoravi Kumaravel, F. Anderson, G. Fitzmaurice, B. Hartmann, and T. Grossman. Loki: Facilitating remote instruction of physical tasks using bi-directional mixed-reality telepresence. In Proceedings of the 32nd Annual… [cited by applicant]
Traceparts. Traceparts, 1990. Retrieved Mar. 8, 2022, from https: //www.traceparts.com/en. [cited by applicant]
Tvori. tvori, Feb. 2022. [cited by applicant]
A. van Dam. Post-wimp user interfaces. Commun. ACM, 40(2):63-67, Feb. 1997. doi: 10.1145/253671.253708. [cited by applicant]
Vreal. vreal, Mar. 2018. [cited by applicant]
C. Y. Wang, M. Sakashita, U. Ehsan, J. Li, and A. S. Won. Again, together: Socially reliving virtual reality experiences when separated. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, p… [cited by applicant]
T. Wang, X. Qian, F. He, X. Hu, Y. Cao, and K. Ramani. GesturAR: An Authoring System for Creating Freehand Interactive Augmented Reality Applications, p. 552-567. Association for Computing Machinery, New York, NY, USA, … [cited by applicant]