IP Library › Granted Patent US 12,682,535
Granted Patent B2
US 12,682,535 · App. 19/026,990 · Granted Jul 14, 2026

System and method for authoring context-aware augmented reality instruction through generative artificial intelligence

Inventors: Karthik Ramani (West Lafayette, IN); Seunggeun Chi (West Lafayette, IN); Hyung-gun Chi (West Lafayette, IN); Rahul Jain (West Lafayette, IN); Jingyu Shi (West Lafayette, IN)
Assignee: Purdue Research Foundation
G06T13/40G06F3/012G06T7/70G06T19/006G06V20/20G09B5/02G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,535
App. No.
19/026,990
Filed
Jan 17, 2025
Granted
Jul 14, 2026
Kind
B2
Art Unit
2626
USPC
345/156
Abstract

A method for generating augmented reality (AR) instructional content is disclosed. The method advantageously provides an AR graphical user interface for generating AR instructional content for performing a task from user-input text descriptions of the task. The method advantageously leverages generative artificial intelligence to enable a code-free and motion-capture-free experience for authoring the AR instructional content, including virtual avatar animations demonstrating performance of the steps of the task. Additionally, the method advantageously overcomes the contextual barrier by enabling the user to author context-aware AR instructions that understand the context and blend physical reality with virtual components.

Claims (42)

1 . A method for generating augmented reality or virtual reality instructional content, the method comprising:

receiving, via at least one input device, a natural language user input from a user describing a task;

generating, with a processor, natural language step-by-step text instructions for performing the task using a first machine learning model, based on the natural language user input, the natural language step-by-step text instructions including an ordered sequence of steps for performing the task, each step including text instructions;

capturing, with at least one sensor, contextual information at least including spatial information relating to an environment in which the task is to be performed, the capturing the contextual information including (i) identifying a first subset of steps from the ordered sequence of steps, and (ii) capturing first contextual information relating to a first area of the environment in which the identified first subset of steps are to be performed, the first contextual information being associated with the first subset of steps; and

generating, with the processor, step-by-step animations of a virtual avatar performing the task using a second machine learning model, based on both the natural language step-by-step text instructions and the contextual information,

wherein the natural language step-by-step text instructions and the step-by-step animations of the virtual avatar are used by an augmented reality device or a virtual reality device to display augmented reality or virtual reality instructional content.

2 . The method according to claim 1 , wherein the at least one input device includes a microphone, the receiving further comprising:

recording, with the microphone, the natural language user input spoken by the user.

3 . The method according to claim 1 , wherein the first machine learning model is a language model configured to receive natural language prompts and generate natural language responses.

4 . The method according to claim 3 , the generating the natural language step-by-step text instructions further comprising:

generating a first natural language prompt based on the natural language user input; and

generating, with the first machine learning model, a first natural language response by inputting the first natural language prompt into the first machine learning model, the first natural language response including the natural language step-by-step text instructions.

5 . The method according to claim 4 , the generating the first natural language prompt further comprising:

forming the first natural language prompt by combining the natural language user input with predefined prompt text configured to prompt the first machine learning model to output the natural language step-by-step text instructions for performing the task that is described in the natural language user input.

6 . The method according to claim 5 , wherein the predefined prompt text includes a list of action labels and the predefined prompt text is configured to prompt the first machine learning model to provide the step-by-step text instructions incorporating action labels from the list of action labels.

7 . The method according to claim 1 further comprising:

modifying the natural language step-by-step text instructions based on a user input received via the at least one input device.

8 . The method according to claim 1 , wherein the at least one sensor includes a camera, the capturing the contextual information further comprising:

capturing images of the environment with the camera; and

determining the spatial information relating to the environment based on the images.

9 . The method according to claim 1 , the capturing the contextual information further comprising:

detecting an object in the environment; and

determining a spatial position of the object within the environment, the contextual information including the spatial position of the object.

10 . The method according to claim 9 , the capturing the contextual information further comprising:

determining a semantic label for the object, the contextual information including the semantic label for the object.

11 . The method according to claim 9 , the capturing the contextual information further comprising:

determining a pose of the object, the contextual information including the pose of the object.

12 . The method according to claim 1 , the capturing the first contextual information further comprising:

capturing a motion trajectory of the user as the user navigates the environment to reach the first area of the environment in which the identified first subset of steps are to be performed, the motion trajectory being associated with the first subset of steps.

13 . The method according to claim 1 , the generating the step-by-step animations further comprising:

generating, for each respective step in the ordered sequence of steps, a respective animation of the virtual avatar performing the respective step.

14 . The method according to claim 13 , the generating the respective animation of the virtual avatar performing the respective step further comprising:

generating the respective animation of the virtual avatar performing the respective step using the second machine learning model based on the text instructions of the respective step and the contextual information,

wherein the respective animation is associated with a particular spatial location within the environment.

15 . The method according to claim 14 , wherein the respective animation includes an interaction with a virtual object by the virtual avatar and the respective animation is associated with a physical object in the environment that corresponds to the virtual object.

16 . The method according to claim 13 , the generating the step-by-step animations further comprising:

forming a continuous animation by combining the respective animations of the virtual avatar performing each respective step for performing the task; and

smoothing transitions between the respective animations of the virtual avatar performing each respective step for performing the task using a temporal smoothing algorithm.

17 . The method according to claim 1 , wherein the second machine learning model is a text-to-motion diffusion model.

18 . The method according to claim 17 , wherein the text-to-motion diffusion model adds an input token to each frame of a motion embedding, the input token being an embedding of text conditions, the text conditions including at least a portion of the natural language step-by-step text instructions.

19 . The method according to claim 1 further comprising:

displaying, in an augmented reality or virtual reality graphical user interface, the step-by-step animations of the virtual avatar, the step-by-step animations of the virtual avatar being superimposed upon the environment in accordance with the contextual information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2025
From: RAMANI, KARTHIK; CHI, SEUNGGEUN; CHI, HYUNG-GUN; JAIN, RAHUL; SHI, JINGYU
To: PURDUE RESEARCH FOUNDATION
Reel/Frame 070064/0243 →
Continuity (2)
Provisional Application 63622351 · Jan 18, 2024
Related Publication 20250238991A1 · Jul 24, 2025
References Cited (137)
US 11200892B1 · Stoops · 2021 [cited by examiner]
US 11232645B1 · Roche · 2022 [cited by examiner]
US 11861778B1 · Donnell · 2024 [cited by examiner]
US 20220076473A1 · Jeon · 2022 [cited by examiner]
US 20230410398A1 · Song · 2023 [cited by examiner]
US 20240257470A1 · Gupta · 2024 [cited by examiner]
US 20240338872A1 · Giovara · 2024 [cited by examiner]
US 20250078377A1 · Kapadia · 2025 [cited by examiner]
US 20250200091A1 · Ali · 2025 [cited by examiner]
US 20250200855A1 · Park · 2025 [cited by examiner]
US 20250225709A1 · Zhang · 2025 [cited by examiner]
Seeliger, A., Weibel, R. P., & Feuerriegel, S. (2024). Context-adaptive visual cues for safe navigation in augmented reality using machine learning. International Journal of Human-Computer Interaction, 40(3), 761-781. [cited by applicant]
Soberl, D. (2023). Mixed reality and deep learning: Augmenting visual information using generative adversarial networks. In Augmented Reality and Artificial Intelligence: The Fusion of Advanced Technologies (pp. 3-29). … [cited by applicant]
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., & Ganguli, S. (Jun. 2015). Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning (pp. 2256-2265). PMLR. [cited by applicant]
Soliman, M., & Al Balushi, M. K. (2023). Unveiling destination evangelism through generative AI tools. Robonomics: The Journal of the Automated Economy, 4, 54. [cited by applicant]
Tang, K., Niu, Y., Huang, J., Shi, J., & Zhang, H. (2020). Unbiased scene graph generation from biased training. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 3716-3725). [cited by applicant]
Unity Technologies. 2023. Unity3D. Retrieved Dec. 9, 2023 from https://unity.com/. [cited by applicant]
Zhang, M., Cai, Z., Pan, L., Hong, F., Guo, X., Yang, L., & Liu, Z. (2022). Motiondiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001. [cited by applicant]
Uriarte-Portillo, A., Ibáñez, M. B., Zatarain-Cabada, R., & Barrón-Estrada, M. L. (2023). Comparison of using an augmented reality learning tool at home and in a classroom regarding motivation and learning outcomes. Mul… [cited by applicant]
Wang, T. (2022). Supporting the Design and Authoring of Pervasive Smart Environments (Doctoral dissertation, Purdue University). [cited by applicant]
Wang, T., Qian, X., He, F., Hu, X., Huo, K., Cao, Y., & Ramani, K. (Oct. 2020). CAPturAR: An augmented reality tool for authoring human-involved context-aware applications. In Proceedings of the 33rd Annual ACM Symposiu… [cited by applicant]
Wang, X., Wang, Y., Shi, Y., Zhang, W., & Zheng, Q. (Oct. 2020). Avatarmeeting: An augmented reality remote interaction system with personalized avatars. In Proceedings of the 28th ACM International Conference on Multim… [cited by applicant]
Wang, X., Yew, A. W. W., Ong, S. K., & Nee, A. Y. (2020). Enhancing smart shop floor management with ubiquitous augmented reality. International Journal of Production Research, 58(8), 2352-2367. [cited by applicant]
Wang, Z., Nguyen, C., Asente, P., & Dorsey, J. (May 2021). Distanciar: Authoring site-specific augmented reality experiences for remote environments. In Proceedings of the 2021 CHI Conference on Human Factors in Computi… [cited by applicant]
Weerasinghe, M., Quigley, A., Pucihar, K. Č., Toniolo, A., Miguel, A., & Kljun, M. (2022). Arigato: Effects of Adaptive Guidance on Engagement and Performance in Augmented Reality Learning Environments. IEEE Transaction… [cited by applicant]
Weidner, F., Boettcher, G., Arboleda, S. A., Diao, C., Sinani, L., Kunert, C., . . . & Raake, A. (2023). A systematic review on the visualization of avatars and agents in ar & vr displayed using head-mounted displays. I… [cited by applicant]
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., & Schmidt, D. C. (2023). A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382. [cited by applicant]
Xu, D., Zhu, Y., Choy, C. B., & Fei-Fei, L. (2017). Scene graph generation by iterative message passing. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5410-5419). [cited by applicant]
Ye, H., & Fu, H. (Apr. 2022). ProGesAR: Mobile AR Prototyping for Proxemic and Gestural Interactions with Real-world IoT Enhanced Spaces. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (… [cited by applicant]
Ye, H., Leng, J., Xiao, C., Wang, L., & Fu, H. (Apr. 2023). Proobjar: Prototyping spatially-aware interactions of smart objects with ar-hmd. In Proceedings of the 2023 CHI Conference on Human Factors in Computing System… [cited by applicant]
Yigitbas, E., Heindörfer, J., & Engels, G. (2019). A context-aware virtual reality first aid training application. In Proceedings of Mensch und Computer 2019 (pp. 885-888). [cited by applicant]
Yin, X., Fan, X., Zhu, W., & Liu, R. (2019). Synchronous AR assembly assistance and monitoring system based on ego-centric vision. Assembly Automation, 39(1), 1-16. [cited by applicant]
Yuan, Y., Chen, X., & Wang, J. (2020). Object-contextual representations for semantic segmentation. In Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, Proceedings, Part VI 16 (pp. 173… [cited by applicant]
Yun, K., Lu, T., & Chow, E. (Apr. 2018). Occluded object reconstruction for first responders with augmented reality glasses using conditional generative adversarial networks. In Pattern Recognition and Tracking XXIX (vo… [cited by applicant]
Zhao, Z., & Ma, X. (Dec. 2018). A compensation method of two-stage image generation for human-ai collaborated in-situ fashion design in augmented reality environment. In 2018 IEEE International Conference on Artificial … [cited by applicant]
Zheng, J., Zheng, Q., Fang, L., Liu, Y., & Yi, L. (2023). Cams: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthesis. In Proceedings of the IEEE/CVF Conference on Computer V… [cited by applicant]
Zhu, Z., Liu, Z., Zhang, Y., Zhu, L., Huang, J., Villanueva, A. M., & Ramani, K. (Apr. 2023). Learniotvr: An end-to-end virtual reality environment providing authentic learning experiences for internet of things. In Pro… [cited by applicant]
Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., & Socher, R. (2019). Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858. [cited by applicant]
Kim, D. Y., Lee, H. K., & Chung, K. (2023). Avatar-mediated experience in the metaverse: The impact of avatar realism on user-avatar relationship. Journal of Retailing and Consumer Services, 73, 103382. [cited by applicant]
Kim, M., Lee, K., Balan, R., & Lee, Y. (Apr. 2023). Bubbleu: Exploring Augmented Reality Game Design with Uncertain AI-based Interaction. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (… [cited by applicant]
Kim, S., Chi, H. G., & Ramani, K. (2021). Object synthesis by learning part geometry with surface and volumetric representations. Computer-Aided Design, 130, 102932. [cited by applicant]
Krings, S., Yigitbas, E., Jovanovikj, I., Sauer, S., & Engels, G. (Jun. 2020). Development framework for context-aware augmented reality applications. In Companion Proceedings of the 12th ACM SIGCHI Symposium on Enginee… [cited by applicant]
Labbe, Y., Manuelli, L., Mousavian, A., Tyree, S., Birchfield, S., Tremblay, J., & Sivic, J. (2022). Megapose: 6d pose estimation of novel objects via render & compare. arXiv preprint arXiv:2212.06870. [cited by applicant]
Lacoche, J., & Villain, E. (Feb. 2022). Prototyping context-aware augmented reality applications for smart environments inside virtual reality. In GRAPP 2022. [cited by applicant]
Lavric, T. (2022). Methodologies and tools for expert knowledge sharing in manual assembly industries by using augmented reality (Doctoral dissertation, Institut Polytechnique de Paris). [cited by applicant]
Lee, B., Sedlmair, M., & Schmalstieg, D. (2023). Design patterns for situated visualization in augmented reality. IEEE Transactions on Visualization and Computer Graphics. [cited by applicant]
Lee, G. A., Kim, G. J., & Billinghurst, M. (2005). Immersive authoring: What you experience is what you get (wyxiwyg). Communications of the ACM, 48(7), 76-81. [cited by applicant]
Lee, J., & Lan, A. (Jun. 2023). Smartphone: Exploring keyword mnemonic with auto-generated verbal and visual cues. In International Conference on Artificial Intelligence in Education (pp. 16-27). Cham: Springer Nature S… [cited by applicant]
Li, W., Li, C., Kim, M., Huang, H., & Yu, L. F. (Apr. 2023). Location-aware adaptation of augmented reality narratives. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (pp. 1-15). [cited by applicant]
Li, Y. L., Zhou, S., Huang, X., Xu, L., Ma, Z., Fang, H. S., & Lu, C. (2019). Transferable interactiveness knowledge for human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision a… [cited by applicant]
Liu, D., Long, C., Zhang, H., Yu, H., Dong, X., & Xiao, C. (2020). Arshadowgan: Shadow generative adversarial network for augmented reality in single light scenes. In Proceedings of the IEEE/CVF conference on computer v… [cited by applicant]
Liu, J. S., Tversky, B., & Feiner, S. (2022). Precueing object placement and orientation for manual tasks in augmented reality. IEEE Transactions on Visualization and Computer Graphics, 28(11), 3799-3809. [cited by applicant]
Liu, S., Toreini, P., & Maedche, A. (2022). Designing Gaze-Aware Attention Feedback for Learning in Mixed Reality. In Proceedings of Mensch und Computer 2022 (pp. 503-508). [cited by applicant]
Liu, V., Vermeulen, J., Fitzmaurice, G., & Matejka, J. (Jul. 2023). 3DALL-E: Integrating text-to-image AI in 3D design workflows. In Proceedings of the 2023 ACM designing interactive systems conference (pp. 1955-1977). [cited by applicant]
Liu, Z., Zhu, Z., Jiang, E., Huang, F., Villanueva, A. M., Qian, X., & Ramani, K. (Apr. 2023). Instrumentar: Auto-generation of augmented reality tutorials for operating digital instruments through recording embodied de… [cited by applicant]
Louie, R., Coenen, A., Huang, C. Z., Terry, M., & Cai, C. J. (Apr. 2020). Novice-AI music co-creation via AI-steering tools for deep generative models. In Proceedings of the 2020 CHI conference on human factors in compu… [cited by applicant]
Luo, Y., Liu, F., She, Y., & Yang, B. (2023). A context-aware mobile augmented reality pet interaction model to enhance user experience. Computer Animation and Virtual Worlds, 34(1), e2123. [cited by applicant]
Lv, Z. (2023). Generative artificial intelligence in the metaverse era. Cognitive Robotics, 3, 208-217. [cited by applicant]
Maio, R., Santos, A., Marques, B., Ferreira, C., Almeida, D., Ramalho, P., & Santos, B. S. (2023). Pervasive Augmented Reality to support logistics operators in industrial scenarios: a shop floor user study on kit assem… [cited by applicant]
Microsoft. 2021. HoloLens2. https://www.microsoft.com/en-us/hololens/hardware. Accessed on Sep. 12, 2023. [cited by applicant]
Monteiro, K., Vatsal, R., Chulpongsatorn, N., Parnami, A., & Suzuki, R. (Apr. 2023). Teachable reality: Prototyping tangible augmented reality with everyday objects by leveraging interactive machine teaching. In Proceed… [cited by applicant]
Morris, A., Guan, J., Lessio, N., & Shao, Y. (Oct. 2020). Toward mixed reality hybrid objects with iot avatar agents. In 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (pp. 766-773). IEEE. [cited by applicant]
Muff, F., & Fill, H. G. (2022). A Framework for Context-Dependent Augmented Reality Applications Using Machine Learning and Ontological Reasoning. In AAAI Spring Symposium: MAKE. [cited by applicant]
Nebeling, M., Lewis, K., Chang, Y. C., Zhu, L., Chung, M., Wang, P., & Nebeling, J. (Apr. 2020). XRDirector: A role-based collaborative immersive authoring system. In Proceedings of the 2020 CHI Conference on Human Fact… [cited by applicant]
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., & Chen, M. (2021). Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741. [cited by applicant]
OpenAI. 2021. ChatGPT: A large-scale generative language model. https://www.openai.com/research/chatgpt. Accessed on Sep. 12, 2023. [cited by applicant]
Pan, X., Zheng, M., Xu, X., & Campbell, A. G. (2021). Knowing your student: Targeted teaching decision support through asymmetric mixed reality collaborative learning. IEEE Access, 9, 164742-164751. [cited by applicant]
Preda, M., & Lavric, T. (2023). Augmented Reality Training in Manufacturing Sectors. In The Digital Twin (pp. 447-496). Cham: Springer International Publishing. [cited by applicant]
Qian, X. (2023). Explore the Design and Authoring of Ai-driven Context-aware Augmented Reality Experiences (Doctoral dissertation, Purdue University). [cited by applicant]
Qian, X., He, F., Hu, X., Wang, T., Ipsita, A., & Ramani, K. (Apr. 2022). Scalar: Authoring semantically adaptive augmented reality experiences in virtual reality. In Proceedings of the 2022 CHI Conference on Human Fact… [cited by applicant]
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., . . . & Sutskever, I. (Jul. 2021). Learning transferable visual models from natural language supervision. In International conference on machine le… [cited by applicant]
Radford, A. (2018). Improving language understanding by generative pre-training. [cited by applicant]
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8), 9. [cited by applicant]
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., . . . & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, … [cited by applicant]
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2), 3. [cited by applicant]
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., & Sutskever, I. (Jul. 2021). Zero-shot text-to-image generation. In International conference on machine learning (pp. 8821-8831). Pmlr. [cited by applicant]
Reddy, V. S. S. S. (2022). An Exploration of the Virtual Digital Twin Capture for Spatial Tasks and its Applications (Master's thesis, Purdue University). [cited by applicant]
Redmon, J. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. 2021. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752[cs.CV]. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patt… [cited by applicant]
Zhao Ruizhi. 2021. Context-Aware AR Scheduling Assistant Using Personalized Avatars. (2021). [cited by applicant]
Sachan, M. S., & Peiris, R. L. (Apr. 2022). Designing augmented reality based interventions to encourage physical activity during virtual classes. In CHI Conference on Human Factors in Computing Systems Extended Abstrac… [cited by applicant]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mo… [cited by applicant]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. 2022. Pho… [cited by applicant]
Sandamini, A., Jayathilaka, C., Pannala, T., Karunanayaka, K., Kumarasinghe, P., & Perera, D. (Nov. 2022). An Augmented Reality-based Fashion Design Interface with Artistic Contents Generated Using Deep Generative Model… [cited by applicant]
Scargill, T. (Oct. 2021). Context-Aware Markerless Augmented Reality for Shared Educational Spaces. In 2021 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct) (pp. 469-472). IEEE. [cited by applicant]
Scargill, T., Chen, Y., Eom, S., Dunn, J., & Gorlatova, M. (Mar. 2022). Environmental, user, and social context-aware augmented reality for supporting personal development and change. In 2022 IEEE conference on virtual … [cited by applicant]
2023.Midjourney. https://www.midjourney.com/Accessed:Aug. 2, 2023. [cited by applicant]
Alblehai, F. M. (2022). Individual experience and engagement in avatar-Mediated environments: the Mediating effect of interpersonal attraction. Journal of Educational Computing Research, 60(4), 986-1007. [cited by applicant]
Athanasiou, N., Petrovich, M., Black, M. J., & Varol, G. (Sep. 2022). Teach: Temporal action composition for 3d humans. In 2022 International Conference on 3D Vision (3DV) (pp. 414-423). IEEE. [cited by applicant]
Daniel Black. 2017. Why can I see my avatar? Embodied visual engagement in the third-person videogame. Games and Culture 12, 2 (2017), 179-199. [cited by applicant]
Bowman, D. A., Gabbard, J., Auerbach, D., Roofigari-Esfahan, N., Britt, K., Ilo, C. I., & Adapa, K. (Mar. 2022). Buildar: a proof-of-concept prototype of intelligent augmented reality in construction. In 2022 IEEE Confe… [cited by applicant]
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., . . . & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877-1901. [cited by applicant]
Cai, Q., Pan, Y., Ngo, C. W., Tian, X., Duan, L., & Yao, T. (2019). Exploring object relation in mean teacher for cross-domain detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit… [cited by applicant]
Cao, J., Liu, X., Su, X., Tarkoma, S., & Hui, P. (Dec. 2021). Context-aware augmented reality with 5G edge. In 2021 IEEE Global Communications Conference (GLOBECOM) (pp. 1-6). IEEE. [cited by applicant]
Cao, Y., Fuste, A., & Heun, V. (Apr. 2022). Mobiletutar: A lightweight augmented reality tutorial system using spatially situated human segmentation videos. In CHI Conference on Human Factors in Computing Systems Extend… [cited by applicant]
Cao, Y., Li, S., Liu, Y., Yan, Z., Dai, Y., Yu, P. S., & Sun, L. (2023). A comprehensive survey of ai-generated content (aigc): A history of generative ai from gan to chatgpt. arXiv preprint arXiv:2303.04226. [cited by applicant]
Cao, Y., Qian, X., Wang, T., Lee, R., Huo, K., & Ramani, K. (Apr. 2020). An exploratory study of augmented reality presence for tutoring machine tasks. In Proceedings of the 2020 CHI conference on human factors in compu… [cited by applicant]
Cao, Y., Wang, T., Qian, X., Rao, P. S., Wadhawan, M., Huo, K., & Ramani, K. (Oct. 2019). GhostAR: A time-space editor for embodied authoring of human-robot collaborative task with augmented reality. In Proceedings of t… [cited by applicant]
Chamola, V., Bansal, G., Das, T. K., Hassija, V., Sai, S., Wang, J., & Niyato, D. (2024). Beyond reality: The pivotal role of generative ai in the metaverse. IEEE Internet of Things Magazine, 7(4), 126-135. [cited by applicant]
Chen, C., Nguyen, C., Hoffswell, J., Healey, J., Bui, T., & Weibel, N. (Oct. 2023). PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences. In Proceedings of the 36… [cited by applicant]
Chen, L., Tang, W., John, N., Wan, T. R., & Zhang, J. J. (2018). Context-aware mixed reality: A framework for ubiquitous interaction. arXiv preprint arXiv:1803.05541. [cited by applicant]
Chen, L., Tang, W., John, N. W., Wan, T. R., & Zhang, J. J. (Feb. 2020). Context-Aware Mixed Reality: A Learning-Based Framework for Semantic-Level Interaction. In Computer Graphics Forum (vol. 39, No. 1, pp. 484-496). [cited by applicant]
Chi, H. G., Chi, S., Chan, S., & Ramani, K. (May 2023). Pose Relation Transformer Refine Occlusions for Human Pose Estimation. In 2023 IEEE International Conference on Robotics and Automation (ICRA) (pp. 6138-6145). IEE… [cited by applicant]
Chidambaram, S., Huang, H., He, F., Qian, X., Villanueva, A. M., Redick, T. S., & Ramani, K. (Jun. 2021). Processar: An augmented reality-based tool to create in-situ procedural 2d/3d ar instructions. In Proceedings of … [cited by applicant]
Chidambaram, S., Reddy, S. S., Rumple, M., Ipsita, A., Villanueva, A., Redick, T., & Ramani, K. (Oct. 2022). EditAR: A Digital Twin Authoring Environment for Creation of AR/VR and Video Instructions from a Single Demons… [cited by applicant]
Christopoulos, A., Pellas, N., Kurczaba, J., & Macredie, R. (2022). The effects of augmented reality-supported instruction in tertiary-level medical education. British Journal of Educational Technology, 53(2), 307-325. [cited by applicant]
Davari, S., Lu, F., & Bowman, D. A. (Mar. 2022). Validating the benefits of glanceable and context-aware augmented reality for everyday information access tasks. In 2022 IEEE Conference on Virtual Reality and 3D User In… [cited by applicant]
Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (Jun. 2009). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255). IEEE. [cited by applicant]
Devlin, J. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. [cited by applicant]
Dhariwal, P., & Nichol, A. (2021). Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34, 8780-8794. [cited by applicant]
Doughty, M., Singh, K., & Ghugre, N. R. (2021). SurgeonAssist-Net: towards context-aware head-mounted display-based augmented reality for surgical guidance. In Medical Image Computing and Computer Assisted Intervention—… [cited by applicant]
Fan, Q., Zhuo, W., Tang, C. K., & Tai, Y. W. (2020). Few-shot object detection with attention-RPN and multi-relation detector. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 40… [cited by applicant]
Fan, Z., Taheri, O., Tzionas, D., Kocabas, M., Kaufmann, M., Black, M. J., & Hilliges, O. (2023). ARCTIC: A dataset for dexterous bimanual hand-object manipulation. In Proceedings of the IEEE/CVF Conference on Computer … [cited by applicant]
Fei, J., Xia, Z., Yu, P., & Xiao, F. (2021). Exposing AI-generated videos with motion magnification. Multimedia Tools and Applications, 80(20), 30789-30802. [cited by applicant]
Blender Foundation. 2023. Blender-Open Source 3D Creation Software. Retrieved Dec. 9, 2023 from https://www.blender.org/. [cited by applicant]
Frizziero, L., Leon-Cardenas, C., Freddi, M., Grassoni, A., & Liverani, A. (2023). Augmented reality applied to design for disassembly assessment for a volumetric pump with rotating cylinder. Production & Manufacturing … [cited by applicant]
Gallardo, A. G., Estrada, M. L. B., Cabada, R. Z., Dalle, M. Z. G., & Portillo, A. U. (Aug. 2022). EstelAR: an Augmented Reality Astronomy learning tool for STEM students. In 2022 IEEE Mexican International Conference o… [cited by applicant]
Epic Games. 2023. Unreal Engine. Retrieved Dec. 9, 2023 from https://www.unrealengine.com/. [cited by applicant]
Gattullo, M., Evangelista, A., Manghisi, V. M., Uva, A. E., Fiorentino, M., Boccaccio, A., . . . & Gabbard, J. L. (2020). Towards next generation technical documentation in augmented reality using a context-aware inform… [cited by applicant]
Gozalo-Brizuela, R., & Garrido-Merchan, E. C. (2023). ChatGPT is not all you need. A State of the Art Review of large Generative AI models. arXiv. arXiv preprint arXiv:2301.04655. [cited by applicant]
Grubert, J., Langlotz, T., Zollmann, S., & Regenbrecht, H. (2016). Towards pervasive augmented reality. [cited by applicant]
Guo, C., Zou, S., Zuo, X., Wang, S., Ji, W., Li, X., & Cheng, L. (2022). Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (p… [cited by applicant]
Gutierrez, F., Verbert, K., & Htun, N. N. (Sep. 2018). PHARA: an augmented reality grocery store assistant. In Proceedings of the 20th International Conference on Human-Computer Interaction with Mobile Devices and Servi… [cited by applicant]
Herskovitz, J., Cheng, Y. F., Guo, A., Sample, A. P., & Nebeling, M. (2022). Xspace: An augmented reality toolkit for enabling spatially-aware distributed collaboration. Proceedings of the ACM on Human-Computer Interact… [cited by applicant]
Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in neural information processing systems, 33, 6840-6851. [cited by applicant]
Hoover, M., Miller, J., Gilbert, S., & Winer, E. (2020). Measuring the performance impact of using the microsoft hololens 1 to provide guided assembly work instructions. Journal of Computing and Information Science in E… [cited by applicant]
Hu, Y., Yuan, M., Xian, K., Samitha Elvitigala, D., & Quigley, A. (2023). Exploring the design space of employing ai-generated content for augmented reality display. arXiv e-prints, arXiv-2303. [cited by applicant]
Huang, G., Qian, X., Wang, T., Patel, F., Sreeram, M., Cao, Y., & Quinn, A. J. (May 2021). Adaptutar: An adaptive tutoring system for machine tasks in augmented reality. In Proceedings of the 2021 CHI Conference on Huma… [cited by applicant]
Huang, Q., Park, J. S., Gupta, A., Bennett, P., Gong, R., Som, S., & Gao, J. (2023). Ark: Augmented reality with knowledge interactive emergent ability. arXiv preprint arXiv:2305.00970. [cited by applicant]
Jeong, H., & Kim, G. J. (Mar. 2023). Table2table: Merging “similar” workspaces and supporting adaptive telepresence demonstration guidance. In 2023 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and… [cited by applicant]
Jo, D., Kim, K. H., & Kim, G. J. (2015). SpaceTime: adaptive control of the teleported avatar for improved AR tele-conference experience. Computer Animation and Virtual Worlds, 26(3-4), 259-269. [cited by applicant]
Johnson, J., Gupta, A., & Fei-Fei, L. (2018). Image generation from scene graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1219-1228). [cited by applicant]
Justus, D., Brennan, J., Bonner, S., & McGough, A. S. (Dec. 2018). Predicting the computational cost of deep learning models. In 2018 IEEE international conference on big data (Big Data) (pp. 3873-3882). IEEE. [cited by applicant]
Kalamkar, D., Mudigere, D., Mellempudi, N., Das, D., Banerjee, K., Avancha, S., . . . & Dubey, P. (2019). A study of BFLOAT16 for deep learning training. arXiv preprint arXiv:1905.12322. [cited by applicant]
Kang, C., Yeom, I., Ashtari, A., Woo, W., & Noh, J. (2023). ARbility: re-inviting older wheelchair users to in-store shopping via wearable augmented reality. Virtual Reality, 27(3), 1919-1936. [cited by applicant]
Karunratanakul, K., Preechakul, K., Suwajanakorn, S., & Tang, S. (2023). Gmd: Controllable human motion synthesis via guided diffusion models. arXiv preprint arXiv:2305.12577, 3. [cited by applicant]