IP Library › Granted Patent US 12,651,418
Granted Patent B2
US 12,651,418 · App. 18/317,851 · Granted Jun 9, 2026

Spatial document system and method

Inventors: Jennifer Healey (San Jose, CA); Tong Sun (San Ramon, CA); Nicholas Rewkowski (San Jose, CA); Nedim Lipka (Campbell, CA); Curtis Wigington (San Jose, CA); Alexa Siu (San Jose, CA)
Assignee: Adobe Inc.
G06T19/006G10L15/08G10L15/22G06T2219/004G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,418
App. No.
18/317,851
Granted
Jun 9, 2026
Kind
B2
Abstract

A computing system captures image data using a camera and captures spatial information using one or more sensors. The computing system receives voice data using a microphone. The computing system analyzes the voice data to identify a keyword. The computing system analyzes the image data and the spatial information to identify an object corresponding to the keyword. The computing system generates text based on the voice data and the keyword. The computing system stores the text in association with the object. The computing system generates and provides output comprising the text linked to the object or a derivative thereof.

Claims (94)

1 . A method comprising:

capturing image data using a camera;

capturing spatial information using one or more sensors;

receiving voice data using a microphone;

analyzing the voice data to identify a keyword;

analyzing the image data and the spatial information to identify an object corresponding to the keyword;

generating text based on the voice data and the keyword;

storing the text in association with the object; and

generating and providing output comprising the text linked to the object or a derivative thereof, wherein generating the output comprises:

capturing second image data;

identifying one or more image markers in the second image data;

traversing stored image data to identify the image markers;

retrieving the generated text based on an association with the stored image data;

retrieving data identifying the object in the stored image data;

aligning a second object in the second image data with the object in the stored image data; and

overlaying the text on the object based on the alignment.

2 . The method of claim 1 , wherein the method is performed by a computing device and storing the text in association with the object comprises:

storing an anchor location with respect to a position of the computing device;

storing information characterizing an intended viewpoint; and

storing the text in association with the anchor location and the information characterizing the intended viewpoint.

3 . The method of claim 1 , further comprising:

establishing a set of weights, each weight corresponding to a type of image, audio, or sensor data based on an accuracy thereof; and

computing a saliency value as a function of input data and respective weights of the set of weights, wherein the weights and the saliency value are further stored in association with the object and the text.

4 . The method of claim 1 , wherein the text is a first text, the method further comprising:

receiving a second text;

identifying a relationship between the second text and the first text or the object; and

updating the stored text in association with the object to comprise the second text and an indication of the relationship.

5 . The method of claim 4 , wherein the first text is received from a first user and the second text is received from a second user.

6 . The method of claim 1 , further comprising:

based on the image data, the spatial information, and the voice data, identifying and tracking one or more floors in a building,

wherein the spatial information comprises GPS data and gyroscope data.

7 . The method of claim 1 , further comprising storing, with the text in association with the object:

metadata specifying a date and a user.

8 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

capturing image data using a camera;

capturing spatial information using one or more sensors;

receiving user input associated with the image data;

analyzing the user input to identify a keyword;

analyzing the image data and the spatial information to identify an object corresponding to the keyword;

generating text based on the user input and the keyword;

storing the text in association with the object; and

generating and providing output comprising the text linked to the object or a derivative thereof, wherein generating the output comprises:

capturing second image data;

identifying one or more image markers in the second image data;

traversing stored image data to identify the image markers;

retrieving the generated text based on an association with the stored image data;

retrieving data identifying the object in the stored image data;

aligning a second object in the second image data with the object in the stored image data; and

overlaying the text on the object based on the alignment.

9 . The system of claim 8 , wherein the operations are performed by a computing device, the operations further comprising:

storing an anchor location with respect to a position of the computing device;

storing information characterizing an intended viewpoint; and

storing the text in association with the anchor location and the information characterizing the intended viewpoint.

10 . The system of claim 8 , the operations further comprising:

establishing a set of weights, each weight corresponding to a type of image, audio, or sensor data based on an accuracy thereof; and

computing a saliency value as a function of input data and respective weights of the set of weights, wherein the weights and the saliency value are further stored in association with the object and the text.

11 . The system of claim 8 , wherein the text is a first text, the operations further comprising:

receiving a second text;

identifying a relationship between the second text and the first text or the object; and

updating the stored text in association with the object to comprise the second text and an indication of the relationship.

12 . The system of claim 11 , wherein:

the first text is received from a first user and the second text is received from a second user.

13 . The system of claim 8 , the operations further comprising:

based on the image data, the spatial information, and the user input, identifying and tracking one or more floors in a building,

wherein the spatial information comprises GPS data and gyroscope data.

14 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

capturing image data using a camera;

capturing spatial information using one or more sensors;

receiving voice data using a microphone;

analyzing the voice data to identify a keyword;

analyzing the image data and the spatial information to identify an object corresponding to the keyword;

generating text based on the voice data and the keyword;

storing the text in association with the object; and

generating and providing output comprising the text linked to the object or a derivative thereof, wherein generating the output comprises:

capturing second image data;

identifying one or more image markers in the second image data;

traversing stored image data to identify the image markers;

retrieving the generated text based on an association with the stored image data;

retrieving data identifying the object in the stored image data;

aligning a second object in the second image data with the object in the stored image data; and

overlaying the text on the object based on the alignment.

15 . The medium of claim 14 , wherein the operations are performed by a computing device, the operations further comprising:

storing an anchor location with respect to a position of the computing device;

storing information characterizing an intended viewpoint; and

storing the text in association with the anchor location and the information characterizing the intended viewpoint.

16 . The medium of claim 14 , the operations further comprising:

establishing a set of weights, each weight corresponding to a type of image, audio, or sensor data based on an accuracy thereof; and

computing a saliency value as a function of input data and respective weights of the set of weights, wherein the weights and the saliency value are further stored in association with the object and the text.

17 . The medium of claim 14 , wherein the text is a first text received from a first user, the operations further comprising:

receiving a second text from a second user;

identifying a relationship between the second text and the first text or the object; and

updating the stored text in association with the object to comprise the second text and an indication of the relationship.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: HEALEY, JENNIFER; SUN, TONG; REWKOWSKI, NICHOLAS; LIPKA, NEDIM; WIGINGTON, CURTIS; SIU, ALEXA
To: ADOBE INC.
Reel/Frame 063648/0001 →
Continuity (1)
Related Publication 20240386675A1 · Nov 21, 2024
References Cited (48)
US 7627556B2 · Liu · 2009 [cited by examiner]
US 11521018B1 · Barzelay · 2022 [cited by examiner]
US 11721333B2 · Lee · 2023 [cited by examiner]
US 20150154232A1 · Ovsjanikov · 2015 [cited by examiner]
US 20170103072A1 · Yuen · 2017 [cited by examiner]
US 20190187479A1 · Nishizawa · 2019 [cited by examiner]
US 20190332657A1 · Jones · 2019 [cited by examiner]
US 20200258517A1 · Park · 2020 [cited by examiner]
US 20210232836A1 · Graefe · 2021 [cited by examiner]
US 20210304451A1 · Fortier · 2021 [cited by examiner]
US 20230055477A1 · Mohajer · 2023 [cited by examiner]
CN 107919127A · 2018 [cited by examiner]
CN 110430356A · 2019 [cited by examiner]
GB 2546368A · 2017 [cited by examiner]
KR 20190115839A · 2019 [cited by examiner]
WO WO2018133307A1 · 2018 [cited by examiner]
“Adding Editable Text”, https://sparkar.facebook.com/ar-studio/learn/articles/2D/editable-text/, (Accessed on Feb. 6, 2023). [cited by applicant]
“AR anchor manager”, AR Foundation, 4.1.13 (unity3d.com), (Accessed on Feb. 6, 2023). [cited by applicant]
“Content Persistence Fundamentals” 2019. Magic Leap. https://ml1-developer.magicleap.com/en-us/learn/guides/content-persistence-fundamentals. (Updated Oct. 23, 2019; Accessed on Mar. 10, 2022). [cited by applicant]
“Creating screen annotations for objects in an AR experience”, https://developer.apple.com/documentation/arkit/content_anchors/creating_screen_annotations_for_objects_in_an_ar_experience, (Accessed on Feb. 6, 2023). [cited by applicant]
“HoloLens.” 2022. Spatial anchors—Mixed Reality | Microsoft Docs. https://docs.microsoft.com/en-us/windows/mixed-reality/design/spatial-anchors (Accessed on Mar. 10, 2023). [cited by applicant]
“Tutorial on voice and text annotations with the HoloLens”, (codeholo.com), (Accessed on Feb. 6, 2023). [cited by applicant]
“Working with Anchor”, ARCore, Google Developers, (Accessed on Feb. 6, 2023). [cited by applicant]
Beck, Stephan, Andre Kunert, Alexander Kulik, and Bernd Froehlich. 2013. Immersive group-to-group telepresence. IEEE transactions on visualization and computer graphics 19, 4 (2013), 616-625. [cited by applicant]
Bertasius, Gedas and Lorenzo Torresani. 2021. Classifying, Segmenting, and Tracking Object Instances in Video with Mask Propagation. arXiv:1912.04573. [cited by applicant]
Breunig, Martin, Patrick Erik Bradley, Markus Jahn, Paul Kuper, Nima Mazroob, Norbert Rösch, Mulhim Al-Doori, Emmanuel Stefanakis, and Mojgan Jadidi. 2020. Geospatial data management research: Progress and future direct… [cited by applicant]
Caarls, Jurjen, Pieter Jonker, and Stelian Persa. 2003. Sensor fusion for augmented reality. In European Symposium on Ambient Intelligence. Springer, 160-176. [cited by applicant]
Caron, Francois, Emmanuel Duflos, Denis Pomorski, and Philippe Vanheeghe. 2006. GPS/IMU data fusion using multisensor Kalman filtering: introduction of contextual aspects. Information fusion 7, 2 (2006), 221-230. [cited by applicant]
Cheah, Thomas CS Cheah and K-W Ng. 2005. A practical implementation of a 3D game engine. In International Conference on Computer Graphics, Imaging and Visualization (CGIV'05). IEEE, 351-358. [cited by applicant]
Gaillard, Jeremy, Adrien Peytavie, and Gilles Gesquière. 2016. A Data Structure for Progressive Visualisation and Edition of Vectorial Geospatial Data. In 3D GeoInfo, vol. 2. 201-209. [cited by applicant]
Guan, Peiyu, Zhiqiang Cao, Erkui Chen, Shuang Liang, Min Tan, and Junzhi Yu. 2020. A real-time semantic visual SLAM approach with points and objects. International Journal of Advanced Robotic Systems 17, 1 (2020), 17298… [cited by applicant]
Jiang, Jade, Michael Tobia, Robert Lawther, Dominic Branchaud, and Tomasz Bednarz. 2020. Double vision: Digital twin applications within extended reality. In ACM SIGGRAPH 2020 Appy Hour. 1-2. [cited by applicant]
Khan, Latif U., Walid Saad, Dusit Niyato, Zhu Han, and Choong Seon Hong. 2021. Digital-Twin-Enabled 6G: Vision, Architectural Trends, and Future Directions. arXiv preprint arXiv:2102.12169 (2021). [cited by applicant]
Klosowski, James T., Martin Held, Joseph SB Mitchell, Henry Sowizral, and Karel Zikan. 1998. Efficient collision detection using bounding volume hierarchies of k-DOPs. IEEE transactions on Visualization and Computer Gra… [cited by applicant]
Li, L. and Michael F. Goodchild. 2011. “Automatically and Accurately Matching Objects in Geospatial Datasets”, Remote Sensing and Spatial Information Sciences,, vol. 38, Part II, pp. 98-103. [cited by applicant]
Mazurek, Patryk, and Tomasz Hachaj. 2021. SLAM-OR: Simultaneous Localization, Mapping and Object Recognition Using Video Sensors Data in Open Environments from the Sparse Points Cloud. Sensors 21, 14. [cited by applicant]
Nicholson, Lachlan, Michael Milford, and Niko Sünderhauf. 2018. QuadricSLAM: Dual Quadrics from Object Detections as Landmarks in Object-oriented SLAM. arXiv:1804.04011 [cs.RO]. [cited by applicant]
Orts-Escolano, Sergio, Christoph Rhemann, Sean Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim, Philip L Davidson, Sameh Khamis, Mingsong Dou, et al. 2016. Holoportation: Virtual 3d teleportation in real-… [cited by applicant]
Pillai, Sudeep and John Leonard. 2015. Monocular SLAM Supported Object Recognition. arXiv:1506.01732 [cs.RO]. [cited by applicant]
Raskar, Ramesh, GregWelch, Matt Cutts, Adam Lake, Lev Stesin, and Henry Fuchs. 1998. The office of the future: A unified approach to image-based modeling and spatially immersive displays. In Proceedings of the 25th annu… [cited by applicant]
Redmon and Farhadi, YOLO 9000: Better, Faster, Stronger, arXiv:1612.08242 (2016). [cited by applicant]
Schall, Gerhard, Daniel Wagner, Gerhard Reitmayr, Elise Taichmann, Manfred Wieser, Dieter Schmalstieg, and Bernhard Hofmann-Wellenhof. 2009. Global pose estimation using multi-sensor fusion for outdoor augmented reality… [cited by applicant]
Sowizral, Henry, 2000. Scene graphs in the new millennium. IEEE Computer Graphics and Applications 20, 1 (2000), 56-67. [cited by applicant]
Wang, Guangting, Chong Luo, Xiaoyan Sun, Zhiwei Xiong, and Wenjun Zeng. 2020. Tracking by Instance Detection: A Meta-Leaming Approach.arXiv:2004.00830. [cited by applicant]
Weng, Xinshuo, Jianren Wang, David Held, and Kris Kitani. 2020. 3D Multi-Object Tracking: A Baseline and New Evaluation Metrics. arXiv:1907.03961. [cited by applicant]
Zhang, Pifu, Jason Gu, Evangelos E Milios, and Peter Huynh. 2005. Navigation with IMU/GPS/digital compass with unscented Kalman filter. In IEEE International Conference Mechatronics and Automation, 2005, vol. 3. IEEE, 1… [cited by applicant]
Zhang, Jun, Mina Henein, Robert Mahony, and Viorela Ila. 2020. VDO-SLAM: A Visual Dynamic Object-aware SLAM System. (2020). arXiv:2005.11052. [cited by applicant]
Zhou, Xingyi, Vladlen Kollun, and Philipp Kr.henbühl. 2020. Tracking Objects as Points, arXiv:2004.01177. [cited by applicant]