IP Library Granted Patent US 12,430,857
Granted Patent B2
US 12,430,857 · App. 18/155,626 · Granted Sep 30, 2025

Neural extension of 3D content in augmented reality environments

Inventors: Hayko Jochen Wilhelm Riemenschneider (Zurich, CH); Erika Varis Doggett (Los Angeles, CA); Evan Matthew Goldberg (Burbank, CA); Shinobu Hattori (Los Angeles, CA); Christopher Richard Schroers (Uster, CH)
Assignee: DISNEY ENTERPRISES, INC.
G06T19/006G06V10/26G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,857
App. No.
18/155,626
Granted
Sep 30, 2025
Kind
B2
Abstract

Generating augmented reality content includes inputting a first layout of a physical space and a first set of anchor content into a machine learning model; generating, via execution of the machine learning model, a first augmented reality view that includes (i) a first portion of the physical space and (ii) an extension of the first set of anchor content across a second portion of the physical space; and causing the first augmented reality view to be outputted in a computing device.

Claims (48)

1. A computer-implemented method for generating augmented reality content, the method comprising:

inputting a first layout of a physical space and a first set of anchor content into a machine learning model, wherein the first set of anchor content is represented within the physical space;

generating, via execution of the machine learning model, a first three-dimensional (3D) volume that includes (i) a first subset of the physical space including the first set of anchor content and (ii) a placement of one or more 3D representations of the first set of anchor content in a second subset of the physical space, wherein:

the placement of the one or more 3D representations is located at a different position within the physical space relative to a position of the first set of anchor content within the first subset of physical space, and

generating the first 3D volume comprises:

applying a first set of neural network layers included in the machine learning model to the first set of anchor content to generate a semantic segmentation of the first set of anchor content, and

applying a second set of neural network layers included in the machine learning model to the first layout, the first set of anchor content, and the semantic segmentation to generate the first 3D volume; and

causing one or more views of the first 3D volume to be outputted within an augmented reality environment provided by in a computing device.

2. The computer-implemented method of claim 1 , wherein generating the first 3D volume comprises:

applying a first set of neural network layers included in the machine learning model to the first set of anchor content to generate the one or more 3D representations; and

applying a second set of neural network layers included in the machine learning model to the one or more 3D representations and the first layout to determine the placement of the one or more 3D representations in the second subset of the physical space.

3. The computer-implemented method of claim 1 , further comprising generating, via execution of the machine learning model, the first layout as a semantic segmentation of sensor data associated with the physical space.

4. The computer-implemented method of claim 3 , wherein the sensor data comprises at least one of an image of the physical space, a point cloud, a mesh, or a depth map.

5. The computer-implemented method of claim 1 , further comprising training the machine learning model based on a set of training layouts, a set of training anchor images, and one or more losses associated with the first 3D volume.

6. The computer-implemented method of claim 5 , wherein the one or more losses comprise a layout loss that is computed based on a representation of the first subset of the physical space in the first 3D volume and a corresponding subset of the physical space.

7. The computer-implemented method of claim 5 , wherein the one or more losses comprise a layout loss that is computed based on the first layout and the placement of the one or more 3D representations of the first set of anchor content within the first 3D volume.

8. The computer-implemented method of claim 1 , wherein the first 3D volume comprises a neural radiance field.

9. The computer-implemented method of claim 1 , wherein the first set of anchor content comprises at least one of an image, a video, or a 3D object.

10. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

inputting a first layout of a physical space and a first set of anchor content into a machine learning model, wherein the first set of anchor content is represented within the physical space;

generating, via execution of the machine learning model, a first three-dimensional (3D) volume that includes (i) a first subset of the physical space including the first set of anchor content and (ii) a placement of one or more 3D representations of the first set of anchor content in a second subset of the physical space, wherein:

the placement of the one or more 3D representations is located at a different position within the physical space relative to a position of the first set of anchor content within the first subset of physical space; and

generating the first 3D volume comprises:

applying a first set of neural network layers included in the machine learning model to the first set of anchor content to generate a semantic segmentation of the first set of anchor content, and

applying a second set of neural network layers included in the machine learning model to the first layout, the first set of anchor content, and the semantic segmentation to generate the first 3D volume; and

causing one or more views of the first 3D volume to be outputted within an augmented reality environment provided by in a computing device.

11. The one or more non-transitory computer-readable media of claim 10 , wherein the instructions further cause the one or more processors to perform the step of applying a set of neural network layers included in the machine learning model to sensor data associated with the physical space to generate the first layout, wherein the first layout includes predictions of objects for regions of the sensor data.

12. The one or more non-transitory computer-readable media of claim 10 , wherein the instructions further cause the one or more processors to perform the steps of:

generating, via execution of the machine learning model, a second 3D volume that includes (i) a third subset of the physical space and (ii) a placement of one or more 3D representations of a second set of anchor content in a fourth subset of the physical space; and

causing one or more views of the second 3D volume to be outputted in the computing device.

13. The one or more non-transitory computer-readable media of claim 12 , wherein the first set of anchor content and the second set of anchor content comprise at least one of different video frames included in a video, depictions of two different scenes, or different sets of 3D objects.

14. The one or more non-transitory computer-readable media of claim 10 , wherein the instructions further cause the one or more processors to perform the step of training the first machine learning model based on one or more losses associated with the first 3D volume.

15. The one or more non-transitory computer-readable media of claim 14 , wherein the one or more losses comprise a similarity loss that is computed based on the first set of anchor content and a rendering of the second subset of the physical space within the 3D volume.

16. The one or more non-transitory computer-readable media of claim 14 , wherein the one or more losses comprise a reconstruction loss that is computed based on the one or more 3D representations of the first set of anchor content generated by the machine learning model and one or more 3D objects.

17. The one or more non-transitory computer-readable media of claim 14 , wherein the one or more losses comprise a segmentation loss that is computed based on a semantic segmentation of the first set of anchor content and a ground truth segmentation associated with the first set of anchor content.

18. The one or more non-transitory computer-readable media of claim 10 , wherein causing the one or more views of the first 3D volume to be outputted in the computing device comprises:

rendering the first 3D volume from the one or more views; and

outputting the one or more views within an augmented reality environment provided by the computing device.

19. A system, comprising:

one or more memories that store instructions, and

one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of:

inputting a first layout of a physical space and a first set of anchor content into a machine learning model, wherein the first set of anchor content is represented within the physical space;

generating, via execution of the machine learning model, a first three-dimensional (3D) volume that includes (i) a first subset of the physical space including the first set of anchor content and (ii) a placement of one or more 3D representations of the first set of anchor content in a second subset of the physical space, wherein:

the placement of the one or more 3D representations is located at a different position within the physical space relative to a position of the first set of anchor content within the first subset of physical space; and

generating the first 3D volume comprises:

applying a first set of neural network layers included in the machine learning model to the first set of anchor content to generate a semantic segmentation of the first set of anchor content, and

applying a second set of neural network layers included in the machine learning model to the first layout, the first set of anchor content, and the semantic segmentation to generate the first 3D volume; and

causing one or more views of the first 3D volume to be outputted within an augmented reality environment provided by in a computing device.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2023
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 062482/0255 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2023
From: RIEMENSCHNEIDER, HAYKO JOCHEN WILHELM; SCHROERS, CHRISTOPHER RICHARD
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 062453/0836 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2023
From: DOGGETT, ERIKA VARIS; GOLDBERG, EVAN MATTHEW; HATTORI, SHINOBU
To: DISNEY ENTERPRISES, INC.
Reel/Frame 062425/0487 →
Continuity (1)
Related Publication 20240242444A1 · Jul 18, 2024
References Cited (35)
US 11210843B1 · Coffey et al. · 2021 [cited by applicant]
US 20120229508A1 · Wigdor et al. · 2012 [cited by applicant]
US 20190371279A1 · Mak · 2019 [cited by applicant]
US 20200111256A1 · Bleyer · 2020 [cited by examiner]
US 20200226736A1 · Kar · 2020 [cited by examiner]
US 20210142497A1 · Pugh et al. · 2021 [cited by applicant]
US 20210173968A1 · Yang et al. · 2021 [cited by applicant]
US 20210201565A1 · Dibra et al. · 2021 [cited by applicant]
US 20210241109A1 · Jie · 2021 [cited by applicant]
US 20210373834A1 · Goldberg et al. · 2021 [cited by applicant]
US 20210383595A1 · Chapman et al. · 2021 [cited by applicant]
JP 2014515130A · 2014 [cited by applicant]
JP 2020042802A · 2020 [cited by applicant]
JP 2021527247A · 2021 [cited by applicant]
JP 2022505775A · 2022 [cited by applicant]
Extended European Search Report for Application No. 24152403.2 dated Jun. 24, 2024. [cited by applicant]
Extended European Search Report for Application No. 24152408.1 dated Jun. 24, 2024. [cited by applicant]
Teterwak et al., “Boundless: Generative Adversarial Networks for Image Extension”, Aug. 19, 2019, pp. 10521-10530. [cited by applicant]
Frueh et al., “Headset Removal for Virtual and Mixed Reality”, Proceedings of SIGGRAPH Talks, DOI: http://dx.doi.org/10.1145/3084363.3085083, Jul. 30-Aug. 3, 2017, 2 pages. [cited by applicant]
Koh et al. “Simple and Effective Synthesis of Indoor 3D Scenes”, The Thirty-Seventh AAAI Conference on Artificial Intelligence , Dec. 1, 2022, pp. 1169-1178. [cited by applicant]
ARCore, “Working with Anchors”, Google Developers, Retrieved from https://developers.google.com/ar/develop/anchors, on Nov. 30, 2022, 5 pages. [cited by applicant]
Azure, “Object Anchors”, Microsoft, Retrieved from https://azure.microsoft.com/en-us/services/object-anchors, on Nov. 30, 2022, 12 pages. [cited by applicant]
Niantic Lightship, “Build the Real-World Metaverse”, Retrieved from https://lightship.dev, on Nov. 30, 2022, 8 pages. [cited by applicant]
“DALL-E 2”, Retrieved from https://openai.com/dall-e-2, on Nov. 30, 2022, 14 pages. [cited by applicant]
Google Research, Brain Team, “Imagen: Text-to-Image Diffusion Models”, Retrieved from https://imagen.research.google, on Nov. 30, 2022, 18 pages. [cited by applicant]
Bastian, Matthias, “The Decoder”, What would Mona Lisa look like with a body? DALL-E 2 has an answer, Retrieved from https://mixed-news.com/en/what-would-mona-lisa-look-like-with-a-body-dall-e-2-has-an-answer/, on Nov. … [cited by applicant]
Sketchfab, “Persistence of Memory 3D”, Retrieved from https://sketchfab.com/3d-models/persistence-of-memory-3d-ffe139656cc241f28d5dc95dd80ccc64, on Nov. 30, 2022, 6 pages. [cited by applicant]
U.S. patent application entitled, “Space and Content Matching for Augmented and Mixed Reality”, U.S. Appl. No. 17/749,005, filed May 19, 2022, 33 pages. [cited by applicant]
U.S. patent application entitled, “Augmented Reality Enhancement of Moving Images”, U.S. Appl. No. 17/887,731, filed Aug. 15, 2022, 35 pages. [cited by applicant]
U.S. patent application entitled, “User Responsive Augmented Reality Enhancement of Moving Images”, U.S. Appl. No. 17/887,742 , filed Aug. 15, 2022, 42 pages. [cited by applicant]
U.S. patent application entitled, “Dynamic Scale Augmented Reality Enhancement of Images”, U.S. Appl. No. 17/887,754, filed Aug. 15, 2022, 40 pages. [cited by applicant]
Non Final Office Action received for U.S. Appl. No. 18/155,635 dated Sep. 5, 2024, 22 pages. [cited by applicant]
Office Action received for Japanese Application No. 2023-209927 dated Jan. 22, 2025, 5 pages. [cited by applicant]
Office Action received for Japanese Application No. 2023-209931 dated Jan. 22, 2025, 5 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 18/155,635 dated Jan. 17, 2025, 18 pages. [cited by applicant]