IP Library Granted Patent US 12682552
Granted Patent B2
US 12682552 · App. 18/476,670 · Granted Jul 14, 2026

AI methods for transforming a text prompt into an immersive volumetric photo or video

Inventors: Lawrence Wayne Neal (Salem, OR); Forrest Briggs (Palo Alto, CA)
Assignee: Stevens Law Group, P.C.
G06T15/10G06F40/40G06T5/77G06T13/20G06T19/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682552
App. No.
18/476,670
Granted
Jul 14, 2026
Kind
B2
Abstract

A text-to-image prompt is processed using a text-to-image machine learning model to obtain a non-immersive (e.g., rectilinear image). The non-immersive image may be enhanced by a superresolution machine learning model and processed with a monoscopic depth estimation model to obtain a depthmap. The non-immersive image and the depthmap may be converted to an immersive projection (e.g., F-theta) and corresponding depth map. The immersive projection may be out-painted. The immersive projection may be used to generate video with simulated camera movement, output on a VR headset, and/or processed to remove a background layer and displayed on an AR headset, or on an holographic glasses-free three-dimensional display.

Claims (71)

1 . A system comprising:

one or more processing devices; and

one or more memory devices operably coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to perform a method comprising:

receiving a rectilinear image;

generating a first depthmap from the rectilinear image using a monocular depth estimation algorithm;

generating an immersive projection of the rectilinear image according to a projection mapping and a second depthmap;

generating a second depthmap corresponding to the immersive projection from the first depthmap;

receiving a text-to-image prompt; and

processing the text-to-image prompt with a text-to-image machine learning model to obtain the rectilinear image.

2 . The system of claim 1 , wherein the immersive projection is one of an F-theta projection or an inflated F-theta projection.

3 . The system of claim 1 , wherein the method further comprises:

alternatively processing the text-to-image prompt, the alternative processing further comprising:

processing the text-to-image prompt with a text-to-image machine learning model to obtain an original image; and

processing the text-to-image prompt with a superresolution machine learning model to obtain the rectilinear image, the rectilinear image having higher resolution than the original image.

4 . The system of claim 1 , wherein the method further comprises out-painting the immersive projection using an out-painting machine learning model.

5 . The system of claim 4 , wherein out-painting the immersive projection comprises:

(a) adjusting a viewing angle with respect to the immersive projection;

(b) generating a rectilinear projection of the immersive projection;

(c) out-painting the rectilinear projection using the out-painting machine learning model;

(d) projecting the rectilinear projection onto the immersive projection; and

(e) repeating (a) to (d) until the immersive projection is completely out-painted.

6 . The system of claim 1 , wherein the rectilinear image is one of a plurality of images.

7 . The system of claim 1 , wherein the method further comprises:

identifying a background in the immersive projection; and

removing the background from the immersive projection.

8 . The system of claim 7 , wherein the method further comprises transmitting a rendering of the immersive projection to an augmented reality display device.

9 . The system of claim 1 , wherein the method further comprises transmitting a rendering of the immersive projection to three-dimensional display device, the three-dimensional display device being any of a virtual reality display device, three-dimensional display device requiring viewing using glasses, three-dimensional display device that does not require viewing using glasses, or holographic three-dimensional display device that does not require viewing using glasses.

10 . The system of claim 1 , wherein the method further comprises:

generating a simulated camera movement; and

generating video frames simulating perception of the immersive projection by a camera traversing the simulated camera movement.

11 . A non-transitory computer-readable medium storing executable code that, when executed by one or more processing devices, causes the one or more processing devices to perform a method comprising:

(a) receiving a text-generation prompt;

(b) processing the text-generation prompt with a large language model (LLM) to obtain a text-to-image prompt;

(c) processing the text-to-image prompt with a text-to-image machine learning model to obtain an image; and

(d) presenting a representation of the image to a source of the text-generation prompt,

wherein the image is a rectilinear image and (d) further comprises generating an immersive projection and depthmap from the rectilinear image.

12 . The non-transitory computer-readable medium of claim 11 , further comprising repeating (a), (b), (c), and (d) with a state of the LLM being updated for each iteration of (b) and (b) being performed according to the state of the LLM.

13 . The non-transitory computer-readable medium of claim 11 , wherein the text-to-image prompt comprises a data object.

14 . The non-transitory computer-readable medium of claim 11 , wherein the text-to-image prompt specifies locations for objects.

15 . The non-transitory computer-readable medium of claim 11 , wherein the rectilinear image is one of a plurality of images.

16 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises:

identifying a background in the immersive projection;

removing the background from the immersive projection; and

transmitting a rendering of the immersive projection to an augmented reality display device.

17 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises transmitting a rendering of the immersive projection to a three-dimensional display device, the three-dimensional display device being any of a virtual reality display device, three-dimensional display device requiring viewing using glasses, three-dimensional display device that does not require viewing using glasses, or holographic three-dimensional display device that does not require viewing using glasses.

18 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises:

generating a simulated camera movement; and

generating video frames simulating perception of the immersive projection by a camera traversing the simulated camera movement.

19 . A system comprising:

one or more processing devices; and

one or more memory devices operably coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to perform a method comprising:

receiving a rectilinear image;

generating a first depthmap from the rectilinear image using a monocular depth estimation algorithm;

generating an immersive projection of the rectilinear image according to a projection mapping and a second depthmap;

generating a second depthmap corresponding to the immersive projection from the first depthmap;

receiving a text-to-image prompt;

processing the text-to-image prompt with a text-to-image machine learning model to obtain an original image; and

processing the text-to-image prompt with a superresolution machine learning model to obtain the rectilinear image, the rectilinear image having higher resolution than the original image.

20 . A system comprising:

one or more processing devices; and

one or more memory devices operably coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to perform a method comprising:

receiving a rectilinear image;

generating a first depthmap from the rectilinear image using a monocular depth estimation algorithm;

generating an immersive projection of the rectilinear image according to a projection mapping and a second depthmap;

generating a second depthmap corresponding to the immersive projection from the first depthmap; and

out-painting the immersive projection using an out-painting machine learning model, wherein out-painting the immersive projection comprises:

(a) adjusting a viewing angle with respect to the immersive projection;

(b) generating a rectilinear projection of the immersive projection;

(c) out-painting the rectilinear projection using the out-painting machine learning model;

(d) projecting the rectilinear projection onto the immersive projection; and

(e) repeating (a) to (d) until the immersive projection is completely out-painted.