IP Library › Granted Patent US 12,688,641
Granted Patent B2
US 12,688,641 · App. 18/528,567 · Granted Jul 21, 2026

Graphic rendering using diffusion models

Inventor: Tony Francis (New York, NY)
Assignee: Dream3D, Inc.
G06T15/00G06V10/764G06V10/7715G06V10/774G06V10/82G06T2210/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,641
App. No.
18/528,567
Filed
Dec 4, 2023
Granted
Jul 21, 2026
Kind
B2
Art Unit
2615
USPC
345/419
Abstract

A system or method for generating computer graphics. One or more three-dimensional (3D) scenes are obtained and rasterized into a first set of two-dimensional (2D) images having a first resolution. Features are extracted from the first set of 2D images, and text prompts are generated based on the features. A diffusion model is applied to the features, and the text prompts to generate a second set of 2D images having a second resolution greater than the first resolution. The diffusion model is trained over a dataset comprising images and corresponding text descriptions to generate an image consistent with a text prompt. The second set of 2D images having the second resolution are caused to be rendered at a client device.

Claims (45)

1 . A method comprising:

receiving, from a client device, a sequence of one or more three-dimensional scenes;

rasterizing the sequence of one or more three-dimensional scenes into a first set of one or more two-dimensional images having a first resolution;

extracting features from the first set of one or more two-dimensional images;

generating one or more text prompts based on the extracted features;

for each of the one or more text prompts,

generating a set of prompt embeddings based on the text prompt; and

generating a set of latent variables based on the features associated with the text prompt;

sampling the first set of one or more two-dimensional images at different times;

applying a diffusion model to 1) the features associated with a first subset of the first set of two-dimensional images and associated with a second subset of the first set of two-dimensional images selected to overlap the first subset and 2) the one or more text prompts to generate a second set of one or more two-dimensional images having a second resolution greater than the first resolution, wherein the diffusion model is trained over a dataset comprising images and corresponding text descriptions to generate an image consistent with a text prompt, wherein the diffusion model includes a two-dimensional Unet configured to receive the extracted features from the first set of one or more two-dimensional images, the sampled first set of one or more two-dimensional images, the set of prompt embeddings, and the set of latent variables to generate the second set of one or more images; and

causing the second set of one or more images to be rendered at a client device.

2 . The method of claim 1 , wherein the features include two-dimensional feature maps extracted by a convolutional neural network.

3 . The method of claim 1 , further comprising:

identifying an object based on the features extracted from the one or more two-dimensional images, wherein the one or more text prompts include a text prompt corresponding to the object.

4 . The method of claim 1 , further comprising:

identifying an environment based on the features extracted from the one or more two-dimensional images, wherein the one or more text prompts include a text prompt corresponding to the environment.

5 . A computer program product comprising a non-transitory computer readable storage medium having instructions encoded thereon that, when executed by one or more processors, cause the one or more processors to:

receive, from a client device, a sequence of one or more three-dimensional scenes;

rasterize the sequence of one or more three-dimensional scenes into a first set of one or more two-dimensional images having a first resolution;

extract features from the first set of one or more two-dimensional images;

generate one or more text prompts based on the extracted features;

for each of the one or more text prompts,

generate a set of prompt embeddings based on the text prompt; and

generate a set of latent variables based on the features associated with the text prompt;

sample the first set of one or more two-dimensional images at different times;

apply a diffusion model to 1) the features associated with a first subset of the first set of two-dimensional images and associated with a second subset of the first set of two-dimensional images selected to overlap the first subset and 2) the one or more text prompts to generate a second set of one or more two-dimensional images having a second resolution greater than the first resolution, wherein the diffusion model is trained over a dataset comprising images and corresponding text descriptions to generate an image consistent with a text prompt, wherein the diffusion model includes a two-dimensional Unet configured to receive the extracted features from the first set of one or more two-dimensional images, the sampled first set of one or more two-dimensional images, the set of prompt embeddings, and the set of latent variables to generate the second set of one or more images; and

cause the second set of one or more images to be rendered at a client device.

6 . The computer program product of claim 5 , wherein the features include two-dimensional feature maps extracted by a convolutional neural network.

7 . The computer program product of claim 5 , wherein the one or more processors are further caused to:

identify an object based on the features extracted from the one or more two-dimensional images, wherein the one or more text prompts include a text prompt corresponding to the object.

8 . The computer program product of claim 5 , wherein the one or more processors are further caused to:

identify an environment based on the features extracted from the one or more two-dimensional images, wherein the one or more text prompts include a text prompt corresponding to the environment.

9 . A computer system comprising:

one or more processors; and

a non-transitory computer readable storage medium having instructions encoded thereon that, when executed by the one or more processors, cause the one or more processors to:

receiving, from a client device, a sequence of one or more three-dimensional scenes;

rasterize the sequence of one or more three-dimensional scenes into a first set of one or more two-dimensional images having a first resolution;

extract features from the first set of one or more two-dimensional images;

generate one or more text prompts based on the extracted features;

for each of the one or more text prompts,

generate a set of prompt embeddings based on the text prompt; and

generate a set of latent variables based on the features associated with the text prompt;

sample the first set of one or more two-dimensional images at different times;

apply a diffusion model to 1) the features associated with a first subset of the first set of two-dimensional images and associated with a second subset of the first set of two-dimensional images selected to overlap the first subset and 2) the one or more text prompts to generate a second set of one or more two-dimensional images having a second resolution greater than the first resolution, wherein the diffusion model is trained over a dataset comprising images and corresponding text descriptions to generate an image consistent with a text prompt, wherein the diffusion model includes a two-dimensional Unet configured to receive the extracted features from the first set of one or more two-dimensional images, the sampled first set of one or more two-dimensional images, the set of prompt embeddings, and the set of latent variables to generate the second set of one or more images; and

cause the second set of one or more images to be rendered at a client device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2023
From: FRANCIS, TONY
To: DREAM3D, INC.
Reel/Frame 065765/0207 →
Continuity (2)
Provisional Application 63430238 · Dec 5, 2022
Related Publication 20240185498A1 · Jun 6, 2024
References Cited (15)
US 12322068B1 · Kim · 2025 [cited by examiner]
US 20210158561A1 · Park et al. · 2021 [cited by applicant]
US 20210374384A1 · Munkberg et al. · 2021 [cited by applicant]
US 20220138455A1 · Nagano et al. · 2022 [cited by applicant]
US 20220284582A1 · Yang et al. · 2022 [cited by applicant]
US 20240005604A1 · Kreis · 2024 [cited by examiner]
US 20240087179A1 · Min · 2024 [cited by examiner]
US 20240169479A1 · Wang · 2024 [cited by examiner]
US 20240185035A1 · Yu · 2024 [cited by examiner]
Gafni, Oran, et al. “Make-a-scene: Scene-based text-to-image generation with human priors.” European conference on computer vision. Cham: Springer Nature Switzerland, 2022. (Year: 2022). [cited by examiner]
Singer, Uriel, et al. “Make-a-video: Text-to-video generation without text-video data.” arXiv preprint arXiv:2209.14792 (2022). (Year: 2022). [cited by examiner]
Tao, Ming, et al. “Df-gan: A simple and effective baseline for text-to-image synthesis.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022. (Year: 2022). [cited by examiner]
Ho, Jonathan, et al. “Video diffusion models.” Advances in neural information processing systems 35 (2022): 8633-8646. (Year: 2022). [cited by examiner]
Zhou, Daquan, et al. “Magicvideo: Efficient video generation with latent diffusion models.” arXiv preprint arXiv:2211.11018 (2022). ( Year: 2022). [cited by examiner]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US23/82366, Mar. 22, 2024, 12 pages. [cited by applicant]