IP Library › Granted Patent US 12,657,667
Granted Patent B2
US 12,657,667 · App. 18/534,243 · Granted Jun 16, 2026

Diversity-preserved domain adaptation using text-to-image diffusion for 3D generative model

Inventors: Se Young Chun (Seoul, KR); Gwanghyun Kim (Seoul, KR)
Assignee: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
G06T5/60G06T5/20G06T5/50G06T7/80G06T19/20G06T2207/20081G06T2207/20084G06T2207/30244G06T2219/2024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,667
App. No.
18/534,243
Granted
Jun 16, 2026
Kind
B2
Abstract

An embodiment of the present disclosure provides a three-dimensional image generation method, which is performed by a server and is capable of domain adaptation, including generating N target images corresponding to a second domain by converting styles of previously collected N source images corresponding to a first domain according to instructions of an input text, selecting only a target image that satisfies a preset condition among the N target images, and generating multiple three-dimensional images corresponding to a specific domain through certain noise data and a preset camera pose parameter by training a three-dimensional generation model, which is previously built, by using the selected target image.

Claims (44)

1 . A three-dimensional image generation method that is performed by a server and is capable of domain adaptation, the three-dimensional image generation method comprising:

generating N target images corresponding to a second domain by converting styles of previously collected N source images corresponding to a first domain according to instructions of an input text;

selecting only a target image that satisfies a preset condition among the N target images; and

generating multiple three-dimensional images corresponding to a specific domain through certain noise data and a preset camera pose parameter by training a three-dimensional generation model, which is previously built, by using the selected target image,

wherein the N source images and the N target images are three-dimensional images each composed of multiple camera viewpoints, and

each of the N target images is converted to an image of at least one style set in the second domain while maintaining an object included in each of the N source images.

2 . The three-dimensional image generation method of claim 1 , further comprising:

before the generating of the N target images, generating the N source images by inputting the certain noise data including an identifier of the first domain and the preset camera pose parameter to the three-dimensional generation model before the training.

3 . The three-dimensional image generation method of claim 1 , wherein

the generating of the N target images includes generating the N target images by stochastically converting the N source images to match multiple styles indicated by the text through a text-to-image diffusion model, which is previously trained, by using data pairs of an image and a text.

4 . The three-dimensional image generation method of claim 1 , wherein

the selecting of only the target image includes filtering a target image of which difference from a style indicated by the text is greater than or equal to a preset threshold among the N target images.

5 . The three-dimensional image generation method of claim 1 , wherein

the selecting of only the target image includes filtering a target image of which a camera viewpoint difference from a corresponding source image is greater than or equal to a preset threshold among the N target images.

6 . The three-dimensional image generation method of claim 1 , wherein the generating of the multiple three-dimensional images includes:

outputting a new three-dimensional image by inputting the certain noise data including an identifier of the first domain and the camera pose parameter to the three-dimensional generation model; and

performing fine tuning of the output three-dimensional image such that an adversarial loss according to a difference between the output three-dimensional image and a target image of the specific domain is reduced.

7 . The three-dimensional image generation method of claim 1 , wherein

the generating of the multiple three-dimensional images includes performing fine tuning of the three-dimensional image output from the three-dimensional generation model through the certain noise data to correspond to the selected target image by selecting at least one of multiple styles set for the specific domain.

8 . The three-dimensional image generation method of claim 1 , wherein

the generating of the multiple three-dimensional images includes generating the multiple three-dimensional images by implementing a three-dimensional embedding space corresponding to the specific domain by mapping previously collected actual two-dimensional images to the three-dimensional embedding space and inputting the mapped two-dimensional images to the three-dimensional generation model.

9 . A three-dimensional image generation server comprising:

a memory storing a program for performing a domain adaptable three-dimensional image generation method; and

a processor configured to execute the program,

wherein the processor includes:

a target image generation unit configured to generate N target images corresponding to a second domain by converting styles of previously collected N source images corresponding to a first domain according to instructions of an input text,

a target image filtering unit configured to select only a target image that satisfies a preset condition among the N target images, and

a domain adaptation unit configured to generating multiple three-dimensional images corresponding to a specific domain through certain noise data and a preset camera pose parameter by training a three-dimensional generation model, which is previously built, by using the selected target image, and

wherein the N source images and the N target images are three-dimensional images each composed of multiple camera viewpoints, and

each of the N target images is converted to an image of at least one style set in the second domain while maintaining an object included in each of the N source images.

10 . The three-dimensional image generation server of claim 9 , wherein

the processor generates the N source images by inputting the certain noise data including an identifier of the first domain and the preset camera pose parameter to the three-dimensional generation model before training by the domain adaptation unit.

11 . The three-dimensional image generation server of claim 9 , wherein

the target image generation unit generates the N target images by stochastically converting the N source images to match multiple styles indicated by the text through a text-to-image diffusion model, which is previously trained, by using data pairs of an image and a text.

12 . The three-dimensional image generation server of claim 9 , wherein

the target image filtering unit filters a target image of which difference from a style indicated by the text is greater than or equal to a preset threshold among the N target images.

13 . The three-dimensional image generation server of claim 9 , wherein

the target image filtering unit filters a target image of which a camera viewpoint difference from a corresponding source image is greater than or equal to a preset threshold among the N target images.

14 . The three-dimensional image generation server of claim 9 , wherein

the domain adaptation unit outputs a new three-dimensional image by inputting the certain noise data including an identifier of the first domain and the camera pose parameter to the three-dimensional generation model, and performs fine tuning of the output three-dimensional image such that an adversarial loss according to a difference between the output three-dimensional image and a target image of the specific domain is reduced.

15 . The three-dimensional image generation server of claim 9 , wherein

the domain adaptation unit performs fine tuning of the three-dimensional image output from the three-dimensional generation model through the certain noise data to correspond to the selected target image by selecting at least one of multiple styles set for the specific domain.

16 . The three-dimensional image generation server of claim 9 , wherein

the domain adaptation unit generates the multiple three-dimensional images by implementing a three-dimensional embedding space corresponding to the specific domain by mapping previously collected actual two-dimensional images to the three-dimensional embedding space and inputting the mapped two-dimensional images to the three-dimensional generation model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2023
From: CHUN, SE YOUNG; KIM, GWANGHYUN
To: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 065817/0535 →
Priority Claims (1)
KR 10-2023-0107395 · Aug 17, 2023 · national
Continuity (1)
Related Publication 20250061545A1 · Feb 20, 2025
References Cited (9)
US 11776180B2 · Xu · 2023 [cited by examiner]
US 20220044352A1 · Liao et al. · 2022 [cited by applicant]
US 20220076374A1 · Li · 2022 [cited by examiner]
US 20240290054A1 · Yin · 2024 [cited by examiner]
CN 116071452A · 2023 [cited by examiner]
KR 102345520B1 · 2021 [cited by applicant]
KR 1020220129828A · 2022 [cited by applicant]
C. Wang, M. Chai, M. He, D. Chen and J. Liao, “Cross-Domain and Disentangled Face Manipulation With 3D Guidance,” in IEEE Transactions on Visualization and Computer Graphics, vol. 29, No. 4, pp. 2053-2066, Apr. 1, 2023.… [cited by examiner]
Gwanghyun Kim et al., DATID-3D: Diversity-Preserved Domain Adaptation Using Text-to-Image Diffusion for 3D Generative Model, arXiv, Nov. 29, 2022, pp. 1-21. [cited by applicant]