IP Library › Granted Patent US 12,664,718
Granted Patent B2
US 12,664,718 · App. 18/279,717 · Granted Jun 23, 2026

Texture completion

Inventors: Jongyoo Kim (Beijing, CN); Jiaolong Yang (Beijing, CN); Xin Tong (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC
G06T15/04G06T7/40G06T7/70G06V10/54G06V20/70G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,718
App. No.
18/279,717
Filed
Aug 31, 2023
Granted
Jun 23, 2026
Kind
B2
Examiner
WANG, YUEHAN
Art Unit
2617
USPC
345/582
Abstract

According to implementations of the present disclosure, there is provided a solution for completing textures of an object. In this solution, a complete texture map of an object is generated from a partial texture map of the object according to a texture generation model. A first prediction on whether a texture of at least one block in the complete texture map is an inferred texture is determined according to a texture discrimination model. A second image of the object is generated based on the complete texture map. A second prediction on whether the first image and the second image are generated images is determined according to an image discrimination model. The texture generation model, the texture and image discrimination models are trained based on the first and second predictions.

Claims (62)

1 . A computer-implemented method comprising:

generating a complete texture map of an object from a partial texture map of the object according to a texture generation model, the partial texture map comprising visible textures in a first image of the object and the complete texture map comprising the visible textures and inferred textures;

determining, for at least one block in the complete texture map, a first prediction on whether the at least one block comprises an inferred texture according to a texture discrimination model;

determining a second prediction on whether the first image and a second image of the object are generated images according to an image discrimination model, the second image generated based on the complete texture map, the second prediction being determined by predicting whether a first set of patches in the first image and a second set of patches in the second image are generated patches according to the image discrimination model; and

training the texture generation model, the texture discrimination model and the image discrimination model based on the first and second predictions by:

determining a first set of labels corresponding to the first set of patches, each of the first set of labels indicating that a respective patch of the first set of patches is not a generated patch,

determining a second set of labels corresponding to the second set of patches based on visibility of textures of the second set of patches in the first image, and

training the texture generation model, the texture discrimination model and the image discrimination model based on the second prediction, the first set of labels and the second set of labels.

2 . The method of claim 1 , wherein training the texture generation model, the texture discrimination model and the image discrimination model comprises:

predicting a position of the at least one block in the complete texture map based on an output of an intermediate layer of the texture discrimination model; and

training the texture generation model, the texture discrimination model and the image discrimination model further based on a difference between the predicted position and an actual position of the at least one block in the complete texture map.

3 . The method of claim 1 , further comprising:

determining a plurality of blocks from the complete texture map, each of the plurality of blocks comprising a plurality of pixels of the complete texture map;

determining, for each of the plurality of blocks, a first ratio of the number of valid pixels to the number of the plurality of pixels, a valid pixel comprising a visible texture in the first image; and

selecting the at least one block from the plurality of blocks based on respective first ratios determined for the plurality of blocks, the first ratio determined for the at least one selected block exceeding a first threshold or being below a second threshold and the first threshold exceeding the second threshold.

4 . The method of claim 1 , wherein determining the second set of labels corresponding to the second set of patches comprises:

generating a visibility mask indicating valid pixels and non-valid pixels in the second image, a valid pixel comprising a visible texture in the first image and a non-valid pixel comprising an inferred texture from the complete texture map; and

determining the second set of labels by performing on the visibility mask a convolution operation associated with the second set of patches.

5 . The method of claim 1 , wherein determining the second set of labels corresponding to the second set of patches comprises:

determining, for a given patch of the second set of patches, a second ratio of the number of valid pixels in the given patch to the number of pixels in the given patch, a valid pixel comprising a visible texture in the first image;

in accordance with a determination that the second ratio exceeds a threshold, assigning a first label to the given patch, the first label indicating that the given patch is not a generated patch; and

in accordance with a determination that the second ratio is below the threshold, assigning a second label to the given patch, the second label indicating that the given patch is a generated patch.

6 . The method of claim 1 , further comprising:

determining a target pose of the object based on a distribution of poses of objects in a plurality of training images comprising the first image; and

generating the second image based on the complete texture map and the target pose.

7 . The method of claim 1 , further comprising:

obtaining a partial texture map of a further object, the partial texture map comprising visible textures in a third image of the further object; and

after training of the texture generation model, generating a complete texture map of the further object from the partial texture map according to the texture generation model, the complete texture map comprising the visible textures in the third image and inferred textures.

8 . An electronic device, comprising:

at least one processing unit; and

at least one memory coupled to the at least one processing unit and having instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising:

generating a complete texture map of an object from a partial texture map of the object according to a texture generation model, the partial texture map comprising visible textures in a first image of the object and the complete texture map comprising the visible textures and inferred textures;

determining, for at least one block in the complete texture map, a first prediction on whether the at least one block comprises an inferred texture according to a texture discrimination model;

determining a second prediction on whether the first image and a second image of the object are generated images according to an image discrimination model, the second image generated based on the complete texture map, the second prediction being determined by predicting whether a first set of patches in the first image and a second set of patches in the second image are generated patches according to the image discrimination model; and

training the texture generation model, the texture discrimination model and the image discrimination model based on the first and second predictions by:

determining a first set of labels corresponding to the first set of patches, each of the first set of labels indicating that a respective patch of the first set of patches is not a generated patch,

determining a second set of labels corresponding to the second set of patches based on visibility of textures of the second set of patches in the first image, and

training the texture generation model, the texture discrimination model and the image discrimination model based on the second prediction, the first set of labels, and the second set of labels.

9 . The device of claim 8 , wherein training the texture generation model, the texture discrimination model and the image discrimination model comprises:

predicting a position of the at least one block in the complete texture map based on an output of an intermediate layer of the texture discrimination model; and

training the texture generation model, the texture discrimination model and the image discrimination model further based on a difference between the predicted position and an actual position of the at least one block in the complete texture map.

10 . The device of claim 8 , wherein the acts further comprises:

determining a plurality of blocks from the complete texture map, each of the plurality of blocks comprising a plurality of pixels of the complete texture map;

determining, for each of the plurality of blocks, a first ratio of the number of valid pixels to the number of the plurality of pixels, a valid pixel comprising a visible texture in the first image; and

selecting the at least one block from the plurality of blocks based on respective first ratios determined for the plurality of blocks, the first ratio determined for the at least one selected block exceeding a first threshold or being below a second threshold and the first threshold exceeding the second threshold.

11 . The device of claim 8 , wherein the acts further comprises:

determining a target pose of the object based on a distribution of poses of objects in a plurality of training images comprising the first image; and

generating the second image based on the complete texture map and the target pose.

12 . The electronic device of claim 8 , the acts further comprising:

obtaining a partial texture map of a further object, the partial texture map comprising visible textures in a third image of the further object; and

after training of the texture generation model, generating a complete texture map of the further object from the partial texture map according to the texture generation model, the complete texture map comprising the visible textures in the third image and inferred textures.

13 . A non-transitory computer storage medium comprising computer-executable instructions which, when executed by a device, cause the device to perform acts comprising:

generating a complete texture map of an object from a partial texture map of the object according to a texture generation model, the partial texture map comprising visible textures in a first image of the object and the complete texture map comprising the visible textures and inferred textures,

determining, for at least one block in the complete texture map, a first prediction on whether the at least one block comprises an inferred texture according to a texture discrimination model;

determining a second prediction on whether the first image and a second image of the object are generated images according to an image discrimination model, the second image generated based on the complete texture map, the second prediction being determined by predicting whether a first set of patches in the first image and a second set of patches in the second image are generated patches according to the image discrimination model; and

training the texture generation model, the texture discrimination model and the image discrimination model based on the first and second predictions by:

determining a first set of labels corresponding to the first set of patches, each of the first set of labels indicating that a respective patch of the first set of patches is not a generated patch,

determining a second set of labels corresponding to the second set of patches based on visibility of textures of the second set of patches in the first image, and

training the texture generation model, the texture discrimination model and the image discrimination model based on the second prediction, the first set of labels, and the second set of labels.

14 . The non-transitory computer storage medium of claim 13 , further comprising computer-executable instructions which, when executed by a device, cause the device to perform acts comprising:

obtaining a partial texture map of a further object, the partial texture map comprising visible textures in a third image of the further object; and

after training of the texture generation model, generating a complete texture map of the further object from the partial texture map according to the texture generation model, the complete texture map comprising the visible textures in the third image and inferred textures.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2023
From: KIM, JONGYOO; YANG, JIAOLONG; TONG, XIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064812/0827 →
Continuity (1)
Related Publication 20240161382A1 · May 16, 2024
References Cited (69)
US 10242484B1 · Cernigliaro · 2019 [cited by examiner]
US 10389994B2 · Graziosi · 2019 [cited by applicant]
US 20200151940A1 · Yu · 2020 [cited by examiner]
CN 110378230A · 2019 [cited by applicant]
CN 111881926A · 2020 [cited by applicant]
WO 2017123163A1 · 2017 [cited by applicant]
WO 2018102700A1 · 2018 [cited by applicant]
Fan et al., Full Face-and-Head 3D Model With Photorealistic Texture, Oct. 27, 2020, Digital Object Identifier 10.1109/Access.2020.3031886, vol. 8, 2020, pp. 210709-210721 (Year: 2020). [cited by examiner]
Office Action received for EP Application No. 21938235.5, mailed on Date Dec. 5, 2023, 3 Pages. [cited by applicant]
Blanz, et al., “A Morphable Model for the Synthesis of 3D Faces”, In Proceedings of the 26th annual conference on Computer graphics and interactive techniques, Jul. 1999, pp. 187-194. [cited by applicant]
Booth, et al., “3D Face Morphable Models “In-the-Wild””, In Proceedings of In IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 48-57. [cited by applicant]
Booth, et al., “A 3D Morphable Model Learnt from 10,000 faces”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 15543-5552. [cited by applicant]
Deng, et al., “Accurate 3D Face Reconstruction with Weakly-Supervised Learning: From Single Image to Image Set”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 16, 2… [cited by applicant]
Deng, et al., “UV-GAN: Adversarial Facial UV Map Completion for Pose-Invariant Face Recognition”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 7093-7102. [cited by applicant]
Fan, et al., “Full Face-and-Head 3D Model With Photorealistic Texture”, In Journal of IEEE Access, vol. 8, Oct. 27, 2020, pp. 210709-210721. [cited by applicant]
Feng, et al., “Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network”, In Proceedings of the European Conference on Computer Vision, Sep. 8, 2018, 18 Pages. [cited by applicant]
Gecer, et al., “GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 1155-116… [cited by applicant]
Gecer, et al., “OSTeC: One-Shot Texture Completion”, In Repository of arxiv code: 2012.15370v1 [cs.CV], Dec. 30, 2020, 19 Pages. [cited by applicant]
Genova, et al., “Unsupervised Training for 3D Morphable Model Regression”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 8377-8386. [cited by applicant]
Guo, et al., “CNN-Based Real-Time Dense Face Reconstruction with Inverse-Rendered Photo-Realistic Face Images”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, Issue, 6, 2018, 16 P… [cited by applicant]
Hassner, et al., “Effective Face Frontalization in Unconstrained Images”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2015, 10 Pages. [cited by applicant]
Ichim, et al., “Dynamic 3D Avatar Creation from Hand-held Video Input”, In Journal of ACM Transactions on Graphics, vol. 34, Issue 4, Jul. 27, 2015, 14 Pages. [cited by applicant]
Iizuka, et al., “Globally and Locally Consistent Image Completion”, In Journal of ACM Transactions on Graphics, vol. 36, Issue 4, Jul. 20, 2017, 14 Pages. [cited by applicant]
Isola, et al., “Image-to-Image Translation with Conditional Adversarial Networks”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 1125-1134. [cited by applicant]
Karra, et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, In Proceedings of in IEEE Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 4401-4410. [cited by applicant]
Karras, et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation”, In Proceedings of 6th International Conference on Learning Representations, Apr. 30, 2018, 26 Pages. [cited by applicant]
Lazova, et al., “360-Degree Textures of People in Clothing from a Single Image”, In Proceedings of International Conference on 3D Vision, Sep. 16, 2019, 11 Pages. [cited by applicant]
Lee, et al., “StyleUV: Diverse and High-quality UV Map Generative Model”, In Repository of arXiv:2011.12893v1, Nov. 25, 2020, 12 Pages. [cited by applicant]
Li, et al., “Recurrent Feature Reasoning for Image Inpainting”, In Proceedings of IEEE International Conference on Computer Vision, Jun. 13, 2020, pp. 7760-7768. [cited by applicant]
Lin, et al., “Cocogan: Generation by Parts via Conditional Coordinating”, In Proceedings of IEEE International Conference on Computer Vision, Oct. 27, 2019, pp. 4512-4521. [cited by applicant]
Lin, et al., “Face Parsing With Rol Tanh-Warping”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 5647-5656. [cited by applicant]
Liu, et al., “Image Inpainting for Irregular Holes Using Partial Convolutions”, In Proceedings of the European Conference on Computer Vision, Sep. 8, 2018, 16 Pages. [cited by applicant]
Mao, et al., “Least Squares Generative Adversarial Networks”, In Proceedings of IEEE International Conference on Computer Vision, Oct. 22, 2017, pp. 2813-2821. [cited by applicant]
Nguyen, et al., “Dual Discriminator Generative Adversarial Nets”, In Proceedings of the 31st International Conference on Neural Information Processing Systems, Sep. 2017, 11 Pages. [cited by applicant]
Odena, et al., “Conditional Image Synthesis with Auxiliary Classifier GANs”, In Proceedings of the 34th International Conference on Machine Learning, Aug. 2017, 10 Pages. [cited by applicant]
Olszewski, Kyle, “Realistic Dynamic Facial Textures from a Single Image using GANs”, In Proceeding of IEEE International Conference on Computer Vision, Oct. 22, 2017, pp. 5439-5448. [cited by applicant]
Park, et al., “SRFeat: Single Image Super-Resolution with Feature Discrimination”, In Proceedings of the European Conference on Computer Vision, Sep. 2018, 17 Pages. [cited by applicant]
Parkhi, et al., “Deep Face Recognition”, In Proceedings of British Machine Vision Conference, Sep. 2015, 12 Pages. [cited by applicant]
Pathak, et al., “Context Encoders: Feature Learning by Inpainting”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 26, 2016, pp. 2536-2544. [cited by applicant]
Pavllo, et al., “Convolutional Generation of Textured 3D Meshes”, In Proceedings of the 34th Conference on Neural Information Processing Systems, Jun. 2020, 13 Pages. [cited by applicant]
Paysan, et al., “A 3D Face Model for Pose and Illumination Invariant Face Recognition”, In Proceedings of the Sixth IEEE International Conference on Advanced Video and Signal Based Surveillance, Sep. 2, 2009, pp. 296-30… [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/CN21/090047”, Mailed Date: Feb. 7, 2022, 9 Pages. [cited by applicant]
Ploumpis, et al., “Combining 3D Morphable Models: A Large scale Face-and-Head Model”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 16, 2019, pp. 10934-10943. [cited by applicant]
Richardson, et al., “3D Face Reconstruction by Learning from Synthetic Data”, In Proceedings of Fourth International Conference on 3D Vision, Oct. 25, 2016, pp. 460-469. [cited by applicant]
Saito, et al., “Photorealistic Facial Texture Inference Using Deep Neural Networks”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 5144-5153. [cited by applicant]
Sengupta, et al., “SfSNet: Learning Shape, Reflectance and Illuminance of Faces in the wild”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, 10 Pages. [cited by applicant]
Wang, et al., “High-Resolution Image Synthesis and Semantic Manipulation With Conditional GANs”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 8798-8807. [cited by applicant]
Wang, et al., “Image Quality Assessment: From Error Visibility to Structural Similarity”, In Journal of IEEE Transactions on Image Processing, vol. 13, Issue 4, Apr. 13, 2004, pp. 600-612. [cited by applicant]
Xu, et al., “Deep 3D Portrait from a Single Image”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 7710-7720. [cited by applicant]
Yamaguchi, et al., “High-Fidelity Facial Reflectance and Geometry Inference from an Unconstrained Image”, In Journal of ACM Transactions on Graphics, vol. 37, Issue 4, Jul. 2018, 14 Pages. [cited by applicant]
Yang, et al., “FaceScape: A Largescale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Apr. 2020, pp. 601-610. [cited by applicant]
Yu, et al., “Free-Form Image Inpainting with Gated Convolution”, In Proceedings of IEEE/CVF International Conference on Computer Vision, Oct. 27, 2019, pp. 4470-4479. [cited by applicant]
Yu, Jiahui, “Generative Image Inpainting with Contextual Attention”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 5505-5514. [cited by applicant]
Yu, et al., “Multi-Scale Context Aggregation by Dilated Convolutions”, In Proceedings of International Conference on Learning Representations, May 2, 2016, 13 Pages. [cited by applicant]
Yuan, et al., “Face De-occlusion using 3D Morphable Model and Generative Adversarial Network”, In Proceedings of IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 10062-10071. [cited by applicant]
Zheng, et al., “Pluralistic Image Completion”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 1438-1447. [cited by applicant]
Kim, et al., “Learning High-Fidelity Face Texture Completion without Complete Face Texture”, In Proceedings of IEEE/CVF International Conference on Computer Vision, Oct. 10, 2021, pp. 13990-13999. [cited by applicant]
Lin, et al., “Towards High-Fidelity 3D Face Reconstruction from In-the-Wild Images using Graph Convolutional Networks”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp… [cited by applicant]
Lattas, et al., “AvatarMe: Realistically Renderable 3D Facial Reconstruction “In-the-Wild””, In Proceedings of IEEE/CVD Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 757-766. [cited by applicant]
Kingma, et al., “ADAM: A Method for Stochastic Optimization”, In Proceedings of 3rd International Conference on Learning Representations, May 7, 2015, 15 Pages. [cited by applicant]
Gross, et al., “Multi-Pie”, In Journal of Image and Vision Computing, vol. 28, Issue 5, May 1, 2010, 21 Pages. [cited by applicant]
Dai, et al., “SGNN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 846-855. [cited by applicant]
Abadi, et al., “TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems”, In Publication of Preliminary White Paper, Nov. 9, 2015, 19 Pages. [cited by applicant]
Extended European search report received for European Application No. 21938235.5, mailed on Jan. 16, 2025, 7 pages. [cited by applicant]
Kim, et al., “Learning High-Fidelity Face Texture Completion Without Complete Face Texture”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 10, 2021, pp. 13990-13999. [cited by applicant]
Xue, et al., “Side Information for Face Completion: a Robust PCA Approach”, Retrieved from: arXiv: 1801.07580v1, Jan. 20, 2018, 15 pages. [cited by applicant]
Yu, et al., “Generative Image Inpainting with Contextual Attention”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 5505-5514. [cited by applicant]
Communication pursuant to Rules 70(2) and 70a(2) received for European Application No. 21938235.5 mailed on Feb. 4, 2025, 1 page. [cited by applicant]
First Office Action Received for Chinese Application No. 202180097514.1, mailed on Mar. 17, 2026, 31 pages. (English translation Provided). [cited by applicant]