IP Library Granted Patent US 12,561,885
Granted Patent B2
US 12,561,885 · App. 18/338,093 · Granted Feb 24, 2026

Method, system, and medium for artificial intelligence-based completion of a 3D image during electronic communication

Inventors: Matǔŝ Kirchmayer (Košice, SK); Gergely Magyar (Vel'ke Kapusany, SK); Mária Virĉíková (Košice, SK); Rudolf Jakša (Košice, SK)
Assignee: MATSUKO s.r.o.
G06T15/04G06T7/11G06T7/55G06T7/60G06T15/10G06T2200/08G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,885
App. No.
18/338,093
Filed
Jun 20, 2023
Granted
Feb 24, 2026
Kind
B2
Examiner
HOANG, PHI
Art Unit
2619
USPC
345/419
Abstract

A method and system for reconstructing a photo-realistic, three-dimensional representation of at least part of a conference participant's head from image data, head alignment data, and depth data. The image data is projected from a world space into an object space using the head alignment data and depth data. In the object space, at least part of the area missing from the image data is completed using a computational model of a person. An artificial neural network is used to reconstruct the texture and depth of the reconstruction as part of the completion. At least the image data and the head alignment data are determined based on two-dimensional or 2.5-dimensional images of the conference participant captured by a camera.

Claims (56)

1 . A method comprising:

obtaining image data, head alignment data, and depth data of at least part of a three-dimensional head comprising a face of a conference participant, wherein at least the image data and head alignment data are determined based on two-dimensional or 2.5-dimensional images of the conference participant captured by a camera;

reconstructing a photo-realistic three-dimensional representation of the at least part of the head from the image data, the head alignment data, and the depth data, wherein reconstructing the representation of the at least part of the head comprises reconstructing an area missing from the image data acquired by the camera,

wherein reconstructing the representation comprises:

projecting the image data from a world space into an object space using the head alignment data and depth data; and

in the object space, completing at least part of the area missing from the image data using a computational model of a person, and

wherein completing at least part of the area missing from the image data comprises applying an artificial neural network to reconstruct a texture and a depth of the at least part of the head.

2 . The method of claim 1 , further comprising obtaining head segmentation data delineating separation of the head from background.

3 . The method of claim 2 , further comprising encoding the head segmentation data as an alpha channel.

4 . The method of claim 3 , further comprising:

(a) obtaining information defining at least one hole respectively representing at least one missing portion in the object space, wherein the information defining the at least one hole is obtained from the head segmentation data, depth data, and head alignment data; and

(b) encoding the information defining the at least one hole in the alpha channel.

5 . The method of claim 1 , wherein the computational model is specific to the conference participant.

6 . The method of claim 1 , wherein the computational model is generated from multiple persons.

7 . The method of claim 1 , further comprising:

(a) identifying at least one of:

(i) frames of the image data that depict the head from an angle beyond an angle limit; and

(ii) frames of the image data that depict parts of the face less than a face threshold; and

(b) removing the identified frames from the image data prior to the completing.

8 . The method of claim 1 , further comprising obtaining a normals channel identifying respective normals for different points on the at least part of the head, and wherein the artificial neural network uses the normals channel during completion.

9 . The method of claim 1 , wherein the reconstructing is performed based on the image data, the head alignment data, and the depth data retrieved from at least one video frame.

10 . The method of claim 9 , wherein the reconstructed three-dimensional representation is determined from a current image frame and a composition frame generated from multiple past image frames, wherein the multiple past image frames are sampled at a constant frequency.

11 . The method of claim 10 , further comprising obtaining head alignment data of the conference participant, and wherein the multiple past image frames are sampled based on the head alignment data.

12 . The method of claim 9 , wherein the reconstructed three-dimensional representation is determined from a current image frame and a composition frame generated from multiple past image frames, wherein the multiple past image frames are irregularly sampled.

13 . The method of claim 1 , further comprising, after the completing, smoothing or removing a back of the head.

14 . The method of claim 1 , further comprising training the artificial neural network prior to the reconstructing, wherein the training comprises:

capturing, using multiple cameras, multiple views of the face of the conference participant from different perspectives to obtain time-synchronized images comprising red, green, blue and depth channels;

stitching the time-synchronized images together to generate three-dimensional training data comprising a complete version of the face of the conference participant; and

training the artificial neural network using the training data.

15 . The method of claim 1 , further comprising training the artificial neural network prior to the reconstructing, wherein the training comprises:

capturing, using a single camera, multiple images of the face of the conference participant, wherein the multiple images are captured from different angles and show differential facial expressions of the conference participant;

stitching the multiple images together to generate three-dimensional training data comprising a complete version of the face of the conference participant; and

training the artificial neural network using the training data.

16 . The method of claim 1 , further comprising training the artificial neural network prior to the reconstructing, wherein the training comprises:

retrieving from storage a recording of a three-dimensional video comprising the face of the conference participant;

stitching different frames of the recording together to generate three-dimensional training data comprising a complete version of the face of the conference participant; and

training the artificial neural network using the training data.

17 . A system comprising:

a network interface;

a processor communicatively coupled to the network interface; and

a non-transitory computer readable medium communicatively coupled to the processor and having stored thereon computer program code that is executable by the processor and that, when executed by the processor, causes the processor to perform a method comprising:

obtaining image data, head alignment data, and depth data of at least part of a three-dimensional head comprising a face of a conference participant, wherein at least the image data and head alignment data are determined based on two-dimensional or 2.5-dimensional images of the conference participant captured by a camera;

reconstructing a photo-realistic three-dimensional representation of the at least part of the head from the image data, the head alignment data, and the depth data, wherein reconstructing the representation of the at least part of the head comprises reconstructing an area missing from the image data acquired by the camera,

wherein reconstructing the representation comprises:

projecting the image data from a world space into an object space using the head alignment data and depth data; and

in the object space, completing at least part of the area missing from the image data using a computational model of a person, and

wherein completing at least part of the area missing from the image data comprises applying an artificial neural network to reconstruct a texture and a depth of the at least part of the head.

18 . The system of claim 17 , further comprising a camera communicatively coupled to the processor, the camera for capturing an image of the conference participant.

19 . The system of claim 17 , further comprising a display device communicatively coupled to the processor, and wherein the method further comprises displaying the reconstructed three-dimensional representation on the display.

20 . A non-transitory computer readable medium having encoded thereon computer program code that is executable by a processor and that, when executed by the processor, causes the processor to perform a method comprising:

obtaining image data, head alignment data, and depth data of at least part of a three-dimensional head comprising a face of a conference participant, wherein at least the image data and head alignment data are determined based on two-dimensional or 2.5-dimensional images of the conference participant captured by a camera;

reconstructing a photo-realistic three-dimensional representation of the at least part of the head from the image data, the head alignment data, and the depth data, wherein reconstructing the representation of the at least part of the head comprises reconstructing an area missing from the image data acquired by the camera,

wherein reconstructing the representation comprises:

projecting the image data from a world space into an object space using the head alignment data and depth data; and

in the object space, completing at least part of the area missing from the image data using a computational model of a person, and

wherein completing at least part of the area missing from the image data comprises applying an artificial neural network to reconstruct a texture and a depth of the at least part of the head.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: VIRCÍKOVÁ, PHD, ING. MÁRIA; KIRCHMAYER, MATÚ¿; JAK¿A, PHD, ING. RUDOLF; MAGYAR, ING. GERGELY
To: MATSUKO S.R.O.
Reel/Frame 064001/0207 →
Continuity (3)
Continuation In Part 17538664 · Nov 30, 2021
Provisional Application 63120061 · Dec 1, 2020
Related Publication 20230334754A1 · Oct 19, 2023
References Cited (46)
US 5500671A · Andersson et al. · 1996 [cited by applicant]
US 6469710B1 · Shum et al. · 2002 [cited by applicant]
US 7302466B1 · Satapathy et al. · 2007 [cited by applicant]
US 9094660B2 · Alregib et al. · 2015 [cited by applicant]
US 9661272B1 · Daniel · 2017 [cited by applicant]
US 9760935B2 · Aarabi et al. · 2017 [cited by applicant]
US 10346893B1 · Duan et al. · 2019 [cited by applicant]
US 11087521B1 · Lombardi et al. · 2021 [cited by applicant]
US 20020135581A1 · Russell · 2002 [cited by examiner]
US 20040066386A1 · Leprevost · 2004 [cited by applicant]
US 20170178306A1 · Le Clerc et al. · 2017 [cited by applicant]
US 20180211444A1 · Shaviv et al. · 2018 [cited by applicant]
US 20180374242A1 · Li et al. · 2018 [cited by applicant]
US 20190213772A1 · Lombardi et al. · 2019 [cited by applicant]
US 20190294103A1 · Hauger et al. · 2019 [cited by applicant]
US 20200021627A1 · Brenes et al. · 2020 [cited by applicant]
US 20200098177A1 · Ni · 2020 [cited by examiner]
US 20200133618A1 · Kim · 2020 [cited by applicant]
US 20200234482A1 · Krokhalev · 2020 [cited by examiner]
US 20200257891A1 · Cole et al. · 2020 [cited by applicant]
US 20210150792A1 · Ulyanov et al. · 2021 [cited by applicant]
US 20210286424A1 · Ivanovitch · 2021 [cited by applicant]
US 20210342983A1 · Lin et al. · 2021 [cited by applicant]
US 20220172424A1 · Vircikova et al. · 2022 [cited by applicant]
KR 1020190112966A · 2019 [cited by applicant]
BNP Paribas Real Estate. “DARE—When Science Fiction Becomes Reality”, Mar. 2019. https://www.youtube.com/watch?v=13ktlkWppVs&feature=youtu.be. [cited by applicant]
Dotson, Kyt. “Spatial, Nreal, Qualcomm join up to deliver 5G-enabled AR collaboration killer app”, Feb. 2020. https://siliconangle.com/2020/02/20/spatial-nreal-qualcomm-join-deliver-5g-enabled-ar-collaboration-killer-ap… [cited by applicant]
Fink, Charlie. “The Trillion Dollar 3D Telepresense Gold Mine”, Nov. 2017. https://www.forbes.com/sites/charliefink/2017/11/20/the-trillion-dollar-3d-telepresence-gold-mine/?sh=256ed6612a72. [cited by applicant]
Goodfellow et al. “Generative Adversarial Nets”, Jun. 2014. [cited by applicant]
He et al. “Deep Residual Learning for Image Recognition”, Dec. 2015. [cited by applicant]
https://spatial.io. [cited by applicant]
https://www.doubleme.me/#holoportal. [cited by applicant]
Iizuka et al. “Globally and Locally Consistent Image Completion”, Jul. 2017. http://iizuka.cs.tsukuba.ac.jp/projects/completion/data/completion_sig2017.pdf. [cited by applicant]
Isola et al. “Image-to-Image Translation with Conditional Adversarial Networks”, Nov. 2018. [cited by applicant]
Jackson et al. “Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression”, Sep. 2017. [cited by applicant]
Krizhevsky et al. “ImageNet Classification with Deep Convolutional Neural Networks”, 2012. [cited by applicant]
Liu et al. “Image Inpainting for Irregular Holes Using Partial Convolutions”, Dec. 2018. [cited by applicant]
Mildenhall et al. “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”, Aug. 2020. [cited by applicant]
Pathak et al. “Context Encoders: Feature Learning by Inpainting”. [cited by applicant]
Ronneberger et al. “U-Net: Convolutional Networks for Biomedical Image Segmentation”, May 2015. [cited by applicant]
Wang et al. “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs”, Aug. 2018. [cited by applicant]
Wang et al. “VCNet: A Robust Approach to Blind Image Inpainting”, Mar. 2020. [cited by applicant]
Yu Deng et al.: “Accurate 3D Face Reconstruction with Weakly Supervised Learning: From Single Image to Image Set”, Arxiv.Org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 20, 20… [cited by applicant]
Sikander Gulbadan et al.: “A Novel Machine Vision-Based 3D Facial Action Unit Identification for Fatigue Detection”, IEEE Transactions On Intelligent Transportation Systems, IEEE, Piscataway, NJ, USA, vol. 22, No. 5, Fe… [cited by applicant]
Applicant: Matsuko s.r.o .; “Method, System, and Medium for Artificial Intelligence-Based Completion of a 3D Image During Electronic Communication”; U.S. Appl. No. 24/183,233; Extended European Search Report; dated Nov.… [cited by applicant]
Cheng Shiyang, et al.; “Faster, Better and More Detailed: 3D Face Reconstruction with Graph Convolutional Networks”; Jan. 26, 2021 (Jan. 26, 2021); Springer, pp. 188-205, XP047577984. [cited by applicant]