IP Library › Granted Patent US 12,664,742
Granted Patent B2
US 12,664,742 · App. 18/224,550 · Granted Jun 23, 2026

Method, system, and medium for enhancing a 3D image during electronic communication

Inventors: Mária Virčiková (Košice, SK); Matúš Kirchmayer (Košice, SK); Rudolf Jakša (Košice, SK)
Assignee: MATSUKO S.R.O.
G06T19/20G06T7/50G06V10/54G06V40/174G06V40/20H04N7/157G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,742
App. No.
18/224,550
Filed
Jul 20, 2023
Granted
Jun 23, 2026
Kind
B2
Examiner
TUNG, KEE M
Art Unit
2611
USPC
345/419
Abstract

A method and system for providing enhanced holographic communication is disclosed. Visual modeling data and verbal communication data of a conference participant in a current time frame are obtained as a current sample. The current sample is encoded into a first vector comprising a plurality of states for the current sample as an encoded result. At least one of the plurality of states is updated by inputting the encoded result to a generative artificial intelligent model as a prompt, and the updated plurality of states are obtained as a predicted result. The visual modeling data is then updated based on the predicted result for rendering a three-dimensional (3D) representation of the conference participant in the current time frame or in one or more future time frames.

Claims (54)

1 . A method comprising:

obtaining visual modeling data and verbal communication data of a conference participant in a current time frame as a current sample;

encoding the current sample into a first vector comprising a plurality of states for the current sample as an encoded result, wherein the plurality of states comprises one or more of texture, shape, eye behavior, mouth and facial expression, head movement, landmark position, speech, or script, and wherein each of the plurality of states is characterized by a plurality of elements;

updating at least one of the plurality of states by inputting the encoded result to a generative artificial intelligence model as a prompt;

selecting the visual modeling data of the conference participant in a plurality of past time frames as past samples;

encoding the past samples into a second plurality of vectors representing the plurality of states for the past samples;

updating at least one of the plurality of states by inputting the second plurality of vectors to the generative artificial intelligence model as additional prompts;

obtaining the updated plurality of states as a predicted result; and

updating the visual modeling data based on the predicted result for rendering a three-dimensional (3D) representation of the conference participant in the current time frame or in one or more future time frames.

2 . The method of claim 1 , wherein updating the visual modeling data comprises inputting the predicted result to a reconstruction or completion artificial neural network model to obtain the updated visual modeling data for rendering the 3D representation of the conference participant in the current time frame or in the one or more future time frames.

3 . The method of claim 1 , wherein updating the visual modeling data comprises decoding the predicted result into the updated visual modeling data for rendering the 3D representation of the conference participant in the current time frame or in the one or more future time frames.

4 . The method of claim 1 , wherein the predicted result has a same format as the encoded result.

5 . The method of claim 1 , wherein the visual modeling data comprises one or more of texture data, shape data, facial expression data, head movement data, or eyes behavior data, and the verbal communication data comprises one or more of audio data or text data.

6 . The method of claim 1 , wherein updating at least one of the plurality of states comprises updating one or more of texture, shape, eye behavior, or mouth and facial expression of the conference participant where a portion of the conference participant's head is obstructed.

7 . The method of claim 1 , wherein the plurality of past time frames are sampled at a constant frequency.

8 . The method of claim 1 , wherein the plurality of past time frames are irregularly sampled.

9 . The method of claim 1 , further comprising, prior to encoding the past samples into the second plurality of vectors:

obtaining at least one of the visual modeling data or the verbal communication data of additional one or more conference participants in the plurality of past time frames; and

adding the at least one of the visual modeling data or the verbal communication data of the additional one or more conference participants in the plurality of past time frames into the past samples.

10 . The method of claim 1 , further comprising, prior to encoding the current sample into a first vector:

obtaining at least one of the visual modeling data or the verbal communication data of additional one or more conference participants in the current time frame; and

adding the at least one of the visual modeling data or the verbal communication data of additional one or more conference participants in the current time frame into the current sample.

11 . The method of claim 1 , further comprising:

encoding the first vector into a sequence of token values; and

updating the encoded result with the sequence of token values, such that the encoded result is compatible with the generative artificial intelligence model being a transformer model.

12 . The method of claim 1 , further comprising:

obtaining a sequence of video frames captured by a camera for the conference participant, the sequence of video frames comprising two-dimensional (2D) image data, depth data, and head alignment data; and

reconstructing, for the current time frame of the sequence of video frames, the visual modeling data by projecting the 2D image data from a world space into an object space using the head alignment data and depth data.

13 . The method of claim 12 , further comprising completing at least one missing part of the reconstructed visual modeling data by a pre-trained artificial neural network.

14 . The method of claim 1 , wherein the 3D representation of the conference participant comprises one or more of a hologram, a parametric human representation, a stereoscopic 3D representation, a volumetric representation, a mesh-based representation, a point-cloud based representation, a radiance field representation, or a hybrid representation.

15 . The method of claim 1 , wherein the generative artificial intelligence model is selected from one of transformer models, Stable Diffusion models, Generative Adversarial Networks, or autoencoders.

16 . The method of claim 1 , wherein the plurality of states comprises eye behavior, and the eye behavior comprises eye-blink behavior, and wherein the generative artificial intelligence model is trained based on the current time frame or a plurality of past training time frames for the conference participant that are capable of being encoded into a plurality of training vectors comprising a plurality of training states, the plurality of training states comprising the eye-blink behavior and at least one of facial expression or speech pattern, such that the trained generative artificial intelligence model is capable of generating predicted future states of eye-blink behavior based on the facial expression or the speech pattern.

17 . The method of claim 16 , wherein the generative artificial intelligence is trained to allow a mapping between the eye-blink behavior and at least one of the facial expression or the speech pattern.

18 . A system comprising:

a network interface;

a processor communicatively coupled to the network interface; and

a non-transitory computer readable medium communicatively coupled to the processor and having stored thereon computer program code that is executable by the processor and that, when executed by the processor, causes the processor to perform a method comprising:

obtaining visual modeling data and verbal communication data of a conference participant in a current time frame as a current sample;

encoding the current sample into a first vector comprising a plurality of states for the current sample as an encoded result, wherein the plurality of states comprises one or more of texture, shape, facial expression, head movement, landmark position, speech, or script, and wherein each of the plurality of states is characterized by a plurality of elements;

updating at least one of the plurality of states by inputting the encoded result to a generative artificial intelligence model as a prompt;

selecting the visual modeling data of the conference participant in a plurality of past time frames as past samples;

encoding the past samples into a second plurality of vectors representing the plurality of states for the past samples;

updating at least one of the plurality of states by inputting the second plurality of vectors to the generative artificial intelligence model as additional prompts;

obtaining the updated plurality of states as a predicted result; and

updating the visual modeling data based on the predicted result for rendering a three-dimensional (3D) representation of the conference participant in the current time frame or in one or more future time frames.

19 . A non-transitory computer readable medium having encoded thereon computer program code that is executable by a processor and that, when executed by the processor, causes the processor to perform a method comprising:

obtaining visual modeling data and verbal communication data of a conference participant in a current time frame as a current sample;

encoding the current sample into a first vector comprising a plurality of states for the current sample as an encoded result, wherein the plurality of states comprises one or more of texture, shape, facial expression, head movement, landmark position, speech, or script, and wherein each of the plurality of states is characterized by a plurality of elements;

updating at least one of the plurality of states by inputting the encoded result to a generative artificial intelligence model as a prompt;

selecting the visual modeling data of the conference participant in a plurality of past time frames as past samples;

encoding the past samples into a second plurality of vectors representing the plurality of states for the past samples;

updating at least one of the plurality of states by inputting the second plurality of vectors to the generative artificial intelligence model as additional prompts;

obtaining the updated plurality of states as a predicted result; and

updating the visual modeling data based on the predicted result for rendering a three-dimensional (3D) representation of the conference participant in the current time frame or in one or more future time frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2023
From: VIRCÍKOVÁ, MÁRIA; KIRCHMAYER, MATÚ¿; JAK¿A, RUDOLF
To: MATSUKO S.R.O.
Reel/Frame 064335/0549 →
Continuity (1)
Related Publication 20250029346A1 · Jan 23, 2025
References Cited (30)
US 9904660B1 · Wang et al. · 2018 [cited by applicant]
US 20190012599A1 · el Kaliouby · 2019 [cited by examiner]
US 20210076002A1 · Peters · 2021 [cited by examiner]
US 20210334935A1 · Grigoriev et al. · 2021 [cited by applicant]
US 20220172424A1 · Vircíkováet al. · 2022 [cited by applicant]
US 20220207810A1 · Nemchinov · 2022 [cited by examiner]
D-ID; “The Digital People Platform”; https://www.d-id.com; retrieved Jul. 20, 2023. [cited by applicant]
Wikipedia the Free Encyclopedia; “Generative pre-trained transformer”; https://en.wikipedia.org/wiki/Generative_pre-trained_transformer; retrieved on Jul. 20, 2023; 11 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “ChatGPT”; https://en.wikipedia.org/wiki/ChatGPT; retrieved on Jul. 20, 2023; 37 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “Vision transformer”; https://en.wikipedia.org/wiki/Vision_transformer; retrieved Jul. 20, 2023; 5 pgs. [cited by applicant]
Arnab, Anurag, et al. “Vivit: A video vision transformer.” Proceedings of the IEEE/CVF international conference on computer vision. 2021; 11 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “Residual neural network”; https://en.wikipedia.org/wiki/Residual_neural_network; retrieved on Jul. 20, 2023; 7 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “ImageNet”; https://en.wikipedia.org/wiki/ImageNet; retrieved on Jul. 20, 2023; 5 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “Autoencoder”; https://en.wikipedia.org/wiki/Autoencoder; retrieved Jul. 20, 2023; 14 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “Principal component analysis”; https://en.wikipedia.org/wiki/Principal_component_analysis; retrieved Jul. 20, 2023; 34 pgs. [cited by applicant]
TensorFlow; “tf.keras.applications.resnet50.ResNet50”; https://www.tensorflow.org/api_docs/python/tf/keras/applications/resnet50/ResNet50; retrieved Jul. 20, 2023; 3 pgs. [cited by applicant]
Phillipi; “pix2pix”; https://github.com/phillipi/pix2pix; retrieved Jul. 20, 2023; 8 pgs. [cited by applicant]
Hong, Wenyi, et al. “Cogvideo: Large-scale pretraining for text-to-video generation via transformers.” arXiv preprint arXiv:2205.15868 (2022); 15 pgs. [cited by applicant]
Khachatryan, L. et al. “Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators”, Mar. 23, 2023; 25 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia' “Self-organizing map”; https://en.wikipedia.org/wiki/Self-organizing_map; retrieved Jul. 20, 2023; 11 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “k-means clustering”; https://en.wikipedia.org/wiki/K-means_clustering; retrieved Jul. 20, 2023; 17 pgs. [cited by applicant]
Wikipedia the Free Encyclopedia; “Feature hashing”; https://en.wikipedia.org/wiki/Feature_hashing; retrieved Jul. 20, 2023; 8 pgs. [cited by applicant]
Inventor: Maria Vircíková, et al.; “Method, System, and Medium for Artificial Intelligence-Based Completion of a 3D Image During Electronic Communication”; U.S. Appl. No. 18/338,093, filed Jun. 20, 2023. [cited by applicant]
Aaron S. Jackson, Adrian Bulat, Vasileios Argyriou, Georgios Tzimiropoulos: “Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression”, ARXIV:1703.07834V2, Sep. 8, 2017 (Sep. 8, 2017). [cited by applicant]
Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton: “ImageNet Classification with Deep Convolutional Neural Networks”, NIPS '12: Proceedings of the 25th International Conference on Neural Information Processing Systems… [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun: “Deep Residual Learning for Image Recognition”, ARXIV:1512.03385V1, Dec. 10, 2015 (Dec. 10, 2015). [cited by applicant]
Olaf Ronneberger, Philipp Fischer, Thomas Brox: “U-Net: Convolutional Networks for Biomedical Image Segmentation”, ARXIV:1505.04597V1, May 18, 2015 (May 18, 2015). [cited by applicant]
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, Alexei A. Efros: “Image-to-Image Translation with Conditional Adversarial Networks”, ARXIV:1611.07004V3, Nov. 26, 2018 (Nov. 26, 2018). [cited by applicant]
Mildenhall, B., et al.: “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”, ARXIV:2003.08934. [cited by applicant]
European Application No. EP24188051; European Search Report dated Dec. 2, 2024; 2 pgs. [cited by applicant]