IP Library Granted Patent US 12,482,163
Granted Patent B2
US 12,482,163 · App. 17/993,025 · Granted Nov 25, 2025

Method, device, and computer program product for processing video

Inventors: Zhisong Liu (Shenzhen, CN); Zijia Wang (WeiFang, CN); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06T13/40G06N3/0464G06N3/08G06T13/205G06T15/04G06T17/00G10L25/57G06T2200/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,163
App. No.
17/993,025
Granted
Nov 25, 2025
Kind
B2
Abstract

Methods, devices and computer program products for processing video are disclosed herein. A method includes: generating, based on a reference image and a first frame of a video comprising an object, a two-dimensional avatar image of the object; and generating a base three-dimensional avatar of the object by performing a three-dimensional transformation on the two-dimensional avatar image and the object in the first frame. The method further includes: generating a three-dimensional avatar video corresponding to the video based on the base three-dimensional avatar and features of the video, the features comprising differences of the object between adjacent frames of the video. This solution enables the generation of a customized three-dimensional avatar video for an object in a video, where the avatar can move in synchronization with the object and retain the unique features of the object, and can provide a more detailed and vivid representation than a two-dimensional avatar.

Claims (67)

1 . A method for processing video, comprising:

generating, based on a reference image and a first frame of a video comprising an object, a two-dimensional avatar image of the object;

generating a base three-dimensional avatar of the object by performing a three-dimensional transformation on the two-dimensional avatar image and the object in the first frame; and

generating a three-dimensional avatar video corresponding to the video based on the base three-dimensional avatar and features of the video, the features comprising image differences of the object between adjacent frames of the video;

wherein the method is implemented using a generation model for three-dimensional avatar videos, and the generation model is trained based on a loss function that includes a plurality of cross-modality loss components including an image-audio loss component, an image-text loss component, and an audio-text loss component; and

wherein the loss function is determined at least in part by:

determining the image-audio loss component as an image-audio contrastive loss function based on image data and audio data;

determining the image-text loss component as an image-text contrastive loss function based on the image data and text data;

determining the audio-text loss component as an audio-text contrastive loss function based on the audio data and the text data; and

determining the loss function based on the image-audio contrastive loss function, the image-text contrastive loss function, and the audio-text contrastive loss function.

2 . The method according to claim 1 , wherein generating the base three-dimensional avatar comprises:

generating a three-dimensional projection representation of the object based on shape, posture, and expression of the object in the first frame and the two-dimensional avatar image.

3 . The method according to claim 2 , wherein generating a three-dimensional projection representation of the object comprises:

generating the three-dimensional projection representation of the object based on the shape, posture, expression, and texture details of the object in the first frame and the two-dimensional avatar image.

4 . The method according to claim 3 , wherein generating the base three-dimensional avatar further comprises:

generating the base three-dimensional avatar based on a camera position at which the first frame was captured, color and lighting information of the two-dimensional avatar image, and the three-dimensional projection representation.

5 . The method according to claim 1 , wherein the features of the video further comprise voice features and text features, and generating the three-dimensional avatar video comprises:

acquiring fused features of the video based on features of the image differences, the voice features of the video, and the text features of the video; and

generating the three-dimensional avatar video based on the base three-dimensional avatar and the fused features.

6 . The method according to claim 1 , wherein the method further comprises:

generating, based on a first frame of a source video comprising a three-dimensional avatar, audio data of the source video, and text data of the source video, a predictive video comprising the three-dimensional avatar; and

determining the loss function for training the generation model based on image data in the source video and the predictive video, the audio data, and the text data.

7 . The method according to claim 6 , further comprising training the generation model by using the loss function.

8 . The method according to claim 6 , further comprising training the generation model by using one or more of the following:

a motion loss function for determining a motion loss of a three-dimensional avatar video generated by the generation model; and

a style loss function for determining a style loss of the three-dimensional avatar video generated by the generation model.

9 . The method according to claim 1 , wherein determining the loss function further comprises:

determining a landmark loss function based on the image data; and

determining the loss function based on the image-audio contrastive loss function, the image-text contrastive loss function, the audio-text contrastive loss function, and the landmark loss function.

10 . A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions, wherein the machine-executable instructions, when executed by a machine, cause the machine to perform actions comprising:

generating, based on a reference image and a first frame of a video comprising an object, a two-dimensional avatar image of the object;

generating a base three-dimensional avatar of the object by performing a three-dimensional transformation on the two-dimensional avatar image and the object in the first frame; and

generating a three-dimensional avatar video corresponding to the video based on the base three-dimensional avatar and features of the video, the features comprising image differences of the object between adjacent frames of the video;

wherein the actions are implemented using a generation model for three-dimensional avatar videos, and the generation model is trained based on a loss function that includes a plurality of cross-modality loss components including an image-audio loss component, an image-text loss component, and an audio-text loss component; and

wherein the loss function is determined at least in part by:

determining the image-audio loss component as an image-audio contrastive loss function based on image data and audio data;

determining the image-text loss component as an image-text contrastive loss function based on the image data and text data;

determining the audio-text loss component as an audio-text contrastive loss function based on the audio data and the text data; and

determining the loss function based on the image-audio contrastive loss function, the image-text contrastive loss function, and the audio-text contrastive loss function.

11 . An electronic device, comprising:

at least one processor; and

memory coupled to the at least one processor, wherein the memory has instructions stored therein which, when executed by the at least one processor, cause the electronic device to perform actions comprising:

generating, based on a reference image and a first frame of a video comprising an object, a two-dimensional avatar image of the object;

generating a base three-dimensional avatar of the object by performing a three-dimensional transformation on the two-dimensional avatar image and the object in the first frame; and

generating a three-dimensional avatar video corresponding to the video based on the base three-dimensional avatar and features of the video, the features comprising image differences of the object between adjacent frames of the video;

wherein the actions are implemented using a generation model for three-dimensional avatar videos, and the generation model is trained based on a loss function that includes a plurality of cross-modality loss components including an image-audio loss component, an image-text loss component, and an audio-text loss component; and

wherein the loss function is determined at least in part by:

determining the image-audio loss component as an image-audio contrastive loss function based on image data and audio data;

determining the image-text loss component as an image-text contrastive loss function based on the image data and text data;

determining the audio-text loss component as an audio-text contrastive loss function based on the audio data and the text data; and

determining the loss function based on the image-audio contrastive loss function, the image-text contrastive loss function, and the audio-text contrastive loss function.

12 . The electronic device according to claim 11 , wherein generating the base three-dimensional avatar comprises:

generating a three-dimensional projection representation of the object based on shape, posture, and expression of the object in the first frame and the two-dimensional avatar image.

13 . The electronic device according to claim 12 , wherein generating a three-dimensional projection representation of the object comprises:

generating the three-dimensional projection representation of the object based on the shape, posture, expression, and texture details of the object in the first frame and the two-dimensional avatar image.

14 . The electronic device according to claim 13 , wherein generating the base three-dimensional avatar further comprises:

generating the base three-dimensional avatar based on a camera position at which the first frame was captured, color and lighting information of the two-dimensional avatar image, and the three-dimensional projection representation.

15 . The electronic device according to claim 11 , wherein the features of the video further comprise voice features and text features, and generating the three-dimensional avatar video comprises:

acquiring fused features of the video based on features of the image differences, the voice features of the video, and the text features of the video; and

generating the three-dimensional avatar video based on the base three-dimensional avatar and the fused features.

16 . The electronic device according to claim 11 , wherein the actions further comprise:

generating, based on a first frame of a source video comprising a three-dimensional avatar, audio data of the source video, and text data of the source video, a predictive video comprising the three-dimensional avatar; and

determining the loss function for training the generation model based on image data in the source video and the predictive video, the audio data, and the text data.

17 . The electronic device according to claim 16 , wherein the actions further comprise training the generation model by using the loss function.

18 . The electronic device according to claim 11 , wherein determining the loss function further comprises:

determining a landmark loss function based on the image data; and

determining the loss function based on the image-audio contrastive loss function, the image-text contrastive loss function, the audio-text contrastive loss function, and the landmark loss function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2022
From: LIU, ZHISONG; WANG, ZIJIA; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 061861/0486 →
Priority Claims (1)
CN 202211296657.8 · Oct 21, 2022 · national
Continuity (1)
Related Publication 20240185494A1 · Jun 6, 2024
References Cited (40)
US 12205577B1 · Kim · 2025 [cited by examiner]
US 20150213604A1 · Li · 2015 [cited by examiner]
US 20180173942A1 · Kim · 2018 [cited by examiner]
US 20190035149A1 · Chen · 2019 [cited by examiner]
US 20200035008A1 · Kwon · 2020 [cited by examiner]
US 20210027511A1 · Shang · 2021 [cited by examiner]
Zheng et al., “Adversarial-Metric Learning for Audio-Visual Cross-Modal Matching”, Jan. 12, 2021, IEEE Transactions on Multimedia, vol. 24, pp. 338-351 (Year: 2021). [cited by examiner]
Yu et al., “CoCa: Contrastive Captioners are Image-Text Foundation Models”, Aug. 2022, Transactions on Machine Learning Research, pp. 1-20 (Year: 2022). [cited by examiner]
Chen et al., “Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning”, Sep. 18-22, 2022, Proc. Interspeech 2022, pp. 2773-2777 (Year: 2022). [cited by examiner]
O. Fried et al., “Text-Based Editing of Talking-head Video,” ACM Transactions on Graphics, vol. 38, No. 4, Jul. 2019, pp. 68:1-68:14. [cited by applicant]
L. Xie et al., “Realistic Mouth-Synching for Speech-Driven Talking Face Using Articulatory Modelling,” IEEE Transactions on Multimedia, vol. 9, No. 3, Apr. 2007, pp. 500-510. [cited by applicant]
E. Cosatto et al., “Lifelike Talking Faces for Interactive Services,” Proceedings of the IEEE, vol. 91, No. 9, Sep. 2003, pp. 1406-1429. [cited by applicant]
A. Wang et al., “Assembling an Expressive Facial Animation System,” Proceedings of the 2007 ACM SIGGRAPH Symposium on Video Games, Aug. 2007, pp. 21-26. [cited by applicant]
A. Shysheya et al., “Textured Neural Avatars,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 2387-2397. [cited by applicant]
S. Ma et al., “Pixel Codec Avatars,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:2104.04638v1, Apr. 9, 2021, 11 pages. [cited by applicant]
M. C. Doukas et al., “HeadGAN: One-Shot Neural Head Synthesis and Editing,” IEEE/CVF International Conference on Computer Vision, Oct. 2021, pp. 14398-14407. [cited by applicant]
Y. Fan et al., “FaceFormer: Speech-Driven 3D Facial Animation with Transformers,” arXiv:2112.05329v4, Mar. 17, 2022, 13 pages. [cited by applicant]
T.-C. Wang et al., “One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 10039-10049. [cited by applicant]
E. Zakharov et al., “Few-Shot Adversarial Learning of Realistic Neural Talking Head Models,” IEEE/CVF International Conference on Computer Vision, arXiv:1905.08233v2, Sep. 25, 2019, 21 pages. [cited by applicant]
K. R. Prajwal et al., “A Lip Sync Expert Is All You Need for Speech to Lip Generation in the Wild,” Proceedings of the 28th ACM International Conference on Multimedia, arXiv:2008.10010v1, Aug. 23, 2020, 10 pages. [cited by applicant]
A. Lahiri et al., “LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from Video using Pose and Lighting Normalization,” arXiv:2106.04185v1, Jun. 8, 2021, 16 pages. [cited by applicant]
J. S. Chung et al., “Out of Time: Automated Lip Sync in the Wild,” Workshop on Multi-view Lip-reading, Asian Conference on Computer Vision, Nov. 2016, 14 pages. [cited by applicant]
X. Huang et al., “Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization,” arXiv:1703.06868v2, Jul. 30, 2017, 11 pages. [cited by applicant]
X. Li et al., “Learning Linear Transformations for Fast Arbitrary Style Transfer,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:1808.04537v1, Aug. 14, 2018, 12 pages. [cited by applicant]
Y. Li et al., “Diversified Texture Synthesis with Feed-Forward Networks,” arXiv:1703.01664v1, Mar. 5, 2017, 11 pages. [cited by applicant]
F. Luan et al., “Deep Photo Style Transfer,” arXiv:1703.07511v3, Apr. 11, 2017, 9 pages. [cited by applicant]
Y. Li et al., “Universal Style Transfer via Feature Transforms,” 31st Conference on Neural Information Processing Systems, Dec. 2017, 11 pages. [cited by applicant]
W. Ma et al., “Block Shuffle: A Method for High-Resolution Fast Style Transfer with Limited Memory,” IEEE Access, arXiv:2008.03706v1, Aug. 9, 2020, 12 pages. [cited by applicant]
D. Y. Park et al., “Arbitrary Style Transfer With Style-Attentional Networks,” arXiv:1812.02342v5, May 23, 2019, 9 pages. [cited by applicant]
T. Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation,” The Sixth International Conference on Learning Representations, arXiv:1710.10196v3, Feb. 26, 2018, 26 pages. [cited by applicant]
T. Karras et al., “Analyzing and Improving the Image Quality of StyleGAN,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 8110-8119. [cited by applicant]
S. Yang et al., “Pastiche Master: Exemplar-Based High-Resolution Portrait Style Transfer,” Conference on Computer Vision and Pattern Recognition, Mar. 24, 2022, 16 pages. [cited by applicant]
T. Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 4401-4410. [cited by applicant]
A. Baevski et al., “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” Proceedings of the 34th International Conference on Neural Information Processing Systems, arXiv:2006.11477v3, Oct. 2… [cited by applicant]
F. Reda et al., “Pytorch Implementation of FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks,” https://github.com/NVIDIA/flownet2-pytorch, Accessed Jul. 21, 2022, 4 pages. [cited by applicant]
M. Loper et al., “SMPL: A Skinned Multi-Person Linear Model,” ACM Transactions on Graphics, vol. 34, No. 6, Nov. 2015, 16 pages. [cited by applicant]
E. Corona et al., “SMPLicit: Topology-aware Generative Model for Clothed People,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 11875-11885. [cited by applicant]
T. Li et al., “Learning a Model of Facial Shape and Expression from 4D Scans,” ACM Transactions on Graphics, vol. 36, No. 6, Nov. 2017, 17 pages. [cited by applicant]
L. Wang et al., “HMM Trajectory-guided Sample Selection for Photo-realistic Talking Head,” Multimedia Tools and Applications, vol. 74, No. 22, Nov. 2015, pp. 9849-9869. [cited by applicant]
K. Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition,” International Conference on Learning Representations, arXiv:1409.1556v6, Apr. 10, 2015, 14 pages. [cited by applicant]