IP Library › Granted Patent US 12,494,010
Granted Patent B2
US 12,494,010 · App. 18/339,341 · Granted Dec 9, 2025

Method, electronic device, and computer program product for video processing

Inventors: Zhisong Liu (Shenzhen, CN); Zijia Wang (WeiFang, CN); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06T13/80G06N3/08G06T11/60G06T13/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,494,010
App. No.
18/339,341
Granted
Dec 9, 2025
Kind
B2
Abstract

A method, an electronic device, and a computer program product for video processing are provided in embodiments of the present disclosure. The method generates an avatar image using a reference image and image data for a first frame in a video stream, and generates an avatar video using the avatar image and image data, audio data, and text data in the video stream. Through this solution, a user-defined avatar video adapted to a user of a real video and actions thereof can be generated more accurately and with high quality.

Claims (81)

1 . A method for video processing, comprising:

acquiring a video stream, the video stream comprising image data, audio data, and text data corresponding to video frames, and the video frames comprising a first frame;

generating a first avatar image using a reference image and image data for the first frame;

obtaining a video integration feature based on the first avatar image, the image data, the audio data, and the text data; and

generating an avatar video corresponding to the video stream based on the first avatar image and the video integration feature;

wherein obtaining the video integration feature based on the first avatar image, the image data, the audio data, and the text data comprises:

converting respective features corresponding to the first avatar image, the image data, the audio data, and the text data to respective vectors in a feature space; and

obtaining the video integration feature based on the respective vectors in the feature space converted from the respective features corresponding to the first avatar image, the image data, the audio data, and the text data.

2 . The method according to claim 1 , wherein obtaining the video integration feature based on the first avatar image, the image data, the audio data, and the text data comprises:

obtaining a first avatar image feature, an image difference feature, an audio feature, and a text feature, wherein the first avatar image feature corresponds to the first avatar image, the image difference feature corresponds to image data of adjacent frames in the video frames, the audio feature corresponds to the audio data, and the text feature corresponds to the text data; and

performing integration processing on the first avatar image feature, the image difference feature, the audio feature, and the text feature to obtain the video integration feature.

3 . The method according to claim 2 , wherein performing integration processing on the first avatar image feature, the image difference feature, the audio feature, and the text feature to obtain the video integration feature comprises:

converting the first avatar image feature, the image difference feature, the audio feature, and the text feature into a first vector, a second vector, a third vector, and a fourth vector in the feature space, respectively;

generating a feature integration vector based on the first vector, the second vector, the third vector, and the fourth vector;

generating a residual vector corresponding to the feature integration vector by using an attention mechanism; and

obtaining the video integration feature based on the feature integration vector and the residual vector.

4 . The method according to claim 1 , wherein the method is implemented by an avatar video generation model.

5 . The method according to claim 4 , further comprising:

obtaining a first loss function based on the avatar video, the audio data, and the text data; and

training the avatar video generation model by using the first loss function.

6 . The method according to claim 5 , wherein obtaining a first loss function based on the avatar video, the audio data, and the text data comprises:

obtaining a video-audio loss function based on the avatar video and the audio data;

obtaining a video-text loss function based on the avatar video and the text data;

obtaining an audio-text loss function based on the audio data and the text data; and

obtaining the first loss function based on the video-audio loss function, the video-text loss function, and the audio-text loss function.

7 . The method according to claim 6 , wherein the video frames further comprise a second frame, and the method further comprises:

obtaining a second loss function based on the first avatar image and a second avatar image for the second frame; and

training the avatar video generation model by using the second loss function.

8 . The method according to claim 7 , further comprising:

obtaining a third loss function based on the image data and the avatar video; and

training the avatar video generation model by using the third loss function.

9 . The method according to claim 8 , wherein the method further comprises:

obtaining a fourth loss function based on the first loss function, the second loss function, and the third loss function; and

training the avatar video generation model by using the fourth loss function.

10 . An electronic device, comprising:

at least one processor; and

at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the electronic device to perform operations comprising:

acquiring a video stream, the video stream comprising image data, audio data, and text data corresponding to video frames, and the video frames comprising a first frame;

generating a first avatar image using a reference image and image data for the first frame;

obtaining a video integration feature based on the first avatar image, the image data, the audio data, and the text data; and

generating an avatar video corresponding to the video stream based on the first avatar image and the video integration feature;

wherein obtaining the video integration feature based on the first avatar image, the image data, the audio data, and the text data comprises:

converting respective features corresponding to the first avatar image, the image data, the audio data, and the text data to respective vectors in a feature space; and

obtaining the video integration feature based on the respective vectors in the feature space converted from the respective features corresponding to the first avatar image, the image data, the audio data, and the text data.

11 . The electronic device according to claim 10 , wherein obtaining the video integration feature based on the first avatar image, the image data, the audio data, and the text data comprises:

obtaining a first avatar image feature, an image difference feature, an audio feature, and a text feature, wherein the first avatar image feature corresponds to the first avatar image, the image difference feature corresponds to image data of adjacent frames in the video frames, the audio feature corresponds to the audio data, and the text feature corresponds to the text data; and

performing integration processing on the first avatar image feature, the image difference feature, the audio feature, and the text feature to obtain the video integration feature.

12 . The electronic device according to claim 11 , wherein performing integration processing on the first avatar image feature, the image difference feature, the audio feature, and the text feature to obtain the video integration feature comprises:

converting the first avatar image feature, the image difference feature, the audio feature, and the text feature into a first vector, a second vector, a third vector, and a fourth vector in the feature space, respectively;

generating a feature integration vector based on the first vector, the second vector, the third vector, and the fourth vector;

generating a residual vector corresponding to the feature integration vector by using an attention mechanism; and

obtaining the video integration feature based on the feature integration vector and the residual vector.

13 . The electronic device according to claim 10 , wherein the operations are implemented by an avatar video generation model.

14 . The electronic device according to claim 13 , wherein the operations further comprise:

obtaining a first loss function based on the avatar video, the audio data, and the text data; and

training the avatar video generation model by using the first loss function.

15 . The electronic device according to claim 14 , wherein obtaining a first loss function based on the avatar video, the audio data, and the text data comprises:

obtaining a video-audio loss function based on the avatar video and the audio data;

obtaining a video-text loss function based on the avatar video and the text data;

obtaining an audio-text loss function based on the audio data and the text data; and

obtaining the first loss function based on the video-audio loss function, the video-text loss function, and the audio-text loss function.

16 . The electronic device according to claim 15 , wherein the video frames further comprise a second frame, and the operations further comprise:

obtaining a second loss function based on the first avatar image and a second avatar image for the second frame; and

training the avatar video generation model by using the second loss function.

17 . The electronic device according to claim 16 , wherein the operations further comprise:

obtaining a third loss function based on the image data and the avatar video; and

training the avatar video generation model by using the third loss function.

18 . The electronic device according to claim 17 , wherein the operations further comprise:

obtaining a fourth loss function based on the first loss function, the second loss function, and the third loss function; and

training the avatar video generation model by using the fourth loss function.

19 . A computer program product that is tangibly stored on a non-transitory computer-readable medium and comprises computer-executable instructions, wherein the computer-executable instructions, when executed by a device, cause the device to perform operations comprising:

acquiring a video stream, the video stream comprising image data, audio data, and text data corresponding to video frames, and the video frames comprising a first frame;

generating a first avatar image using a reference image and image data for the first frame;

obtaining a video integration feature based on the first avatar image, the image data, the audio data, and the text data; and

generating an avatar video corresponding to the video stream based on the first avatar image and the video integration feature;

wherein obtaining the video integration feature based on the first avatar image, the image data, the audio data, and the text data comprises:

converting respective features corresponding to the first avatar image, the image data, the audio data, and the text data to respective vectors in a feature space; and

obtaining the video integration feature based on the respective vectors in the feature space converted from the respective features corresponding to the first avatar image, the image data, the audio data, and the text data.

20 . The computer program product according to claim 19 , wherein obtaining the video integration feature based on the first avatar image, the image data, the audio data, and the text data comprises:

obtaining a first avatar image feature, an image difference feature, an audio feature, and a text feature, wherein the first avatar image feature corresponds to the first avatar image, the image difference feature corresponds to image data of adjacent frames in the video frames, the audio feature corresponds to the audio data, and the text feature corresponds to the text data; and

performing integration processing on the first avatar image feature, the image difference feature, the audio feature, and the text feature to obtain the video integration feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: LIU, ZHISONG; WANG, ZIJIA; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 064027/0177 →
Priority Claims (1)
CN 202211020897.5 · Aug 24, 2022 · national
Continuity (1)
Related Publication 20240070956A1 · Feb 29, 2024
References Cited (34)
Wang et al.; “Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion;” Jul. 20, 2021; arXiv:2107.09293v1; pp. 1-8 (Year: 2021). [cited by examiner]
L. Yu et al.; “Multimodal Learning for Temporally Coherent Talking Face Generation With Articulator Synergy,” Jun. 25, 2021; in IEEE Transactions on Multimedia, vol. 24, pp. 2950-2962 (Year: 2021). [cited by examiner]
O. Fried et al., “Text-Based Editing of Talking-head Video,” ACM Transactions on Graphics, vol. 38, No. 4, Jul. 2019, pp. 68:1-68:14. [cited by applicant]
L. Xie et al., “Realistic Mouth-Synching for Speech-Driven Talking Face Using Articulatory Modelling,” IEEE Transactions on Multimedia, vol. 9, No. 3, Apr. 2007, pp. 500-510. [cited by applicant]
E. Cosatto et al., “Lifelike Talking Faces for Interactive Services,” Proceedings of the IEEE, vol. 91, No. 9, Sep. 2003, pp. 1406-1429. [cited by applicant]
A. Wang et al., “Assembling an Expressive Facial Animation System,” Proceedings of the 2007 ACM SIGGRAPH Symposium on Video Games, Aug. 2007, pp. 21-26. [cited by applicant]
A. Shysheya et al., “Textured Neural Avatars,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 2387-2397. [cited by applicant]
S. Ma et al., “Pixel Codec Avatars,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:2104.04638v1, Apr. 9, 2021, 11 pages. [cited by applicant]
M. C. Doukas et al., “HeadGAN: One-Shot Neural Head Synthesis and Editing,” IEEE/CVF International Conference on Computer Vision, Oct. 2021, pp. 14398-14407. [cited by applicant]
Y. Fan et al., “FaceFormer: Speech-Driven 3D Facial Animation with Transformers,” arXiv:2112.05329v4, Mar. 17, 2022, 13 pages. [cited by applicant]
T.-C. Wang et al., “One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing, ” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 10039-10049. [cited by applicant]
E. Zakharov et al., “Few-Shot Adversarial Learning of Realistic Neural Talking Head Models,” IEEE/CVF International Conference on Computer Vision, arXiv:1905.08233v2, Sep. 25, 2019, 21 pages. [cited by applicant]
K. R. Prajwal et al., “A Lip Sync Expert Is All You Need for Speech to Lip Generation in the Wild,” Proceedings of the 28th ACM International Conference on Multimedia, arXiv:2008.10010v1, Aug. 23, 2020, 10 pages. [cited by applicant]
A. Lahiri et al., “LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from Video using Pose and Lighting Normalization,” arXiv:2106.04185v1, Jun. 8, 2021, 16 pages. [cited by applicant]
J. S. Chung et al., “Out of Time: Automated Lip Sync in the Wild,” Workshop on Multi-view Lip-reading, Asian Conference on Computer Vision, Nov. 2016, 14 pages. [cited by applicant]
X. Huang et al., “Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization,” arXiv:1703.06868v2, Jul. 30, 2017, 11 pages. [cited by applicant]
X. Li et al., “Learning Linear Transformations for Fast Arbitrary Style Transfer,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:1808.04537v1, Aug. 14, 2018, 12 pages. [cited by applicant]
Y. Li et al., “Diversified Texture Synthesis with Feed-Forward Networks,” arXiv:1703.01664v1, Mar. 5, 2017, 11 pages. [cited by applicant]
F. Luan et al., “Deep Photo Style Transfer,” arXiv:1703.07511v3, Apr. 11, 2017, 9 pages. [cited by applicant]
Y. Li et al., “Universal Style Transfer via Feature Transforms,” 31st Conference on Neural Information Processing Systems, Dec. 2017, 11 pages. [cited by applicant]
W. Ma et al., “Block Shuffle: A Method for High-Resolution Fast Style Transfer with Limited Memory,” IEEE Access, arXiv:2008.03706v1, Aug. 9, 2020, 12 pages. [cited by applicant]
D. Y. Park et al., “Arbitrary Style Transfer With Style-Attentional Networks,” arXiv:1812.02342v5, May 23, 2019, 9 pages. [cited by applicant]
T. Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation,” The Sixth International Conference on Learning Representations, arXiv:1710.10196v3, Feb. 26, 2018, 26 pages. [cited by applicant]
T. Karras et al., “Analyzing and Improving the Image Quality of StyleGAN,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 8110-8119. [cited by applicant]
S. Yang et al., “Pastiche Master: Exemplar-Based High-Resolution Portrait Style Transfer,” Conference on Computer Vision and Pattern Recognition, Mar. 24, 2022, 16 pages. [cited by applicant]
T. Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 4401-4410. [cited by applicant]
A. Baevski et al., “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” Proceedings of the 34th International Conference on Neural Information Processing Systems, arXiv:2006.11477v3, Oct. 2… [cited by applicant]
F. Reda et al., “Pytorch Implementation of FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks,” https://github.com/NVIDIA/flownet2-pytorch, Accessed Jul. 21, 2022, 4 pages. [cited by applicant]
M. Loper et al., “SMPL: A Skinned Multi-Person Linear Model,” ACM Transactions on Graphics, vol. 34, No. 6, Nov. 2015, 16 pages. [cited by applicant]
E. Corona et al., “SMPLicit: Topology-aware Generative Model for Clothed People,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, p. 11875-11885. [cited by applicant]
T. Li et al., “Learning a Model of Facial Shape and Expression from 4D Scans,” ACM Transactions on Graphics, vol. 36, No. 6, Nov. 2017, 17 pages. [cited by applicant]
L. Wang et al., “HMM Trajectory-guided Sample Selection for Photo-realistic Talking Head,” Multimedia Tools and Applications, vol. 74, No. 22, Nov. 2015, pp. 9849-9869. [cited by applicant]
K. Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition,” International Conference on Learning Representations, arXiv:1409.1556v6, Apr. 10, 2015, 14 pages. [cited by applicant]
U.S. Appl. No. 17/993,025 filed in the name of Zhisong Liu et al. on Nov. 23, 2022, and entitled “Method, Device, and Computer Program Product for Processing Video.”. [cited by applicant]