IP Library Granted Patent US 12,477,120
Granted Patent B2
US 12,477,120 · App. 18/484,134 · Granted Nov 18, 2025

Keypoints based video compression

Inventors: Zhao Wang (Beijing, CN); Bolin Chen (Hong Kong, HK); Yan Ye (San Diego, CA); Shiqi Wang (Hong Kong, HK)
Assignee: Alibaba Damo (Hangzhou) Technology Co., Ltd.
H04N19/154G06T3/18G06T3/20G06T3/60H04N19/169
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,477,120
App. No.
18/484,134
Granted
Nov 18, 2025
Kind
B2
Abstract

Methods and apparatuses are provided for compressing video data based on keypoint features. An exemplary video compression method includes: receiving a video sequence; encoding one or more pictures of the video sequence; and generating a bitstream; wherein the encoding includes: representing a first picture by a first set of keypoints and a second set of keypoints, the second set comprising less keypoints than the first set; and compressing the video sequence based on the first set and second set of keypoints.

Claims (45)

1 . A video compression method, comprising:

receiving a video sequence;

encoding one or more pictures of the video sequence, wherein the encoding comprises:

representing a first picture by a first set of keypoints and a second set of keypoints, the second set comprising less keypoints than the first set, the first set of keypoints comprising canonical keypoints extracted from a second picture; and

compressing the video sequence based on the first set and second set of keypoints; and

generating a bitstream associated with the compressed video sequence.

2 . The video compression method of claim 1 , wherein the second set of keypoints is a subset of the first set of keypoints, and representing the first picture by the first set of keypoints and the second set of keypoints comprises:

transforming the first set of keypoints;

determining a deformation to the second set of keypoints; and

representing the first picture based on the transformed first set of keypoints and the deformation to the second set of keypoints.

3 . The video compression method of claim 2 , wherein transforming the first set of keypoints comprises performing at least one of a rotation or a translation of the first set of keypoints.

4 . The video compression method of claim 1 , wherein the first and second pictures comprises a representation of a face, and the canonical keypoints are independent of a pose and an expression of the face.

5 . The video compression method of claim 1 , wherein the first picture comprises a representation of a face, and representing the first picture by the first set of keypoints and the second set of keypoints comprises:

representing a pose of the face based on the first set of keypoints; and

representing an expression of the face based on the second set of keypoints.

6 . The video compression method of claim 1 , wherein the encoding further comprises:

generating, based on the first and second sets of keypoints, a reconstructed picture associated with the first picture; and

determining a quality loss of the reconstructed picture, based on a weighted combination of at least two models, wherein the at least two models are trained on different training datasets.

7 . The video compression method of claim 6 , wherein at least one of the training datasets comprises face pictures.

8 . A video decoding method, comprising:

receiving a bitstream; and

decoding, using coded information of the bitstream, one or more pictures, wherein the decoding comprises:

decoding warping data associated with a first set of keypoints and a second set of keypoints, the second set comprising less keypoints than the first set, the first set of keypoints comprising canonical keypoints extracted from a second picture; and

generating a first picture based on the first set of keypoints, the second set of keypoints, and the warping data.

9 . The video decoding method of claim 8 , wherein the second set of keypoints is a subset of the first set of keypoints.

10 . The video decoding method of claim 8 , wherein generating the first picture based on the first set of keypoints, the second set of keypoints, and the warping data comprises:

warping one or more features of a second picture, based on the warping data; and

generating the first picture based on the warped one or more features.

11 . The video decoding method of claim 10 , wherein the first and second pictures comprises a representation of a face, and the canonical keypoints are independent of a pose and an expression of the face.

12 . The video decoding method of claim 8 , wherein the first picture comprises a representation of a face, and generating the first picture based on the first set of keypoints, the second set of keypoints, and the warping data comprises:

generating a pose of the face based on the first set of keypoints; and

generating an expression of the face based on the second set of keypoints.

13 . The video decoding method of claim 8 , further comprising:

generating, based on the first and second sets of keypoints, a reconstructed picture associated with the first picture; and

determining a quality loss of the reconstructed picture, based on a weighted combination of at least two models, wherein the at least two models are trained on different training datasets.

14 . The video decoding method of claim 13 , wherein at least one of the training datasets comprises face pictures.

15 . A method of storing a bitstream of a video, the method comprising:

generating a first set of keypoints and a second set of keypoints representing a first picture, the second set comprising less keypoints than the first set, the first set of keypoints comprising canonical keypoints extracted from a second picture;

generating warping data associated with the first and second sets of keypoints, wherein the first set of keypoints, the second set of keypoints, and the warping data are used for generating the first picture;

generating a bitstream comprising the warping data; and

storing the bitstream in a non-transitory computer readable storage medium.

16 . The method of claim 15 , wherein the second set of keypoints is a subset of the first set of keypoints, and the first picture is represented by a combination of:

a transformation of the first set of keypoints; and

a deformation to the second set of keypoints.

17 . The method of claim 16 , wherein the transformation of the first set of keypoints comprises at least one of a rotation or a translation of the first set of keypoints.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA DAMO (HANGZHOU) TECHNOLOGY CO., LTD.
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 075529/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2024
From: WANG, ZHAO; CHEN, BOLIN; YE, YAN; WANG, SHIQI
To: ALIBABA DAMO (HANGZHOU) TECHNOLOGY CO., LTD.
Reel/Frame 066199/0480 →
Continuity (2)
Provisional Application 63379783 · Oct 17, 2022
Related Publication 20240129487A1 · Apr 18, 2024
References Cited (52)
US 5633956A · Burl · 1997 [cited by examiner]
US 5909224A · Fung · 1999 [cited by examiner]
US 6493043B1 · Bollmann · 2002 [cited by examiner]
US 6642968B1 · Ford · 2003 [cited by examiner]
US 8559513B2 · Demos · 2013 [cited by examiner]
US 11556580B1 · Ramnath · 2023 [cited by examiner]
US 11729417B2 · Bordes · 2023 [cited by examiner]
US 11792419B2 · Park · 2023 [cited by examiner]
US 11895314B2 · Paluri · 2024 [cited by examiner]
US 20150271514A1 · Yoshikawa et al. · 2015 [cited by applicant]
US 20220038720A1 · Hashimoto · 2022 [cited by examiner]
US 20220103860A1 · Demyanov et al. · 2022 [cited by applicant]
US 20220312028A1 · Li · 2022 [cited by examiner]
CN 114170334A · 2022 [cited by applicant]
CN 114556941A · 2022 [cited by applicant]
EP 3110153A1 · 2016 [cited by applicant]
Ahmed et al., “Discrete cosine transform,” IEEE transactions on Computers, vol. 100, No. 1, pp. 90-93, 1974. [cited by applicant]
Balle et al., End-to-End Optimized Image Compression, ICLR, 2017, 27 pages. [cited by applicant]
Balle et al., “Density modeling of images using a generalized normalization transformation,” ICLR, 14 pages, 2016. [cited by applicant]
Blanz et al., “A morphable model for the synthesis of 3d faces,” in Proceedings of the 26th annual conference on Computer graphics and interactive techniques, 1999, pp. 187-194. [cited by applicant]
Bross et al., “Overview of the Versatile Video Coding (VVC) Standard and Its Applications,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 10, pp. 3736-3764, 2023. [cited by applicant]
Chen et al., “Beyond key-point coding: Temporal evolution inference with compact feature representation for talking face video compression,” in Proceedings of the IEEE Data Compression Conference, 10 pages, 2022. [cited by applicant]
Ding et al., “Image Quality Assessment: Unifying Structure and Texture Similarity,” IEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 5, May 2022. [cited by applicant]
Farneback, “Two-frame motion estimation based on polynomial expansion,” Image Analysis, pp. 363-370, 2003. [cited by applicant]
Feng et al., “A generative compression framework for low bandwidth video conference,” in ICME Workshop, 6 pages 2021. [cited by applicant]
Goodfellow et al., “Generative adversarial nets,” Advances in neural information processing systems, vol. 27, 9 pages, 2014. [cited by applicant]
Hu et al., “FVC: A New Framework towards Deep Video Compression in Feature Space,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1502-1511. [cited by applicant]
Jia et al., “Content-Aware Convolutional Neural Network for In-Loop Filtering in High Efficiency Video Coding,” IEEE Transactions on Image Processing, vol. 28, No. 7, pp. 3343-3356 (2019). [cited by applicant]
Kingma et al., “Auto-encoding variational bayes,” in Proceedings of the 2nd International Conference on Learning Representations (ICLR), 2014, p. 14. [cited by applicant]
Koufakis et al., “Very low bit rate face video compression using linear combination of 2d face views and principal components analysis,” Image and Vision computing, vol. 17, No. 14, pp. 1031-1051, 1999. [cited by applicant]
Li et al., “Efficient, Multiple-Line-Based Intra Prediction for HEVC,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, No. 4, pp. 947-957, 2016. [cited by applicant]
Lin et al., “M-LVC: multiple frames prediction for learned video compression,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 3546-3554. [cited by applicant]
Lopez et al., “Head pose computation for very low bitrate video coding,” in International Conference on Computer Analysis of Images and Patterns. Springer, 1995, pp. 440-447. [cited by applicant]
Lu et al., “Content adaptive and error propagation aware deep video compression,” European Conference on Computer Vision. Springer, 2020, pp. 456-472. [cited by applicant]
Lu et al. “DVC: An end-to-end deep video compression framework,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 11006-11015, 2019. [cited by applicant]
Ma et al. “Convolutional neural network-based arithmetic coding for hevc intra-predicted residues,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 7, pp. 1901-1916, 2019. [cited by applicant]
Oquab et al., “Low bandwidth video-chat compression using deep generative models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 2388-2397. [cited by applicant]
Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention, 8 pages, 2015. [cited by applicant]
Siarohin et al., “First Order Motion Model for Image Animation,” Advances in Neural Information Processing Systems, vol. 32, pp. 7137-7147, 2019. [cited by applicant]
Simonyan et al., “Very deep convolutional networks for large scale image recognition,” arXiv preprint, arXiv: 1409, 1556 (2015). [cited by applicant]
Sullivan et al., “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, pp. 1649-1668 (2012). [cited by applicant]
Torres et al., “A proposal for high compression of faces in video sequences using adaptive eigenspaces,” in International Conference on image Processing. IEEE, 2002, vol. 1. [cited by applicant]
Wang et al., “One-shot free-view neural talking-head synthesis for video conferencing,” in Proceedings of the IEEE/CVFConference on Computer Vision and Pattern Recognition, 2021, pp. 10039-10049. [cited by applicant]
Wang et al., “Few-shot video-to-video synthesis.,” in NeurIPS, 2019, 14 pages. [cited by applicant]
Wiegand et al., “Overview of the H.264/AVC Video Coding Standard,” IEEE Transactions on Circuitds and Systems for Video Technology, vol. 13, No. 7, Jul. 2003, pp. 560-576. [cited by applicant]
Wiles et al., “X2face: A network for controlling face generation using images, audio, and pose codes,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 670-686. [cited by applicant]
Yang et al., “Learning for video compression with hierarchical quality and recurrent enhancement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 6628-6637. [cited by applicant]
Yang et al., “Learning for video compression with recurrent auto-encoder and recurrent probability model,” IEEE Journal of Selected Topics in Signal Processing, vol. 15, No. 2, pp. 388-401, 2020. [cited by applicant]
Zakharov et al., “Few-shot adversarial learning of realistic neural talking head models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019,pp. 9459-9468. [cited by applicant]
Zhao et al., Enhanced Bi-Prediction with convolutional neural network for high-efficiency video coding, IEEE Transactions on Circuits and Systems for Video Technology, vol. 29, No. 11, pp. 3291-3301, 2018. [cited by applicant]
Zhu et al., “Generative adversarial network-based intra prediction for video coding,” IEEE transactions on multimedia, vol. 22, No. 1, pp. 45-58, 2019. [cited by applicant]
PCT International Search Report and Written Opinion mailed Dec. 21, 2023 issued in corresponding International Application No. PCT/CN2023/124825 (7 pgs.). [cited by applicant]