IP Library Patent Application 19096523
Patent Application
App. No. 19/096,523

PLENO-GENERATION FACE VIDEO COMPRESSION FRAMEWORK FOR GENERATIVE FACE VIDEO COMPRESSION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/096,523
Abstract

Methods and systems implement a pleno-generation face video compression framework with bandwidth intelligence for generative models and compression. Heterogeneous-granularity facial description regularizes long-term dependencies between video frames and compensates for motion estimation errors caused by compact representations of motion information. A generative decoder reconstructs heterogeneous-granularity visual representations, providing auxiliary visual signals for attention-based recalibration of a GFVC-reconstructed face signal. A coarse-to-fine generation strategy avoids error accumulation. High efficiency for heterogeneous-granularity signal compression is achieved by two different entropy-based signal compression methods: heterogeneous-granularities feature representation from the key-reference frame as hyperpriors to optimize the entropy model for compressing heterogeneous-granularity feature from subsequent inter frames, and a feature difference operation for heterogeneous-granularities feature representation between key-reference and subsequent inter frames, such that the entropy model only compresses heterogeneous-granularities feature residual for redundancy reduction. Mixed-model dataset generation and training and model-specific dataset generation and training are also provided.

Claims (76)

1 . A computing system, comprising:

one or more processors, and

a computer-readable storage medium communicatively coupled to the one or more processors, the computer-readable storage medium storing computer-readable instructions executable by the one or more processors that, when executed by the one or more processors, perform associated operations comprising:

reconstructing a plurality of reconstructed inter frames of a video sequence by inputting a plurality of original inter frames to a generative face video compression (“GFVC”) model;

extracting an original auxiliary facial signal from the plurality of original inter frames;

extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames;

predicting a reconstructed auxiliary facial signal from a quantized auxiliary facial signal, based on a difference between the model-generated auxiliary facial signal and the original auxiliary facial signal; and

boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal.

2 . The computing system of claim 1 , wherein extracting the original auxiliary facial signal from the plurality of original inter frames comprises:

downsampling the plurality of original inter frames; and

transforming the plurality of original inter frames to a high-dimensional face feature map; and extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises:

downsampling the plurality of reconstructed inter frames; and

transforming the plurality of reconstructed inter frames to a high-dimensional face feature map.

3 . The computing system of claim 1 , wherein extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises:

selecting a higher or lower granularity of the original auxiliary facial signal based on higher or lower bitstream bandwidth.

4 . The computing system of claim 1 , wherein the reconstructed auxiliary facial signal is predicted based further on a Gaussian distribution comprising entropy parameters.

5 . The computing system of claim 4 , wherein the entropy parameters are conditioned upon:

a hyperprior comprising the model-generated auxiliary facial signal; and

a causal context of the quantized auxiliary facial signal.

6 . The computing system of claim 1 , wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal comprises:

transforming the reconstructed inter frames into facial features;

transforming the reconstructed auxiliary facial signal into signal features having a same feature dimensionality as the facial features;

performing linear projection upon the facial features and the signal features to yield latent feature maps; and

inputting the latent feature maps into an attention layer to yield fused attention features.

7 . The computing system of claim 6 , wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal further comprises:

inputting the attention features to a coarse face generator U-Net decoder to yield coarsely enhanced inter frames;

learning a motion estimation field and a facial occlusion map by concatenating a reconstructed key-reference frame and the coarsely enhanced inter frames; and

applying the motion estimation field and the facial occlusion map to multi-scale spatial features derived from the reconstructed key-reference frame.

8 . A method, comprising:

reconstructing a plurality of reconstructed inter frames of a video sequence by inputting a plurality of original inter frames to a generative face video compression (“GFVC”) model;

extracting an original auxiliary facial signal from the plurality of original inter frames;

extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames;

predicting a reconstructed auxiliary facial signal based on a difference between the model-generated auxiliary facial signal and the original auxiliary facial signal; and boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal.

9 . The method of claim 8 , wherein extracting the original auxiliary facial signal from the plurality of original inter frames comprises:

downsampling the plurality of original inter frames; and

transforming the plurality of original inter frames to a high-dimensional face feature map; and extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises:

downsampling the plurality of reconstructed inter frames; and

transforming the plurality of reconstructed inter frames to a high-dimensional face feature map.

10 . The method of claim 8 , wherein extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises:

selecting a higher or lower granularity of the original auxiliary facial signal based on higher or lower bitstream bandwidth.

11 . The method of claim 8 , wherein the reconstructed auxiliary facial signal is predicted based further on a Gaussian distribution comprising entropy parameters.

12 . The method of claim 11 , wherein the entropy parameters are conditioned upon:

a hyperprior comprising the model-generated auxiliary facial signal; and

a causal context of the quantized auxiliary facial signal.

13 . The method of claim 8 , wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal comprises:

transforming the reconstructed inter frames into facial features;

transforming the reconstructed auxiliary facial signal into signal features having a same feature dimensionality as the facial features;

performing linear projection upon the facial features and the signal features to yield latent feature maps; and

inputting the latent feature maps into an attention layer to yield fused attention features.

14 . The method of claim 13 , wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal further comprises:

inputting the attention features to a coarse face generator U-Net decoder to yield coarsely enhanced inter frames;

learning a motion estimation field and a facial occlusion map by concatenating a reconstructed key-reference frame and the coarsely enhanced inter frames; and

applying the motion estimation field and the facial occlusion map to multi-scale spatial features derived from the reconstructed key-reference frame.

15 . One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform operations comprising:

reconstructing a plurality of reconstructed inter frames of a video sequence by inputting a plurality of original inter frames to a generative face video compression (“GFVC”) model;

extracting an original auxiliary facial signal from the plurality of original inter frames;

extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames;

predicting a reconstructed auxiliary facial signal based on a difference between the model-generated auxiliary facial signal and the original auxiliary facial signal; and

boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal.

16 . The non-transitory computer-readable media of claim 15 , wherein extracting the original auxiliary facial signal from the plurality of original inter frames comprises:

downsampling the plurality of original inter frames; and

transforming the plurality of original inter frames to a high-dimensional face feature map; and extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises:

downsampling the plurality of reconstructed inter frames; and

transforming the plurality of reconstructed inter frames to a high-dimensional face feature map.

17 . The non-transitory computer-readable media of claim 15 , wherein extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises:

selecting a higher or lower granularity of the original auxiliary facial signal based on higher or lower bitstream bandwidth.

18 . The non-transitory computer-readable media of claim 15 , wherein the reconstructed auxiliary facial signal is predicted based further on a Gaussian distribution comprising entropy parameters.

19 . The non-transitory computer-readable media of claim 15 , wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal comprises:

transforming the reconstructed inter frames into facial features;

transforming the reconstructed auxiliary facial signal into signal features having a same feature dimensionality as the facial features;

performing linear projection upon the facial features and the signal features to yield latent feature maps; and

inputting the latent feature maps into an attention layer to yield fused attention features.

20 . The non-transitory computer-readable media of claim 19 , wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal further comprises:

inputting the attention features to a coarse face generator U-Net decoder to yield coarsely enhanced inter frames;

learning a motion estimation field and a facial occlusion map by concatenating a reconstructed key-reference frame and the coarsely enhanced inter frames; and

applying the motion estimation field and the facial occlusion map to multi-scale spatial features derived from the reconstructed key-reference frame.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA (CHINA) CO., LTD.
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 075529/0567 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2025
From: LIAO, RU-LING; CHEN, JIE; YE, YAN; CHEN, BOLIN; WANG, SHIQI
To: ALIBABA (CHINA) CO., LTD.
Reel/Frame 070856/0578 →