IP Library › Granted Patent US 12,530,588
Granted Patent B2
US 12,530,588 · App. 17/971,546 · Granted Jan 20, 2026

Generative video compression with a transformer-based discriminator

Inventors: Pengli Du (Santa Clara, CA); Ying Liu (Santa Clara, CA); Nam Ling (Santa Clara, CA); Yongxiong Ren (San Jose, CA); Lingzhi Liu (San Jose, CA)
Assignees: Beijing Dajia Internet Information Technology Co., Ltd.; Santa Clara University
G06N3/084H04N19/521H04N19/573H04N19/91
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,588
App. No.
17/971,546
Granted
Jan 20, 2026
Kind
B2
Abstract

A method, an apparatus, and a non-transitory computer-readable storage medium for video compression using a generative adversarial network (GAN) are provided. The method includes obtaining, by a generator of the GAN, a reconstructed target frame based on a reference frame and a raw target frame to be reconstructed; concatenating, by a transformer-based discriminator of the GAN, the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data; determining, by the transformer-based discriminator of the GAN, whether the paired data is real or fake to guide reconstruction of the raw target frame; and determining a generator loss and a transformer-based discriminator loss, and performing gradient back propagation and updating network parameters of the GAN based on the generator loss and the transformer-based discriminator loss.

Claims (91)

1 . A method for video compression, performed by a terminal using a generative adversarial network (GAN), comprising:

obtaining, by a generator of the GAN, a reconstructed target frame based on a reference frame and a raw target frame to be reconstructed;

concatenating, by a transformer-based discriminator of the GAN, the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data, wherein the transformer-based discriminator is configured to model long-distance dependencies across the reference frame, the raw target frame and the reconstructed target frame;

determining, by the transformer-based discriminator of the GAN, whether the paired data is real or fake to guide reconstruction of the raw target frame, wherein the reconstruction of the raw target frame comprises encoding and decoding of target frames;

determining a generator loss and a transformer-based discriminator loss, and performing gradient back propagation and updating network parameters of the GAN based on the generator loss and the transformer-based discriminator loss to obtain a trained GAN for the terminal; and

compressing, by the trained GAN of the terminal, a video stream for communication, storage, or processing.

2 . The method for video compression of claim 1 , wherein obtaining the reconstructed target frame based on the reference frame and the raw target frame to be reconstructed further comprises:

obtaining the reference frame and the raw target frame to be reconstructed;

obtaining, by a motion estimation network of the generator, an estimated motion based on the reference frame and the raw target frame;

encoding, by a motion encoder network of the generator, the estimated motion to obtain an encoded motion, quantizing the encoded motion into a quantized encoded motion, and converting the quantized encoded motion into a bit stream with entropy encoding;

decoding, by a motion decoder network of the generator, the bit stream with entropy decoding, dequantizing and decoding the bit stream to obtain a decoded motion, and warping the decoded motion with the reference frame to obtain a warped target frame; and

concatenating the warped target frame, the reference frame and a reconstructed motion together as a tensor, and obtaining a predicted targe frame by feeding the tensor into a motion compensation convolutional neural network.

3 . The method for video compression of claim 2 , further comprising:

subtracting the predicted target frame from the raw target frame to obtain a residue;

obtaining, by a residue encoder network of the generator, an encoded residue by feeding the residue into the residue encoder network, quantizing the encoded residue into a quantized encoded residue, and converting the quantized encoded residue into a residual bit stream with entropy encoding; and

decoding and dequantizing, by a residue decoder network of the generator, the residual bit stream to obtain a reconstructed residue, and adding the reconstructed residue to the predicted target frame to obtain a reconstructed target frame.

4 . The method for video compression of claim 1 , wherein the raw target frame to be reconstructed comprises a plurality of raw target frames to be constructed, and the plurality of raw target frames to be constructed are generated sequentially.

5 . The method for video compression of claim 2 , wherein obtaining the reference frame and the raw target frame to be reconstructed comprises:

obtaining a compressed intra I frame and a plurality of raw frames as inputs to the generator;

setting the compressed intra I frame as a first reference frame to generate a first reconstructed target P frame; and

setting the first reconstructed target P frame as a second reference frame to generate a second reconstructed target P frame.

6 . The method for video compression of claim 1 , wherein concatenating the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data further comprises:

concatenating, by the transformer-based discriminator, a quantized encoded motion and a quantized encoded residue together to feed into a Spatial Feature Extractor (SFE) to obtain an extracted feature, and concatenating the extracted feature and an estimated flow to form a condition; and

concatenating the raw target frame, the raw reference frame and the condition as a true data, and concatenating the generated target frame, the generated reference frame and the condition as a fake data; and

obtaining the paired data comprising the true data and the fake data.

7 . The method for video compression of claim 1 , wherein determining whether the paired data is real or fake further comprises:

feeding the paired data into a feature extraction convolutional neural network, and flattening the extract feature to obtain a flattened feature;

obtaining a transformed feature by feeding the flattened feature into a transformer block; and

determining whether the transformed feature is real or fake by feeding the transformed feature into a multi-layer perceptron head and a sigmoid activation function.

8 . The method for video compression of claim 1 , wherein determining the generator loss and the transformer-based discriminator loss further comprises:

determining the generator loss for reconstructing decoded frames by determining five terms, wherein the five terms comprise an adversarial loss term, a distortion loss term, a feature matching loss term, an entropy loss term, and a perceptual loss term; and

determining the discriminator loss based on last-layer discriminator probability obtained from both a reconstructed target frame and a raw target frame.

9 . The method for video compression of claim 8 , further comprising:

determining the adversarial loss term based on the reconstructed target frame;

determining the distortion loss term based on a mean squared error (MSE) between the raw target frame and the reconstructed target frame;

determining the feature matching loss term based on a MSE between discriminator features extracted from three scales of the reconstructed target frame and the raw target frame;

determining the entropy loss term based on an estimated entropy of a quantized encoded motion and a residue; and

determining the perceptual loss term based on a summed MSE between a true data feature and a fake data feature extracted from five different layers.

10 . An apparatus for video compression, for use in a terminal, comprising:

one or more processors; and

a memory configured to store a generative adversarial network (GAN) comprising a generator and a transformer-based discriminator, the GAN being executable by the one or more processors,

wherein the one or more processors, upon execution of the instructions, are configured to:

obtain a reconstructed target frame based on a reference frame and a raw target frame to be reconstructed;

concatenate the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data, wherein the transformer-based discriminator is configured to model long-distance dependencies across the reference frame, the raw target frame and the reconstructed target frame;

determine whether the paired data is real or fake to guide reconstruction of the raw target frame, wherein the reconstruction of the raw target frame comprises encoding and decoding of target frames;

determine a generator loss and a transformer-based discriminator loss, and perform gradient back propagation and update network parameters of the GAN based on the generator loss and the transformer-based discriminator loss to obtain a trained GAN for the terminal; and

compressing, by the trained GAN of the terminal, a video stream for communication, storage, or processing.

11 . The apparatus for video compression of claim 10 , wherein the one or more processors are further configured to:

obtain the reference frame and the raw target frame to be reconstructed;

obtain an estimated motion based on the reference frame and the raw target frame;

encode the estimated motion to obtain an encoded motion, quantize the encoded motion into a quantized encoded motion, and convert the quantized encoded motion into a bit stream with entropy encoding;

decode the bit stream with entropy decoding, dequantize and decode the bit stream to obtain a decoded motion, and warp the decoded motion with the reference frame to obtain a warped target frame; and

concatenate the warped target frame, the reference frame and a reconstructed motion together as a tensor, and obtain a predicted targe frame by feeding the tensor into a motion compensation convolutional neural network.

12 . The apparatus for video compression of claim 11 , wherein the one or more processors are further configured to:

subtract the predicted target frame from the raw target frame to obtain a residue;

obtain an encoded residue by feeding the residue into the residue encoder network, quantize the encoded residue into a quantized encoded residue, and convert the quantized encoded residue into a residual bit stream with entropy encoding; and

decode and dequantize the residual bit stream to obtain a reconstructed residue, and add the reconstructed residue to the predicted target frame to obtain a reconstructed target frame.

13 . The apparatus for video compression of claim 10 , wherein the one or more processors are further configured to:

concatenate a quantized encoded motion and a quantized encoded residue together to feed into a Spatial Feature Extractor (SFE) to obtain an extracted feature, and concatenate the extracted feature and an estimated flow to form a condition; and

concatenate the raw target frame, the raw reference frame and the condition as a true data, and concatenate the generated target frame, the generated reference frame and the condition as a fake data; and

obtain the paired data comprising the true data and the fake data.

14 . The apparatus for video compression of claim 10 , wherein the one or more processors are further configured to:

feed the paired data into a feature extraction convolutional neural network, and flatten the extract feature to obtain a flattened feature;

obtain a transformed feature by feeding the flattened feature into a transformer block; and

determine whether the transformed feature is real or fake by feeding the transformed feature into a multi-layer perceptron head and a sigmoid activation function.

15 . The apparatus for video compression of claim 10 , wherein the one or more processors are further configured to:

determine the generator loss for reconstructing decoded frames by determining five terms, wherein the five terms comprise an adversarial loss term, a distortion loss term, a feature matching loss term, an entropy loss term, and a perceptual loss term; and

determine the discriminator loss based on last-layer discriminator probability obtained from both a reconstructed target frame and a raw target frame.

16 . A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by a terminal having one or more computer processors, cause the one or more computer processors of the terminal to perform acts comprising:

obtaining, by a generator of a generative adversarial network (GAN), a reconstructed target frame based on a reference frame and a raw target frame to be reconstructed;

concatenating, by a transformer-based discriminator of the GAN, the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data, wherein the transformer-based discriminator is configured to model long-distance dependencies across the reference frame, the raw target frame and the reconstructed target frame;

determining, by the transformer-based discriminator of the GAN, whether the paired data is real or fake to guide reconstruction of the raw target frame, wherein the reconstruction of the raw target frame comprises encoding and decoding of target frames;

determining a generator loss and a transformer-based discriminator loss, and performing gradient back propagation and updating network parameters of the GAN based on the generator loss and the transformer-based discriminator loss to obtain a trained GAN for the terminal; and

compressing, by the trained GAN of the terminal, a video stream for communication, storage, or processing.

17 . The non-transitory computer-readable storage medium for of claim 16 , wherein obtaining the reconstructed target frame based on the reference frame and the raw target frame to be reconstructed further comprises:

obtaining the reference frame and the raw target frame to be reconstructed;

obtaining, by a motion estimation network of the generator, an estimated motion based on the reference frame and the raw target frame;

encoding, by a motion encoder network of the generator, the estimated motion to obtain an encoded motion, quantizing the encoded motion into a quantized encoded motion, and converting the quantized encoded motion into a bit stream with entropy encoding;

decoding, by a motion decoder network of the generator, the bit stream with entropy decoding, dequantizing and decoding the bit stream to obtain a decoded motion, and warping the decoded motion with the reference frame to obtain a warped target frame; and

concatenating the warped target frame, the reference frame and a reconstructed motion together as a tensor, and obtaining a predicted targe frame by feeding the tensor into a motion compensation convolutional neural network.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the instructions cause the one or more processors to perform acts further comprising:

subtracting the predicted target frame from the raw target frame to obtain a residue;

obtaining, by a residue encoder network of the generator, an encoded residue by feeding the residue into the residue encoder network, quantizing the encoded residue into a quantized encoded residue, and converting the quantized encoded residue into a residual bit stream with entropy encoding; and

decoding and dequantizing, by a residue decoder network of the generator, the residual bit stream to obtain a reconstructed residue, and adding the reconstructed residue to the predicted target frame to obtain a reconstructed target frame.

19 . The non-transitory computer-readable storage medium of claim 16 , wherein concatenating the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data further comprises:

concatenating, by the transformer-based discriminator, a quantized encoded motion and a quantized encoded residue together to feed into a Spatial Feature Extractor (SFE) to obtain an extracted feature, and concatenating the extracted feature and an estimated flow to form a condition; and

concatenating the raw target frame, the raw reference frame and the condition as a true data, and concatenating the generated target frame, the generated reference frame and the condition as a fake data; and

obtaining the paired data comprising the true data and the fake data.

20 . The non-transitory computer-readable storage medium of claim 16 , wherein determining the generator loss and the transformer-based discriminator loss further comprises:

determining the generator loss for reconstructing decoded frames by determining five terms, wherein the five terms comprise an adversarial loss term, a distortion loss term, a feature matching loss term, an entropy loss term, and a perceptual loss term; and

determining the discriminator loss based on last-layer discriminator probability obtained from both a reconstructed target frame and a raw target frame.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2026
From: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
To: BEIJING TRANSTREAMS TECHNOLOGY CO., LTD.
Reel/Frame 074656/0721 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2025
From: KWAI INC.
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 073219/0688 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2022
From: DU, PENGLI; LIU, YING; LING, NAM
To: SANTA CLARA UNIVERSITY
Reel/Frame 061504/0486 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2022
From: REN, YONGXIONG; LIU, LINGZHI
To: KWAI INC.
Reel/Frame 061504/0489 →
Continuity (1)
Related Publication 20240185075A1 · Jun 6, 2024
References Cited (28)
US 20200021873A1 · Swaminathan · 2020 [cited by examiner]
Feng, Runsen, et al. “Versatile learned video compression.” arXiv preprint arXiv:2111.03386 (2021) (Year: 2021). [cited by examiner]
Wang, Chaoyue, et al. “Evolutionary generative adversarial networks.” IEEE Transactions on Evolutionary Computation 23.6 (2019): 921-934. (Year: 2019). [cited by examiner]
T. Wiegand, G. J. Sullivan, G. Bjontegaard, and A. Luthra, “Overview of the h.264/avc video coding standard,” IEEE Trans. Circuits Syst. Video Technol, vol. 13, No. 7, pp. 560-576, Aug. 2003, (17p). [cited by applicant]
G. J. Sullivan, J. R. Ohm, W. J. Han, and T. Wiegand, “Overview of the high efficiency video coding (hevc) standard,” IEEE Trans. Circuits Syst. Video Technol., vol. 22, No. 12, pp. 1649-1668, Dec. 2012, (20p). [cited by applicant]
B. Bross, Y. K. Wang, Y. Ye, S. Liu, J. Chen, G. J. Sul-livan, and J. R. Ohm, “Overview of the versatile video coding (vvc) standard and its applications,” IEEE Trans. Circuits Syst. Video Technol., vol. 31, No. 10, pp.… [cited by applicant]
C. Y. Wu, N. Singhal, and P. Krahenbuhl, “Video com-pression through image interpolation,” in Proc. Eur. Conf. Comput. Vision, Aug. 2018, pp. 416-431, (18p). [cited by applicant]
A. Habibian, T. V. Rozendaal, J. M. Tomczak, , and T. S. Cohen, “Video compression with rate-distortion autoen-coders,” in Proc. Int. Conf. Comput. Vision, Oct. 2019, (10p). [cited by applicant]
G.Lu,W.Ouyang,D.Xu,X.Zhang,C.Cai,andZ.Gao, “Dvc: An end-to-end deep video compression frame-work,” in Proc. IEEE/CVF Conf. Comput. Vision and Pattern Recognit., Jun. 2019, pp. 11006-11015, (10p). [cited by applicant]
R. Yang, F. Mentzer, L. V. Gool, and R. Timofte, “Learning for video compression with hierarchical qual-ity and recurrent enhancement,” in Proc. IEEE/CVF Conf. Comput. Vision and Pattern Recognit., Jun. 2020, pp. 6628-6… [cited by applicant]
R. Yang, F. Mentzer, L. V. Gool, and R. Timofte, “Learning for video compression with recurrent auto-encoder and recurrent probability model,” IEEE Trans. Selected Topics in Signal Process., vol. 15, No. 2, pp. 388-401,… [cited by applicant]
E. Agustsson, M. Tschannen, F. Mentzer, R. Timofte, and L. V. Gool, “Generative adversarial networks for extreme learned image compression,” in Proc. Int. Conf. Comput. Vision, Oct. 2019, pp. 221-231, (11p). [cited by applicant]
S. Iwai, T. Miyazaki, Y. Sugaya, and S. Omachi, “Fidelity-controllable extreme image compression with generative adversarial networks,” in Proc. IEEE/CVF Conf. Comput. Vision and Pattern Recognit., Jan. 2021, pp. 8235-8… [cited by applicant]
R. Yang, R. Timofte, and L. V. Gool, “Perceptual video compression with recurrent conditional gan,” in Pro-cessings of the International Joint Conference on Artifi-cial Intelligence (IJCAI), 2022, (8p). [cited by applicant]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst., Dec. 2017, pp. 5998-6008, (11p). [cited by applicant]
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Min-derer, G. Heigold, S. Gelly, and J. Uszkoreit, “An image is worth 16x16 words: Transformers for image recogni-tion at… [cited by applicant]
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kir-illov, , and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. Eur. Conf. Comput. Vision, Aug. 2020, pp. 213-229, (17p). [cited by applicant]
F. Yang, H. Yang, J. Fu, H. Lu, and B. Guo, “Learning texture transformer network for image super-resolution,” in Proc. IEEE/CVF Conf. Comput. Vision and Pattern Recognit., Jun. 2020, pp. 5791-5800, (10p). [cited by applicant]
J. Johnson, A. Alahi, and F. Li., “Perceptual losses for real-time style transfer and super-resolution,” in Proc. Eur. Conf. Comput. Vision, Oct. 2016, pp. 694-711. [cited by applicant]
F. Bellard, “Bpg image format,” https://bellard.org/bpg/, 2018, (3p). [cited by applicant]
A. Ranjan and M. J. Black, “Optical flow estimation using a spatial pyramid network,” in Proc. IEEE/CVF Conf. Comput. Vision and Pattern Recognit., Jun. 2017, pp. 4161-4170, (10p). [cited by applicant]
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proc. Adv. Neural Inf. Process. Syst., Dec. 2014, pp. 2672-2680, (9p). [cited by applicant]
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. Int. Conf. on Learning Representations, May 2015, (14p). [cited by applicant]
J. Ballé, D. Minnen, S. Singh, S. J. Hwang, and N. John-ston, “Variational image compression with a scale hyperprior,” in Proc. Int. Conf. on Learning Representa-tions, May 2018, (20p). [cited by applicant]
T. Xue, B. Chen, J. Wu, D. Wei, and W. T. Freeman, “Video enhancement with task-oriented flow,” IEEE Trans. Int. Comput. Vision, vol. 127, No. 8, pp. 1106-1125, Aug. 2019, (20p). [cited by applicant]
M. Heusel, H. Ramsauer, T. Unterthiner, B. Unterthiner, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Proc. Adv. Neural Inf. Process. Syst., 2017, (12p). [cited by applicant]
M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton, “Demystifying mmd gans,” in Proc. Int. Conf. on Learning Representations, May 2018, (36p). [cited by applicant]
F. Mentzer, G. D. Toderici, M. Tschannen, and E. Agustsson, “High-fidelity generative image compression,” Adv. Neural Inf. Process. Syst., vol. 33, Jan. 2020, (12p). [cited by applicant]
Cited By (1)
US 12,657,928