IP Library Granted Patent US 12,689,762
Granted Patent B2
US 12,689,762 · App. 18/281,844 · Granted Jul 21, 2026

Temporal structure-based conditional convolutional neural networks for video compression

Inventors: Fabien Racape (San Francisco, CA); Jean Begaint (Menlo Park, CA); Simon Feltman (Sunnyvale, CA); Akshay Pushparaja (San Jose, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N19/537H04N19/177H04N19/184H04N19/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,689,762
App. No.
18/281,844
Granted
Jul 21, 2026
Kind
B2
Abstract

Video encoding and decoding is implemented with auto encoders using luminance information to derive motion information for chrominance prediction. In one embodiment conditional convolutions are used to encode motion flow information. A current condition, for example, GOP structure, is used as input to a succession of fully connected layers to implement the conditional convolution. In a related embodiment, more than one reference frame is used to encode motion flow information.

Claims (46)

1 . A method, comprising:

receiving an input to a conditional convolution layer, wherein the input comprises a concatenated tensor of a current block and at least one reference block, and wherein the conditional convolution layer is at least one layer in a series of fully connected convolution layers;

processing at least one conditional convolution on the input, wherein the at least one conditional convolution is processed based at least on data representative of a current condition, and wherein the current condition comprises temporal-structure information associated with a group of pictures (GOP) structure;

encoding motion flow using an output from the at least one conditional convolution; and

generating a bitstream comprising the encoded motion flow.

2 . The method of claim 1 , wherein the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure indicates at least one of: a framerate, a number of frames separating a first reference frame and a second reference frame, a distance between a current frame and a reference frame, a display order indicating whether the reference frame is before or after the current frame, a number of frames used to predict the current frame, a prediction direction between the current block and the reference block, or a content type.

3 . The method of claim 1 , wherein the GOP structure comprises a content type, and wherein the content type is at least one of: gaming, VR360 content, or screen content.

4 . The method of claim 1 , wherein the current condition is indicated by a one-hot encoded vector over a plurality of temporal conditions.

5 . The method of claim 1 , wherein the method further comprises:

indicating, in the bitstream, the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure.

6 . The method of claim 1 , further comprising:

determining a scaling set associated with the conditional convolution layer, wherein the scaling set is determined from the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure; and

processing the at least one conditional convolution on the input further by applying, via the conditional convolution layer, a weight and the scaling set to the input.

7 . The method of claim 6 , wherein the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure indicates a framerate and a number of frames separating a first reference frame and a second reference frame, and wherein the scaling set is determined based on the framerate and the number of frames.

8 . The method of claim 1 , further comprising:

determining a bias vector associated with the conditional convolution layer, wherein the bias vector is determined from the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure; and

processing the at least one conditional convolution on the input further by applying, via the conditional convolution layer, a weight and the bias vector to the input.

9 . The method of claim 8 , wherein the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure indicates a framerate and a number of frames separating a first reference frame and a second reference frame, and wherein the bias vector is determined based on the framerate and the number of frames.

10 . An encoding device, comprising:

a processor, configured to:

receive an input to a conditional convolution layer, wherein the input comprises a concatenated tensor of a current block and at least one reference block, and wherein the conditional convolution layer is at least one layer in a series of fully connected convolution layers;

process at least one conditional convolution on the input, wherein the at least one conditional convolution is processed based at least on data representative of a current condition, wherein the current condition comprises temporal-structure information associated with a group of pictures (GOP) structure;

encode motion flow using an output from the at least one conditional convolution; and

generate a bitstream comprising the encoded motion flow.

11 . A method, comprising:

entropy decoding a bitstream comprising motion flow data associated with a current block;

processing at least one conditional deconvolution on the motion flow data via a conditional deconvolution layer to generate a reconstructed residue, wherein the conditional deconvolution layer is at least one layer in a series of fully connected layers, and wherein the at least one conditional deconvolution is processed based at least on data representative of a current condition, wherein the current condition comprises temporal-structure information associated with a group of pictures (GOP) structure; and

decoding the current block based on the reconstructed residue and a prediction of the current block.

12 . The method of claim 11 , wherein the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure indicates at least one of: a framerate, a number of frames separating a first reference frame and a second reference frame, a distance between a current frame and a reference frame, a display order indicating whether the reference frame is before or after the current frame, a number of frames used to predict the current frame, a prediction direction between the current block and the reference block, or a content type.

13 . The method of claim 11 , wherein the GOP structure comprises a content type, and wherein the content type is at least one of: gaming, VR360 content, or screen content.

14 . The method of claim 11 , wherein the current condition is indicated by a one-hot encoded vector over a plurality of temporal conditions.

15 . The method of claim 11 , wherein the method further comprises:

receiving, in the bitstream, an indication of the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure.

16 . The method of claim 11 , further comprising:

determining a scaling set associated with the conditional deconvolution layer, wherein the scaling set is determined from the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure; and

processing the at least one conditional deconvolution on the motion flow data further by applying, via the conditional deconvolution layer, a weight and the scaling set to the motion flow data.

17 . The method of claim 16 , wherein the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure indicates a framerate and a number of frames separating a first reference frame and a second reference frame, and wherein the scaling set is determined based on the framerate and the number of frames.

18 . The method of claim 11 , further comprising:

determining a bias vector associated with the conditional deconvolution layer, wherein the bias vector is determined from the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure; and

processing the at least one conditional deconvolution on the motion flow data further by applying, via the conditional deconvolution layer, a weight and the bias vector to the motion flow data.

19 . The method of claim 18 , wherein the data representative of the current condition that comprises the temporal-structure information associated with the GOP structure indicates a framerate and a number of frames separating a first reference frame and a second reference frame, and wherein the bias vector is determined based on the framerate and the number of frames.

20 . A decoding device, comprising:

a processor, configured to:

entropy decode a bitstream comprising motion flow data associated with a current block;

process at least one conditional deconvolution on the motion flow data via a conditional deconvolution layer to generate a reconstructed residue, wherein the conditional deconvolution layer is at least one layer in a series of fully connected layers, and wherein the at least one conditional deconvolution is processed based at least on data representative of a current condition, wherein the current condition comprises temporal-structure information associated with a group of pictures (GOP) structure; and

decode the current block based on the reconstructed residue and a prediction of the current block.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 065166 FRAME: 0925. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 16, 2023
From: RACAPE, FABIEN; BEGAINT, JEAN; FELTMAN, SIMON; PUSHPARAJA, AKSHAY
To: VID SCALE, INC.
Reel/Frame 066018/0722 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2023
From: RACAPE, FABIEN; BEGAINT, JEAN; FELTMAN, SIMON; PUSHPARAJA, AKSHAY
To: PRITCHETT, BRIAN
Reel/Frame 065166/0925 →
Continuity (2)
Provisional Application 63162791 · Mar 18, 2021
Related Publication 20240187640A1 · Jun 6, 2024
References Cited (8)
US 20160366415A1 · Liu · 2016 [cited by examiner]
US 20190327484A1 · Grange · 2019 [cited by examiner]
US 20220129740A1 · Yang · 2022 [cited by examiner]
WO 2020016857 · 2020 [cited by applicant]
Agustsson, et al., Scale-Space Flow for End-to-End Optimized Video Compression, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, pp. 8503-8512. [cited by applicant]
Park et al., Deep Predictive Video Compression Using Mode-Selective Uni- and Bi-Directional Predictions Based on Multi-Frame Hypothesis, IEEE Access, vol. 9, Dec. 21, 2020, pp. 72-85. [cited by applicant]
Yang et al., Learning for Video Compression with Hierarchical Quality and Recurrent Enhancement, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, pp. 1-14. [cited by applicant]
Li et al., AHG11: Updated Information on Inter-Prediction Coding Tool With Deep Neural Network, 21. JVET Meeting, Jan. 6, 2021-Jan. 15, 2021, Teleconference, (The Joint Video Exploration Team of ISO/IEC JTC1/SC29/WG11 a… [cited by applicant]