IP Library Granted Patent US 12,556,720
Granted Patent B2
US 12,556,720 · App. 18/033,697 · Granted Feb 17, 2026

Learned video compression and connectors for multiple machine tasks

Inventors: Fabien Racape (San Francisco, CA); Lahiru Dulanjana Hewa Gamage (Davis, CA); Jean Begaint (Menlo Park, CA); Simon Feltman (Sunnyvale, CA)
Assignee: INTERDIGITAL VC HOLDINGS, INC.
H04N19/186H04N19/176H04N19/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,556,720
App. No.
18/033,697
Granted
Feb 17, 2026
Kind
B2
Abstract

A processing module, or connector, adapts an output of a codec, or a decoded output, to a form suitable for an alternate task. In one embodiment, the output of a codec is used for a machine task and the connector adapts this output to a form suitable for a video display. In another embodiment, metadata accompanies the codec output, which can instruct the connector how to adapt the codec output for an alternate task. In other embodiments, the processing module performs averaging over a N×M window, or convolution.

Claims (39)

1 . A method, comprising:

coding a tensor comprising a video portion using a learned transform comprising a first set of constraints to generate a latent map, wherein metadata comprising trained parameters is included in said latent map;

quantizing and entropy coding the latent map;

entropy decoding the latent map;

processing the entropy decoded latent map under a second set of task-specific constraints, distinct from compression-related residual constraints, the processing comprising applying a convolutional neural network trained to perform semantic or perceptual analysis of video content; and

performing a task using the trained parameters with the processed latent map, the task comprising semantic or perceptual video processing distinct from bitstream reconstruction or residual prediction, including at least one of anomaly detection, video enhancement, object recognition, or augmented reality overlay.

2 . The method of claim 1 , wherein the second set of constraints comprises semantic segmentation constraints trained using pixel-level semantic labels of the video portion.

3 . The method of claim 1 , wherein the task comprises video super-resolution or perceptual enhancement, distinct from residual coding or entropy reconstruction.

4 . The method of claim 3 , wherein said processing further comprises weighting of samples.

5 . The method of claim 1 , wherein said processing comprises a temporal convolution across multiple consecutive frames in addition to spatial convolution across pixel locations.

6 . The method of claim 1 , wherein said processing comprises generating an augmented reality overlay or annotation of detected objects in the video portion.

7 . The method of claim 1 , wherein the second set of constraints is dynamically adapted based on feedback from the semantic or perceptual task output.

8 . A non-transitory computer readable medium containing data content generated according to the method of claim 1 , for playback using a processor.

9 . An apparatus, comprising:

a processor, configured to:

code a tensor comprising a video portion using a learned transform comprising a first set of constraints to generate a latent map, wherein metadata comprising trained parameters is included in said latent map;

quantizing and entropy coding the latent map;

entropy decoding the latent map;

process the entropy decoded latent map under a second set of task-specific constraints, distinct from compression-related residual constraints, the processing comprising applying a convolutional neural network trained to perform semantic or perceptual analysis; and

perform a task using the trained parameters with the processed latent map, the task comprising semantic or perceptual video processing distinct from bitstream reconstruction or residual prediction.

10 . The apparatus of claim 9 , further comprising a hardware accelerator configured to execute the convolutional neural network under the second set of semantic constraints.

11 . The apparatus of claim 9 , wherein the processor is further configured to store in memory distinct parameter sets for compression constraints and for semantic constraints, and selectively apply said sets to the latent map.

12 . The apparatus of claim 9 , wherein the processor is further configured to adapt the latent map for anomaly detection in video surveillance.

13 . The apparatus of claim 9 , wherein the processor is further configured to output a semantically enhanced video stream distinct from a decoded reconstruction of the original video stream.

14 . A method, comprising:

decoding a video bitstream comprising a tensor comprising a video portion using a learned transform comprising a first set of constraints, wherein metadata comprising trained parameters is included in said tensor; and

processing the decoded video bitstream with a two-dimensional convolution on each color component at each pixel location under a second set of semantic constraints, distinct from compression residual coding, the processing adapting the decoded video bitstream for a subsequent semantic or perceptual task including anomaly detection, super-resolution, video quality enhancement, or object classification.

15 . A non-transitory computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of claim 14 .

16 . The method of claim 14 , wherein said processing operates on a window of samples centered at each sample location.

17 . The method of claim 14 , wherein said processing further comprises object recognition using convolutional features extracted under the second set of constraints.

18 . An apparatus, comprising:

a processor, configured to:

decode a video bitstream comprising a tensor comprising a video portion using a learned transform comprising a first set of constraints, wherein metadata comprising trained parameters is included in said tensor; and

process the decoded video bitstream with a two-dimensional convolution on each color component at each pixel location under a second set of semantic constraints distinct from compression residual coding, the processing adapting the decoded video bitstream for a subsequent semantic or perceptual task distinct from reconstruction or prediction coding.

19 . A device comprising:

an apparatus according to claim 18 ; and

at least one of (i) an antenna configured to receive a signal, the signal including the video block, (ii) a band limiter configured to limit the received signal to a band of frequencies that includes the video block, and (iii) a display configured to display an output representative of a video block.

20 . The apparatus of claim 18 , wherein said processing comprises convolution on each color component.

21 . The apparatus of claim 18 , wherein said processing comprises classification of video content for subsequent recommendation or retrieval tasks.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2023
From: RACAPE, FABIEN; HEWA GAMAGE, LAHIRU DULANJANA; BEGAINT, JEAN; FELTMAN, SIMON
To: VID SCALE, INC.
Reel/Frame 063684/0368 →
Continuity (2)
Provisional Application 63109498 · Nov 4, 2020
Related Publication 20230370622A1 · Nov 16, 2023
References Cited (10)
US 20190273948A1 · Yin · 2019 [cited by examiner]
US 20200151559A1 · Karras · 2020 [cited by examiner]
US 20200196024A1 · Hwang · 2020 [cited by examiner]
US 20200304802A1 · Habibian · 2020 [cited by examiner]
US 20200327702A1 · Wang · 2020 [cited by examiner]
EP 3706046 · 2020 [cited by applicant]
Minnen et al., Joint Autoregressive and Hierarchical Priors for Learned Image Compression, Advances in Neural Information Processing Systems 31, arXiv:1809.02736v1 (cs.CV), pp. 1-22, Sep. 8, 2018. [cited by applicant]
Liu, et al., DSIC: Deep Stereo Image Compression, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 9, 2019. [cited by applicant]
Balle, et al., Variational Image Compression with a Scale Hyperprior, ArXiv180201436 Cs Eess Math, May 2018,. [cited by applicant]
Chen et al., CNN Feature Coding for Cloud-Based Visual Analysis, 129. MPEG Meeting, Jan. 13, 2020-Jan. 17, 2020, Brussels, (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. m52162, Jan. 7, 2020. [cited by applicant]