IP Library Patent Application 18033693
Patent Application
App. No. 18/033,693

LEARNED VIDEO COMPRESSION FRAMEWORK FOR MULTIPLE MACHINE TASKS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/033,693
Abstract

Processing of a compressed representation of a video signal is optimized for multiple tasks, such as object detection, viewing of displayed video, or other machine tasks. In one embodiment, multiple analysis stages and a single synthesis is performed as part of a coding/decoding operation with training of an encoder side analysis and, optionally, a corresponding machine task. In another embodiment, multiple synthesis operations are performed on the decoding side, so that respective analysis, synthesis, and task stages are optimized. Other embodiments comprise feeding decoded feature maps to tasks, predictive coding, and using hyperprior-based models.

Claims (33)

1 . A method, comprising:

generating a plurality of tensors of feature maps from multiple analyses of at least one image portion; and

encoding said plurality of tensors into a bitstream, wherein said bitstream comprises a number of layers in the bitstream or dependency flags for layers to point at reference tensors.

2 . An apparatus, comprising:

a processor, configured to:

generate a plurality of tensors of feature maps from multiple analyses of at least one image portion; and

encode said plurality of tensors into a bitstream, wherein said bitstream comprises a number of layers in the bitstream or dependency flags for layers to point at reference tensors.

3 . A method, comprising:

decoding a bitstream to generate multiple feature maps, wherein said bitstream comprises a number of layers in the bitstream or dependency flags for layers to point at reference tensors; and

processing the multiple feature maps using at least one synthesizer to generate outputs for multiple tasks.

4 . An apparatus, comprising:

a processor, configured to:

decode a bitstream to generate multiple feature maps, wherein said bitstream comprises a number of layers in the bitstream or dependency flags for layers to point at reference tensors; and

process the multiple feature maps using at least one synthesizer to generate outputs for multiple tasks.

5 . The method of claim 1 , wherein each tensor of feature maps is input to a different synthesis stage, for performing a given task.

6 . The method of claim 5 , wherein said different synthesis stages are optimized for said given task.

7 . The apparatus of claim 2 , wherein tensors are compressed using predictive coding.

8 . The apparatus of claim 7 , wherein predictive coding comprises transmitting encoded residuals between different tensors.

9 . The method of claim 3 , wherein one synthesizer is used for a viewing task.

10 . The apparatus of claim 2 , wherein said bitstream comprises multi-view video coding.

11 . (canceled)

12 . A device comprising:

an apparatus according to claim 2 ; and

at least one of (i) an antenna configured to receive a signal, the signal including the video block, (ii) a band limiter configured to limit the received signal to a band of frequencies that includes the video block, and (iii) a display configured to display an output representative of a video block.

13 . A non-transitory computer readable medium containing data content generated according to claim 1 , for playback using a processor.

14 . (canceled)

15 . A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of claim 3 .

16 . The method of claim 3 , wherein each tensor of feature maps is input to a different synthesis stage, for performing a given task.

17 . The method of claim 5 , wherein said different synthesis stages are optimized for said given task.

18 . The apparatus of claim 4 , wherein tensors are compressed using predictive coding.

19 . The apparatus of claim 7 , wherein predictive coding comprises transmitting encoded residuals between different tensors.

20 . The apparatus of claim 4 , wherein one synthesizer is used for a viewing task.

21 . The apparatus of claim 4 , wherein said bitstream comprises multi-view video coding.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2023
From: RACAPE, FABIEN; HEWA GAMAGE, LAHIRU DULANJANA; PUSHPARAJA, AKSHAY; BEGAINT, JEAN
To: VID SCALE, INC.
Reel/Frame 063684/0200 →