IP Library Patent Application 19121914
Patent Application
App. No. 19/121,914

SYSTEMS AND METHODS FOR FRAME AND REGION TRANSFORMATIONS WITH SUPERRESOLUTION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/121,914
Abstract

Systems and methods for encoding and decoding video for machine consumption in which superresolution is used to upsample regions of interest at a decoder are provided. A decoder receives an encoded bitstream that has at least one packed frame with at least one region of interest therein being downsampled prior to encoding. A superresolution processor selectively performs upsampling of the identified regions. Parameters related to downsampling may be signaled in the bitstream and used to identify regions of interest to upsample. The superresolution processor may include a neural network trained on a dataset of images that have similar statistical properties to the pictures that are anticipated by a machine system receiving the decoded bitstream and may be trained based on a loss function in a task network based on that dataset.

Claims (22)

1 . A decoder for decoding an encoded bitstream having a packed frame of regions of interest, at least one region of interest being downsampled prior to encoding, the decoder comprising:

a video decoder, the video decoder receiving the encoded bitstream and decompressing the bitstream to identify regions of interest and region parameters therefrom;

a superresolution processing and region unpacking module, the superresolution processing and region unpacking module:

identifying decompressed regions of interest that were downsampled prior to encoding and to be upsampled using superresolution;

applying said downsampled regions to a superresolution network to perform upsampling; and

arranging the decoded regions of interest and upsampled regions of interest in a reconstructed frame with size, position and orientation corresponding to the original frame.

2 . The decoder of claim 1 , wherein parameters related to regions downsampled prior to encoding are signaled in the bitstream and the superresolution processing and region unpacking module use said parameters to identify regions of interest to upsample.

3 . The decoder of claim 2 , wherein the parameters include an indicator of one of a plurality of region types and wherein the superresolution network applies upsampling specific to the indicated region type

4 . The decoder of claim 1 , wherein the superresolution processing and region unpacking module incudes a superrsolution network comprising a trained neural network.

5 . The docoder of claim 4 , wherein the superresolution network is trained on a dataset of images that have similar statistical properties to the pictures that are anticipated by a machine system processing the decoded bitstream.

6 . The decoder of claim 5 , wherein the superresolution network is trained to optimize object detection in the pictures anticipated by the machine system.

7 . The decoder of claim 5 , wherein the superresolution network is trained based on a loss function in a task network.

8 . An encoder for video for machine consumption comprising:

a region detector module, the region detector module receiving the source video and identifying regions of interest therein, the regions of interest being defined in part by coordinates of a bounding box in the frame of source video;

a region transform module, the region transform module performing downsampling on at least one region of interest, wherein the downsampling is based in part on transform parameters from a machine system coupled to a decoder with a superresolution network for upsampling;

a region packing module, the region packing module receiving the set of regions of interest and transformed regions of interest and arranging the regions of interest into a packed frame in which pixels outside the regions of interest are substantially excluded; and

a video encoder receiving the packed frame and coordinates of the regions of interest and encoding the packed frame and region parameters into a coded bitstream.

9 . The encoder of claim 8 , wherein the region of interest comprise a plurality of object types and wherein the transform properties are specific to an object type.

10 . The encoder of claim 8 , wherein the transform properties are based at least in part on a superresolution network in a decoder which is trained by a task network to optimize a loss function.

11 . The encoder of claim 10 , wherein the superresolution network is trained on a dataset of images that have similar statistical properties to the pictures that are anticipated by a machine system.

12 . The encoder of claim 8 , wherein downsampling is determined on a region-by-region basis.

13 . The encoder of claim 12 , wherein downsampling is characterized by downsampling parameters applied by the region transform module for a region of interest and wherein the downsampling parameters are signaled in the coded bitstream.