IP Library › Granted Patent US 11,386,583
Granted Patent B2
US 11,386,583 · App. 16/875,038 · Granted Jul 12, 2022

Image coding apparatus, probability model generating apparatus and image decoding apparatus

Inventors: Sihan Wen (Beijing, CN); Jing Zhou (Beijing, CN); Zhiming Tan (Beijing, CN)
Assignee: FUJITSU LIMITED
G06T9/002G06F17/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,386,583
App. No.
16/875,038
Granted
Jul 12, 2022
Kind
B2
Abstract

Embodiments of this disclosure provide an image coding apparatus, a probability model generating apparatus and an image decoding apparatus. A processor is to perform feature extraction on an input image to obtain first feature maps of N channels; to perform feature extraction on the input image with a size of the input image being adjusted K times, to respectively obtain second feature maps of N channels; and to concatenate the first feature maps of the K×N channels with the second feature maps of K×N channels to output a concatenated feature maps of channels. Hence, features of images may be accurately extracted and more competitive latent representations may be obtained.

Claims (29)

1. An apparatus, comprising:

a processor to couple to a memory and to,

perform feature extraction on an input image to obtain first feature maps of N channels;

perform feature extraction on the input image with a size of the input image being adjusted K times, to respectively obtain second feature maps of K×N channels;

concatenate the first feature maps of the N channels with the second feature maps of K×N channels to output concatenated feature maps of channels;

assign weights to the concatenated feature maps of channels; and

perform down-dimension processing on the weighted concatenated feature maps of channels to obtain feature maps of M channels and output the feature maps of M channels.

2. The apparatus according to claim 1 , wherein to obtain the first feature maps the processor is to configure a plurality of inception processors, each inception processor being sequentially connected, which perform the feature extraction on the input image or perform feature extraction on feature maps from a preceding inception processor to obtain global information and high-level information of the input image.

3. The apparatus according to claim 2 , wherein each inception processor of the inception processors is to:

perform the feature extraction on the input image or perform feature extraction on the feature maps from the preceding inception processor by using different convolution kernels and identical numbers of channels, to respectively obtain the first feature maps of N channels;

perform a pooling by down-dimension processing on the input image or on the feature map from the preceding inception processor to obtain the first feature maps of the N channels;

concatenate the first feature maps of the N channels with the first feature maps of the N channels from the pooling to obtain feature maps of 4N channels; and

perform down-dimension processing on the feature maps of the 4N channels to obtain the first feature maps of the N channels.

4. The apparatus according to claim 1 , wherein to obtain the second features maps of K×N channels, the processor is to:

adjust the size of the input image; and

perform feature extraction on the input image with the size being adjusted to obtain the second feature maps of the K×N channels.

5. The apparatus according to claim 4 , to adjust the size of the input image, the processor is to configure one or more groups of size adjusting processors, size adjusting processors of different groups performing size adjustment on the input image by using different scales, and performing feature extraction on the input image with the size being adjusted by using different convolution kernels.

6. An apparatus, comprising:

a processor to couple to a memory and to,

perform feature extraction on output of a hyper decoder to obtain multi-scale auxiliary information;

concatenate a latent representation of an input image from an arithmetic decoder with the multi-scale auxiliary information; and

decode output from the concatenator to obtain a reconstructed image of the input image.

7. The apparatus according to claim 6 , wherein to perform the feature extraction, the processor is to use dilated convolution kernels of different dilation ratios and identical numbers of channels to obtain the multi-scale auxiliary information.

8. An apparatus, comprising:

a processor to couple to a memory and to,

perform feature extraction on output of a hyper decoder to obtain multi-scale auxiliary information;

obtain information indicating content-based prediction by taking a latent representation of an input image from a quantizer as input; and

process the information indicating the content-based prediction and the multi-scale auxiliary information to obtain a predicted probability model,

wherein to perform the feature extraction, the processor is to use dilated convolution kernels of different dilation ratios and identical numbers of channels to obtain the multi-scale auxiliary information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2020
From: WEN, SIHAN; ZHOU, JING; TAN, ZHIMING
To: FUJITSU LIMITED
Reel/Frame 052701/0732 →
Priority Claims (1)
CN 201910429870.3 · May 22, 2019 · national
Continuity (1)
Related Publication 20200372686A1 · Nov 26, 2020
Cited By (1)
US 12,361,734