IP Library Granted Patent US 12,651,381
Granted Patent B2
US 12,651,381 · App. 17/816,832 · Granted Jun 9, 2026

End-to-end optimization of adaptive spatial resampling towards machine vision

Inventors: Shurun Wang (Beijing, CN); Zhao Wang (Beijing, CN); Yan Ye (San Diego, CA); Shiqi Wang (Kowloon Tong, HK)
Assignee: Alibaba Innovation Private Limited
G06T9/002G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,381
App. No.
17/816,832
Granted
Jun 9, 2026
Kind
B2
Abstract

A computer-implemented method for training spatial resampling modules includes: down-sampling, by a down-sampling module, an input image data to generate a down-sampled image data; up-sampling, by an up-sampling module, the down-sampled image data to generate a first up-sampled image data; analyzing, by a plurality of analysis models corresponding to a plurality of tasks, the first up-sampled image data; and training the down-sampling module based on a loss function associated with the plurality of analysis models according to the input image data and the first up-sampled image data.

Claims (30)

1 . A computer-implemented method for spatial resampling, comprising:

performing an instance segmentation to an image to be analyzed;

selecting a resampling factor from a plurality of resampling factor candidates based on an area of object regions calculated according to the instance segmentation;

down-sampling, by a down-sampling module, the image to be analyzed based on the selected resampling factor for resampling the image to generate a down-sampled image data; and

compressing, by an encoder, the down-sampled image data to obtain a quantized and compressed bitstream, wherein the bitstream is decoded by a decoder to obtain a reconstructed image data that is up-sampled, by an up-sampling module, based on the selected resampling factor to generate an up-sampled image data.

2 . The computer-implemented method of claim 1 , further comprising:

up-sampling, by an up-sampling module, the down-sampled image data based on the selected resampling factor to generate an up-sampled image data.

3 . The computer-implemented method of claim 1 , wherein the selecting the resampling factor comprises:

selecting the resampling factor based on a width and a height of the image, and the area of object regions calculated by an instance segmentation network performing the instance segmentation.

4 . The computer-implemented method of claim 1 , further comprising:

skipping the down-sampling in response to the resampling factor being 100 percent when a portion parameter calculated based on the area of object regions is lower than or equal to a threshold value.

5 . The computer-implemented method of claim 1 , wherein the down-sampling module is trained based on a loss function associated with a plurality of analysis models.

6 . A computer-implemented method for spatial resampling, comprising:

decoding, by a decoder, a quantized and compressed bitstream to obtain reconstructed image data and a resampling factor for resampling an image; and

up-sampling, by an up-sampling module, the reconstructed image data based on the resampling factor to generate an up-sampled image data, wherein the resampling factor is selected from a plurality of resampling factor candidates based on an area of object regions calculated according to instance segmentation of the image, wherein the bitstream is obtained by compressing, by an encoder, a down-sampled image data.

7 . The computer-implemented method of claim 6 , wherein the resampling factor is selected based on a width and a height of the image, and the area of object regions calculated by an instance segmentation network performing the instance segmentation.

8 . The computer-implemented method of claim 6 , further comprising:

skipping the up-sampling in response to the resampling factor being 100 percent when a portion parameter calculated based on the area of object regions is lower than or equal to a threshold value.

9 . The computer-implemented method of claim 6 , wherein the up-sampling module is trained based on a loss function associated with a plurality of analysis models.

10 . A method for processing a bitstream, comprising:

receiving a bitstream comprising coded data associated with an input image, wherein the bitstream is obtained by compressing, by an encoder, a down-sampled image data;

decoding, by a decoder, the bitstream to obtain reconstructed image data; and

up-sampling, by an up-sampling module, the reconstructed image data based on a resampling factor to generate an up-sampled image data, wherein the resampling factor is selected from a plurality of resampling factor candidates based on an area of object regions calculated according to instance segmentation of the image.

11 . The method of claim 10 , wherein the bitstream comprises an index representing the resampling factor.

12 . The method of claim 10 , wherein the resampling factor is selected based on a width and a height of the input image, and an area of object regions calculated by an instance segmentation network performing the instance segmentation.

13 . The method of claim 10 , further comprising:

skipping the up-sampling in response to the resampling factor being 100 percent when a portion parameter calculated based on the area of object regions is lower than or equal to a threshold value.

14 . The method of claim 10 , wherein the up-sampling module is trained based on a loss function associated with a plurality of analysis models during a training stage.

15 . The method of claim 14 , wherein the bitstream is provided by compressing a down-sampled image data generated by down-sampling the input image by a down-sampling module trained based on the same loss function associated with the plurality of analysis models.

16 . The method of claim 14 , wherein the loss function comprises a contour loss function, a plurality of feature map distortions respectively associated with the analysis models, a plurality of analysis loss functions respectively associated with the analysis models, or any combinations thereof.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2024
From: ALIBABA SINGAPORE HOLDING PRIVATE LIMITED
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 066348/0252 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2022
From: WANG, SHURUN; WANG, ZHAO; YE, YAN; WANG, SHIQI
To: ALIBABA SINGAPORE HOLDING PRIVATE LIMITED
Reel/Frame 061232/0524 →
Continuity (1)
Related Publication 20240046527A1 · Feb 8, 2024
References Cited (26)
US 20220086469A1 · Sarwer · 2022 [cited by examiner]
Adini et al., “Context-enabled learning in the human visual system,” Nature, vol. 415, 2002, pp. 790-793. [cited by applicant]
An et al., “Block partitioning structure for next generation video coding,” International Telecommunications Union, 2015, 8 pages. [cited by applicant]
Balle, et al., “Variational Image Compression with a Scale Hyperprior,” International Conference on Learning Representations, 2018, 23 pages. [cited by applicant]
Balle et al., End-to-End Optimized Image Compression, ICLR, 2017, 27 pages. [cited by applicant]
Bross et al., Developments in International Video Coding Standardization After AVC, With an Overview of Versatile Video Coding (VVC), Proceedings of the IEEE, vol. 109, No. 9, Sep. 2021, pp. 1463-1493. [cited by applicant]
Chao, et al., “A Novel Rate Control Framework for SIFT/SURF Feature Preservation in H.264/AVC Video Compression,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 25, No. 6, Jun. 2015, pp. 958-972. [cited by applicant]
Chen et al., “Contour Loss: Boundary-Aware Learning for Salient Object Segmentation,” arXiv, 2019, 12 pages. [cited by applicant]
Duan et al., “Overview of the MPEG-CDVS Standard,” IEEE Transactions on Image Processing, vol. 25, No., 1, Jan. 2016, pp. 179-194. [cited by applicant]
Duan et al., Compact Descriptors for Video Analysis: The Emerging MPEG Standard, IEEE Computer Society, 2018, pp. 44-54. [cited by applicant]
Duan et al., “Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics,” IEEE Transactions on Image Processing, vol. 29, 2020, pp. 8680-8695. [cited by applicant]
Jiang et al., “An End-to-End Compression Framework Based on Convolutional Neural Networks,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, No. 10, Oct. 2018, pp. 3007-3018. [cited by applicant]
Li et al., “λ Domain Rate Control Algorithm for High Efficiency Video Coding,” IEEE Transactions on Image Processing, vol. 23, No. 9, Sep. 2014, pp. 3841-3854. [cited by applicant]
Lin et al., “Adaptive Downsampling to Improve Image Compression at Low Bit Rates,” IEEE Transactions on Image Processing, vol. 15, No. 9, Sep. 2006, pp. 2513-2521. [cited by applicant]
Liu et al., “Compressive Sampling-Based Image Coding for Resource-Deficient Visual Communication,” IEEE Transactions on Image Processing, vol. 25, No. 6, Jun. 2016, pp. 2844-2855. [cited by applicant]
Liu et al., “CNN-Basesd DCT-Like Transform for Image Compression,” Springer, 2018, pp. 61-72. [cited by applicant]
Minnen et al., “Joint Autoregressive and Hierarchical Priors for Learned Image Compression,” 32nd Conference on Neural Information Processing System, 2018, 10 pages. [cited by applicant]
Mohan et al., “Internet of Video Things in 2030: A World with Many Cameras,” IEEE Xplore, 2017, 4 pages. [cited by applicant]
Pfaff et al., “CE3: Affine linear weighted intra prediction (CE3-4.1, CE3-4.2),” JVET-N0217, 14 [cited by applicant]
Rabbani et al., “JPEG2000: Image compression fundamentals, standards and practice,” Journal of Electronic Imaging, 2002, 11(2): 286. [cited by applicant]
Sullivan et al., “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, pp. 1649-1668 (2012). [cited by applicant]
Toderici et al., “Variable Rate Image Compression with Recurrent Neural Networks,” ICLR, 2016, 12 pages. [cited by applicant]
Wallace et al., “The JPEG Still Picture Compression Standard,” IEEE Transactions on Consumer Electronics, vol. 38, No. 1, Feb. 1992, 17 pages. [cited by applicant]
Wang et al., “Extended Coding Unit Partitioning for Future Video Coding,” IEEE Transactions on Image Processing, vol. 29, 2020, pp. 2931-2946. [cited by applicant]
Wiegand et al., “Overview of the H.264/AVC Video Coding Standard,” IEEE Transactions on Circuitds and Systems for Video Technology, vol. 13, No. 7, Jul. 2003, pp. 560-576. [cited by applicant]
Zhao et al., “Mode-dependent non separable secondary transform,” ITU-T 2015. [cited by applicant]