IP Library Granted Patent US 12,439,094
Granted Patent B2
US 12,439,094 · App. 18/478,680 · Granted Oct 7, 2025

Pre-analysis based image compression methods

Inventors: Shurun Wang (Beijing, CN); Zhao Wang (Hangzhou, CN); Yan Ye (San Diego, CA); Shiqi Wang (Hong Kong, HK)
Assignee: Alibaba Damo (Hangzhou) Technology Co., Ltd.
H04N19/85H04N19/119H04N19/14H04N19/172H04N19/176H04N19/42H04N19/124
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,439,094
App. No.
18/478,680
Granted
Oct 7, 2025
Kind
B2
Abstract

The present disclosure provides pre-analysis based methods for adaptively compressing image data consumed by machine vision tasks. An exemplary method includes: receiving a video sequence; encoding one or more input pictures associated with the video sequence; and generating a bitstream, wherein the encoding includes: performing instance segmentation of an input picture, to generate one or more segment masks; combining the one or more segment masks to generate a merged mask; extracting, from the input picture, a region comprising the merged mask; and compressing image data representing the extracted region.

Claims (64)

1. An image data processing method, comprising:

receiving a video sequence;

encoding one or more input pictures associated with the video sequence; and

generating a bitstream,

wherein the encoding comprises:

performing instance segmentation of an input picture, to generate one or more segment masks;

combining the one or more segment masks to generate a merged mask;

extracting, from the input picture, a region comprising the merged mask; and

compressing image data representing the extracted region, wherein the extracted region comprises one or more pixels adjacent to the merged mask.

2. The method of claim 1 , wherein extracting, from the input picture, the region comprising the merged mask comprises:

in the input picture, identifying one or more blocks not overlapping with the merged mask; and

extracting, from the input picture, a region not including the one or more blocks.

3. The method of claim 2 , wherein identifying the one or more blocks using a slide window, and a size of the slide window is determined based on at least one of:

a quantization parameter (QP) for compressing the input picture, or

a definition of the input picture.

4. The method of claim 1 , wherein extracting, from the input picture, the region comprising the merged mask comprises:

dilating the merged mask to determine a first region extending from a boundary of the merged mask; and

extracting, from the input picture, the first region and the dilated mask.

5. The method of claim 4 , wherein dilating the merged mask by applying a kernel to the merged mask, and a size of the kernel is determined based on at least one of:

a quantization parameter (QP) for compressing the input picture, or

a definition of the input picture.

6. The method of claim 1 , wherein extracting, from the input picture, the region comprising the merged mask comprises:

dilating the merged mask to determine a first region extending from a boundary of the merged mask;

in the input picture, identifying one or more blocks not overlapping with the first region or the merged mask; and

extracting, from the input picture, a region not including the one or more blocks.

7. The method of claim 1 , wherein extracting, from the input picture, the region comprising the merged mask further comprises:

in the extracted region, determining a second region not overlapping with the merged mask; and

blurring the second region.

8. The method of claim 1 , wherein extracting, from the input picture, the region comprising the merged mask further comprises:

determining, based on a quantization parameter (QP) for compressing the input picture, whether to blur at least a part of the extracted region; and

in response to the QP being greater than a predetermined threshold, blurring the part of the extracted region.

9. The method of claim 8 , wherein the part of the extracted region does not overlap with the merged mask.

10. The method of claim 1 , wherein the instance segmentation is performed by a convolutional neural network (CNN).

11. A non-transitory computer readable storage medium storing a bitstream generated by receiving a video sequence, encoding the video sequence to generate coded information included in the bitstream, and transmit the bitstream, wherein the encoding comprises:

performing instance segmentation of an input picture, to generate one or more segment masks;

combining the one or more segment masks to generate a merged mask;

extracting, from the input picture, a region comprising the merged mask; and

compressing image data representing the extracted region, to generate the bitstream, wherein the extracted region comprises one or more pixels adjacent to the merged mask.

12. The medium of claim 11 , wherein extracting, from the input picture, the region comprising the merged mask comprises:

in the input picture, identifying one or more blocks not overlapping with the merged mask; and

extracting, from the input picture, a region not including the one or more blocks.

13. The medium of claim 12 , wherein identifying the one or more blocks using a slide window, and a size of the slide window is determined based on at least one of:

a quantization parameter (QP) for compressing the input picture, or

a definition of the input picture.

14. The medium of claim 11 , wherein extracting, from the input picture, the region comprising the merged mask comprises:

dilating the merged mask to determine a first region extending from a boundary of the merged mask; and

extracting, from the input picture, the first region and the merged mask.

15. The medium of claim 14 , wherein dilating the merged mask by applying a kernel to the merged mask, and a size of the kernel is determined based on at least one of:

a quantization parameter (QP) for compressing the input picture, or

a definition of the input picture.

16. The medium of claim 11 , wherein extracting, from the input picture, the region comprising the merged mask comprises:

dilating the merged mask to determine a first region extending from a boundary of the merged mask;

in the input picture, identifying one or more blocks not overlapping with the first region or the merged mask; and

extracting, from the input picture, a region not including the one or more blocks.

17. The medium of claim 11 , wherein extracting, from the input picture, the region comprising the merged mask further comprises:

in the extracted region, determining a second region not overlapping with the merged mask; and

blurring the second region.

18. An image data processing apparatus, comprising:

a memory storing a set of instructions; and

one or more processors configured to execute the set of instructions to cause the apparatus to perform operations comprising:

performing instance segmentation of an input picture, to generate one or more segment masks;

combining the one or more segment masks to generate a merged mask;

extracting, from the input picture, a region comprising the merged mask; and

compressing image data representing the extracted region, wherein the extracted region comprises one or more pixels adjacent to the merged mask.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA DAMO (HANGZHOU) TECHNOLOGY CO., LTD.
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 075529/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2023
From: WANG, SHURUN; WANG, ZHAO; YE, YAN; WANG, SHIQI
To: ALIBABA DAMO (HANGZHOU) TECHNOLOGY CO., LTD.
Reel/Frame 065247/0502 →
Continuity (2)
Provisional Application 63378888 · Oct 10, 2022
Related Publication 20240121445A1 · Apr 11, 2024
References Cited (28)
CN 103002289A · 2013 [cited by applicant]
CN 109409371A · 2019 [cited by applicant]
WO 2021097595A1 · 2021 [cited by applicant]
WO 2022127865A1 · 2022 [cited by applicant]
Adini et al., “Context-enabled leaming in the human visual system,” Nature, vol. 415, 2002, pp. 790-793. [cited by applicant]
An et al., “Block partitioning structure for next generation video coding,” International Telecommunications Union, 2015, 8 pages. [cited by applicant]
Balle, et al., “Variational Image Compression with a Scale Hyperprior,” International Conference on Learning Representations, 2018, 23 pages. [cited by applicant]
Balle et al., End-to-End Optimized Image Compression, ICLR, 2017, 27 pages. [cited by applicant]
Bross et al., Developments in International Video Coding Standardization After AVC, With an Overview of Versatile Video Coding (VVC), Proceedings of the IEEE, vol. 109, No. 9, Sep. 2021, pp. 1463-1493. [cited by applicant]
Brouard, Olivier, “Pre-analysis of video for its advanced coding. Application to the HDTV coding in H.264 streams,” Université de Nantes, 48 pages, 2010. [cited by applicant]
Chao, et al., “A Novel Rate Control Framework for SIFT/SURF Feature Preservation in H.264/AVC Video Compression,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 25, No. 6, Jun. 2015, pp. 958-972. [cited by applicant]
Duan et al, “Overview of the MPEG-CDVS Standard,” IEEE Transactions on Image Processing, vol. 25, No. 1, Jan. 2016, pp. 179-194. [cited by applicant]
Duan et al., Compact Descriptors for Video Analysis: The Emerging MPEG Standard, IEEE Computer Society, 2018, pp. 44-54. [cited by applicant]
Duan et al., “Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics,” IEEE Transactions on Image Processing, vol. 29, 2020, pp. 8680-8695. [cited by applicant]
Garcia-Lucas et al., “Acceleration of the integer motion estimation in JEM through pre-analysis techniques,” J. Supercomput, 75:1203-1214, 2019. [cited by applicant]
Li et al., “λ Domain Rate Control Algorithm for High Efficiency Video Coding,” IEEE Transactions on Image Processing, vol. 23, No. 9, Sep. 2014, pp. 3841-3854. [cited by applicant]
Liu et al., “CNN-Basesd DCT-Like Transform for Image Compression,” Springer, 2018, pp. 61-72. [cited by applicant]
Minnen et al., “Joint Autoregressive and Hierarchical Priors for Learned Image Compression,” 32 [cited by applicant]
Mohan et al., “Internet of Video Things in 2030: A World with Many Cameras,” IEEE Xplore, 2017, 4 pages. [cited by applicant]
Pfaff et al., “CE3: Affine linear weighted intra prediction (CE3-4.1, CE3-4.2),” JVET-N0217, 14 [cited by applicant]
Rabbani et al., “JPEG2000: Image compression fundamentals, standards and practice,” Journal of Electronic Imaging, 2002, 11(2): 286. [cited by applicant]
Sullivan et al., “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, pp. 1649-1668 (2012). [cited by applicant]
Toderici et al., “Variable Rate Image Compression with Recurrent Neural Networks,” ICLR, 2016, 12 pages. [cited by applicant]
Wallace et al., “The JPEG Still Picture Compression Standard,” IEEE Transactions on Consumer Electronics, vol. 38, No. 1, Feb. 1992, 17 pages. [cited by applicant]
Wang et al., “Attention-Based Dual-Scale CNN In-Loop Filter for Versatile Video Coding,” IEEE Access, vol. 7, 145214-145226, 2019. [cited by applicant]
Yokoyama et al., “A Rate Control Method With Pre-Analysis For Real-Time MPEG-2 Video Coding,” 2001 International Conference on Image Processing, IEEE, pp. 514-517, 2001. [cited by applicant]
Zhao et al., “Mode-dependent non-separable secondary transform,” ITU-T SG16/Q6 Doc. COM16-C1044, 5 pages, 2015. [cited by applicant]
PCT International Search Report and Written Opinion mailed Jan. 13, 2024, issued in corresponding International Application No. PCT/CN2023/123861 (7 pgs.). [cited by applicant]