IP Library › Granted Patent US 12,739,385
Granted Patent B2
US 12,739,385 · App. 18/618,551 · Granted Sep 15, 2026

Methods and non-transitory computer readable storage medium for spatial resampling towards machine vision

Inventors: Shurun Wang (Beijing, CN); Yan Ye (San Diego, CA)
Assignee: Alibaba Innovation Private Limited
H04N19/132H04N19/172H04N19/186H04N19/436
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,385
App. No.
18/618,551
Granted
Sep 15, 2026
Kind
B2
Abstract

A method of encoding a video sequence into a bitstream. The method includes receiving a video sequence; performing a plurality of convolutions on an input image data of the video sequence in YUV format; wherein performing the plurality of convolutions includes performing a first stage convolution on the input image data, wherein the first stage convolution comprises a first convolution and a second convolution that are provided in parallel; performing a second stage convolution on a channel-wise concatenation result of an output of the first convolution and an output of the second convolution; performing a third stage convolution on an output of the second stage convolution; and obtaining an output image data based on an output of the third stage convolution; and encoding the output image data for generating the bitstream.

Claims (63)

1 . A method of decoding a bitstream to output one or more pictures for a video stream, the method comprising:

receiving a bitstream; and

decoding, using coded information of the bitstream, one or more pictures comprising a down-sampled image data in YUV format; and

performing a plurality of convolutions on the down-sampled image data, wherein performing the plurality of convolutions comprises:

performing a first stage convolution on the down-sampled image data, wherein the first stage convolution comprises a first convolution and a second convolution provided in parallel;

performing a second stage convolution on a channel-wise concatenation result of an output of the first convolution and an output of the second convolution;

performing a third stage convolution on an output of the second stage convolution, wherein the third stage convolution comprises a fourth convolution and a fifth convolution provided in parallel;

performing a bicubic interpolation on the down-sampled image data to obtain a bicubic interpolation result; and

performing an element-wise addition to an output of third stage convolution and the bicubic interpolation result to obtain an up-sampled image data;

wherein performing the element-wise addition to the output of third stage convolution and the bicubic interpolation result to obtain the up-sampled image data, further comprises:

performing a first element-wise addition to an output of the fourth convolution and a first bicubic interpolation result of a Y component of the down-sampled image data to obtain a Y component of the up-sampled image data; and

performing a second element-wise addition to an output of the fifth convolution and a second bicubic interpolation result of a channel-wise concatenation result of a U component and a V component of the down-sampled image data to obtain a channel-wise concatenation result of a U component and a V component of the up-sampled image data.

2 . The method according to claim 1 , wherein performing the first stage convolution on the image data comprises:

performing the first convolution on a Y component of the down-sampled image data; and

performing the second convolution on a channel-wise concatenation result of a U component and a V component of the down-sampled image data.

3 . The method according to claim 1 , wherein the second stage convolution comprises a series of convolutions.

4 . The method according to claim 1 , wherein performing the third stage convolution on the output of the second stage convolution further comprises:

performing the fourth convolution on the output of the second stage convolution; and

performing the fifth convolution on the output of the second stage convolution.

5 . The method according to claim 1 , wherein a set of parameters of each of the plurality of convolutions comprises: an input channel number, an output channel number, a kernel size, a stride, and a padding size.

6 . The method according to claim 5 , wherein the set of parameters is determined based on a type of a YUV format and a number of the plurality of convolutions.

7 . The method according to claim 1 , wherein a Rectified Linear Unit (ReLU) is applied to each convolution in the first stage convolution and the second stage convolution as an activation function.

8 . An apparatus for image data processing, comprising:

a memory configured to store instructions; and

one or more processors configured to execute the instructions to cause the apparatus to perform operations comprising:

performing a plurality of convolutions on a down-sampled image data in YUV format, wherein performing the plurality of convolutions comprises:

performing a first stage convolution on the down-sampled image data, wherein the first stage convolution comprises a first convolution and a second convolution provided in parallel;

performing a second stage convolution on a channel-wise concatenation result of an output of the first convolution and an output of the second convolution;

performing a third stage convolution on an output of the second stage convolution, wherein the third stage convolution comprises a fourth convolution and a fifth convolution provided in parallel;

performing a bicubic interpolation on the down-sampled image data to obtain a bicubic interpolation result; and

performing an element-wise addition to an output of third stage convolution and the bicubic interpolation result to obtain an up-sampled image data;

wherein performing the element-wise addition to the output of third stage convolution and the bicubic interpolation result to obtain the up-sampled image data, further comprises:

performing a first element-wise addition to an output of the fourth convolution and a first bicubic interpolation result of a Y component of the down-sampled image data to obtain a Y component of the up-sampled image data; and

performing a second element-wise addition to an output of the fifth convolution and a second bicubic interpolation result of a channel-wise concatenation result of a U component and a V component of the down-sampled image data to obtain a channel-wise concatenation result of a U component and a V component of the up-sampled image data.

9 . The apparatus according to claim 8 , wherein performing the first stage convolution on the down-sampled image data comprises:

performing the first convolution on a Y component of the down-sampled image data; and

performing the second convolution on a channel-wise concatenation result of a U component and a V component of the down-sampled image data.

10 . The apparatus according to claim 8 , wherein the second stage convolution comprises a series of convolutions.

11 . The apparatus according to claim 8 , wherein performing the third stage convolution on the output of the second stage convolution further comprises:

performing the fourth convolution on the output of the second stage convolution; and

performing the fifth convolution on the output of the second stage convolution.

12 . The apparatus according to claim 8 , wherein a set of parameters of each of the plurality of convolutions comprises: an input channel number, an output channel number, a kernel size, a stride, and a padding size.

13 . The apparatus according to claim 12 , wherein the set of parameters is determined based on a type of a YUV format and a number of the plurality of convolutions.

14 . The apparatus according to claim 8 , wherein a Rectified Linear Unit (ReLU) is applied to each convolution in the first stage convolution and the second stage convolution as an activation function.

15 . A non-transitory computer readable medium that stores a set of instructions that is executable by one or more processors of an apparatus to cause the apparatus to perform operations comprising:

performing a plurality of convolutions on a down-sampled image data in YUV format, wherein performing the plurality of convolutions comprises:

performing a first stage convolution on the down-sampled image data, wherein the first stage convolution comprises a first convolution and a second convolution provided in parallel;

performing a second stage convolution on a channel-wise concatenation result of an output of the first convolution and an output of the second convolution;

performing a third stage convolution on an output of the second stage convolution, wherein the third stage convolution comprises a fourth convolution and a fifth convolution provided in parallel;

performing a bicubic interpolation on the down-sampled image data to obtain a bicubic interpolation result; and

performing an element-wise addition to an output of third stage convolution and the bicubic interpolation result to obtain an up-sampled image data;

wherein performing the element-wise addition to the output of third stage convolution and the bicubic interpolation result to obtain the up-sampled image data, further comprises:

performing a first element-wise addition to an output of the fourth convolution and a first bicubic interpolation result of a Y component of the down-sampled image data to obtain a Y component of the up-sampled image data; and

performing a second element-wise addition to an output of the fifth convolution and a second bicubic interpolation result of a channel-wise concatenation result of a U component and a V component of the down-sampled image data to obtain a channel-wise concatenation result of a U component and a V component of the up-sampled image data.

16 . The non-transitory computer readable medium according to claim 15 , wherein performing the first stage convolution on the down-sampled image data comprises:

performing the first convolution on a Y component of the down-sampled image data; and

performing the second convolution on a channel-wise concatenation result of a U component and a V component of the down-sampled image data.

17 . The non-transitory computer readable medium according to claim 15 , wherein the second stage convolution comprises a series of convolutions.

18 . The non-transitory computer readable medium according to claim 15 , wherein performing the third stage convolution on the output of the second stage convolution further comprises:

performing the fourth convolution on the output of the second stage convolution; and

performing the fifth convolution on the output of the second stage convolution.

19 . The non-transitory computer readable medium according to claim 15 , wherein a set of parameters of each of the plurality of convolutions comprises: an input channel number, an output channel number, a kernel size, a stride, and a padding size.

20 . The non-transitory computer readable medium according to claim 19 , wherein the set of parameters is determined based on a type of a YUV format and a number of the plurality of convolutions.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
CHANGE OF NAME Recorded Aug 5, 2026
From: SIM IP 5 LLC
To: VDT IP PROTECTION LLC
Reel/Frame 076110/0492 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2024
From: WANG, SHUNRUN; YE, YAN
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 068696/0819 →
Continuity (2)
Provisional Application 63495369 · Apr 11, 2023
Related Publication 20240357118A1 · Oct 24, 2024
References Cited (18)
US 20220086463A1 · Coban et al. · 2022 [cited by applicant]
US 20230116285A1 · Cui · 2023 [cited by examiner]
US 20240007658A1 · Mao · 2024 [cited by examiner]
US 20250259275A1 · Xie · 2025 [cited by examiner]
CN 111145290A · 2020 [cited by applicant]
CN 113962882A · 2022 [cited by applicant]
CN 114781622A · 2022 [cited by applicant]
KR 20220077893A · 2022 [cited by applicant]
WO 2022126120A1 · 2022 [cited by applicant]
Gao et al., “Response to VCM Call for Proposals from Tencent—an End-to-end Learning based Solution,” International Organisation for Standardisation Organisation Internationale de Normalisation ISO/IEC JTC 1/SC 29/WG 2 M… [cited by applicant]
Jiang et al., “An End-to-End Compression Framework Based on Convolutional Neural Networks,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, No. 10, Oct. 2018, pp. 3007-3018. [cited by applicant]
Lin et al., “Adaptive Downsampling to Improve Image Compression at Low Bit Rates,” IEEE Transactions on Image Processing, vol. 15, No. 9, Sep. 2006, pp. 2513-2521. [cited by applicant]
Liu et al., “Compressive Sampling-Based Image Coding for Resource-Deficient Visual Communication,” IEEE Transactions on Image Processing, vol. 25, No. 6, Jun. 2016, pp. 2844-2855. [cited by applicant]
Liu et al., “[VCM] Response to VCM Call for Proposals from Tencent and Wuhan University—an ECM based solution,” International Organisation for Standardisation Organisation Internationale de Normalisation ISO/IEC JTC 1/S… [cited by applicant]
Liu et al. “[VCM] response to VCM call for proposals—an EVC based solution,” International Organisation for Standardisation Organisation Internationale de Normalisation ISO/IEC JTC 1/SC 29/WG 2 MPEG Technical Requiremen… [cited by applicant]
Sun et al., “Learned Image Downscaling for Upscaling Using Content Adaptive Resampler,” IEEE Transactions on Image Processing, vol. 29, pp. 4027-4040, 2020. [cited by applicant]
Wang et al., “[VCM] Video Coding for Machines CfP Response from Alibaba and City University of Hong Kong,” International Organisation for Standardisation Organisation Internationale de Normalisation ISO/IEC JTC 1/SC 29/… [cited by applicant]
PCT International Search Report and Written Opinion mailed Jun. 24, 2024, issued in corresponding International Application No. PCT/CN2024/087334 (7 pgs.). [cited by applicant]