IP Library Granted Patent US 12,003,728
Granted Patent B2
US 12,003,728 · App. 17/816,825 · Granted Jun 4, 2024

Methods and systems for temporal resampling for multi-task machine vision

Inventors: Shurun Wang (Beijing, CN); Zhao Wang (Beijing, CN); Yan Ye (San Diego, CA); Shiqi Wang (Kowloon Tong, HK)
Assignee: Alibaba Innovation Private Limited
H04N19/132H04N19/137H04N19/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,003,728
App. No.
17/816,825
Granted
Jun 4, 2024
Kind
B2
Abstract

A method for temporal resampling for multi-task machine vision is provided. The method includes receiving a bitstream of a video sequence after temporal resampling; and constructing a target frame from the bitstream using a frame construction model.

Claims (40)

1. A method for temporal resampling for multi-task machine vision, comprising:

receiving a bitstream of a video sequence after temporal resampling; and

constructing a target frame from the bitstream using a frame construction model, wherein an input of the frame construction model is a channel-wise concatenation of an average residual of the target frame and a prediction of the target frame, the average residual of the target frame is obtained based on resampled frames and a resampling factor.

2. The method according to claim 1 , wherein the average residual is obtained using the resampling factor and a frame from the video sequence.

3. The method according to claim 1 , wherein the prediction is obtained using the average residual.

4. The method according to claim 1 , wherein the frame construction model is trained with feature maps, the feature maps being extracted from corresponding one or more machine analysis models for the target frame and the constructed frame of the target frame, respectively.

5. A method for temporal resampling for multi-task machine vision, comprising:

receiving a bitstream of a video sequence after temporal resampling; and

constructing a target frame from the bitstream using a frame construction model; wherein the frame construction model is trained with feature maps, the feature maps being extracted from corresponding one or more machine analysis models for the target frame and the constructed frame of the target frame, respectively, and one or more loss functions of corresponding one or more machine analysis models are obtained with label information.

6. The method according to claim 5 , wherein a total loss function is determined using contour loss, feature map distortion of the one or more feature maps, and the one or more loss functions.

7. The method according to claim 6 , wherein the loss function of each of the one or more machine analysis model is determined by a mean squared difference between a machine analysis model feature map of the target frame and a machine analysis model feature map of the constructed frame of the machine analysis model.

8. An apparatus for temporal resampling for multi-task machine vision, the apparatus comprising:

a memory figured to store instructions; and

one or more processors configured to execute the instructions to cause the apparatus to perform:

receiving a bitstream of a video sequence after temporal resampling; and

constructing a target frame from the bitstream using a frame construction model wherein an input of the frame construction model is a channel-wise concatenation of an average residual of the target frame and a prediction of the target frame, the average residual of the target frame is obtained based on resampled frames and a resampling factor.

9. The apparatus according to claim 8 , wherein the average residual is obtained using the resampling factor and a frame from the video sequence.

10. The apparatus according to claim 8 , wherein the prediction is obtained using the average residual.

11. The apparatus according to claim 8 , wherein the frame construction model is trained with feature maps, the feature maps being extracted from corresponding one or more machine analysis models for the target frame and the constructed frame of the target frame, respectively.

12. An apparatus for temporal resampling for multi-task machine vision, the apparatus comprising:

a memory configured to store instructions; and

one or more processors configured to execute the instructions to cause the apparatus to perform:

receiving a bitstream of a video sequence after temporal resampling; and

constructing a target frame from the bitstream using a frame construction model; wherein the frame construction model is trained with feature maps, the feature maps being extracted from corresponding one or more machine analysis models for the target frame and the constructed frame of the target frame, respectively; and one or more loss functions of corresponding one or more machine analysis models are obtained with label information.

13. The apparatus according to claim 12 , wherein a total loss function is determined using contour loss, feature map distortion of the one or more feature maps, and the one or more loss functions.

14. A method for temporal resampling for multi-task machine vision, comprising:

receiving a video sequence;

determining a moving complexity of the video sequence, wherein determining the moving complexity further comprises:

calculating an average mean absolute difference (MAD) for a testing point with a range;

sampling all the testing points across the video sequence equally with an interval; and

obtaining the moving complexity of the video sequence based on the average MAD and the interval; and

determining whether to perform temporal resampling based on the moving complexity.

15. The method according to claim 14 , wherein determining whether to perform the temporal resampling based on the moving complexity further comprises:

not performing the temporal resampling when the moving complexity is greater than or equal to a first threshold; or

performing the temporal resampling when the moving complexity is less than a second threshold.

16. The method according to claim 15 , when the moving complexity is less than the first threshold and greater than or equal to the second threshold, the method further comprises:

determining a quantization parameter associated with the video sequence; and

determining whether to perform the temporal resampling based on the quantization parameter.

17. The method according to claim 16 , wherein determining whether to perform the temporal resampling based on the quantization parameter further comprises:

performing the temporal resampling when the quantization parameter is equal to or greater than a preset value.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2024
From: ALIBABA SINGAPORE HOLDING PRIVATE LIMITED
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 066348/0252 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2022
From: WANG, SHURUN; WANG, ZHAO; YE, YAN; WANG, SHIQI
To: ALIBABA SINGAPORE HOLDING PRIVATE LIMITED
Reel/Frame 061191/0851 →
Continuity (1)
Related Publication 20240048709A1 · Feb 8, 2024