IP Library Granted Patent US 12,387,097
Granted Patent B2
US 12,387,097 · App. 18/111,756 · Granted Aug 12, 2025

Efficient video processing via temporal progressive learning

Inventors: Peng Wang (Los Angleles, CA); Heng Wang (Los Angeles, CA); Xianhang Li (Los Angeles, CA); Xinyu Li (Los Angeles, CA)
Assignee: Lemon Inc.
G06N3/08G06N3/0455G06V10/751G06V10/771G06V10/7715G06V10/82G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,097
App. No.
18/111,756
Granted
Aug 12, 2025
Kind
B2
Abstract

Systems and methods for performing temporal progressive learning for video processing are provided herein. Some examples include receiving a video that includes a plurality of frames, extracting a first subset of frames from the plurality of frames, and inputting the first subset of frames into a model that includes an encoder and a decoder. The examples further include comparing a first output of the model to the first subset of frames and updating the encoder, thereby training the encoder, and extracting a second subset of frames from the plurality of frames. The second subset of frames includes a number of frames that is larger than a number of frames in the first subset of frames. The examples further include inputting the second subset of frames into the model, comparing a second output of the model to the second subset of frames and updating the encoder, thereby further training the encoder.

Claims (52)

1. A method of performing temporal progressive learning for video processing, the method comprising:

receiving a video stream comprising a plurality of frames;

extracting a first subset of frames from the plurality of frames;

inputting the first subset of frames into a model, wherein the model includes an encoder and a decoder;

comparing a first output of the model to the first subset of frames and updating the encoder based on the comparison, thereby training the encoder;

extracting a second subset of frames from the plurality of frames, the second subset of frames comprising a number of frames that is larger than a number of frames in the first subset of frames;

inputting the second subset of frames into the model;

comparing a second output of the model to the second subset of frames and updating the encoder based on the comparison, thereby further training the encoder; and

providing the encoder.

2. The method of claim 1 , wherein the model is a masked auto encoder (MAE) model.

3. The method of claim 2 , wherein each frame in the first and second subsets of frames are randomly masked, prior to being input into the MAE model.

4. The method of claim 1 , wherein the first and second subsets of frames are randomly selected from the plurality of frames.

5. The method of claim 1 , wherein the second subset of frames comprises twice as many frames as the first subset of frames.

6. The method of claim 1 , further comprising, prior to providing the model:

extracting a third subset of frames from the plurality of frames, the third subset of frames being randomly selected from the plurality of frames and comprising a number of frames that is larger than the number of frames in the second subset of frames;

inputting the third subset of frames into the model; and

comparing a third output of the model to the third subset of frames and updating the encoder based on the comparison, thereby further training the encoder.

7. The method of claim 6 , wherein the third subset of frames comprises twice as many frames as the second subset of frames.

8. The method of claim 1 , wherein each sequence of the extracting, the inputting, and the comparing define a respective stage, and wherein the number of frames extracted in the subset of frames for each stage are determined based on a total number of stages and a computational budget.

9. A system for performing temporal progressive learning for video processing, the system comprising:

a processor;

memory storing instructions that, when executed by the processor, cause the system to perform a set of operations, the set of operations comprising:

receiving a video stream comprising a plurality of frames;

extracting a first subset of frames from the plurality of frames;

inputting the first subset of frames into a model, wherein the model includes an encoder and a decoder;

comparing a first output of the model to the first subset of frames and updating the encoder based on the comparison, thereby training the encoder;

extracting a second subset of frames from the plurality of frames, the second subset of frames comprising a number of frames that is larger than a number of frames in the first subset of frames;

inputting the second subset of frames into the model;

comparing a second output of the model to the second subset of frames and updating the encoder based on the comparison, thereby further training the encoder; and

providing the encoder.

10. The system of claim 9 , wherein the model is a masked auto encoder (MAE) model.

11. The system of claim 10 , wherein each frame in the first and second subsets of frames are randomly masked, prior to being input into the MAE model.

12. The system of claim 9 , wherein the first and second subsets of frames are randomly selected from the plurality of frames.

13. The system of claim 9 , wherein the second subset of frames comprises twice as many frames as the first subset of frames.

14. The system of claim 9 , further comprising, prior to providing the model:

extracting a third subset of frames from the plurality of frames, the third subset of frames being randomly selected from the plurality of frames and comprising a number of frames that is larger than the number of frames in the second subset of frames;

inputting the third subset of frames into the model; and

comparing a third output of the model to the third subset of frames and updating the encoder based on the comparison, thereby further training the encoder.

15. The system of claim 14 , wherein the third subset of frames comprises twice as many frames as the second subset of frames.

16. The system of claim 9 , wherein each sequence of the extracting, the inputting, and the comparing define a respective stage, and wherein the number of frames extracted in the subset of frames for each stage are determined based on a total number of stages and a computational budget.

17. One or more computer readable non-transitory storage media embodying software that is operable when executed, by at least one processor of a device, to:

receive a video stream comprising a plurality of frames;

extract a first subset of frames from the plurality of frames;

input the first subset of frames into a model, wherein the model includes an encoder and a decoder;

compare a first output of the model to the first subset of frames and updating the encoder based on the comparison, thereby training the encoder;

extract a second subset of frames from the plurality of frames, the second subset of frames comprising a number of frames that is larger than a number of frames in the first subset of frames;

input the second subset of frames into the model;

compare a second output of the model to the second subset of frames and updating the encoder based on the comparison, thereby further training the encoder; and

provide the encoder.

18. The one or more computer readable non-transitory storage media of claim 17 , wherein the model is a masked auto encoder (MAE) model.

19. The one or more computer readable non-transitory storage media of claim 18 , wherein each frame in the first and second subsets of frames are randomly masked, prior to being input into the MAE model.

20. The one or more computer readable non-transitory storage media of claim 17 , wherein the first and second subsets of frames are randomly selected from the plurality of frames.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2026
From: WANG, PENG; WANG, HENG; LI, XIANHANG; LI, XINYU
To: BYTEDANCE INC.
Reel/Frame 074085/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2026
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 074085/0375 →
Continuity (1)
Related Publication 20230206067A1 · Jun 29, 2023
References Cited (43)
US 11238093B2 · Ayush · 2022 [cited by examiner]
US 11386680B2 · Maity · 2022 [cited by examiner]
US 11967121B2 · Takagi · 2024 [cited by examiner]
US 12045728B2 · Ketyko · 2024 [cited by examiner]
US 20180174050A1 · Holt · 2018 [cited by examiner]
US 20190108399A1 · Escorcia · 2019 [cited by examiner]
US 20200342643A1 · Gouws · 2020 [cited by examiner]
US 20210004677A1 · Menick · 2021 [cited by examiner]
US 20210081673A1 · Lai · 2021 [cited by examiner]
US 20210201460A1 · Gong · 2021 [cited by examiner]
US 20210295160A1 · Park · 2021 [cited by examiner]
US 20210303935A1 · Ma · 2021 [cited by examiner]
US 20220108136A1 · Son · 2022 [cited by examiner]
US 20220156584A1 · Fersch · 2022 [cited by examiner]
US 20220180189A1 · Adrian · 2022 [cited by examiner]
US 20220180199A1 · Xu · 2022 [cited by examiner]
US 20220245932A1 · Gauerhof · 2022 [cited by examiner]
US 20220309342A1 · Borgohain · 2022 [cited by examiner]
US 20230005251A1 · Take · 2023 [cited by examiner]
US 20230046066A1 · Bulat · 2023 [cited by examiner]
US 20230061517A1 · Yang · 2023 [cited by examiner]
US 20230082050A1 · Li · 2023 [cited by examiner]
US 20230177639A1 · Chee · 2023 [cited by examiner]
US 20230262237A1 · Mitra · 2023 [cited by examiner]
US 20230274527A1 · Chen · 2023 [cited by examiner]
US 20230298219A1 · Galpin · 2023 [cited by examiner]
US 20230368520A1 · Goldin · 2023 [cited by examiner]
US 20230412796A1 · Liu · 2023 [cited by examiner]
US 20240064318A1 · Letunovskiy · 2024 [cited by examiner]
US 20240095878A1 · Chen · 2024 [cited by examiner]
US 20240127794A1 · Seo · 2024 [cited by examiner]
US 20240161445A1 · Fujiwaka · 2024 [cited by examiner]
US 20240203123A1 · Dimitriou · 2024 [cited by examiner]
US 20240220856A1 · Schreiber · 2024 [cited by examiner]
US 20240267542A1 · Henry · 2024 [cited by examiner]
US 20240273944A1 · Chen · 2024 [cited by examiner]
US 20240312208A1 · Ghosh · 2024 [cited by examiner]
US 20240320971A1 · Thumpudi · 2024 [cited by examiner]
US 20240338572A1 · Gillian · 2024 [cited by examiner]
US 20250046071A1 · Goldin · 2025 [cited by examiner]
US 20250086471A1 · Bubeck · 2025 [cited by examiner]
US 20250104210A1 · Pu · 2025 [cited by examiner]
US 20250118434A1 · Lee · 2025 [cited by examiner]