IP Library › Granted Patent US 12,223,719
Granted Patent B2
US 12,223,719 · App. 17/548,824 · Granted Feb 11, 2025

Apparatus and method for prediction of video frame based on deep learning

Inventors: Kun Fan (Seoul, KR); Chung-In Joung (Seoul, KR); Seungjun Baek (Seoul, KR); Seunghwan Byun (Seoul, KR)
Assignee: Korea University Research and Business Foundation
G06V20/46G06N3/08G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,719
App. No.
17/548,824
Granted
Feb 11, 2025
Kind
B2
Abstract

An apparatus and a method of predicting a video frame are provided. The apparatus includes a level encoder configured to extract and learn at least one feature from a video frame, a feature learning unit configured to learn based on the at least one feature or transmit predicted feature data corresponding to the at least one feature, and a level decoder configured to obtain and learn a predicted video frame based on the predicted feature data.

Claims (58)

1. An apparatus for predicting a video frame, the apparatus comprising:

a level encoder configured to extract and learn at least one feature from a video frame;

a feature learning unit configured to learn based on the at least one feature or transmit predicted feature data corresponding to the at least one feature; and

a level decoder configured to obtain and learn a predicted video frame based on the predicted feature data,

wherein the level encoder receives first to (T−1)th video frames, respectively, and extracts at least one feature from each of the first to (T−1)th video frames, where “T” includes a natural number equal to or greater than 2,

wherein the feature learning unit is trained based on at least one feature extracted from each of the first to (T−1)th video frames,

wherein the level encoder receives the T-th video frame,

wherein the level decoder obtains a (T+1)th predicted video frame corresponding to the T th video frame,

wherein the level encoder receives the (T+1)th predicted video frame, and

wherein the level decoder obtains a (T+2)th predicted video frame corresponding to the (T+1)th predicted video frame.

2. The apparatus of claim 1 , wherein the level encoder includes a first-level encoder to an N-th level encoder, and

wherein each of the first-level encoder to the N-th level encoder extracts features of different levels from the video frame where “N” is a natural number equal to or greater than 2.

3. The apparatus of claim 2 , wherein the feature learning unit includes a first feature learning unit to an N-th feature learning unit corresponding to each of the first-level encoder to the N-th level encoder, and

wherein each of the first feature learning unit to the N-th feature learning unit receives each feature of the different levels, obtains and transmit predicted feature data corresponding to each feature of the different levels.

4. The apparatus of claim 3 , wherein the level decoder includes a first-level decoder to an N-th level decoder corresponding to each of the first-level encoder to the N-th level encoder or corresponding to each of the first feature learning unit to the N-th feature learning unit, and

wherein the first-level decoder to the N-th level decoder receive each of the predicted feature data, respectively, and generate a predicted video frame by using the predicted feature data.

5. The apparatus of claim 1 , wherein at least one of the level encoder and the level decoder is based on at least one of a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and a deep belief neural network (DBN).

6. The apparatus of claim 1 , wherein the feature learning unit is based on a long short term memory (LSTM).

7. A method of predicting a video frame, the method comprising:

extracting at least one feature from a video frame;

obtaining predicted feature data trained based on the at least one feature or corresponding to the at least one feature; and

obtaining a predicted video frame based on the predicted feature data,

wherein the extracting of the at least one feature from the video frame includes:

receiving first to (T−1)th video frames, respectively, and extracting at least one feature from each of the first to (T−1)th video frames, where “T” includes a natural number of 2 or more,

wherein the training based on the at least one feature includes:

training based on at least one feature extracted from each of the first to (T−1)th video frames,

wherein the extracting of the at least one feature from the video frame includes:

receiving a T-th video frame and extracting at least one feature from the T-th video frame,

wherein the obtaining of the predicted video frame based on the predicted feature data includes:

obtaining a (T+1) th predicted video frame corresponding to the T th video frame based on the predicted feature data.

8. The method of claim 7 , wherein the extracting of the at least one feature from the video frame includes:

extracting, by the first-level encoder to the N-th level encoder, features of different levels from each other from the video frame, where “N” includes a natural number equal to or greater than 2.

9. The method of claim 8 , wherein the obtaining of the predicted feature data corresponding to the at least one feature includes:

receiving, by a first feature learning unit to an N-th feature learning unit corresponding to the first-level encoder to the N-th level encoder, the features of the different levels, respectively; and

obtaining, by the first feature learning unit to the N-th feature learning unit, the predicted feature data corresponding to each of the features of the different levels, respectively, and transmitting the predicted feature data to a next frame processing process.

10. The method of claim 9 , wherein the obtaining of the predicted video frame based on the predicted feature data includes:

receiving each of the predicted feature data by a first-level decoder to an N-th level decoder, respectively; and

generating, by the first-level decoder to the N-th level decoder, the predicted video frame by using the predicted feature data, and

wherein the first-level decoder to the N-th level decoder correspond to each of the first-level encoder to the N-th level encoder, or corresponding to each of the first feature learning unit to the N-th feature learning unit.

11. The method of claim 7 , further comprising:

receiving the (T+1)th predicted video frame and obtaining a (T+2)th predicted video frame corresponding to the (T+1)th predicted video frame.

12. A method of predicting a video frame, the method comprising:

extracting at least one feature from a video frame;

obtaining predicted feature data trained based on the at least one feature or corresponding to the at least one feature; and

obtaining a predicted video frame based on the predicted feature data,

wherein the extracting of the at least one feature from the video frame includes:

receiving first to (T−1)th video frames, respectively, and extracting at least one feature from each of the first to (T−1)th video frames, where “T” includes a natural number of 2 or more, and

wherein the training based on the at least one feature includes:

training based on at least one feature extracted from each of the first to (T−1)th video frames,

wherein the extracting of the at least one feature from the video frame includes:

extracting, by the first-level encoder to the N-th level encoder, features of different levels from each other from the video frame, where “N” includes a natural number equal to or greater than 2,

wherein the obtaining of the predicted feature data corresponding to the at least one feature includes:

receiving, by a first feature learning unit to an N-th feature learning unit corresponding to the first-level encoder to the N-th level encoder, the features of the different levels, respectively; and

obtaining, by the first feature learning unit to the N-th feature learning unit, the predicted feature data corresponding to each of the features of the different levels, respectively, and transmitting the predicted feature data to a next frame processing process, and

wherein the obtaining of the predicted video frame based on the predicted feature data includes:

receiving each of the predicted feature data by a first-level decoder to an N-th level decoder, respectively; and

generating, by the first-level decoder to the N-th level decoder, the predicted video frame by using the predicted feature data, and

wherein the first-level decoder to the N-th level decoder correspond to each of the first-level encoder to the N-th level encoder, or corresponding to each of the first feature learning unit to the N-th feature learning unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2021
From: FAN, KUN; JOUNG, CHUNG-IN; BAEK, SEUNGJUN; BYUN, SEUNGHWAN
To: KOREA UNIVERSITY RESEARCH AND BUSINESS FOUNDATION
Reel/Frame 058370/0168 →
Priority Claims (2)
KR 10-2020-0173072 · Dec 11, 2020 · national
KR 10-2020-0186716 · Dec 29, 2020 · national
Continuity (1)
Related Publication 20220189171A1 · Jun 16, 2022
References Cited (7)
US 10911775B1 · Zhu · 2021 [cited by examiner]
US 20180253640A1 · Goudarzi · 2018 [cited by examiner]
US 20210064925A1 · Shih · 2021 [cited by examiner]
US 20210168395A1 · Cricri · 2021 [cited by examiner]
Fan, Kun et al. “Sequence-to-Sequence Video Prediction by Learning Hierarchical Representations” Applied Sciences 10, No. 22: 8288. 2020, https://doi.org/10.3390/app10228288, (14 pages in English). [cited by applicant]
Mukherjee, Subham, et al. “Predicting Video-Frames Using Encoder-Convlstm Combination.” ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019. pp. 1-5. [cited by applicant]
Papadomanolaki, Maria, et al. “Detecting Urban Changes With Recurrent Neural Networks From Multitemporal Sentinel-2 Data.” IGARSS 2019-2019 IEEE international geoscience and remote sensing symposium. IEEE, arXiv:1910.07… [cited by applicant]