IP Library Granted Patent US 12,664,607
Granted Patent B2
US 12,664,607 · App. 18/658,694 · Granted Jun 23, 2026

Method and apparatus for generating video intermediate frame

Inventors: Xin Jin (Nanjing, CN); Ban Chen (Nanjing, CN); Longhai Wu (Nanjing, CN); Jie Chen (Nanjing, CN); Jayoon Koo (Suwon-si, KR); Cheulhee Hahm (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T3/18G06V10/761G06V10/82G06V20/48
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,607
App. No.
18/658,694
Filed
May 8, 2024
Granted
Jun 23, 2026
Kind
B2
Art Unit
2667
USPC
382/100
Abstract

A method for generating a video intermediate frame includes performing a warp operation on a plurality of target video frames, based on a bidirectional optical flow between the plurality of target video frames, to obtain a plurality of pictures, determining a similarity between sub-pictures corresponding to image sub-regions in the plurality of pictures, predicting, based on the similarity, a network depth of a frame synthesis network matched with a corresponding image sub-region, the network depth increasing as the similarity decreases, performing synthetic processing, using the frame synthesis network, on corresponding sub-pictures of each of the image sub-regions, based on the network depth matched with the corresponding image sub-region, to obtain a plurality of images, and splicing the plurality of images to obtain intermediate frames of the plurality of target video frames, based on each of the image sub-regions.

Claims (50)

1 . A method for generating a video intermediate frame, the method comprising:

performing a warp operation on a plurality of target video frames, based on a bidirectional optical flow between the plurality of target video frames, to obtain a plurality of pictures;

determining a similarity between sub-pictures corresponding to image sub-regions in the plurality of pictures;

predicting, based on the similarity, a network depth of a frame synthesis network matched with a corresponding image sub-region, the network depth increasing as the similarity decreases;

performing synthetic processing on corresponding sub-pictures of each of the image sub-regions using the frame synthesis network having the network depth matched with the corresponding image sub-region, to obtain a plurality of images; and

splicing the plurality of images to obtain intermediate frames of the plurality of target video frames, based on each of the image sub-regions.

2 . The method of claim 1 , wherein the performing of the warp operation comprises:

performing a forward-warping operation on the plurality of target video frames.

3 . The method of claim 1 , wherein a size of each of the image sub-regions is less than or equal to 64×64 dots-per-inch (dpi).

4 . The method of claim 1 , wherein the predicting of the network depth comprises:

predicting the network depth of the frame synthesis network based on the plurality of pictures and using a fully convolutional network (FCN).

5 . The method of claim 1 , wherein the frame synthesis network comprises a U-Net network.

6 . The method of claim 1 , wherein the performing of the synthetic processing comprises:

performing synthetic processing on the corresponding sub-pictures of a first image sub-region of the image sub-regions using the frame synthesis network with a first network depth matched to the first image sub-region; and

performing synthetic processing on the corresponding sub-pictures of a second image sub-region of the image sub-regions using the frame synthesis network with a second network depth matched to the second image sub-region,

wherein a first similarity of the corresponding sub-pictures of the first image sub-region is greater than a second similarity of the corresponding sub-pictures of the second image sub-region, and

wherein the first network depth is less than the second network depth.

7 . The method of claim 1 , wherein the performing of the warp operation comprises performing the warp operation using an optical flow estimation network,

wherein the predicting of the network depth comprises predicting the network depth of the frame synthesis network using a fully convolutional network (FCN), and

wherein the method further comprises integrating at least one of the optical flow estimation network, the FCN, and the frame synthesis network into an intermediate frame generation model.

8 . The method of claim 7 , further comprising:

training the intermediate frame generation model,

wherein the training of the intermediate frame generation model comprises:

providing sample video frames as input data to the intermediate frame generation model;

generating intermediate frames corresponding to the sample video frames;

calculating loss function values based on the intermediate frames corresponding to the sample video frames; and

adjusting, based on the loss function values, at least one network parameter of at least one of the optical flow estimation network, the FCN, and the frame synthesis network.

9 . A device for generating a video intermediate frame, comprising:

one or more processors;

a memory storing instructions that, when executed by the one or more processors, cause the device to:

perform a warp operation on a plurality of target video frames, based on a bidirectional optical flow between the plurality of target video frames, to obtain a plurality of pictures;

determine a similarity between sub-pictures corresponding to image sub-regions in the plurality of pictures;

predict, based on the similarity, a network depth of a frame synthesis network matched with a corresponding image sub-region, the network depth increasing as the similarity decreases;

perform synthetic processing on corresponding sub-pictures of each of the image sub-regions using the frame synthesis network having the network depth matched with the corresponding image sub-region, to obtain a plurality of images; and

splice the plurality of images to obtain intermediate frames of the plurality of target video frames, based on each of the image sub-regions.

10 . The device of claim 9 , wherein the instructions, when executed by the one or more processors, further cause the device to:

perform a forward-warping operation on the plurality of target video frames.

11 . The device of claim 9 , wherein a size of each of the image sub-regions is less than or equal to 64×64 dots-per-inch (dpi).

12 . The device of claim 9 , wherein the instructions, when executed by the one or more processors, further cause the device to:

predict the network depth of the frame synthesis network based on the plurality of pictures and using a fully convolutional network (FCN).

13 . The device of claim 9 , wherein the frame synthesis network comprises a U-Net network.

14 . The device of claim 9 , wherein the instructions, when executed by the one or more processors, further cause the device to:

perform synthetic processing on the corresponding sub-pictures of a first image sub-region of the image sub-regions using the frame synthesis network with a first network depth matched to the first image sub-region; and

perform synthetic processing on the corresponding sub-pictures of a second image sub-region of the image sub-regions using the frame synthesis network with a second network depth matched to the second image sub-region,

wherein a first similarity of the corresponding sub-pictures of the first image sub-region is greater than a second similarity of the corresponding sub-pictures of the second image sub-region, and

wherein the first network depth is less than the second network depth.

15 . The device of claim 9 , wherein the instructions, when executed by the one or more processors, further cause the device to:

perform the warp operation using an optical flow estimation network;

predict the network depth of the frame synthesis network using a fully convolutional network (FCN), and

integrate at least one of the optical flow estimation network, the FCN, and the frame synthesis network into an intermediate frame generation model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: JIN, XIN; CHEN, BAN; WU, LONGHAI; CHEN, JIE; KOO, JAYOON; HAHM, CHEULHEE
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 067365/0995 →
Priority Claims (1)
CN 202310745901.2 · Jun 21, 2023 · national
Continuity (2)
Continuation PCTIB2024052668 · Mar 20, 2024
Related Publication 20240428371A1 · Dec 26, 2024
References Cited (36)
US 10812825B2 · Liu · 2020 [cited by examiner]
US 11122238B1 · van Amersfoort et al. · 2021 [cited by applicant]
US 11363271B2 · Li et al. · 2022 [cited by applicant]
US 12185023B2 · Jin et al. · 2024 [cited by applicant]
US 20020048395A1 · Harman · 2002 [cited by examiner]
US 20190138889A1 · Jiang et al. · 2019 [cited by applicant]
US 20200035362A1 · Abou Shousha · 2020 [cited by examiner]
US 20210281867A1 · Golinski · 2021 [cited by examiner]
US 20220092795A1 · Liu et al. · 2022 [cited by applicant]
US 20220284552A1 · Yang et al. · 2022 [cited by applicant]
US 20220375030A1 · Choi et al. · 2022 [cited by applicant]
US 20220383573A1 · Schroers et al. · 2022 [cited by applicant]
US 20230007240A1 · Li et al. · 2023 [cited by applicant]
US 20230214458A1 · Marsden · 2023 [cited by examiner]
US 20230345020A1 · Yu · 2023 [cited by examiner]
US 20240029196A1 · Barragán del Rey · 2024 [cited by examiner]
US 20240146868A1 · Zhang et al. · 2024 [cited by applicant]
CN 114071223A · 2022 [cited by applicant]
CN 112104830B · 2022 [cited by applicant]
CN 115065796A · 2022 [cited by applicant]
KR 102242343B1 · 2021 [cited by applicant]
KR 102467673B1 · 2022 [cited by applicant]
KR 1020220158598A · 2022 [cited by applicant]
WO 2021085757A1 · 2021 [cited by applicant]
WO 2022099313A1 · 2022 [cited by applicant]
WO 2022267957A1 · 2022 [cited by applicant]
International Search Report (PCT/ISA/210) issued on Jun. 19, 2024 by the International Searching Authority in International Application No. PCT/IB2024/052668. [cited by applicant]
Written Opinion (PCT/ISA/237) issued on Jun. 19, 2024 by the International Searching Authority in International Application No. PCT/IB2024/052668. [cited by applicant]
Niklaus, Simon et al., “Softmax Splatting for Video Frame Interpolation”, CVPR, 2020, pp. 5437-5446. (10 pages total). [cited by applicant]
Choi, Myungsub et al., “Channel Attention Is All You Need for Video Frame Interpolation”, The Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-20), 2020, pp. 10663-10671. (9 pages total). [cited by applicant]
Choi, Myungsub et al., “Motion-Aware Dynamic Architecture for Efficient Frame Interpolation”, ICCV, 2021, pp. 13839-13848. (10 pages total). [cited by applicant]
Xue, Tianfan et al., “Video Enhancement with Task-Oriented Flow”, International Journal of Computer Vision, arXiv:1711.09078v3 [cs.CV], Nov. 10, 2019. (20 pages total). [cited by applicant]
Lee, Hyeongmin et al., “AdaCof: Adaptive Collaboration of Flows for Video Frame Interpolation”, CVPR, 2020, pp. 5316-5325. (10 pages total). [cited by applicant]
Park, Junheum et al., “Asymmetric Bilateral Motion Estimation for Video Frame Interpolation”, ICCV, 2021, pp. 14539-14548. (10 pages total). [cited by applicant]
Communication dated Feb. 3, 2026 issued by the European Patent Office in European Patent Application No. 24825403.9. [cited by applicant]
Yizeng Han et al., “Dynamic Neural Networks: A Survey”, arXiv:2102.04906v2 [cs.CV], Feb. 10, 2021, XP 081877781, pp. 1-20 (20 pages total). [cited by applicant]