IP Library › Granted Patent US 11,948,280
Granted Patent B2
US 11,948,280 · App. 16/860,754 · Granted Apr 2, 2024

System and method for multi-frame contextual attention for multi-frame image and video processing using deep neural networks

Inventors: Mostafa El-Khamy (San Diego, CA); Ryan Szeto (Ann Arbor, MI); Jungwon Lee (San Diego, CA)
Assignee: Samsung Electronics Co., Ltd
G06T5/005G06T5/50G06T7/246G06T2207/10016G06T2207/20084G06T2207/20182
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,948,280
App. No.
16/860,754
Granted
Apr 2, 2024
Kind
B2
Abstract

A method and system for multi-frame contextual attention are provided. The method includes obtaining a reference frame to be processed, identifying context frames with respect to the reference frame, and producing a refined reference frame by processing the obtained reference frame based on the context frames.

Claims (27)

1. A method for multi-frame contextual attention, the method comprising:

obtaining a reference frame from a video, the reference frame having an occluded region;

identifying context frames from the video with respect to the reference frame, the context frames comprising a first set of one or more consecutive neighboring frames immediately prior to the reference frame and a second set of one or more consecutive neighboring frames immediately subsequent to the reference frame, wherein a first number of frames in the first set is equal to a second number of frames in the second set, and the first number and the second number have an inversely proportional relationship to a scene changing speed in the video;

producing a coarse prediction of the reference frame using at least the reference frame to fill a hole of the occluded region based on filtering; and

producing a refined reference frame by processing the coarse prediction of the reference frame and the context frames.

2. The method of claim 1 , wherein producing the refined reference frame further includes producing a coarse prediction of each context frame based on convolutional filtering.

3. The method of claim 2 , wherein producing the refined reference frame further includes processing the coarse predictions of each context frame and the coarse prediction of the reference frame in a multi-frame contextual attention stream.

4. The method of claim 1 , wherein producing the refined reference frame further includes processing the coarse prediction of the reference frame and the identified context frames in a multi-frame contextual attention stream.

5. The method of claim 1 , wherein the identified context frames include non-masked regions in feature maps corresponding to multiple frames.

6. The method of claim 1 , further comprising detecting the amount of motion in the video.

7. The method of claim 1 , wherein the reference frame is occluded by an object, and further comprising:

removing the object producing the hole in the reference frame,

wherein processing the coarse prediction of the reference frame and the context frames includes filling the hole in the obtained reference frame.

8. A system for multi-frame contextual attention, the system comprising:

a memory; and

a processor configured to:

obtain a reference frame from a video, the reference frame having an occluded region;

identify context frames from the video with respect to the reference frame, the context frames comprising a first set of one or more consecutive neighboring frames immediately prior to the reference frame and a second set of one or more consecutive neighboring frames immediately subsequent to the reference frame, wherein a first number of frames in the first set is equal to a second number of frames in the second set, and the first number and the second number have an inversely proportional relationship to a scene changing speed in the video; and

produce a coarse prediction of the reference frame using at least the reference frame to fill a hole of the occluded region through filtering;

produce a refined reference frame by processing the coarse prediction of the reference frame and the context frames.

9. The system of claim 8 , wherein producing the refined reference frame further includes producing a coarse prediction of each context frame based on convolutional filtering.

10. The system of claim 9 , wherein the processor is further configured to produce the refined reference frame by processing the coarse predictions of each context frame and the coarse prediction of the reference frame in a multi-frame contextual attention stream.

11. The system of claim 8 , wherein the processor is further configured to produce the refined reference frame by processing the coarse prediction of the reference frame and the identified context frames in a multi-frame contextual attention stream.

12. The system of claim 8 , wherein the identified context frames include non-masked regions in feature maps corresponding to multiple frames.

13. The system of claim 8 , wherein the processor is further configured to detect the amount of motion in the video.

14. The system of claim 8 , wherein the reference frame is occluded by an object, and

wherein the processor is further configured to remove the object producing a hole in the reference frame, and processing the coarse prediction of the reference frame and the context frames includes filling the hole in the obtained reference frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2020
From: EL-KHAMY, MOSTAFA; SZETO, RYAN; LEE, JUNGWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 052989/0700 →
Continuity (2)
Provisional Application 62960867 · Jan 14, 2020
Related Publication 20210217145A1 · Jul 15, 2021