IP Library Granted Patent US 11,640,714
Granted Patent B2
US 11,640,714 · App. 16/852,647 · Granted May 2, 2023

Video panoptic segmentation

Inventors: Joon-Young Lee (San Jose, CA); Sanghyun Woo (Daejeon, KR); Dahun Kim (Daejeon, KR)
Assignee: ADOBE INC.
G06K9/629G06K9/6256G06K9/6292G06N3/08G06V20/46G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,640,714
App. No.
16/852,647
Granted
May 2, 2023
Kind
B2
Abstract

Systems and methods for panoptic video segmentation are described. A method may include identifying a target frame and a reference frame from a video, generating target features for the target frame and reference features for the reference frame, combining the target features and the reference features to produce fused features for the target frame, generating a feature matrix comprising a correspondence between objects from the reference features and objects from the fused features; and generating panoptic segmentation information for the target frame based on the feature matrix.

Claims (57)

1. A method for image processing, comprising:

identifying a target frame from a first time in a video and a reference frame from a second time of the video, wherein the first time is different than the second time;

generating target features for the target frame from the first time and reference features for the reference frame from the second time;

combining the target features for the target frame from the first time and the reference features for the reference frame from the second time to produce fused features for the target frame;

generating a feature matrix comprising a correspondence between objects from the reference features and objects from the fused features; and

generating panoptic segmentation information for the target frame based on the feature matrix.

2. The method of claim 1 , further comprising:

combining a plurality of input features to produce the target features, wherein each of the plurality of input features has a different resolution, and wherein the target features have a same resolution as the target frame.

3. The method of claim 1 , further comprising:

aligning the reference features with the target features, wherein the fused features are combined based on the aligned reference features.

4. The method of claim 1 , wherein:

the combining of the target features and the reference features comprises applying a spatial-temporal attention module to the target features and the reference features.

5. The method of claim 1 , further comprising:

identifying an object order for the objects from the fused features based on the feature matrix, wherein the panoptic segmentation information is based at least in part on the object order.

6. The method of claim 1 , further comprising:

identifying a bounding box for each of the objects from the fused features.

7. The method of claim 1 , further comprising:

classifying each of the objects from the fused features.

8. The method of claim 1 , further comprising:

generating a pixel mask for each of the objects from the fused features.

9. The method of claim 1 , further comprising:

classifying each pixel of the target frame based on the fused features.

10. The method of claim 1 , wherein:

the panoptic segmentation information comprises classification information and instance information for each pixel of the target frame.

11. The method of claim 1 , wherein:

the panoptic segmentation information is generated based on an object order for the objects from the fused features, an object classification for each of the objects from the fused features, a pixel mask for each of the objects from the fused features, and a pixel classification for each pixel of the target frame.

12. The method of claim 1 , further comprising:

sampling a plurality of frames from the video; and

generating the panoptic segmentation information for each of the plurality of frames.

13. A method for training an artificial neural network (ANN) for video segmentation, comprising:

identifying a training set comprising a plurality of video clips and original panoptic segmentation information for each of the plurality of video clips;

identifying a target frame from a first time and a reference frame from a second time for each of the plurality of video clips, wherein the first time is different than the second time;

generating target features for the target frame from the first time and reference features for the reference frame from the second time;

combining the target features for the target frame from the first time and the reference features for the reference frame from the second time to produce fused features for the target frame;

generating a feature matrix comprising a correspondence between objects from the reference features and objects from the fused features;

generating predicted panoptic segmentation information for the target frame based on the feature matrix;

comparing the predicted panoptic segmentation information to the original panoptic segmentation information; and

updating the ANN based on the comparison.

14. The method of claim 13 , further comprising:

sampling a plurality of frames from each of the plurality of video clips; and

generating panoptic segmentation information for each of the plurality of frames.

15. The method of claim 13 , wherein:

the predicted panoptic segmentation information is generated based on an object order for the objects from the fused features, an object classification for each of the objects from the fused features, a pixel mask for each of the objects from the fused features, and a pixel classification for each pixel of the target frame.

16. The method of claim 13 , further comprising:

applying a spatial-temporal attention module to the target features and the reference features.

17. An apparatus for image processing, comprising:

an encoder configured to generate target features for a target frame from a first time in a video and reference features for a reference frame from a second time of the video, wherein the first time is different than the second time;

a fusion component configured to combine the target features for the target frame from the first time and the reference features for the reference frame from the second time to produce fused features for the target frame, wherein the fusion component comprises a spatial-temporal attention module configured to combine the target features and the reference features;

a track head configured to generate a feature matrix comprising a correspondence between objects from the reference features and objects from the fused features;

a semantic head configured to classify each pixel of the target frame based on the fused features; and

a segmentation component configured to generate panoptic segmentation information for the target frame based on the feature matrix and the classification of each pixel of the target frame.

18. The apparatus of claim 17 , further comprising:

a bounding box head configured to identify a bounding box for each of the objects from the fused features.

19. The apparatus of claim 17 , further comprising:

a mask head configured to generate a pixel mask for each of the objects from the fused features.

20. The apparatus of claim 17 , wherein:

the track head is further configured to classify each object from the fused features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2020
From: LEE, JOON-YOUNG; WOO, SANGHYUN; KIM, DAHUN
To: ADOBE INC.
Reel/Frame 052437/0720 →
Continuity (1)
Related Publication 20210326638A1 · Oct 21, 2021