IP Library › Granted Patent US 11,741,611
Granted Patent B2
US 11,741,611 · App. 17/584,988 · Granted Aug 29, 2023

Cyclical object segmentation neural networks

Inventor: Ning Xu (San Jose, CA)
Assignee: Adobe Inc.
G06T7/10G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,741,611
App. No.
17/584,988
Granted
Aug 29, 2023
Kind
B2
Abstract

Introduced here are computer programs and associated computer-implemented techniques for training and then applying computer-implemented models designed for segmentation of an object in the frames of video. By training and then applying the segmentation model in a cyclical manner, the errors encountered when performing segmentation can be eliminated rather than propagated. In particular, the approach to segmentation described herein allows the relationship between a reference mask and each target frame for which a mask is to be produced to be explicitly bridged or established. Such an approach ensures that masks are accurate, which in turn means that the segmentation model is less prone to distractions.

Claims (41)

1. A system comprising:

a memory device storing a network-based model; and

at least one processer configured to cause the system to:

generate a first mask by applying the network-based model to a given digital image based on a reference digital image and a reference mask, wherein the reference mask defines, in the reference digital image, a boundary of an object;

generate a second mask by applying the network-based model to the reference digital image based on the given digital image and the first mask; and

modifying the network-based model based on a comparison of the second mask and the reference mask.

2. The system of claim 1 , wherein:

the network-based model is parameterized by weights; and

modifying the network-based model comprises changing at least one of the weights to account for differences between the second mask and the reference mask.

3. The system of claim 1 , wherein:

the reference digital image and the given digital image are frames in a video comprised of multiple frames; and

the at least one processor is configured to iteratively train the network-based model multiple times in succession, each time with a different frame of the multiple frames serving as the given digital image.

4. The system of claim 3 , wherein the at least one processor is configured to iteratively train the network-based model based on analyses pairs of frames in a forward temporal order and a backward temporal order.

5. The system of claim 3 , wherein the at least one processor is configured to iteratively train the network-based model by reversing a temporal order of the multiple frames.

6. The system of claim 3 , wherein the at least one processor is configured to construct a frame set that is comprised of the multiple frames and a mask set that is comprised of masks produced for the multiple frames.

7. The system of claim 1 , wherein modifying the network-based model based on the comparison of the second mask and the reference mask comprises computing a metric indicative of a correspondence between normalized pixel values at corresponding coordinates of the second mask and the reference mask.

8. The system of claim 1 , wherein the network-based model is parameterized by weights that are modified at runtime as the network-based model is applied to frames of a video of which the reference digital image and the given digital image are a part.

9. A non-transitory computer-readable medium with instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising:

generating a first mask by applying a network-based model to a given digital image based on a reference digital image and a reference mask, wherein the reference mask defines, in the reference digital image, a boundary of an object;

generating a second mask by applying the network-based model to the reference digital image based on the given digital image and the first mask; and

generating a third mask on a comparison of the second mask and the reference mask.

10. The non-transitory computer-readable medium of claim 9 , wherein the operations, except are performed a predetermined number of times to iteratively improve quality of masks produced by the network-based model for the reference digital image.

11. The non-transitory computer-readable medium of claim 9 , wherein generating the second mask comprises generating a predicted boundary of the object in the reference digital image.

12. The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise determining an Intersection over Union score to quantify similarity between the second mask and the reference mask to make the comparison.

13. The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise evaluating a cross-entropy loss score to quantify similarity between the second mask and the reference mask to make the comparison.

14. The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise:

causing display of the reference digital image on an interface; and

recording, based on input received through the interface on which the reference digital image is displayed, a definition of the reference mask.

15. A method comprising:

generating a first mask by applying a segmentation model to a given frame based on a reference frame and a reference mask, wherein the reference mask defines, in the reference frame, a boundary of an object;

generating a second mask by applying the segmentation model to the reference frame based on the given frame and the first mask; and

modifying the segmentation model based on a comparison of the second mask and the reference mask.

16. The method of claim 15 , wherein:

the segmentation model is parameterized by weights; and

modifying the segmentation model comprises changing at least one of the weights to account for differences between the second mask and the reference mask.

17. The method of claim 15 , wherein:

the reference frame and the given frame are frames in a video comprised of multiple frames; and

the method further comprises iteratively training the segmentation model multiple times in succession, each time with a different frame of the multiple frames serving as the given frame.

18. The method of claim 15 , wherein modifying the segmentation model based on the comparison of the second mask and the reference mask comprises determining an Intersection over Union score to quantify similarity between the second mask and the reference mask to make the comparison.

19. The method of claim 15 , wherein modifying the segmentation model based on the comparison of the second mask and the reference mask comprises evaluating a cross-entropy loss score to quantify similarity between the second mask and the reference mask to make the comparison.

20. The method of claim 15 , wherein the segmentation model is parameterized by weights that are modified at runtime as the segmentation model is applied to frames of a video of which the reference frame and the given frame are a part.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2022
From: XU, NING
To: ADOBE INC.
Reel/Frame 058779/0445 →
Continuity (3)
Continuation 17086012 · Oct 30, 2020
Provisional Application 63088327 · Oct 6, 2020
Related Publication 20220148183A1 · May 12, 2022