IP Library Granted Patent US 11,676,283
Granted Patent B2
US 11,676,283 · App. 17/660,361 · Granted Jun 13, 2023

Iteratively refining segmentation masks

Inventors: Zichuan Liu (San Jose, CA); Wentian Zhao (San Jose, CA); Shitong Wang (San Jose, CA); He Qin (San Jose, CA); Yumin Jia (San Jose, CA); Yeojin Kim (Seoul, KR); Xin Lu (Mountain View, CA); Jen-Chan Chien (Saratoga, CA)
Assignee: Adobe Inc.
G06T7/11G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,676,283
App. No.
17/660,361
Granted
Jun 13, 2023
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that generate refined segmentation masks for digital visual media items. For example, in one or more embodiments, the disclosed systems utilize a segmentation refinement neural network to generate an initial segmentation mask for a digital visual media item. The disclosed systems further utilize the segmentation refinement neural network to generate one or more refined segmentation masks based on uncertainly classified pixels identified from the initial segmentation mask. To illustrate, in some implementations, the disclosed systems utilize the segmentation refinement neural network to redetermine whether a set of uncertain pixels corresponds to one or more objects depicted in the digital visual media item based on low-level (e.g., local) feature values extracted from feature maps generated for the digital visual media item.

Claims (76)

1. A method comprising:

receiving a digital visual media item portraying an object;

generating, utilizing a segmentation refinement neural network, an initial segmentation mask for the digital visual media item;

determining, utilizing the segmentation refinement neural network, a set of pixels having an associated classification from the initial segmentation mask; and

generating a refined segmentation mask for the digital visual media item by utilizing the segmentation refinement neural network to iteratively refine the initial segmentation mask by iteratively refining the set of pixels having the associated classification from the initial segmentation mask.

2. The method of claim 1 , wherein:

determining, utilizing the segmentation refinement neural network, the set of pixels having the associated classification from the initial segmentation mask comprises determining, utilizing the segmentation refinement neural network, uncertain pixels having an uncertain classification from the initial segmentation mask; and

utilizing the segmentation refinement neural network to iteratively refine the initial segmentation mask comprises utilizing the segmentation refinement neural network to iteratively refine the uncertain pixels from the initial segmentation mask.

3. The method of claim 2 , wherein determining, utilizing the segmentation refinement neural network, the uncertain pixels from the initial segmentation mask comprises:

generating, utilizing the segmentation refinement neural network, an uncertainty map comprising uncertainty scores for pixels of the initial segmentation mask; and

determining the uncertain pixels from the initial segmentation mask based on the uncertainty scores of the uncertainty map.

4. The method of claim 1 , wherein generating the refined segmentation mask for the digital visual media item by utilizing the segmentation refinement neural network to iteratively refine the initial segmentation mask comprises:

generating a first refined segmentation mask for the digital visual media item by utilizing the segmentation refinement neural network to refine the initial segmentation mask; and

generating a second refined segmentation mask for the digital visual media item by utilizing the segmentation refinement neural network to refine the first refined segmentation mask.

5. The method of claim 1 ,

further comprising generating, utilizing the segmentation refinement neural network, a set of feature maps corresponding to the digital visual media item,

wherein generating, utilizing the segmentation refinement neural network, the initial segmentation mask for the digital visual media item comprises generating, utilizing the segmentation refinement neural network, the initial segmentation mask based on the set of feature maps.

6. The method of claim 5 , wherein generating, utilizing the segmentation refinement neural network, the set of feature maps corresponding to the digital visual media item comprises utilizing the segmentation refinement neural network to:

generate an initial feature map from the digital visual media item;

generate an additional initial feature map from the initial feature map; and

generate a final feature map from the additional initial feature map.

7. The method of claim 6 , wherein:

generating the initial feature map comprises generating a first set of low-level feature values corresponding to local attributes of the digital visual media item;

generating the additional initial feature map comprises generating a second set of low-level feature values corresponding to the local attributes of the digital visual media item, the second set of low-level feature values comprising a higher level of feature values than the first set of low-level feature values; and

generating the final feature map comprises generating a set of high-level feature values corresponding to global attributes of the digital visual media item.

8. The method of claim 1 , wherein:

generating, utilizing the segmentation refinement neural network, the initial segmentation mask for the digital visual media item comprises generating, utilizing the segmentation refinement neural network the initial segmentation mask indicating a location and a rough shape of the object portrayed in the digital visual media item; and

generating the refined segmentation mask for the digital visual media item by utilizing the segmentation refinement neural network to iteratively refine the initial segmentation mask comprises generating the refined segmentation mask by utilizing the segmentation refinement neural network to iteratively recapture local details associated with the object portrayed in the digital visual media item.

9. The method of claim 1 , further comprising:

providing the digital visual media item and one or more modification elements for display via a graphical user interface of a computing device;

receiving, via the graphical user interface, a user interaction with the one or more modification elements, wherein generating the initial segmentation mask and the refined segmentation mask is in response to receiving the user interaction;

modifying the digital visual media item utilizing the refined segmentation mask; and

providing the modified digital visual media item for display via the graphical user interface.

10. A non-transitory computer-readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

receiving a digital visual media item comprising a plurality of pixels, the digital visual media item portraying an object;

generating, utilizing a segmentation refinement neural network, an initial segmentation mask that classifies pixels from the plurality of pixels as corresponding to the object or not corresponding to the object;

determining uncertain pixels that have an associated uncertainty that the uncertain pixels have been classified correctly within the initial segmentation mask; and

generating a refined segmentation mask for the digital visual media item by utilizing the segmentation refinement neural network to iteratively refine the uncertain pixels.

11. The non-transitory computer-readable medium of claim 10 , wherein utilizing the segmentation refinement neural network to iteratively refine the uncertain pixels comprises utilizing the segmentation refinement neural network to iteratively reclassify at least a subset of uncertain pixels from the uncertain pixels as corresponding to the object or not corresponding to the object.

12. The non-transitory computer-readable medium of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

generating, utilizing the segmentation refinement neural network, an uncertainty map comprising uncertainty scores for pixels of the initial segmentation mask; and

determining a ranking for the uncertainty scores of the uncertainty map,

wherein determining the uncertain pixels comprises determining the uncertain pixels based on the ranking of the uncertainty scores from the uncertainty map.

13. The non-transitory computer-readable medium of claim 12 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

generating, utilizing the segmentation refinement neural network, a feature map from the digital visual media item; and

combining the feature map and the initial segmentation mask to generate a combined map, wherein generating, utilizing the segmentation refinement neural network, the uncertainty map comprises generating, utilizing the segmentation refinement neural network, the uncertainty map from the combined map.

14. The non-transitory computer-readable medium of claim 10 ,

wherein generating the refined segmentation mask for the digital visual media item comprises generating the refined segmentation mask for a video frame of a digital video feed; and

further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

generating, utilizing the segmentation refinement neural network, a plurality of additional refined segmentation masks for a plurality of additional video frames of the digital video feed; and

modifying the digital video feed utilizing the refined segmentation mask and the plurality of additional refined segmentation masks.

15. A system comprising:

one or more memory components; and

one or more processing devices coupled to the one or more memory components, the one or more processing devices to perform operations comprising utilizing a segmentation refinement neural network to:

extract, from a digital visual media item, a set of feature maps comprising feature values corresponding to the digital visual media item;

generate, utilizing at least one feature map from the set of feature maps, an initial segmentation mask that classifies pixels of the digital visual media item as corresponding to an object portrayed in the digital visual media item or not corresponding to the object;

determine, utilizing the at least one feature map, uncertain pixels that have an associated uncertainty that the uncertain pixels have been classified correctly within the initial segmentation mask; and

generate, utilizing the at least one feature map, a refined segmentation mask for the digital visual media item by iteratively refining the uncertain pixels.

16. The system of claim 15 , wherein the one or more processing devices:

extract the set of feature maps by extracting, from the digital visual media item, a set of initial feature maps comprising low-level feature values corresponding to the digital visual media item; and

generate, utilizing the at least one feature map, the refined segmentation mask for the digital visual media item by iteratively refining the uncertain pixels by, for a given iteration:

extracting, from the set of initial feature maps, a subset of low-level feature values that correspond to the uncertain pixels; and

refining the uncertain pixels utilizing the subset of low-level feature values.

17. The system of claim 16 , wherein the one or more processing devices:

extract the set of feature maps by extracting, from the digital visual media item, a final feature map comprising high-level feature values corresponding to the digital visual media item; and

generate, utilizing the at least one feature map, the refined segmentation mask for the digital visual media item by iteratively refining the uncertain pixels by, for the given iteration:

extracting, from the final feature map, a subset of high-level feature values that correspond to the uncertain pixels; and

refining the uncertain pixels utilizing the subset of high-level feature values.

18. The system of claim 17 , wherein the one or more processing devices generate, utilizing the at least one feature map, the refined segmentation mask for the digital visual media item by iteratively refining the uncertain pixels by, for the given iteration:

extracting, from the final feature map, an additional subset of high-level feature values that correspond to certain pixels that have an associated certainty that the certain pixels have been classified correctly within the initial segmentation mask; and

refining the uncertain pixels utilizing the additional subset of high-level feature values.

19. The system of claim 18 , wherein the one or more processing devices generate, utilizing the at least one feature map, the refined segmentation mask for the digital visual media item by iteratively refining the uncertain pixels by, for the given iteration:

generating a convolutional output from the subset of low-level feature maps utilizing a convolutional layer;

combining the convolutional output with the subset of high-level feature values and the additional subset of high-level feature values to generate a combined map; and

refining the uncertain pixels utilizing the combined map.

20. The system of claim 15 , wherein the one or more processing devices generate, utilizing the at least one feature map, the refined segmentation mask for the digital visual media item by iteratively refining the uncertain pixels by iteratively generating a plurality of refined segmentation masks, where one or more refined segmentation masks from the plurality of refined segmentation masks comprises a reclassification of at least a subset of the uncertain pixels based on a previously generated refined segmentation mask.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2022
From: LIU, ZICHUAN; ZHAO, WENTIAN; WANG, SHITONG; QIN, HE; JIA, YUMIN; KIM, YEOJIN; LU, XIN; CHIEN, JEN-CHAN
To: ADOBE INC.
Reel/Frame 059685/0642 →
Continuity (2)
Continuation 16988408 · Aug 7, 2020
Related Publication 20220245824A1 · Aug 4, 2022
Cited By (1)
US 12,494,036