IP Library › Granted Patent US 11,972,569
Granted Patent B2
US 11,972,569 · App. 17/158,527 · Granted Apr 30, 2024

Segmenting objects in digital images utilizing a multi-object segmentation model framework

Inventors: Brian Price (San Jose, CA); David Hart (Orem, UT); Zhihong Ding (Fremont, CA); Scott Cohen (Sunnyvale, UT)
Assignee: Adobe Inc.
G06T7/11G06T3/4046G06T7/174G06T7/187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,972,569
App. No.
17/158,527
Filed
Jan 26, 2021
Granted
Apr 30, 2024
Kind
B2
Art Unit
2672
USPC
382/173
Abstract

The present disclosure relates to a multi-model object segmentation system that provides a multi-model object segmentation framework for automatically segmenting objects in digital images. In one or more implementations, the multi-model object segmentation system utilizes different types of object segmentation models to determine a comprehensive set of object masks for a digital image. In various implementations, the multi-model object segmentation system further improves and refines object masks in the set of object masks utilizing specialized object segmentation models, which results in more improved accuracy and precision with respect to object selection within the digital image. Further, in some implementations, the multi-model object segmentation system generates object masks for portions of a digital image otherwise not captured by various object segmentation models.

Claims (72)

1. A non-transitory computer-readable medium storing executable instructions that, when executed by a processing device, cause the processing device to perform operations comprising:

generating a first set of object masks for a digital image comprising a plurality of objects utilizing a first object segmentation model;

generating a second set of object masks for the digital image utilizing a second object segmentation model;

detecting an overlap between a first object mask from the first set of object masks and a second object mask from the second set of object masks;

merging the overlapping first and second object masks to generate a combined object mask for the digital image;

generating a third set of object masks for the digital image that comprises the combined object mask and non-overlapping object masks from one or more of the first set of object masks or the second set of object masks;

determining that an object mask of the third set of object masks corresponds to a specialist object segmentation neural network;

generating, utilizing the specialist object segmentation neural network, an updated object mask for the object mask of the third set of object masks;

replacing the object mask of the third set of object masks in the third set of object masks with the updated object mask; and

based on detecting a selection request of a target object in the digital image, providing an object mask of the target object from the third set of object masks for the digital image.

2. The non-transitory computer-readable medium of claim 1 , wherein detecting the overlap between the first object mask and the second object mask comprises determining that a first set of pixels included in the first object mask overlaps a second set of pixels included in the second object mask by at least a pixel overlap threshold amount.

3. The non-transitory computer-readable medium of claim 2 , wherein detecting the overlap between the first object mask and the second object mask comprises matching object classification labels associated with the first object mask and the second object mask.

4. The non-transitory computer-readable medium of claim 1 , wherein merging the overlapping first and second object masks comprises:

identifying a set of non-overlapping pixels in the first object mask from the first set of object masks that is non-overlapping with the second object mask from the second set of object masks; and

generating the combined object mask by adding the set of non-overlapping pixels to the pixels included in the second object mask.

5. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the processing device, cause the processing device to refine the combined object mask by utilizing an object mask machine-learning model to improve segmentation of a corresponding object in the digital image.

6. The non-transitory computer-readable medium of claim 1 ,

wherein generating the first set of object masks for the digital image utilizing the first object segmentation model comprises utilizing a first object segmentation neural network that segments known object classes within the digital image; and

wherein generating the second set of object masks for the digital image utilizing the second object segmentation model comprises utilizing a second object segmentation neural network that segments semantic objects within the digital image.

7. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the processing device, cause the processing device to classify each segmented object in the digital image corresponding to the first set of object masks.

8. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the processing device, cause the processing device to:

determine that the object mask of the third set of object masks corresponds to the specialist object segmentation neural network based on a classification label assigned to the object mask of the third set of object masks; and

based on determining that the object mask of the third set of object masks corresponds to the specialist object segmentation neural network, provide the object mask of the third set of object masks and the digital image to the specialist object segmentation neural network to generate the updated object mask.

9. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the processing device, cause the processing device to:

determine that an object mask of the third set of object masks corresponds to a partial-object segmentation neural network based on a classification label assigned to the object mask of the third set of object masks;

based on determining that the object mask of the third set of object masks corresponds to the partial-object segmentation neural network, provide the object mask of the third set of object masks and the digital image to the partial-object segmentation neural network to generate a plurality of partial-object masks within the object mask of the third set of object masks; and

update the third set of object masks to add the plurality of partial-object masks as sub-masks within the object mask of the third set of object masks.

10. The non-transitory computer-readable medium of claim 9 , further comprising instructions that, when executed by the processing device, cause the processing device to:

identify an object mask hierarchy between the plurality of partial-object masks and the object mask of the third set of object masks;

determine that the selection request of the target object in the digital image corresponds to the object mask of the third set of object masks and a partial-object mask of the plurality of partial-object masks; and

in response to detecting the selection request of the target object in the digital image, indicate that the object mask of the third set of object masks and the partial-object mask correspond to the target object.

11. A method comprising:

generating a first set of object masks for a digital image comprising a plurality of objects utilizing a first object segmentation model;

generating a second set of object masks for the digital image utilizing a second object segmentation model;

detecting an overlap between a first object mask from the first set of object masks and a second object mask from the second set of object masks;

merging the overlapping first and second object masks to generate a combined object mask for the digital image;

generating a third set of object masks for the digital image that comprises the combined object mask and non-overlapping object masks from one or more of the first set of object masks or the second set of object masks;

determining that an object mask of the third set of object masks corresponds to a partial-object segmentation neural network;

generating, utilizing the partial-object segmentation neural network, a plurality of partial-object masks within the object mask of the third set of object masks;

updating the third set of object masks to add the plurality of partial-object masks as sub-masks within the object mask of the third set of object masks; and

based on detecting a selection request of a target object in the digital image, providing an object mask of the target object from the third set of object masks for the digital image.

12. The method of claim 11 , wherein detecting the overlap between the first object mask and the second object mask comprises determining that a first set of pixels included in the first object mask overlaps a second set of pixels included in the second object mask by at least a pixel overlap threshold amount.

13. The method of claim 11 , wherein detecting the overlap between the first object mask and the second object mask comprises matching object classification labels associated with the first object mask and the second object mask.

14. The method of claim 11 , wherein merging the overlapping first and second object masks comprises:

identifying a set of non-overlapping pixels in the first object mask from the first set of object masks that is non-overlapping with the second object mask from the second set of object masks; and

generating the combined object mask by adding the set of non-overlapping pixels to the pixels included in the second object mask.

15. The method of claim 11 , further comprising:

determining that an object mask of the third set of object masks corresponds to a specialist object segmentation neural network based on a classification label assigned to the object mask of the third set of object masks;

based on determining that the object mask of the third set of object masks corresponds to the specialist object segmentation neural network, providing the object mask of the third set of object masks and the digital image to the specialist object segmentation neural network to generate an updated object mask; and

replacing the object mask of the third set of object masks in the third set of object masks with the updated object mask.

16. A system comprising:

a memory component; and

one or more processing devices coupled to the memory component, the one or more processing devices to perform operations comprising:

generating a first set of object masks for a digital image comprising a plurality of objects utilizing a first object segmentation model;

generating a second set of object masks for the digital image utilizing a second object segmentation model;

detecting an overlap between a first object mask from the first set of object masks and a second object mask from the second set of object masks;

merging the overlapping first and second object masks to generate a combined object mask for the digital image;

generating a third set of object masks for the digital image that comprises the combined object mask and non-overlapping object masks from one or more of the first set of object masks or the second set of object masks;

determining that an object mask of the third set of object masks corresponds to a specialist object segmentation neural network;

generating, utilizing the specialist object segmentation neural network, an updated object mask for the object mask of the third set of object masks;

replacing the object mask of the third set of object masks in the third set of object masks with the updated object mask; and

based on detecting a selection request of a target object in the digital image, providing an object mask of the target object from the third set of object masks for the digital image.

17. The system of claim 16 , wherein the operations further comprise refining the combined object mask by utilizing an object mask machine-learning model to improve segmentation of a corresponding object in the digital image.

18. The system of claim 16 , wherein generating the first set of object masks for the digital image utilizing the first object segmentation model comprises utilizing a first object segmentation neural network that segments known object classes within the digital image, and wherein generating the second set of object masks for the digital image utilizing the second object segmentation model comprises utilizing a second object segmentation neural network that segments semantic objects within the digital image.

19. The system of claim 16 , wherein the operations further comprise:

determining that an object mask of the third set of object masks corresponds to a partial-object segmentation neural network based on a classification label assigned to the object mask of the third set of object masks;

based on determining that the object mask of the third set of object masks corresponds to the partial-object segmentation neural network, providing the object mask of the third set of object masks and the digital image to the partial-object segmentation neural network to generate a plurality of partial-object masks within the object mask of the third set of object masks; and

updating the third set of object masks to add the plurality of partial-object masks as sub-masks within the object mask of the third set of object masks.

20. The system of claim 19 , wherein the operations further comprise:

identifying an object mask hierarchy between the plurality of partial-object masks and the object mask of the third set of object masks;

determining that the selection request of the target object in the digital image corresponds to the object mask of the third set of object masks and a partial-object mask of the plurality of partial-object masks; and

in response to detecting the selection request of the target object in the digital image, indicating that the object mask of the third set of object masks and the partial-object mask correspond to the target object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2021
From: PRICE, BRIAN; DING, ZHIHONG; COHEN, SCOTT; HART, DAVID
To: ADOBE INC.
Reel/Frame 055034/0848 →
Continuity (1)
Related Publication 20220237799A1 · Jul 28, 2022
Cited By (2)
US 12,437,512 US 12,450,748