IP Library › Granted Patent US 12,749,289
Granted Patent B2
US 12,749,289 · App. 18/423,037 · Granted Sep 29, 2026

Generating labels for segments of a digital image using a mask-aware classification neural network

Inventors: Jason Wen Yong Kuen (Santa Clara, CA); Cristina Isabel Gonzalez Osorio (Bogota, CO); Jing Shi (Rochester, NY); Jiuxiang Gu (College Park, MD); Kushal Kafle (Boston, MA); Yuqian Zhou (Urbana, IL); Zijun Wei (San Jose, CA)
Assignee: Adobe Inc.
G06V10/764G06T7/10G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,289
App. No.
18/423,037
Granted
Sep 29, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media that generate segment labels for image segments of a digital image that have been determined via deep segmentation. For instance, in some embodiments, the disclosed systems generate, using a segment classification neural network, an image embedding for a digital image portraying a plurality of image segments. Additionally, the disclosed systems determine, using the segment classification neural network, masked segment embeddings for the plurality of image segments of the digital image based on the image embedding and a plurality of masks corresponding to the plurality of image segments. Based on the masked segment embeddings, the disclosed systems use the segment classification neural network to determine segment labels for the plurality of image segments.

Claims (47)

1 . A computer-implemented method comprising:

generating, using a segment classification neural network, an image embedding for a digital image portraying a plurality of image segments;

determining, using the segment classification neural network, masked segment embeddings for the plurality of image segments of the digital image based on the image embedding and a plurality of masks corresponding to the plurality of image segments by:

generating a set of masked multi-scale features for an image segment from the plurality of image segments by executing a mask-based pooling operation using the image embedding and a mask corresponding to the image segment; and

determining a masked segment embedding for the image segment by combining features from the set of masked multi-scale features; and

determining, using the segment classification neural network, segment labels for the plurality of image segments based on the masked segment embeddings.

2 . The computer-implemented method of claim 1 , further comprising:

generating the plurality of masks for the digital image using a segmentation neural network; and

providing the plurality of masks with the digital image as input to the segment classification neural network.

3 . The computer-implemented method of claim 1 , wherein determining the masked segment embeddings for the plurality of image segments of the digital image based on the image embedding and the plurality of masks comprises determining a masked segment embedding for an image segment of the digital image by applying a mask for the image segment to the image embedding to prevent incorporating features represented in the image embedding that are unassociated with the image segment.

4 . The computer-implemented method of claim 1 , wherein executing the mask-based pooling operations comprises executing one or more mask-based pooling operations via one or more mask-based pooling layers of the segment classification neural network.

5 . The computer-implemented method of claim 1 , wherein combining the features from the set of masked multi-scale features comprises averaging the features from the set of masked multi-scale features.

6 . The computer-implemented method of claim 5 , wherein generating the set of masked multi-scale features for the image segment comprises generating the set of masked multi-scale features to include one or more features of an additional image segment from the plurality of image segments.

7 . The computer-implemented method of claim 1 , wherein determining the segment labels for the plurality of image segments comprises determining the segment labels from an open-vocabulary label set.

8 . The computer-implemented method of claim 1 , wherein determining the masked segment embeddings using the segment classification neural network comprises determining the masked segment embeddings using the segment classification neural network having parameters determined using weak-alignment data that includes training images and training image captions, and strong-alignment data that includes additional training images and training segment labels.

9 . A system comprising:

one or more memory devices; and

one or more processors configured to cause the system to:

receive a digital image portraying a plurality of image segments;

generate, using a class-agnostic segmentation neural network, a plurality of masks for the digital image by generating a mask for each image segment of the plurality of image segments; and

determine segment labels for the plurality of image segments of the digital image by using a segment classification neural network to:

generate an image embedding for the digital image by generating a plurality of multi-scale features from the digital image using an image encoder;

generate a masked segment embedding for each image segment of the digital image based on the image embedding and a corresponding mask from the plurality of masks by applying the corresponding mask to the plurality of multi-scale features generated from the digital image; and

determine a segment label for each image segment of the digital image based on a corresponding masked segment embedding.

10 . The system of claim 9 , wherein the one or more processors are further configured to cause the system to determine parameters for the segment classification neural network via training iterations that incorporate weak-alignment data having training images and training image captions and additional training iterations that incorporate strong-alignment data having additional training images and training segment labels.

11 . The system of claim 10 , wherein the one or more processors are configured to cause the system to determine the parameters for the segment classification neural network via the training iterations and the additional training iterations by:

providing the training images and masks corresponding to the training images as input to the segment classification neural network during the training iterations; and

providing the additional training images and additional masks corresponding to the additional training images as input to the segment classification neural network during the additional training iterations.

12 . The system of claim 11 , wherein the one or more processors are further configured to cause the system to:

generate the masks from the training images using a segmentation neural network; and

generate the additional masks from the additional training images using the segmentation neural network.

13 . The system of claim 10 , wherein the one or more processors are configured to cause the system to determine the parameters for the segment classification neural network via the additional training iterations that incorporate the strong-alignment data by providing the additional training images and bounding boxes or point locations corresponding to image segments of the additional training images as input to the segment classification neural network during the additional training iterations.

14 . The system of claim 9 , wherein applying the corresponding mask to the plurality of multi-scale features generated from the digital image comprises executing at least one pooling operation using the plurality of multi-scale features and the corresponding mask.

15 . The system of claim 9 , wherein the one or more processors are configured to cause the system to determine the segment label for each image segment of the digital image by:

generating a first segment label for a first image segment corresponding to an object portrayed in the digital image;

generating a second segment label for a second image segment corresponding to a background of the digital image; and

generating a third segment label for a third image segment corresponding to a foreground of the digital image.

16 . A non-transitory computer-readable medium storing executable instructions that, when executed by a processing device, cause the processing device to perform operations comprising:

generating an image embedding for a digital image using a segment classification neural network having parameters determined based on weak-alignment data that includes training images and training image captions, and strong-alignment data that includes additional training images and training segment labels;

determining, using the segment classification neural network, a masked segment embedding for each image segment portrayed within the digital image based on the image embedding and a corresponding mask by:

generating masked multi-scale features by applying the corresponding mask to the image embedding via a mask-based pooling operation; and

determining the masked segment embedding based on the masked multi-scale features; and

determining, using the segment classification neural network, a segment label for each image segment portrayed within the digital image based on a masked segment embedding determined for the image segment.

17 . The non-transitory computer-readable medium of claim 16 , wherein determining the masked segment embedding for each image segment comprises determining, for each image segment, the masked segment embedding having features of the digital image that represent a context surrounding the image segment within the digital image.

18 . The non-transitory computer-readable medium of claim 16 , wherein determining the masked segment embedding based on the masked multi-scale features comprises determining the masked segment embedding by combining features from the masked multi-scale features.

19 . The non-transitory computer-readable medium of claim 18 , wherein determining the masked segment embedding based on the masked multi-scale features comprises determining the masked segment embedding by merging the masked multi-scale features.

20 . The non-transitory computer-readable medium of claim 16 , wherein determining the segment label for each image segment comprises determining a plurality of segment labels for a plurality of objects portrayed in the digital image, each object of the plurality of objects corresponding to a mask generated for the digital image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2024
From: KUEN, JASON WEN YONG; OSORIO, CRISTINA ISABEL GONZALEZ; SHI, JING; GU, JIUXIANG; KAFLE, KUSHAL; ZHOU, YUQIAN; WEI, ZIJUN
To: ADOBE INC.
Reel/Frame 066252/0125 →
Continuity (1)
Related Publication 20250245964A1 · Jul 31, 2025
References Cited (7)
US 20160358337A1 · Dai · 2016 [cited by examiner]
US 20220230321A1 · Zhao · 2022 [cited by examiner]
US 20240153093A1 · Xu · 2024 [cited by examiner]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, Ross Girshick. “Segment Anything.” IEEE Conference On … [cited by applicant]
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, Diana Marculescu. “Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP.” Proceedings of the IEEE/CVF Confere… [cited by applicant]
Lu Qi, Jason Kuen, Yi Wang, Jiuxiang Gu, Hengshuang Zhao, Zhe Lin, Philip Torr, Jiaya Jia. “Open-World Entity Segmentation.” IEEE Conference On Computer Vision And Pattern Recognition (CVPR) 2021. arXiv:2107.14228. 14 P… [cited by applicant]
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, Saining Xie. “A ConvNet for the 2020s.” CVPR 2022. arXiv:2201.03545. 15 Pages. Jan. 10, 2022. [cited by applicant]