IP Library › Granted Patent US 12,475,565
Granted Patent B2
US 12,475,565 · App. 18/056,987 · Granted Nov 18, 2025

Amodal instance segmentation using diffusion models

Inventors: Jianming Zhang (Fremont, CA); Qing Liu (Santa Clara, CA); Yilin Wang (Sunnyvale, CA); Zhe Lin (Clyde Hill, WA); Bowen Zhang (Adelaide, AU)
Assignee: ADOBE INC.
G06T7/10G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,565
App. No.
18/056,987
Granted
Nov 18, 2025
Kind
B2
Abstract

Systems and methods for instance segmentation are described. Embodiments include identifying an input image comprising an object that includes a visible region and an occluded region that is concealed in the input image. A mask network generates an instance mask for the input image that indicates the visible region of the object. A diffusion model then generates a segmentation mask for the input image based on the instance mask. The segmentation mask indicates a completed region of the object that includes the visible region and the occluded region.

Claims (51)

1 . A method comprising:

identifying an input image comprising an object that includes a visible region and an occluded region, wherein the occluded region is concealed in the input image;

generating an instance mask for the input image using a first machine learning model, wherein the instance mask indicates the visible region of the object; and

generating a segmentation mask for the input image by identifying a noise map with random noise and denoising the noise map based on the instance mask using a second machine learning model comprising a diffusion model, wherein the segmentation mask indicates a completed region that includes the visible region and the occluded region.

2 . The method of claim 1 , further comprising:

encoding the input image to obtain image features; and

decoding the image features to obtain the instance mask.

3 . The method of claim 2 , further comprising:

decoding the image features to obtain an occlusion mask, wherein the segmentation mask is based on the occlusion mask.

4 . The method of claim 2 , further comprising:

decoding the image features to obtain a plurality of class agnostic values indicating object presence.

5 . The method of claim 2 , further comprising:

generating an occluded image based on the input image, wherein the segmentation mask is based on the occluded image.

6 . The method of claim 1 , further comprising:

identifying the noise map including random noise, wherein the segmentation mask is generated based on the noise map.

7 . The method of claim 1 , wherein the segmentation mask comprises a plurality of overlapping regions corresponding to a plurality of objects in the image, respectively.

8 . A method comprising:

identifying a training image and a ground truth segmentation mask for the training image, wherein the training image comprising an object that includes a visible region and an occluded region and the ground truth segmentation mask indicates a completed region of the object that includes the visible region and the occluded region;

generating an instance mask for the training image using a first machine learning model comprising a mask network, wherein the instance mask indicates the visible region of the object;

generating a predicted segmentation mask for the training image by identifying a noise map with random noise and denoising the noise map based on the instance mask using a second machine learning model comprising a diffusion model, wherein the predicted segmentation mask indicates the completed region;

comparing the predicted segmentation mask to the ground truth segmentation mask; and

training the mask network by updating the parameters of the mask network based on the comparison.

9 . The method of claim 8 , further comprising:

training the diffusion model by updating parameters of the diffusion model based on the comparison.

10 . The method of claim 8 , wherein:

the diffusion model is pretrained before training the mask network.

11 . The method of claim 8 , further comprising:

encoding the training image to obtain image features; and

decoding the image features to obtain the instance mask.

12 . The method of claim 11 , further comprising:

decoding the image features to obtain an occlusion mask, wherein the predicted segmentation mask is based on the occlusion mask.

13 . The method of claim 11 , further comprising:

identifying a positive supervision region corresponding to the occluded region; and

identifying a negative supervision region corresponding to a region that does not correspond to the object, wherein the mask network is trained based on the positive supervision region and the negative supervision region.

14 . The method of claim 13 , further comprising:

identifying a neutral supervision region corresponding to a candidate occlusion region, wherein the mask network is not trained based on the neutral supervision region.

15 . An apparatus comprising:

a processor;

a memory containing instructions executable by the processor;

a first machine learning model configured to generate an instance mask for an input image comprising an object that includes a visible region and an occluded region, wherein the instance mask indicates the visible region of the object; and

a second machine learning model comprising a diffusion model configured to generate a segmentation mask for the input image by identifying a noise map with random noise and denoising the noise map based on the instance mask, wherein the segmentation mask indicates a complete region of the object that includes the visible region and the occluded region.

16 . The apparatus of claim 15 , wherein:

the first machine learning model includes an image encoder configured to encode the input image to obtain image features.

17 . The apparatus of claim 16 , wherein:

the first machine learning model comprises a decoder including a first head configured to generate the instance mask and a second head configured to generate an occlusion mask indicating a candidate occlusion region.

18 . The apparatus of claim 17 , wherein:

the decoder comprises a third head configured to generate class agnostic values indicating object presence.

19 . The apparatus of claim 17 , wherein:

the encoder comprises a convolutional neural network, and the decoder comprises a transformer network.

20 . The apparatus of claim 15 , wherein:

the diffusion model comprises a U-Net architecture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2022
From: ZHANG, JIANMING; LIU, QING; WANG, YILIN; LIN, ZHE; ZHANG, BOWEN
To: ADOBE INC.
Reel/Frame 061826/0918 →
Continuity (1)
Related Publication 20240169541A1 · May 23, 2024
References Cited (64)
US 20060291001A1 · Sung · 2006 [cited by examiner]
US 20210150751A1 · Liu · 2021 [cited by examiner]
US 20210279503A1 · Qi · 2021 [cited by examiner]
US 20210304422A1 · Yu · 2021 [cited by examiner]
US 20220148284A1 · Kim · 2022 [cited by examiner]
US 20230026278A1 · Vianello · 2023 [cited by examiner]
US 20230095092A1 · Xiao · 2023 [cited by examiner]
US 20230103638A1 · Saharia · 2023 [cited by examiner]
US 20230289971A1 · Back · 2023 [cited by examiner]
US 20240161250A1 · Balaji · 2024 [cited by examiner]
Alexe, et al, “Measuring the Objectness of Image Windows”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, No. 11, 2012, pp. 2189-2202, 14 pages. [cited by applicant]
Arbeláez, et al, “Contour Detection and Hierarchical Image Segmentation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. (33) 2011, 898-916, 20 pages. [cited by applicant]
Bao, et al, “BEIT: BERT Pre-Training of Image Transformers”, arXiv preprint arXiv:2106.08254v2 [cs.CV] Sep. 3, 2022, 18 pages. [cited by applicant]
Carion, et al, “End-to-End Object Detection with Transformers”, arXiv preprint arXiv:2005.12872v3 [cs.CV] May 28, 2020, 26 pages. [cited by applicant]
Cheng, et al, “Masked-attention Mask Transformer for Universal Image Segmentation”, arXiv preprint arXiv:2112.01527v3 [cs.CV] Jun. 15, 2022, 20 pages. [cited by applicant]
Cheng, et al, “Per-Pixel Classification is Not All You Need for Semantic Segmentation”, arXiv preprint arXiv:2107.06278v2 [cs.CV] Oct. 31, 2021, 17 pages. [cited by applicant]
Connor, et al, “Representing Closed Transformation Paths in Encoded Network Latent Space”, arXiv preprint strucarXiv:1912.02644v1 [stat.ML] Dec. 5, 2019, 10 pages. [cited by applicant]
Croitoru, et al, “Diffusion Models in Vision: A Survey”, arXiv:2209.04747v2 [cs.CV] Oct. 6, 2022, 22 pages. [cited by applicant]
Dai, et al, “Convolutional Feature Masking for Joint Object and Stuff Segmentation”, arXiv preprint arXiv:1412.1283v4 [cs.CV] Apr. 2, 2015, 10 pages. [cited by applicant]
Dalal, et al, “Histograms of Oriented Gradients for Human Detection”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05), Jun. 20-25, 2005, 8 pages. [cited by applicant]
Donahue, et al, “DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition”, arXiv preprint arXiv:1310.1531v1 [cs.CV] Oct. 6, 2013, 10 pages. [cited by applicant]
Dosovitskiy, et al, “An Image is Worth 16X16 Words: Transformers for Image Recognition at Scale”, arXiv preprint arXiv:2010.11929v2 [cs.CV] Jun. 3, 2021, 22 pages. [cited by applicant]
Follmann, et al, “Learning to See the Invisible: End-to-End Trainable Amodal Instance Segmentation”, arXiv preprint arXiv:1804.08864v1 [cs.CV] Apr. 24, 2018, 21 pages. [cited by applicant]
Follmann, et al, “Oriented Boxes for Accurate Instance Segmentation”, arXiv preprint arXiv:1911.07732v1 [cs.CV] Nov. 18, 2019, 8 pages. [cited by applicant]
Fu, et al, “IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things”, arXiv preprint arXiv:1906.06597v1 [cs.CV] Jun. 15, 2019, 13 pages. [cited by applicant]
Gansbeke, et al, “Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals”, arXiv preprint arXiv:2102.06191v3 [cs.CV] Aug. 3, 2021, 14 pages. [cited by applicant]
Girshick, et al, “Fast R-CNN”, arXiv preprint arXiv:1504.08083v2 [cs.CV] Sep. 27, 2015, 9 pages. [cited by applicant]
Girshick, et al, “Region-based Convolutional Networks for Accurate Object Detection and Segmentation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, No. 1, pp. 142-158, Jan. 1, 2016, 16 pages. [cited by applicant]
Girshick, et al, “Rich feature hierarchies for accurate object detection and semantic segmentation”, arXiv preprint arXiv:1311.2524v5 [cs.CV] Oct. 22, 2014, 21 pages. [cited by applicant]
Gao, et al, “Res2Net: A New Multi-scale Backbone Architecture”, arXiv preprint arXiv:1904.01169v3 [cs.CV] Jan. 27, 2021, 11 pages. [cited by applicant]
Hariharan, et al, “Simultaneous Detection and Segmentation”, arXiv preprint arXiv:1407.1808v1 [cs.CV] Jul. 7, 2014, 16 pages. [cited by applicant]
He, et al, “Deep Residual Learning for Image Recognition”, arXiv preprint arXiv:1512.03385v1 [cs.CV] Dec. 10, 2015, 12 pages. [cited by applicant]
He, et al, “Mask R-CNN”, arXiv preprint arXiv:1703.06870v1 [cs.CV] Mar. 20, 2017, 9 pages. [cited by applicant]
He, et al, “Mask R-CNN”, arXiv preprint arXiv:1703.06870v3 [cs.CV] Jan. 24, 2018, 12 pages. [cited by applicant]
He, et al, “Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition”, arXiv preprint arXiv:1406.4729v4 [cs.CV] Apr. 23, 2015, 14 pages. [cited by applicant]
Heitz, et al, “Learning Spatial Context: Using Stuff to Find Things”, Computer Vision—ECCV 2008, ECCV 2008. Lecture Notes in Computer Science, vol. 5302. Springer, 14 pages. [cited by applicant]
Ho, et al, “Classifier-Free Diffusion Guidance”, arXiv preprint arXiv:2207.12598v1 [cs.LG] Jul. 26, 2022, 14 pages. [cited by applicant]
Jiang, et al, “All Tokens Matter: Token Labeling for Training Better Vision Transformers”, arXiv preprint arXiv:2104.10858v3 [cs.CV] Jun. 9, 2021, 16 pages. [cited by applicant]
Kingma, et al, “Auto-Encoding Variational Bayes”, arXiv preprint arXiv:1312.6114v10 [stat.ML] May 1, 2014, 14 pages. [cited by applicant]
Kirillov, et al, “Panoptic Segmentation”, arXiv preprint arXiv:1801.00868v3 [cs.CV] Apr. 10, 2019, 10 pages. [cited by applicant]
LeCun, Y, et al., “Backpropagation Applied to Handwritten Zip Code Recognition”, Neural Computation, vol. 1, Issue: 4, Dec. 1989, 11 pages. [cited by applicant]
Li, et al, “Amodal Instance Segmentation”, arXiv preprint arXiv:1604.08202v2 [cs.CV] Aug. 17, 2016, 23 pages. [cited by applicant]
Li, et al, “Lecture 1: Welcome to CS231n”, Stanford University, Apr. 4, 2017, 48 pages. [cited by applicant]
Liu, et al, “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows”, arXiv preprint arXiv:2103.14030v2 [cs.CV] Aug. 17, 2021, 14 pages. [cited by applicant]
Long, et al, “Fully Convolutional Networks for Semantic Segmentation”, arXiv preprint arXiv:1411.4038v2 [cs.CV] Mar. 8, 2015, 10 pages. [cited by applicant]
Lowe, David G., “Distinctive Image Features from Scale-Invariant Keypoints”, International Journal of Computer Vision 60(2), 91-110, Jan. 2004, Kluwer Academic Publishers, 20 pages. [cited by applicant]
Luo, Calvin, “Understanding Diffusion Models: A Unified Perspective”, arXiv preprint arXiv:2208.11970v1 [cs.LG] Aug. 25, 2022, 23 pages. [cited by applicant]
Meng, et al, “Sdedit: Guided Image Synthesis and Editing With Stochastic Differential Equations”, arXiv preprint arXiv:2108.01073v2 [cs.CV] Jan. 5, 2022, 33 pages. [cited by applicant]
Qi, et al, “Amodal Instance Segmentation with KINS Dataset”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2019, 10 pages. [cited by applicant]
Ramesh, et al, “Zero-Shot Text-to-Image Generation”, arXiv preprint arXiv:2102.12092v1 [cs.CV] Feb. 24, 2021, 20 pages. [cited by applicant]
Redmon, et al, “YOLOv3: An Incremental Improvement”, arXiv preprint arXiv:1804.02767v1 [cs.CV] Apr. 8, 2018, 6 pages. [cited by applicant]
Redmon, et al, “YOLO9000: Better, Faster, Stronger”, arXiv preprint arXiv:1612.08242v1 [cs.CV] Dec. 25, 2016, 9 pages. [cited by applicant]
Redmon, et al, “You Only Look Once: Unified, Real-Time Object Detection”, arXiv preprint arXiv:1506.02640v5 [cs.CV] May 9, 2016, 10 pages. [cited by applicant]
Ren, et al, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, arXiv preprint arXiv:1506.01497v3 [cs.CV] Jan. 6, 2016, 14 pages. [cited by applicant]
Ronneberger, et al, “U-Net: Convolutional Networks for Biomedical Image Segmentation”, arXiv preprint arXiv:1505.04597v1 [cs.CV] May 18, 2015, 8 pages. [cited by applicant]
Ryoo, et al, “TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?”, arXiv preprint arXiv:2106.11297v4 [cs.CV] Apr. 3, 2022, 16 pages. [cited by applicant]
Sohl-Dickstein, et al, “Deep Unsupervised Learning using Nonequilibrium Thermodynamics”, arXiv preprint arXiv:1503.03585v8 [cs.LG] Nov. 18, 2015, 18 pages. [cited by applicant]
Szegedy, et al, “Deep Neural Networks for Object Detection”, Advances in Neural Information Processing Systems 26 (NIPS), Dec. 5-10, 2013, 9 pages. [cited by applicant]
Tian, et al, “Conditional Convolutions for Instance Segmentation”, arXiv preprint arXiv:2003.05664v4 [cs.CV] Jul. 26, 2020, 18 pages. [cited by applicant]
Wang, et al, “SOLO: Segmenting Objects by Locations”, arXiv preprint arXiv:1912.04488v3 [cs.CV] Jul. 19, 2020, 19 pages. [cited by applicant]
Wu, et al, “Logit-based Uncertainty Measure in Classification”, arXiv preprint arXiv:2107.02845v1 [cs.LG] Jul. 6, 2021, 14 pages. [cited by applicant]
Xiao, et al, “Amodal Segmentation Based on Visible Region Segmentation and Shape Prior”, arXiv preprint arXiv:2012.05598v2 [cs.CV] Dec. 19, 2020, 9 pages. [cited by applicant]
Xie, et al, “SimMIM: a Simple Framework for Masked Image Modeling”, arXiv preprint arXiv:2111.09886 [cs.CV] Nov. 18, 2021, 13 pages. [cited by applicant]
Zhu, et al, “Semantic Amodal Segmentation”, arXiv preprint arXiv:1509.01329v2 [cs.CV] Dec. 14, 2016, 11 pages. [cited by applicant]