IP Library › Granted Patent US 12,488,555
Granted Patent B2
US 12,488,555 · App. 18/215,551 · Granted Dec 2, 2025

Efficient object segmentation

Inventors: Zichuan Liu (San Jose, CA); Xin Lu (Saratoga, CA); Mingyuan Wu (Champaign, IL)
Assignee: Adobe Inc.
G06V10/26G06T7/11G06V10/82G06T2207/20104
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,555
App. No.
18/215,551
Filed
Jun 28, 2023
Granted
Dec 2, 2025
Kind
B2
Art Unit
2665
USPC
382/173
Abstract

In implementations of systems for efficient object segmentation, a computing device implements a segment system to receive a user input specifying coordinates of a digital image. The segment system computes receptive fields of a machine learning model based on the coordinates of the digital image. The machine learning model is trained on training data to generate segment masks for objects depicted in digital images. The segment system processes a portion of a feature map of the digital image using the machine learning model based on the receptive fields. A segment mask is generated for an object depicted in the digital image based on processing the portion of the feature map of the digital image using the machine learning model.

Claims (34)

1 . A method comprising:

receiving, by a processing device, a user input based on a single interaction relative to a user interface, the user input specifying coordinates of a digital image;

computing, by the processing device, receptive fields by a machine learning model based on the coordinates of the digital image, the machine learning model trained on training data to generate segment masks based on one or more objects depicted in training digital images;

processing, by the processing device, a portion of a feature map of the digital image using the machine learning model based on the receptive fields; and

generating, by the processing device, a segment mask for an object depicted in the digital image based on processing the portion of the feature map of the digital image using the machine learning model.

2 . The method as described in claim 1 , wherein the feature map of the digital image is a multi-level feature pyramid generated by processing the digital image using a feature pyramid network.

3 . The method as described in claim 1 , wherein the portion of the feature map of the digital image represents a portion of the digital image that includes the coordinates of the digital image and the object depicted in the digital image.

4 . The method as described in claim 1 , wherein the receptive fields are computed using receptive field tracing.

5 . The method as described in claim 1 , wherein the receptive fields are computed for nodes of layers of the machine learning model.

6 . The method as described in claim 5 , wherein the receptive fields are computed from a lowest layer of the layers to a highest layer of the layers.

7 . The method as described in claim 5 , wherein the layers perform layer operations including at least one of activation, convolution, pooling, normalization, or interpolation.

8 . The method as described in claim 1 , wherein the portion of the feature map of the digital image corresponds to a union of the receptive fields.

9 . The method as described in claim 1 , wherein the coordinates of the digital image are included in the object depicted in the digital image.

10 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

receiving a user input defining input coordinates of a pixel of a digital image, the pixel included in an object depicted in the digital image;

computing receptive fields for nodes of layers of a machine learning model based on the input coordinates of the pixel;

processing a portion of a feature map of the digital image-using the machine learning model based on the receptive fields for the nodes of the layers, the portion of the feature map corresponding to a union of the receptive fields; and

generating a segment mask for the object based on processing the portion of the feature map of the digital image using the machine learning model.

11 . The system as described in claim 10 , wherein the receptive fields for the nodes are computed from a lowest layer of the layers to a highest layer of the layers.

12 . The system as described in claim 10 , wherein the receptive fields for the nodes of the layers are computed using receptive field tracing.

13 . The system as described in claim 10 , wherein the feature map of the digital image is a multi-level feature pyramid generated by processing the digital image using a feature pyramid network.

14 . The system as described in claim 10 , wherein the portion of the feature map of the digital image corresponds to a union of the receptive fields for the nodes of the layers.

15 . The system as described in claim 10 , wherein the coordinates of the digital image are included in the object depicted in the digital image.

16 . The system as described in claim 10 , wherein the user input is generated based on a single interaction relative to a user interface.

17 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a user input specifying coordinates of a digital image, the coordinates are included in an object depicted in the digital image;

computing receptive fields of a machine learning model based on the coordinates of the digital image, the machine learning model trained on training data to generate segment masks for training objects depicted in training digital images;

processing a portion of a feature map of the digital image using the machine learning model based on the receptive fields; and

generating a segment mask for the object depicted in the digital image based on processing the portion of the feature map of the digital image using the machine learning model.

18 . The non-transitory computer-readable storage medium as described in claim 17 , wherein the receptive fields are computed using receptive field tracing.

19 . The non-transitory computer-readable storage medium as described in claim 17 , wherein the user input is generated based on a single interaction relative to a user interface.

20 . The non-transitory computer-readable storage medium as described in claim 17 , wherein the feature map of the digital image is a multi-level feature pyramid generated by processing the digital image using a feature pyramid network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2023
From: LIU, ZICHUAN; LU, XIN; WU, MINGYUAN
To: ADOBE INC.
Reel/Frame 064198/0449 →
Continuity (1)
Related Publication 20250005884A1 · Jan 2, 2025
References Cited (62)
US 10325371B1 · Kim · 2019 [cited by examiner]
US 11410315B2 · Homayounfar · 2022 [cited by examiner]
US 11587234B2 · Zhao · 2023 [cited by examiner]
US 12347005B2 · Smith · 2025 [cited by examiner]
US 20200167943A1 · Kim · 2020 [cited by examiner]
US 20200327334A1 · Goren · 2020 [cited by examiner]
US 20200349711A1 · Duke · 2020 [cited by examiner]
US 20210042928A1 · Takeda · 2021 [cited by examiner]
US 20210166400A1 · Goel · 2021 [cited by examiner]
US 20210383534A1 · Tadross · 2021 [cited by examiner]
US 20210390700A1 · Lee · 2021 [cited by examiner]
US 20220044365A1 · Zhang · 2022 [cited by examiner]
US 20220156943A1 · Zhang · 2022 [cited by examiner]
US 20220222832A1 · Fu · 2022 [cited by examiner]
US 20220383505A1 · Homayounfar · 2022 [cited by examiner]
US 20230204424A1 · Ranganathan · 2023 [cited by examiner]
US 20230289969A1 · Li · 2023 [cited by examiner]
US 20240078797A1 · Azarian Yazdi · 2024 [cited by examiner]
US 20240104831A1 · Lin · 2024 [cited by examiner]
US 20240169542A1 · Borse · 2024 [cited by examiner]
US 20240355018A1 · Aggarwal · 2024 [cited by examiner]
CA 3157919A1 · 2021 [cited by examiner]
WO WO2024053846A1 · 2024 [cited by examiner]
WO WO2024112452A1 · 2024 [cited by examiner]
Araujo, André, et al., “Computing Receptive Fields of Convolutional Neural Networks”, Google Research Perception Labs Google Research [retrieved Jan. 31, 2023]. Retrieved from the Internet <https://distill.pub/2019/comp… [cited by applicant]
Boykov, Yuri Y, et al., “Interactive Graph Cuts for Optimal Boundary & Region Segmentation of Objects in N-D Images”, Proceedings of International Conference on Computer Vision, Vancouver, Canada, vol. 1 [retrieved Jan.… [cited by applicant]
Cao, Jiale , et al., “SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation”, Cornell University arXiv, arXiv.org [retrieved Jan. 31, 2023]. Retrieved from the Internet <https://arxiv.… [cited by applicant]
Chen, Xi , et al., “FocalClick: Towards Practical Interactive Image Segmentation”, Cornell University arXiv, arXiv.org [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2204.02574.pdf>., Apr.… [cited by applicant]
Cheng, Bowen , et al., “Masked-attention Mask Transformer for Universal Image Segmentation”, Cornell University arXiv, arXiv.org [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2112.01527.p… [cited by applicant]
Cheng, Bowen , et al., “Pointly-Supervised Instance Segmentation”, Cornell University arXiv, arXiv.org [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2104.06404.pdf>., Jun. 15, 2022, 14 Pag… [cited by applicant]
Grady, Leo , “Random Walks for Image Segmentation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, No. 11 [retrieved Feb. 1, 2023]. Retrieved from the Internet <http://vision.cse.psu.edu/people… [cited by applicant]
Gupta, Agrim , et al., “LVIS: A Dataset for Large Vocabulary Instance Segmentation”, Cornell University arXiv, arXiv.org [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1908.03195.pdf>., Se… [cited by applicant]
He, Kaiming , et al., “Deep Residual Learning for Image Recognition”, Cornell University arXiv Preprint, arXiv.org [retrieved Aug. 11, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1512.03385.pdf>., Dec. 10,… [cited by applicant]
He, Kaiming , et al., “Mask R-CNN”, Cornell University arXiv, arXiv.org [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1703.06870.pdf>., Jan. 24, 2018, 12 pages. [cited by applicant]
Hu, Jie , et al., “ISTR: End-to-End Instance Segmentation with Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2105.00637.pdf>., May 6, 202… [cited by applicant]
Ioffe, Sergey , et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, Proceedings of the 32nd International Conference on Machine Learning, PMLR [retrieved Mar. 16, 2023… [cited by applicant]
Jang, Won-Dong , et al., “Interactive Image Segmentation via Backpropagating Refinement Scheme”, IEEE/CVF [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://openaccess.thecvf.com/content_CVPR_2019/papers/Ja… [cited by applicant]
Laradji, Issam , et al., “Proposal-Based Instance Segmentation With Point Supervision”, IEEE International Conference on Image Processing [retrieved Feb. 17, 2023]. Retrieved from the Internet <https://ieeexplore.ieee.o… [cited by applicant]
Lee, Youngwan , et al., “CenterMask: Real-Time Anchor-Free Instance Segmentation”, IEEE/CVF Conference on Computer Vision and Pattern Recognition [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://openaccess… [cited by applicant]
Lin, Tsung-Yi , et al., “Feature Pyramid Networks for Object Detection”, Cornell University arXiv, arXiv.org [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1612.03144.pdf>., Apr. 19, 2017, … [cited by applicant]
Lin, Tsung-Yi , et al., “Focal Loss for Dense Object Detection”, Cornell University arXiv, arXiv.org [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1708.02002.pdf>., Oct. 2017, 10 pages. [cited by applicant]
Lin, Zheng , et al., “Interactive Image Segmentation With First Click Attention”, IEEE/CVF Conference on Computer Vision and Pattern Recognition [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://openaccess… [cited by applicant]
Lin, Tsung-Yi , et al., “Microsoft COCO: Common Objects in Context”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1405.0312.pdf>., Feb. 21, 2015, 15 P… [cited by applicant]
Liu, Qin , et al., “PseudoClick: Interactive Image Segmentation with Click Imitation”, Cornell University arXiv, arXiv.org [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2207.05282.pdf>., J… [cited by applicant]
Long, Jonathan , et al., “Fully Convolutional Networks for Semantic Segmentation”, Cornell University arXiv, arXiv.org [retrieved Aug. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1411.4038.pdf>., Nov. 1… [cited by applicant]
Luo, Wenjie , et al., “Understanding the Effective Receptive Field in Deep Convolutional Neural Networks”, Cornell University arXiv, arXiv.org [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf… [cited by applicant]
Mortensen, Eric , et al., “Intelligent Scissors for Image Composition”, Proceedings of the 22nd annual conference on Computer graphics and interactive techniques [retrieved Feb. 1, 2023]. Retrieved from the Internet <ht… [cited by applicant]
Mortensen, Eric N, et al., “Interactive Segmentation with Intelligent Scissors”, Graphical Models and Image Processing, 60 [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://citeseerx.ist.psu.edu/document?re… [cited by applicant]
Najibi, Mahyar , et al., “AutoFocus: Efficient Multi-Scale Inference”, Cornell University arXiv, arXiv.org [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1812.01600.pdf>., Aug. 1, 2019, 11 … [cited by applicant]
Rother, Carsten , et al., ““GrabCut”: interactive foreground extraction using iterated graph cuts”, ACM Transactions on Graphics, vol. 23, No. 3 [retrieved Feb. 1, 2023]. Retrieved from the Internet <https://www.cs.jhu.… [cited by applicant]
Sofiiuk, Konstantin , et al., “F-BRS: Rethinking Backpropagating Refinement for Interactive Segmentation”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition [retrieved Mar. 15, 2023]. Ret… [cited by applicant]
Sofiiuk, Konstantin , et al., “Reviving Iterative Training with Mask Guidance for Interactive Segmentation”, Cornell University arXiv, arXiv.org [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://arxiv.org/… [cited by applicant]
Tang, Chufeng , et al., “Active Pointly-Supervised Instance Segmentation”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2207.11493.pdf>., Oct. 2022, 22… [cited by applicant]
Tian, Zhi , et al., “Conditional Convolutions for Instance Segmentation”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2003.05664.pdf>., Jul. 26, 2020,… [cited by applicant]
Tian, Zhi , et al., “FCOS: A Simple and Strong Anchor-free Object Detector”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2006.09214.pdf>., Oct. 12, 20… [cited by applicant]
Wang, Xinlong , “SOLOv2: Dynamic and Fast Instance Segmentation”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2003.10152.pdf>., Oct. 23, 2020, 17 Page… [cited by applicant]
Xie, Enze , et al., “PolarMask: Single Shot Instance Segmentation with Polar Representation”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1909.13226.p… [cited by applicant]
Xu, Ning , et al., “Deep Interactive Object Selection”, Cornell University arXiv, arXiv.org [retrieved Mar. 15, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1603.04042.pdf>., Mar. 13, 2016, 9 pages. [cited by applicant]
Yang, Chenhongyi , et al., “QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://ar… [cited by applicant]
Yu, Fisher , et al., “BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1805.04687… [cited by applicant]
Yu, Xiaodong , et al., “SOIT: Segmenting Objects with Instance-Aware Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2112.11037.pdf>., Dec.… [cited by applicant]
Zhou, Chong , “Yolact++ Better Real-Time Instance Segmentation”, University of California, Davis ProQuest Dissertations Publishing [retrieved Jan. 31, 2023]. Retrieved from the Internet <https://www.proquest.com/openvie… [cited by applicant]