IP Library Granted Patent US 12,632,941
Granted Patent B2
US 12,632,941 · App. 18/458,778 · Granted May 19, 2026

Identifying salient regions based on multi-resolution partitioning

Inventors: Sriram Ravindran (Milpitas, CA); Debraj Debashish Basu (Sunnyvale, CA)
Assignee: Adobe Inc.
G06T5/75G06V10/46G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,941
App. No.
18/458,778
Granted
May 19, 2026
Kind
B2
Abstract

In implementation of techniques for generating salient regions based on multi-resolution partitioning, a computing device implements a salient object system to receive a digital image including a salient object. The salient object system generates a first mask for the salient object by partitioning the digital image into salient and non-salient regions. The salient object system also generates a second mask for the salient object that has a resolution that is different than the first mask by partitioning a resampled version of the digital image into salient and non-salient regions. Based on the first mask and the second mask, the salient object system generates an indication of a salient region of the digital image using a machine learning model. The salient object system then displays the indication of the salient region in a user interface.

Claims (37)

1 . A method comprising:

receiving, by a processing device, a digital image including a salient object;

generating, by the processing device, a first mask identifying a location of the salient object by partitioning the digital image into salient and non-salient regions;

generating, by the processing device, a second mask for the salient object that has a resolution that is different than the first mask by partitioning a resampled version of the digital image into salient and non-salient regions;

generating, by the processing device, an indication of a salient region of the digital image based on the location of the salient object of the first mask and the resolution of the second mask using a machine learning model; and

displaying, by the processing device, the indication of the salient region in a user interface.

2 . The method of claim 1 , wherein the first mask and the second mask are generated simultaneously.

3 . The method of claim 1 , wherein the machine learning model comprises at least one of a Self-Distillation with No Labels (DINO) model, a transformer encoder, an upsampling network, or a Graph Total Variation Regularizer model that uses guided reconstruction to optimize the indication of the salient region.

4 . The method of claim 1 , wherein the first mask is trained to minimize a normalized cut based on an adjacency matrix.

5 . The method of claim 1 , wherein the second mask is trained to replicate the first mask when downsampled to a resolution of the first mask.

6 . The method of claim 1 , wherein the resampled version of the digital image is an upsampled version of the digital image.

7 . The method of claim 1 , wherein the first mask is generated using a Self-Distillation with No Labels (DINO) model to generate embeddings for patches of the digital image.

8 . The method of claim 7 , wherein the embeddings for the patches of the digital image are processed by a transformer encoder.

9 . The method of claim 1 , further comprising regularizing the first mask and the second mask using Graph Total Variation (GTV) that discourages island pixels at low resolution and encourages the second mask to follow contours in the digital image.

10 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a digital image including a salient object;

generating a first mask identifying a location of the salient object by partitioning the digital image into salient and non-salient regions;

generating a second mask for the salient object that has a resolution that is different than the first mask by partitioning a resampled version of the digital image into salient and non-salient regions;

generating an indication of a salient region of the digital image based on the location of the salient object of the first mask and the resolution of the second mask using a machine learning model; and

displaying the indication of the salient region in a user interface.

11 . The non-transitory computer-readable storage medium of claim 10 , wherein the first mask and the second mask are generated simultaneously.

12 . The non-transitory computer-readable storage medium of claim 10 , wherein the machine learning model comprises at least one of a Self-Distillation with No Labels (DINO) model, a transformer encoder, an upsampling network, or a Graph Total Variation Regularizer model that uses guided reconstruction to optimize the indication of the salient region.

13 . The non-transitory computer-readable storage medium of claim 10 , wherein the first mask is trained to minimize a normalized cut based on an adjacency matrix and the second mask is trained to replicate the first mask when downsampled to a resolution of the first mask.

14 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

receiving a digital image including a salient object;

generating a first mask identifying a location of the salient object by partitioning the digital image into salient and non-salient regions;

generating a second mask for the salient object that has a resolution that is different than the first mask by partitioning a resampled version of the digital image into salient and non-salient regions;

generating an indication of a salient region of the digital image based on the location of the salient object of the first mask and the resolution of the second mask using a machine learning model; and

displaying the indication of the salient region in a user interface.

15 . The system of claim 14 , wherein the first mask and the second mask are generated simultaneously.

16 . The system of claim 14 , wherein the machine learning model comprises at least one of a Self-Distillation with No Labels (DINO) model, a transformer encoder, an upsampling network, or a Graph Total Variation Regularizer model that uses guided reconstruction to optimize the indication of the salient region.

17 . The system of claim 14 , wherein the first mask is trained to minimize a normalized cut based on an adjacency matrix.

18 . The system of claim 14 , wherein the first mask is generated using a Self-Distillation with No Labels (DINO) model to generate embeddings for patches of the digital image.

19 . The system of claim 18 , wherein the embeddings for the patches of the digital image are processed by a transformer encoder.

20 . The method of claim 1 , wherein the generating the salient region of the digital image further comprises co-optimizing the first mask with the second mask.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2023
From: RAVINDRAN, SRIRAM; BASU, DEBRAJ DEBASHISH
To: ADOBE INC.
Reel/Frame 064757/0013 →
Continuity (1)
Related Publication 20250078220A1 · Mar 6, 2025
References Cited (66)
US 10521699B2 · Bremer · 2019 [cited by examiner]
US 11282208B2 · Cohen · 2022 [cited by examiner]
US 11430084B2 · Stent · 2022 [cited by examiner]
US 12020400B2 · Hsieh · 2024 [cited by examiner]
US 20160012293A1 · Mate · 2016 [cited by examiner]
US 20190340462A1 · Pao · 2019 [cited by examiner]
US 20200074589A1 · Stent · 2020 [cited by examiner]
US 20200143194A1 · Hou · 2020 [cited by examiner]
“The Pascal Visual Object Classes Challenge 2007”, Pascal Network [retrieved Oct. 5, 2023]. Retrieved from the Internet <http://host.robots.ox.ac.uk/pascal/VOC/voc2007/>., Apr. 7, 2007, 6 pages. [cited by applicant]
“Visual Object Classes Challenge 2012 (VOC2012)”, Pascal Network [retrieved Oct. 5, 2023]. Retrieved from the Internet <http://host.robots.ox.ac.uk/pascal/VOC/voc2012/>, Feb. 20, 2012, 11 pages. [cited by applicant]
Aflalo, Amit et al., “DeepCut: Unsupervised Segmentation using Graph Neural Networks Clustering”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2212.05… [cited by applicant]
Allard, William K. , “Total variation regularization for image denoising, i. geometric theory”, SIAM J. Math. Anal, vol. 39, No. 4 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://services.math.duke.edu/˜… [cited by applicant]
Barron, Jonathan et al., “The Fast Bilateral Solver”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1511.03296.pdf>., Jul. 22, 2016, 50 pages. [cited by applicant]
Bielsk, Adam et al., “MOVE: unsupervised movable object segmentation and detection”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2210.07920.pdf>., Oc… [cited by applicant]
Borji, Ali et al., “Salient object detection: A survey”, Computational Visual Media, vol. 5, No. 2 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://link.springer.com/article/10.1007/s41095-019-0149-9>., J… [cited by applicant]
Caron, Mathilde et al., “Emerging Properties in Self-Supervised Vision Transformers”, Cornell University arXiv, arXiv.org [retrieved Aug. 10, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2104.14294.pdf>., M… [cited by applicant]
Caron, Mathilde et al., “Unsupervised learning of visual features by contrasting cluster assignments”, Advances in Neural Information Processing Systems 33 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https:/… [cited by applicant]
Chen, Mickaël et al., “Unsupervised Object Segmentation by Redrawing”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1905.13539.pdf>., Nov. 29, 2019, 1… [cited by applicant]
Chen, Xinlei et al., “Improved Baselines with Momentum Contrastive Learning”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2003.04297.pdf)>., Mar. 9, … [cited by applicant]
Cheng, Bowen et al., “Per-Pixel Classification is Not All You Need for Semantic Segmentation”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2107.06278… [cited by applicant]
De Lutio, Riccardo et al., “Guided Super-Resolution as Pixel-to-Pixel Transformation”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <http://128.84.21.203/pdf/1904.01501>., A… [cited by applicant]
De Lutio, Riccardo et al., “Learning Graph Regularisation for Guided Super-Resolution”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2203.14297.pdf>.,… [cited by applicant]
Deng, Jia et al., “ImageNet: A Large-Scale Hierarchical Image Database”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://www-c… [cited by applicant]
Devlin, Jacob et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding”, Cornell University, arXiv Preprint, arXiv.org [retrieved on Jun. 14, 2023]. Retrieved from the Internet <https://… [cited by applicant]
Dosovitskiy, Alexey et al., “An image is worth 16×16 words: Transformers for image recognition at scale.”, Cornell University arXiv, arXiv.org [retrieved Jun. 29, 2023]. Retrieved from the Internet <https://arxiv.org/pd… [cited by applicant]
Estrela, Vania V. et al., “Total Variation Applications in Computer Vision”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/ftp/arxiv/papers/1603/1603.09599… [cited by applicant]
Gatti, Alice et al., “Deep Learning and Spectral Embedding for Graph Partitioning”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2110.08614.pdf>., Dec… [cited by applicant]
Gatti, Alice et al., “Graph Partitioning and Sparse Matrix Ordering using Reinforcement Learning and Graph Neural Networks”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <ht… [cited by applicant]
Hamilton, Mark et al., “Unsupervised semantic segmentation by distilling feature correspondences”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2203.0… [cited by applicant]
He, Kaiming et al., “Masked Autoencoders Are Scalable Vision Learners”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2111.06377.pdf>., Dec. 19, 2021, … [cited by applicant]
Hou, Xianxu et al., “FEAT: Face Editing with Attention”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2202.02713.pdf>., Feb. 6, 2022, 14 Pages. [cited by applicant]
Kingma, Diederik P. et al., “Adam: A Method for Stochastic Optimization”, Cornell University, arXiv Preprint, arXiv.org [retrieved Aug. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1412.6980.pdf>., Jan. … [cited by applicant]
Krahenbuhl, Philipp et al., “Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials”, Advances in Neural Information Processing Systems 24 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https… [cited by applicant]
Lafferty, John et al., “Conditional random fields: Probabilistic models for segmenting and labeling sequence data”, Department of Computer and Information Science, University of Pennsylvania [retrieved Jun. 14, 2023]. R… [cited by applicant]
Lin, Tsung-Yi et al., “Microsoft COCO: Common Objects in Context”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1405.0312.pdf>., Feb. 21, 2015, 15 Pag… [cited by applicant]
Liu, Jiaming et al., “Image restoration using total variation regularized deep image prior”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1810.12864.p… [cited by applicant]
Liu, Tie et al., “Learning to Detect a Salient Object”, IEEE Transactions on Pattern Analysis and Machine Intelligence vol. 33, No. 2 [retrieved Jun. 14, 2023]. Retrieved from the Internet <http://mmlab.ie.cuhk.edu.hk/a… [cited by applicant]
Melas-Kyriazi, Luke et al., “Deep Spectral Methods: A Surprisingly Strong Baseline for Unsupervised Semantic Segmentation and Localization”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from … [cited by applicant]
Melas-Kyriazi, Luke et al., “Finding an Unsupervised Image Segmenter in Each of Your Deep Generative Model”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/… [cited by applicant]
Ortega, Antonio et al., “Graph Signal Processing: Overview, Challenges and Applications”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1712.00468.pdf>… [cited by applicant]
Qin, Xuebin et al., “U 2-Net: Going Deeper with Nested U-Structure for Salient Object Detection”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2005.09… [cited by applicant]
Revanur, Ambareesh et al., “CoralStyleCLIP: Co-optimized Region and Layer Selection for Image Editing”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2… [cited by applicant]
Savarese, Pedro , “Information-Theoretic Segmentation by Inpainting Error Maximization”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2012.07287.pdf>.… [cited by applicant]
Shi, Jianbo et al., “Normalized Cuts and Image Segmentation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 22, No. 8 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://www.cs.swarthm… [cited by applicant]
Shi, Jianping et al., “Hierarchical Image Saliency Detection on Extended CSSD”, IEEE Transactions on Pattern Analysis and Machine Intelligence vol. 38, No. 4 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https… [cited by applicant]
Shin, Gyungin et al., “NamedMask: Distilling Segmenters from Complementary Foundation Models”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2209.11228… [cited by applicant]
Shin, Gyungin et al., “Unsupervised Salient Object Detection with Spectral Cluster Voting”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2203.12614.pd… [cited by applicant]
Simeoni, Oriane et al., “Localizing objects with self-supervised transformers and no labels”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2109.14279.… [cited by applicant]
Siméoni, Oriane et al., “Unsupervised Object Localization: Observing the Background to Discover Objects”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf… [cited by applicant]
Tang, Meng et al., “Normalized Cut Loss for Weakly-supervised CNN Segmentation”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1804.01346.pdf>., Apr. 4… [cited by applicant]
Vo, Huy , “Toward Unsupervised, Multi-Object Discovery in Large-Scale Image Collections”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2007.02662.pdf>… [cited by applicant]
Vo, Van Huy et al., “Large-Scale Unsupervised Object Discovery”, Advances in Neural Information Processing Systems 34 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://proceedings.neurips.cc/paper_files/pa… [cited by applicant]
Voynov, Andrey et al., “Object Segmentation Without Labels with Large-Scale Generative Models”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2006.0498… [cited by applicant]
Voynov, Andrey et al., “Unsupervised Discovery of Interpretable Directions in the GAN Latent Space”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2002… [cited by applicant]
Vu, Huy et al., “Unrolling of deep graph total variation for image denoising”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2010.11290.pdf>., Mar. 24,… [cited by applicant]
Wang, Lijun et al., “Learning to Detect Salient Objects with Image-level Supervision”, IEEE Conference on Computer Vision and Pattern Recognition [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://openacces… [cited by applicant]
Wang, Peng et al., “Salient object detection for searched web images via global saliency”, IEEE Conference on Computer Vision and Pattern Recognition [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://jingd… [cited by applicant]
Wang, Wenguan et al., “Salient Object Detection in the Deep Learning Era: An In-depth Survey”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1904.09146… [cited by applicant]
Wang, Xinlong , “FreeSOLO: Learning to Segment Objects without Annotations”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)[retrieved Jun. 14, 2023]. Retrieved from the Internet <https://authors.l… [cited by applicant]
Wang, Xinlong et al., “SOLO: segmenting objects by locations”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1912.04488.pdf]>., Jul. 19, 2020, 19 Pages. [cited by applicant]
Wang, Yangtao et al., “Tokencut: Segmenting objects in images and videos with self-supervised transformer and normalized cut”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <… [cited by applicant]
Wei, Xiu-Shen et al., “Unsupervised Object Discovery and Co-Localization by Deep Descriptor Transforming”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pd… [cited by applicant]
Wu, Z. et al., “An optimal graph theoretic approach to data clustering: theory and its application to image segmentation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 15, No. 11 [retrieved Jun. … [cited by applicant]
Yang, Chuan et al., “Saliency detection via graph-based manifold ranking”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://www.… [cited by applicant]
Yun, Yi Ke et al., “Selfreformer: Self-refined network with transformer for salient object detection”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/22… [cited by applicant]
Zhou, Yuan et al., “Salient Object Detection via Fuzzy Theory and Object-Level Enhancement”, IEEE Transactions on Multimedia, vol. 21, No. 1 [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://doi.org/10.110… [cited by applicant]