IP Library › Granted Patent US 12,646,176
Granted Patent B2
US 12,646,176 · App. 18/170,336 · Granted Jun 2, 2026

Transformer-based image segmentation on mobile devices

Inventors: Jingyuan Liu (Santa Clara, CA); Qing Liu (Santa Clara, CA); Jimei Yang (Merced, CA); Yuhong Wu (Sammamish, WA); Su Chen (San Jose, CA)
Assignee: Adobe Inc.
G06T7/11G06V10/267G06V10/7715G06V10/82G06V20/70G06T2207/20021G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,176
App. No.
18/170,336
Granted
Jun 2, 2026
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating segmentation masks for a digital visual media item. In particular, in one or more embodiments, the disclosed systems generate, utilizing a neural network encoder, high-level features of a digital visual media item. Further, the disclosed systems generate, utilizing the neural network encoder, low-level features of the digital visual media item. In some implementations, the disclosed systems generate, utilizing a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features. Moreover, the disclosed systems generate, utilizing the neural network decoder, a refined segmentation mask of the digital visual media item from the initial segmentation mask and the high-level features.

Claims (50)

1 . A computer-implemented method comprising:

generating, utilizing a neural network encoder, high-level features of a digital visual media item;

generating, utilizing the neural network encoder, low-level features of the digital visual media item;

generating, utilizing a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features; and

generating, utilizing the neural network decoder, a refined segmentation mask of the digital visual media item by passing the initial segmentation mask and the high-level features through the neural network decoder.

2 . The computer-implemented method of claim 1 , wherein the neural network encoder comprises a transformer block; and further comprising:

down-sampling a portion of the low-level features to match a dimension of a lower-level set of the low-level features; and

generating a transformed set of features by passing the down-sampled portion of the low- level features through the transformer block.

3 . The computer-implemented method of claim 1 , wherein generating the initial segmentation mask of the digital visual media item from the low-level features comprises decoding the low-level features, without the high-level features.

4 . The computer-implemented method of claim 3 , further comprising:

up-sampling a portion of the low-level features to match a dimension of a higher-level set of the low-level features; and

combining the up-sampled portion of the low-level features with the higher-level set of the low-level features utilizing a series sum operation.

5 . The computer-implemented method of claim 1 , wherein generating the refined segmentation mask of the digital visual media item from the initial segmentation mask and the high-level features comprises decoding the initial segmentation mask and the high-level features by:

generating, utilizing a multilayer perceptron, refined high-level features from the initial segmentation mask and the high-level features;

up-sampling the refined high-level features; and

combining the up-sampled refined high-level features.

6 . The computer-implemented method of claim 5 , wherein combining the up-sampled refined high-level features comprises utilizing a concatenation operation.

7 . The computer-implemented method of claim 5 , wherein combining the up-sampled refined high-level features comprises utilizing a series sum operation.

8 . The computer-implemented method of claim 1 , further comprising generating, utilizing a feature refinement head, a feature refinement segmentation mask of the digital visual media item.

9 . The computer-implemented method of claim 1 , wherein at least one of generating the initial segmentation mask or generating the refined segmentation mask comprises segmenting a portrait into a plurality of semantic regions.

10 . The computer-implemented method of claim 1 , wherein generating the initial segmentation mask and generating the refined segmentation mask are performed on a mobile device.

11 . A non-transitory computer-readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating, utilizing a neural network encoder, high-level features of a digital visual media item;

generating, utilizing the neural network encoder, low-level features of the digital visual media item;

generating, utilizing a first segmentation head of a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features; and

generating a refined segmentation mask of the digital visual media item by processing the initial segmentation mask and the high-level features utilizing a second segmentation head of the neural network decoder.

12 . The non-transitory computer-readable medium of claim 11 , wherein the neural network encoder comprises a transformer block; and wherein the operations further comprise:

down-sampling a portion of the low-level features to match a dimension of a lower-level set of the low-level features; and

generating a transformed set of features by passing the down-sampled portion of the low-level features through the transformer block.

13 . The non-transitory computer-readable medium of claim 11 , wherein generating the initial segmentation mask of the digital visual media item from the low-level features comprises decoding the low-level features without the high-level features.

14 . The non-transitory computer-readable medium of claim 11 , wherein the instructions cause a processor of a mobile device to perform the operations.

15 . A system comprising:

at least one processor; and

at least one memory device coupled to the at least one processor that causes the system to perform operations comprising:

generating, utilizing a neural network encoder, high-level features of a digital visual media item;

generating, utilizing the neural network encoder, low-level features of the digital visual media item;

generating, utilizing a neural network decoder, an initial segmentation mask of the digital visual media item from the low-level features; and

generating, utilizing the neural network decoder, a refined segmentation mask of the digital visual media item by processing the initial segmentation mask and the high-level features through the neural network decoder.

16 . The system of claim 15 , wherein the neural network encoder comprises a transformer block; and further comprising:

down-sampling a portion of the low-level features to match a dimension of a lower-level set of the low-level features; and

generating a transformed set of features by passing the down-sampled portion of the low-level features through the transformer block.

17 . The system of claim 15 , wherein generating the initial segmentation mask of the digital visual media item from the low-level features comprises decoding the low-level features, without the high-level features.

18 . The system of claim 17 , further comprising:

up-sampling a portion of the low-level features to match a dimension of a higher-level set of the low-level features; and

combining the up-sampled portion of the low-level features with the higher-level set of the low-level features utilizing a series sum operation.

19 . The system of claim 15 , wherein generating the refined segmentation mask of the digital visual media item from the initial segmentation mask and the high-level features comprises decoding the initial segmentation mask and the high-level features by:

generating, utilizing a multilayer perceptron, refined high-level features from the initial segmentation mask and the high-level features;

up-sampling the refined high-level features; and

combining the up-sampled refined high-level features.

20 . The system of claim 19 , wherein combining the up-sampled refined high-level features comprises utilizing a concatenation operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2023
From: LIU, JINGYUAN; LIU, QING; YANG, JIMEI; WU, YUHONG; CHEN, SU
To: ADOBE INC.
Reel/Frame 062724/0828 →
Continuity (1)
Related Publication 20240281978A1 · Aug 22, 2024
References Cited (23)
US 20210365717A1 · Cao · 2021 [cited by examiner]
US 20220044407A1 · Liu · 2022 [cited by examiner]
US 20220198209A1 · Spears · 2022 [cited by examiner]
US 20230334813A1 · Spears · 2023 [cited by examiner]
CN 113920099A · 2022 [cited by examiner]
CN 114743103A · 2022 [cited by examiner]
CN 114882599A · 2022 [cited by examiner]
Xu, X.—“MulTNet: A Multi-Scale Transformer Network for Marine Image Segmentation toward Fishing”—Sensors Sep. 23, 2022—pp. 1-17 (Year: 2022). [cited by examiner]
Zhang, G.—“RefineMask: Towards High-Quality Instance Segmentation with Fine-Grained Features”—CVPR 2021—pp. 6861-6869 (Year: 2021). [cited by examiner]
Li, Z.—“MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation”—arXiv—Nov. 25, 2021—pp. 1-10 (Year: 2021). [cited by examiner]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transfo… [cited by applicant]
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for mobilenetv3. In Int. Conf. Comput. Vis., pp. 1314-1324, 2019. [cited by applicant]
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017. [cited by applicant]
Deng, Jia, et al. “Imagenet: A large-scale hierarchical image database.” 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009. [cited by applicant]
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In IEEE Conf. Comput. Vis. Pattern Recog., pp. 4510-4520, 2018. [cited by applicant]
Sachin Mehta and Mohammad Rastegari. Mobilevit: Light-weight, general-purpose, and mobile-friendly vision trans-former. arXiv preprint arXiv:2110.02178, 2021. [cited by applicant]
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. Int. Conf. Comput. V… [cited by applicant]
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. arXiv:2104.13840, 2021. [cited by applicant]
Xie, Enze, et al. “SegFormer: Simple and efficient design for semantic segmentation with transformers.” Advances in Neural Information Processing Systems 34 (2021): 12077-12090. [cited by applicant]
Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Xiaoyi Dong, Lu Yuan, and Zicheng Liu. Mobile-former: Bridging mobilenet and transformer. arXiv preprint arXiv:2108.05895, 2021. [cited by applicant]
Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang. Hrformer: High-resolution transformer for dense prediction. arXiv preprint arXiv:2110.09408, 2021. [cited by applicant]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. Int. Conf. Comput. Vis., 2021. [cited by applicant]
Zhang, Wenqiang, et al. “TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. [cited by applicant]