IP Library › Granted Patent US 12,430,934
Granted Patent B2
US 12,430,934 · App. 18/316,617 · Granted Sep 30, 2025

Utilizing implicit neural representations to parse visual components of subjects depicted within visual content

Inventors: Mausoom Sarkar (New Delhi, IN); Nikitha S R (Namakkal, IN); Mayur Hemani (Noida, IN); Rishabh Jain (Gurgaon, IN); Balaji Krishnamurthy (Noida, IN)
Assignee: Adobe Inc.
G06V20/70G06T3/4046G06T7/11G06V10/46G06V10/764G06V10/7715G06V10/774G06V10/82G06V40/171G06V40/172G06T2207/20021G06T2207/20076G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,934
App. No.
18/316,617
Granted
Sep 30, 2025
Kind
B2
Abstract

This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that utilize a local implicit image function neural network to perform image segmentation with a continuous class label probability distribution. For example, the disclosed systems utilize a local-implicit-image-function (LIIF) network to learn a mapping from an image to its semantic label space. In some instances, the disclosed systems utilize an image encoder to generate an image vector representation from an image. Subsequently, in one or more implementations, the disclosed systems utilize the image vector representation with a LIIF network decoder that generates a continuous probability distribution in a label space for the image to create a semantic segmentation mask for the image. Moreover, in some embodiments, the disclosed systems utilize the LIIF-based segmentation network to generate segmentation masks at different resolutions without changes in an input resolution of the segmentation network.

Claims (43)

1. A computer-implemented method comprising:

generating, utilizing an image encoder, an image vector representation from an image depicting a subject;

utilizing a local implicit image function neural network to generate a continuous class label probability distribution for one or more class labels from the image vector representation; and

creating a semantic segmentation mask for the image comprising one or more labeled semantic regions based on the continuous class label probability distribution.

2. The computer-implemented method of claim 1 , further comprising utilizing the continuous class label probability distribution for the one or more class labels to determine a class prediction for a coordinate between pixels of the image.

3. The computer-implemented method of claim 1 , further comprising:

generating an unfolded image vector representation from the image vector representation utilizing an unfolding operation;

generating a reduced channel image vector representation by utilizing one or more multilayer perceptron decoders to reduce channels of the unfolded image vector representation; and

generating the continuous class label probability distribution for the one or more class labels from the reduced channel image vector representation.

4. The computer-implemented method of claim 1 , further comprising:

generating a global pool feature vector from the image vector representation utilizing global pooling; and

generating the continuous class label probability distribution for the one or more class labels from the global pool feature vector.

5. The computer-implemented method of claim 1 , wherein the image encoder comprises:

a residual block comprising instance normalization layers and convolution layers; and

strided convolution layers between one or more residual blocks.

6. The computer-implemented method of claim 1 , further comprising:

selecting a plurality of upsample coordinates for generating the semantic segmentation mask at an upsampled resolution; and

creating the semantic segmentation mask by generating semantic label predictions at the plurality of upsample coordinates utilizing the continuous class label probability distribution.

7. The computer-implemented method of claim 1 , further comprising utilizing the semantic segmentation mask to edit the one or more labeled semantic regions in the image.

8. The computer-implemented method of claim 1 , further comprising learning parameters of the local implicit image function neural network utilizing an edge-aware loss using a ground truth image with edges for known semantic regions of the ground truth image.

9. A system comprising:

a memory component comprising an image depicting a subject, a convolutional image encoder, and a local implicit image function neural network; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

generating, utilizing the convolutional image encoder, an image vector representation from the image depicting the subject;

generating, utilizing the local implicit image function neural network, a continuous class label probability distribution for one or more class labels from the image vector representation;

selecting a plurality of upsample coordinates for generating a semantic segmentation mask at an upsampled resolution; and

creating a semantic segmentation mask by generating semantic label predictions at the plurality of upsample coordinates utilizing the continuous class label probability distribution.

10. The system of claim 9 , wherein the operations further comprise generating the image depicting the subject by downsampling a higher resolution image depicting the subject.

11. The system of claim 9 , wherein the operations further comprise utilizing the continuous class label probability distribution for the one or more class labels to determine a class prediction for an upsample coordinate from the plurality of upsample coordinates between pixels of the image.

12. The system of claim 9 , wherein the operations further comprise generating, utilizing the local implicit image function neural network, the continuous class label probability distribution by utilizing a reduced channel image vector representation based on the image vector representation and a global pool feature vector based on the image vector representation.

13. The system of claim 9 , wherein the convolutional image encoder comprises a residual block comprising instance normalization layers and convolution layers.

14. The system of claim 9 , wherein the semantic segmentation mask comprises one or more labeled semantic regions based on the semantic label predictions at the plurality of upsample coordinates and wherein the operations further comprise utilizing the semantic segmentation mask to edit the one or more labeled semantic regions in a higher resolution image of the image.

15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating, utilizing an image encoder, an image vector representation from an image depicting a human face;

generating, utilizing a local implicit image function neural network, a continuous class label probability distribution for one or more facial feature labels; and

creating a semantic segmentation mask for the image comprising one or more labeled facial feature regions based on the continuous class label probability distribution.

16. The non-transitory computer-readable medium of claim 15 , wherein the one or more facial feature labels comprise an eye label, a nose label, a lips label, a skin label, an eyebrows label, a teeth label, and a hair label.

17. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise utilizing the continuous class label probability distribution for the one or more facial feature labels to determine a facial feature prediction for a coordinate between pixels of the image.

18. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise generating, utilizing the local implicit image function neural network, the continuous class label probability distribution based on:

a reduced channel image vector representation generated utilizing one or more multilayer perceptron decoders with the image vector representation; and

a global pool feature vector generated utilizing global pooling on the image vector representation.

19. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise creating the semantic segmentation mask at an upsampled resolution by generating semantic label predictions at a plurality of upsample coordinates utilizing the continuous class label probability distribution.

20. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise utilizing the semantic segmentation mask to edit the one or more labeled facial feature regions in the image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: SARKAR, MAUSOOM; S R, NIKITHA; HEMANI, MAYUR; JAIN, RISHABH; KRISHNAMURTHY, BALAJI
To: ADOBE INC.
Reel/Frame 063627/0115 →
Continuity (1)
Related Publication 20240378912A1 · Nov 14, 2024
References Cited (81)
US 11971955B1 · Chakraborty · 2024 [cited by examiner]
US 20150279113A1 · Knorr · 2015 [cited by examiner]
US 20200302225A1 · Dutta · 2020 [cited by examiner]
US 20210012576A1 · Riegler · 2021 [cited by examiner]
US 20210209837A1 · Chen · 2021 [cited by examiner]
US 20230101653A1 · Matsumura · 2023 [cited by examiner]
US 20230245495A1 · Ninh · 2023 [cited by examiner]
US 20230298272A1 · Ezhov · 2023 [cited by examiner]
US 20230306600A1 · Zhang · 2023 [cited by examiner]
US 20230377093A1 · Djelouah · 2023 [cited by examiner]
US 20230410339A1 · Sawarkar · 2023 [cited by examiner]
US 20240029203A1 · Xu · 2024 [cited by examiner]
US 20240062365A1 · Shen · 2024 [cited by examiner]
US 20240070809A1 · Xu · 2024 [cited by examiner]
US 20240089580A1 · Nomura · 2024 [cited by examiner]
US 20240249434A1 · Hong · 2024 [cited by examiner]
US 20240265676A1 · Janousková · 2024 [cited by examiner]
US 20240412319A1 · Singh · 2024 [cited by examiner]
Sarkar, M. “Parameter Efficient Local Implicit Image Function Network for Face Segmentation” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 20970-20980, pp. 1-11. (Year: 2023). [cited by examiner]
Chen, Y. “Learning Continuous Image Representation with Local Implicit Image Function” Computer Vision and Pattern Recognition (cs.CV); Machine Learning, pp. 1-11. (Year: 2021). [cited by examiner]
Aaron S Jackson, Michel Valstar, and Georgios Tzimiropoulos. A cnn cascade for landmark guided semantic part segmentation. In European Conference on Computer Vision, pp. 143-155. Springer, 2016. [cited by applicant]
Abdallah Dib, Cédric Thébault, Junghyun Ahn, Philippe-Henri Gosselin, Christian Theobalt, and Louis Chevallier. Towards high fidelity monocular face reconstruction with rich reflectance using self-supervised learning an… [cited by applicant]
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. [cited by applicant]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language superv… [cited by applicant]
Alexandros Lattas, Stylianos Moschoglou, Baris Gecer, Stylianos Ploumpis, Vasileios Triantafyllou, Abhijeet Ghosh, and Stefanos Zafeiriou. Avatarme: Realistically renderable 3d facial reconstruction “in-the-wild”. In Pr… [cited by applicant]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16×16 words: Transfo… [cited by applicant]
Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recognition wor… [cited by applicant]
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. [cited by applicant]
Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. Maskgan: Towards diverse and interactive facial image manipulation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. [cited by applicant]
Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. Maskgan: Towards diverse and interactive facial image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5549-5558… [cited by applicant]
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. [cited by applicant]
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022, 2016. [cited by applicant]
Eduard Ramon, Gil Triginer, Janna Escur, Albert Pumarola, Jaime Garcia, Xavier Giro-i Nieto, and Francesc Moreno-Noguer. H3d-net: Few-shot high-fidelity 3d head reconstruction. In Proceedings of the IEEE/CVF Internation… [cited by applicant]
Eric R Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In Proceedings of the IEEE/CVF conference on compute… [cited by applicant]
Frederick Ira Parke. A parametric model for human faces. Technical report, Utah Univ Salt Lake City Dept of Computer Science, 1974. [cited by applicant]
Gusi Te, Wei Hu, Yinglu Liu, Hailin Shi, and Tao Mei. Agrnet: Adaptive graph representation learning and reasoning for face parsing. IEEE Transactions on Image Processing, 30:8236-8250, 2021. [cited by applicant]
Gusi Te, Yinglu Liu, Wei Hu, Hailin Shi, and Tao Mei. Edgeaware graph representation learning and reasoning for face parsing. In European Conference on Computer Vision, pp. 258-274. Springer, 2020. [cited by applicant]
Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re… [cited by applicant]
He Zhang, Benjamin S Riggan, Shuowen Hu, Nathaniel J Short, and Vishal M Patel. Synthesis of high-quality visible faces from polarimetric thermal faces using generative adversarial networks. International Journal of Com… [cited by applicant]
Huawei Wei, Shuang Liang, and Yichen Wei. 3d dense face alignment via graph convolution networks. arXiv preprint arXiv:1904.05562, 2019. [cited by applicant]
Ira Kemelmacher-Shlizerman. Transfiguring portraits. 35(4), 2016. [cited by applicant]
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on compu… [cited by applicant]
Jiaxiang Shang, Tianwei Shen, Shiwei Li, Lei Zhou, Mingmin Zhen, Tian Fang, and Long Quan. Self-supervised monocular 3d face reconstruction by occlusion-aware multiview geometry consistency. In European Conference on Co… [cited by applicant]
Jinpeng Lin, Hao Yang, Dong Chen, Ming Zeng, Fang Wen, and Lu Yuan. Face parsing with Rol tanh-warping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5654-5663, 2019. [cited by applicant]
Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3431-3440, 2015. [cited by applicant]
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE conference on computer vision and p… [cited by applicant]
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF conference on computer vision an… [cited by applicant]
Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview neural surface reconstruction by disentangling geometry and appearance. Advances in Neural Information Processing… [cited by applicant]
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. of the… [cited by applicant]
Michael Zollhofer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt. State of the art on monocular 3d face reconstruction, tracking, and … [cited by applicant]
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognitio… [cited by applicant]
Petr Kellnhofer, Lars C Jebe, Andrew Jones, Ryan Spicer, Kari Pulli, and Gordon Wetzstein. Neural lumigraph rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4287-4297,… [cited by applicant]
Ping Luo, Xiaogang Wang, and Xiaoou Tang. Hierarchical face parsing via deep learning. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 2480-2487, 2012. [cited by applicant]
Qi Zheng, Jiankang Deng, Zheng Zhu, Ying Li, and Stefanos Zafeiriou. Decoupled multi-task learning with cyclical selfregulation for face parsing. In Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Sifei Liu, Jianping Shi, Ji Liang, and Ming Hsuan Yang. Face parsing via recurrent propagation. In British Machine Vision Conference 2017, BMVC 2017, British Machine Vision Conference 2017, BMVC 2017. BMVA Press, 2017. [cited by applicant]
Sifei Liu, Jimei Yang, Chang Huang, and Ming-Hsuan Yang. Multi-objective convolutional learning for face labeling. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3451-3459, 2015. [cited by applicant]
Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh. Neural vols. Learning dynamic renderable volumes from images. arXiv preprint arXiv:1906.07751, 2019. [cited by applicant]
Stylianos Ploumpis, Evangelos Ververas, Eimear O'Sullivan, Stylianos Moschoglou, Haoyang Wang, Nick Pears, William AP Smith, Baris Gecer, and Stefanos Zafeiriou. Towards a complete 3d morphable model of the human head. … [cited by applicant]
Tarun Yenamandra, Ayush Tewari, Florian Bernard, Hans-Peter Seidel, Mohamed Elgharib, Daniel Cremers, and Christian Theobalt. i3dmm: Deep implicit 3d morphable model of human heads. In Proceedings of the IEEE/CVF Confer… [cited by applicant]
Thomas Gerig, Andreas Morel-Forster, Clemens Blumer, Bernhard Egger, Marcel Luthi, Sandro Schonborn, and Thomas Vetter. Morphable face models—an open framework. In 2018 13th IEEE International Conference on Automatic Fa… [cited by applicant]
Tianchu Guo, Youngsung Kim, Hui Zhang, Deheng Qian, ByungIn Yoo, Jingtao Xu, Dongqing Zou, Jae-Joon Han, and Changkyu Choi. Residual encoder decoder network and adaptive prior for face parsing. In Proceedings of the AAA… [cited by applicant]
Vincent Sitzmann, Julien N.P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. In Proc. NeurIPS, 2020. [cited by applicant]
Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Proceedings of the 26th annual conference on Computer graphics and interactive techniques, pp. 187-194, 1999. [cited by applicant]
Vuong Le, Jonathan Brandt, Zhe Lin, Lubomir Bourdev, and Thomas S Huang. Interactive facial feature localization. In ECCV. Springer, 2012. [cited by applicant]
Xiangtai Li, Ansheng You, Zhen Zhu, Houlong Zhao, Maoke Yang, Kuiyuan Yang, and Yunhai Tong. Semantic flow for fast and accurate scene parsing. In ECCV, 2020. [cited by applicant]
Xiaoguang Tu, Jian Zhao, Zihang Jiang, Yao Luo, Mei Xie, Yang Zhao, Linxiao He, Zheng Ma, and Jiashi Feng. Joint 3d face reconstruction and dense face alignment from a single image with 2d- assisted self-supervised lear… [cited by applicant]
Xinyu Ou, Si Liu, Xiaochun Cao, and Hefei Ling. Beauty emakeup: A deep makeup transfer system. In Proceedings of the 24th ACM International Conference on Multimedia, p. 701-702, New York, NY, USA, 2016. Association for … [cited by applicant]
Yao Feng, Fan Wu, Xiaohu Shao, Yanfeng Wang, and Xi Zhou. Joint 3d face reconstruction and dense alignment with position map regression network. In Proceedings of the European conference on computer vision (ECCV), pp. 5… [cited by applicant]
Yao Feng, Haiwen Feng, Michael J Black, and Timo Bolkart. Learning an animatable detailed 3d face model from in-the-wild images. ACM Transactions on Graphics (ToG), 40(4):1-13, 2021. [cited by applicant]
Yijun Li, Sifei Liu, Jimei Yang, and Ming-Hsuan Yang. Generative face completion. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3911-3919, 2017. [cited by applicant]
Yiming Lin, Jie Shen, Yujiang Wang, and Maja Pantic. Roi tanh-polar transformer network for face parsing in the wild, 2021. [cited by applicant]
Yinbo Chen, Sifei Liu, and Xiaolong Wang. Learning continuous image representation with local implicit image function. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8628-8638,… [cited by applicant]
Yinglin Zheng, Hao Yang, Ting Zhang, Jianmin Bao, Dongdong Chen, Yangyu Huang, Lu Yuan, Dong Chen, Ming Zeng, and Fang Wen. General facial representation learning in a visual-linguistic manner. In Proceedings of the IEE… [cited by applicant]
Yinglu Liu, Hailin Shi, Hao Shen, Yue Si, Xiaobo Wang, and Tao Mei. A new dataset and boundary—attention semantic segmentation for face parsing. In AAAI, pp. 11637-11644, 2020. [cited by applicant]
Yisu Zhou, Xiaolin Hu, and Bo Zhang. Interlinked convolutional neural networks for face parsing. In Advances in Neural Networks—ISNN 2015, pp. 222-231. Springer International Publishing, 2015. [cited by applicant]
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. In Proceedings of the IEEE/CVF Conference on Compu… [cited by applicant]
Yuval Nirkin, Iacopo Masi, Anh Tran Tuan, Tal Hassner, and Gerard Medioni. On face segmentation, face swapping, and face perception. In 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG… [cited by applicant]
Yuval Nirkin, Yosi Keller, and Tal Hassner. Fsgan: Subject agnostic face swapping and reenactment. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 7184-7193, 2019. [cited by applicant]
Z. Shu, E. Yumer, S. Hadap, K. Sunkavalli, E. Shechtman, and D. Samaras. Neural face editing with intrinsic image disentangling. In Computer Vision and Pattern Recognition, 2017. CVPR 2017. IEEE Conference on. IEEE, 201… [cited by applicant]
Zhen Wei, Si Liu, Yao Sun, and Hefei Ling. Accurate facial image parsing at real-time speed. IEEE Transactions on Image Processing, 28(9):4659-4670, 2019. [cited by applicant]
Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5939-5948, 2019. [cited by applicant]
Cited By (1)
US 12,646,253