IP Library › Granted Patent US 12,614,407
Granted Patent B2
US 12,614,407 · App. 17/552,857 · Granted Apr 28, 2026

Generating segmentation masks for objects in digital videos using pose tracking data

Inventors: Seoung Wug Oh (San Jose, CA); Miran Heo (Seoul, KR); Joon-Young Lee (Miliptas, CA)
Assignee: Adobe Inc.
G06V40/103G06T7/10G06T7/70G06V10/40G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,407
App. No.
17/552,857
Granted
Apr 28, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that generate joint-based segmentation masks for digital objects portrayed in digital videos. In particular, in one or more embodiments, the disclosed systems utilize a video masking model having a pose tracking neural network and a segmentation neural network to generate the joint-based segmentation masks. To illustrate, in some embodiments, the disclosed systems utilize the pose tracking neural network to identify a set of joints of the digital object across the frames of the digital video. The disclosed systems further utilize the segmentation neural network to generate joint-based segmentation masks for the video frames that portray the object using the identified joints. In some cases, the segmentation neural network includes a multi-layer perceptron mixer layer for mixing visual features propagated via convolutional layers.

Claims (57)

1 . A method comprising:

generating a down-sampled frame corresponding to a frame of a digital video;

determining, using a global pose tracking neural network and from the down-sampled frame, a set of joint coordinates corresponding to a digital object portrayed in the frame of the digital video;

generating, using the set of joint coordinates, a joint heat map corresponding to the digital object portrayed in the frame of the digital video, the joint heat map having values that distinguish between joint points corresponding to the set of joint coordinates and areas of the frame of the digital video that surround the joint points;

determining, using one or more convolutional layers of a local segmentation neural network, visual features from the joint heat map and the frame of the digital video;

generating, using at least one multi-perceptron mixer layer of the local segmentation neural network, mixer encodings that mix the visual features across feature channels by generating the mixer encodings from the visual features using layer normalization and an activation function of the at least one multi-perceptron mixer layer; and

generating, using the local segmentation neural network and from the mixer encodings, a joint-based segmentation mask that corresponds to the digital object portrayed in the frame.

2 . The method of claim 1 , wherein determining, using the global pose tracking neural network and the down-sampled frame, the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video comprises determining the set of joint coordinates by using the global pose tracking neural network to globally analyze the down-sampled frame and an additional down-sampled frame corresponding to a preceding frame of the digital video.

3 . The method of claim 1 , wherein generating the joint-based segmentation mask that corresponds to the digital object comprises generating the joint-based segmentation mask from a bounding box associated with the digital object for the frame of the digital video.

4 . The method of claim 1 , wherein generating, using the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask comprises generating, using one or more additional convolutional layers of the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask.

5 . The method of claim 3 , further comprising:

determining the bounding box associated with the digital object for the frame of the digital video based on the set of joint coordinates; and

generating, from the frame of the digital video, a cropped frame that includes the digital object using the bounding box,

wherein generating the joint-based segmentation mask from the bounding box associated with the digital object comprises generating the joint-based segmentation mask from the cropped frame.

6 . The method of claim 1 ,

further comprising extracting, from the digital video, the frame of the digital video and a preceding frame of the digital video,

wherein determining the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video comprises determining the set of joint coordinates using the frame of the digital video and the preceding frame of the digital video.

7 . The method of claim 1 , wherein generating the joint heat map using the set of joint coordinates comprises centering a Gaussian distribution at each joint point associated with the set of joint coordinates.

8 . The method of claim 1 , further comprising modifying the frame of the digital video utilizing the joint-based segmentation mask.

9 . A non-transitory computer-readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating a down-sampled frame corresponding to a frame of a digital video;

determining, using a global pose tracking neural network and from the down-sampled frame, a set of joint coordinates corresponding to a digital object portrayed in the frame of the digital video;

generating, using the set of joint coordinates, a joint heat map corresponding to the digital object portrayed in the frame of the digital video, the joint heat map having values that distinguish between joint points corresponding to the set of joint coordinates and areas of the frame of the digital video that surround the joint points;

determining, using one or more convolutional layers of a local segmentation neural network, visual features from the joint heat map and the frame of the digital video;

generating, using at least one multi-perceptron mixer layer of the local segmentation neural network, mixer encodings that mix the visual features across feature channels by generating the mixer encodings from the visual features using layer normalization and an activation function of the at least one multi-perceptron mixer layer; and

generating, using the local segmentation neural network and from the mixer encodings, a joint-based segmentation mask that corresponds to the digital object portrayed in the frame.

10 . The non-transitory computer-readable medium of claim 9 , wherein:

determining, utilizing the global pose tracking neural network and the down-sampled frame, the set of joint coordinates comprises determining the set of joint coordinates by globally analyzing the down-sampled frame using the global pose tracking neural network; and

determining the visual features from the joint heat map and the frame of the digital video using the one or more convolutional layers of the local segmentation neural network comprises determining the visual features by locally analyzing a portion of the frame of the digital video that contains the digital object based on the joint heat map using the one or more convolutional layers of the local segmentation neural network.

11 . The non-transitory computer-readable medium of claim 10 , wherein:

the operations further comprise generating, from the frame of the digital video, a cropped frame that includes the digital object based on the set of joint coordinates; and

determining the visual features by locally analyzing the portion of the frame of the digital video that contains the digital object based on the joint heat map utilizing the one or more convolutional layers of the local segmentation neural network comprises determining the visual features based on the cropped frame and the joint heat map utilizing the one or more convolutional layers of the local segmentation neural network.

12 . The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise determining, utilizing the global pose tracking neural network, a bounding box associated with the digital object and a tracking identifier that distinguishes the digital object from other digital objects portrayed in the frame of the digital video.

13 . The non-transitory computer-readable medium of claim 9 , wherein determining, utilizing the global pose tracking neural network, the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video comprises determining, utilizing the global pose tracking neural network, the set of joint coordinates associated with a set of pre-determined human joints corresponding to a person portrayed in the frame of the digital video.

14 . The non-transitory computer-readable medium of claim 9 , wherein generating, using the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask comprises generating, using one or more additional convolutional layers of the local segmentation neural network and from the mixer encodings, the joint-based segmentation mask.

15 . The non-transitory computer-readable medium of claim 10 , wherein the operations further comprise:

determining, utilizing the global pose tracking neural network, an additional set of joint coordinates corresponding to an additional digital object portrayed in the frame of the digital video; and

generating, utilizing the local segmentation neural network, an additional joint-based segmentation mask corresponding to the additional digital object based on the additional set of joint coordinates.

16 . The non-transitory computer-readable medium of claim 10 , wherein the operations further comprise:

determining intermediate feature values associated with the digital object portrayed in the frame of the digital video utilizing the global pose tracking neural network; and

providing the intermediate feature values from the global pose tracking neural network to the local segmentation neural network for generating the joint-based segmentation mask via a skip link.

17 . A system comprising:

at least one memory device; and

at least one server device configured to cause the system to:

generate a down-sampled frame corresponding to a frame of a digital video;

determine, utilizing a global pose tracking neural network and from the down-sampled frame, a set of joint coordinates corresponding to a digital object portrayed in the frame of the digital video;

generate, using the set of joint coordinates, a joint heat map corresponding to the digital object portrayed in the frame of the digital video, the joint heat map having values that distinguish between joint points corresponding to the set of joint coordinates and areas of the frame of the digital video that surround the joint points;

determine, utilizing an encoder of a segmentation neural network, mixer encodings for the digital object based on the joint heat map, the encoder comprising:

a plurality of convolutional layers that determine a set of features using the joint heat map; and

a multi-layer perceptron mixer layer that determines the mixer encodings by mixing the set of features using layer normalization and an activation function; and

generate, utilizing a decoder of the segmentation neural network, a joint-based segmentation mask for the digital object portrayed in the frame of the digital video based on the mixer encodings.

18 . The system of claim 17 , wherein the at least one server device is configured to cause the system to determine, utilizing the global pose tracking neural network, the set of joint coordinates corresponding to the digital object portrayed in the frame of the digital video by determining, utilizing the global pose tracking neural network, the set of joint coordinates based on the down-sampled frame and an additional down-sampled frame that corresponds to a preceding frame of the digital video that portrays the digital object.

19 . The system of claim 17 , wherein the at least one server device is further configured to cause the system to:

determine intermediate feature values associated with the digital object utilizing the multi-layer perceptron mixer layer of the encoder; and

provide the intermediate feature values from the encoder of the segmentation neural network to the decoder of the segmentation neural network for generating the joint-based segmentation mask via a skip link.

20 . The system of claim 17 , wherein the at least one server device is configured to cause the system to generate, utilizing the decoder of the segmentation neural network, the joint-based segmentation mask based on the mixer encodings by:

generating, using one or more convolutional layers of the decoder and from the mixer encodings, the joint-based segmentation mask.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: OH, SEOUNG WUG; HEO, MIRAN; LEE, JOON-YOUNG
To: ADOBE INC.
Reel/Frame 058408/0022 →
Continuity (1)
Related Publication 20230196817A1 · Jun 22, 2023
References Cited (76)
US 10961170B2 · Hermans · 2021 [cited by examiner]
US 11417061B1 · Feng · 2022 [cited by examiner]
US 12014472B1 · Mirhosseini · 2024 [cited by examiner]
US 20110267344A1 · Germann · 2011 [cited by examiner]
US 20120142479A1 · Serrarens · 2012 [cited by applicant]
US 20180232887A1 · Lin · 2018 [cited by examiner]
US 20200272812A1 · Wang · 2020 [cited by examiner]
US 20200327334A1 · Goren · 2020 [cited by examiner]
US 20200327465A1 · Baek · 2020 [cited by examiner]
US 20200349722A1 · Schmid · 2020 [cited by examiner]
US 20210201052A1 · Ranga · 2021 [cited by examiner]
US 20210216759A1 · Asayama · 2021 [cited by examiner]
US 20220148453A1 · Lin · 2022 [cited by examiner]
US 20220156943A1 · Zhang · 2022 [cited by examiner]
US 20220375211A1 · Tolstikhin · 2022 [cited by examiner]
US 20230154091A1 · Cho · 2023 [cited by examiner]
US 20230336758A1 · Ikonin · 2023 [cited by examiner]
US 20240303859A1 · Nakamura · 2024 [cited by examiner]
US 20240350010A1 · Krueger · 2024 [cited by examiner]
CN 111179281A · 2020 [cited by examiner]
CN 113177946A · 2021 [cited by applicant]
CN 113610102A · 2021 [cited by applicant]
WO 2021186223A1 · 2021 [cited by applicant]
Examination Report as received in GB application 2214974.4 dated May 5, 2023. [cited by applicant]
Mykhaylo Andriluka, Umar Iqbal, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, and Bernt Schiele. Posetrack: A benchmark for human pose estimation and tracking. In CVPR, 2018. [cited by applicant]
Ali Athar, Sabarinath Mahadevan, Aljosa Osep, Laura Leal-Taixe, and Bastianan Leibe. Stem-seg: Spatio-temporal embeddings for instance segmentation in videos. In ECCV, 2020. [cited by applicant]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. [cited by applicant]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In ICLR, 2015. [cited by applicant]
Gedas Bertasius and Lorenzo Torresani. Classifying, segmenting, and tracking object instances in video with mask propagation. In CVPR, 2020. [cited by applicant]
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? arXiv preprint arXiv:2102.05095, 2021. [cited by applicant]
Daniel Bolya, Chong Zhou, Fanyi Xiao, and Yong Jae Lee. Yolact: Real-time instance segmentation. In ICCV, 2019. [cited by applicant]
Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In CVPR, 2018. [cited by applicant]
Jiale Cao, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, and Ling Shao. Sipmask: Spatial information preservation for fast image and video instance segmentation. In ECCV, 2020. [cited by applicant]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. [cited by applicant]
Liang-Cheih Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. In IEEE TPAMI, 201… [cited by applicant]
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018. [cited by applicant]
Bowen Cheng, Ross Girshick, Piotr Dollar, Alexander C Berg, and Alexander Kirillov. Boundary iou: Improving object-centric image segmentation evaluation. In CVPR, 2021. [cited by applicant]
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. ICLR, 2020. [cited by applicant]
Rohit Girdhar, Georgia Gkioxari, Lorenzo Torresani, Manohar Paluri, and Du Tran. Detect-and-track: Efficient pose estimation in videos. In CVPR, 2018. [cited by applicant]
Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, 2017. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. [cited by applicant]
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus).arXiv preprint arXiv:1606.08415, 2016. [cited by applicant]
Sukjun Hwang, Miran Heo, Seoung Wug Oh, and Seon Joo Kim. Video instance segmentation using inter-frame communication transformers.arXiv preprint arXiv:2106.03299, 2021. [cited by applicant]
Umar Iqbal, Anton Milan, and Juergen Gall. Posetrack: Joint multi-person pose estimation and tracking. InCVPR, 2017. [cited by applicant]
Sheng Jin, Wentao Liu, Wanli Ouyang, and Chen Qian. Multi-person articulated tracking with spatial and temporal embeddings. In CVPR, 2019. [cited by applicant]
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset.arXiv preprint arXiv:170… [cited by applicant]
Harold W Kuhn. The hungarian method for the assignment problem. InNaval research logistics quarterly, 1955. [cited by applicant]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. [cited by applicant]
Dongfang Liu, Yiming Cui, Wenbo Tan, and Yingjie Chen. Sg-net: Spatial granularity network for one-stage video instance segmentation. InCVPR, 2021. [cited by applicant]
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InICLR, 2019. [cited by applicant]
David G Lowe. Distinctive image features from scale-invariant keypoints. 2004. [cited by applicant]
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In3DV, 2016. [cited by applicant]
Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim. Video object segmentation using space-time memory networks. InICCV, 2019. [cited by applicant]
Yaadhav Raaj, Haroon Idrees, Gines Hidalgo, and Yaser Sheikh. Efficient online multi-person 2d pose tracking with recurrent spatio-temporal affinity fields. InCVPR, 2019. [cited by applicant]
Michael Snower, Asim Kadav, Farley Lai, and Hans Peter Graf. 15 keypoints is all you need. InCVPR, 2020. [cited by applicant]
Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In ICCV,2019. [cited by applicant]
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, et al. Mlp-mixer: An all-mlp architecture for vision.arXiv … [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InNeurIPS, 2017. [cited by applicant]
Paul Voigtlaender, Michael Krause, Aljosa Osep, Jonathon Luiten, Berin Balachandar Gnana Sekar, Andreas Geiger, and Bastian Leibe. Mots: Multi-object tracking and segmentation. In CVPR, 2019. [cited by applicant]
Huiyu Wang, Yukun Zhu, Bradley Green, Hartwig Adam, 1026 Alan Yuille, and Liang-Chieh Chen. Axial-deeplab: Stand-alone axial-attention for panoptic segmentation. In ECCV, 1028 2020. [cited by applicant]
Manchen Wang, Joseph Tighe, and Davide Modolo. Combining detection and tracking for human pose estimation in videos. In CVPR, pp. 11088-11096, 2020. [cited by applicant]
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018. [cited by applicant]
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In CVPR, 1037 2020. [cited by applicant]
Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang, Ming Yang, and Junsong Yuan. Track to detect and segment: An online multi-object tracker. arXiv preprint arXiv:2103.08808, 2021. [cited by applicant]
Bin Xiao, Haiping Wu, and Yichen Wei. Simple baselines for human pose estimation and tracking. In ECCV, 2018. [cited by applicant]
Yuliang Xiu, Jiefeng Li, Haoyu Wang, Yinghong Fang, and Cewu Lu. Pose flow: Efficient online pose tracking. arXiv preprint arXiv:1802.00977, 2018. [cited by applicant]
Zhenbo Xu, Wei Zhang, Xiao Tan, Wei Yang, Huan Huang, Shilei Wen, Errui Ding, and Liusheng Huang. Segment as points for efficient online multi-object tracking and segmentation. In ECCV, 2020. [cited by applicant]
Linjie Yang, Yuchen Fan, and Ning Xu. Video instance segmentation. In ICCV, 2019. [cited by applicant]
Shusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. Crossover learning for fast online video instance segmentation. arXiv preprint arXiv:2104.05970, 2021. [cited by applicant]
Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell. Deep layer aggregation. In CVPR, 2018. [cited by applicant]
Xingyi Zhou, Vladlen Koltun, and Philipp Krahenbuhl. Tracking objects as points. In ECCV, 2020. [cited by applicant]
Xingyi Zhou, Dequan Wang, and Philipp Krahenbuhl. Objects as points. arXiv preprint arXiv:1904.07850, 2019. [cited by applicant]
Rempe et al., in Contact and Human Dynamics from Monocular Video, ECCV 2020 available at https://arxiv.org/abs/2007.11678. [cited by applicant]
Examination Report as received in GB application 2214974.4 dated Aug. 7, 2024. [cited by applicant]
Examination Report as received in GB application 2214974.4 dated Nov. 4, 2025. [cited by applicant]
Office Action as received in Chinese Application 2022-11165862.2 dated Nov. 26, 2025. [cited by applicant]