IP Library › Granted Patent US 12,223,661
Granted Patent B2
US 12,223,661 · App. 17/735,728 · Granted Feb 11, 2025

System for automatic object mask and hotspot tracking

Inventors: Lu Zhang (Dalian, CN); Jianming Zhang (Campbell, CA); Zhe Lin (Fremont, CA); Radomir Mech (Mountain View, CA)
Assignee: ADOBE INC.
G06T7/20G06F3/013G06F3/04845G06N3/02G06T7/0012G06T11/60G06T2207/10016G06T2207/20084G06T2207/30041
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,661
App. No.
17/735,728
Granted
Feb 11, 2025
Kind
B2
Abstract

Systems and methods provide editing operations in a smart editing system that may generate a focal point within a mask of an object for each frame of a video segment and perform editing effects on the frames of the video segment to quickly provide users with natural video editing effects. An eye-gaze network may produce a hotspot map of predicted focal points in a video frame. These predicted focal points may then be used by a gaze-to-mask network to determine objects in the image and generate an object mask for each of the detected objects. This process may then be repeated to effectively track the trajectory of objects and object focal points in videos. Based on the determined trajectory of an object in a video clip and editing parameters, the editing engine may produce editing effects relative to an object for the video clip.

Claims (52)

1. A method comprising:

generating a mask of an object in a video segment based on an initial frame of the video segment and a first focal region associated with the object in the initial frame;

for each frame of a plurality of frames of the video segment, generating, based on the frame and the mask of the object, using a neural network, a corresponding focal region associated with the mask of the object that tracks the object through the plurality of frames; and

executing an operation based on the corresponding focal region for one or more frames of the plurality of frames.

2. The method of claim 1 , further comprising generating the mask using a masking network.

3. The method of claim 1 , further comprising generating the first focal region, using a focal region prediction network and based on the initial frame of the video segment, by extracting spatial features of the initial frame during an encoding process and decoding the spatial features.

4. The method of claim 1 , wherein the neural network comprises a first branch configured to generate the corresponding focal region and a second branch configured to generate a corresponding mask of the object in the frame.

5. The method of claim 1 , further comprising generating the mask by:

receiving, at a decoder block of a masking network, spatial information extracted by an encoder block of the masking network via a skip connection; and

integrating the spatial information with a representation of the first focal region.

6. The method of claim 1 , further comprising generating the mask of the object by:

extracting, by an encoder block of a masking network, image features of the initial frame;

receiving the image features at a decoder branch via a first skip connection from the encoder block;

receiving a representation of the first focal region at the decoder branch via a second skip connection;

generating an integrated input that integrates the image features and the representation of the first focal region; and

generating, by the decoder branch, the mask of the object based on the integrated input.

7. The method of claim 1 , further comprising generating the corresponding focal region for the frame by: generating a template image of the object by cropping the initial frame based on the mask, and feeding the template image and the frame into the neural network.

8. A computer system comprising:

one or more hardware processors and memory configured to provide computer program instructions to the one or more hardware processors;

a segmentation and hotspot module configured to use the one or more hardware processors and the memory to:

generate a mask of an object in a video segment based on an initial frame of the video segment and a first focal region associated with the object in the initial frame; and

for each frame of a plurality of frames of the video segment, generate, based on the frame and the mask of the object, using a neural network, a corresponding focal region associated with the mask of the object that tracks the object through the plurality of frames; and

a crop suggestion module configured to use the one or more hardware processors and the memory to execute an operation based on the corresponding focal region for one or more frames of the plurality of frames.

9. The computer system of claim 8 , further comprising a masking network configured to generate the mask.

10. The computer system of claim 8 , further comprising a focal region prediction network configured to generate the first focal region based on the initial frame of the video segment by extracting spatial features of the initial frame during an encoding process and decoding the spatial features.

11. The computer system of claim 8 , wherein the neural network comprises a first branch to generate the corresponding focal region and a second branch configured to generate a corresponding mask of the object in the frame.

12. The computer system of claim 8 , further comprising a masking network configured to generate the mask by:

receiving, at a decoder block of the masking network, spatial information extracted by an encoder block of the masking network via a skip connection; and

integrating the spatial information with a representation of the first focal region.

13. The computer system of claim 8 , further comprising a masking network configured to generate the mask by:

extracting, by an encoder block of the masking network, image features of the initial frame;

receiving the image features at a decoder branch via a first skip connection from the encoder block;

receiving a representation of the first focal region at the decoder branch via a second skip connection;

generating an integrated input that integrates the image features and the representation of the first focal region; and

generating, by the decoder branch, the mask of the object based on the integrated input.

14. The computer system of claim 8 , wherein the segmentation and hotspot module is further configured to generate the corresponding focal region for the frame based on the mask by generating a template image of the object by cropping the initial frame based on the mask, and feeding the template image and the frame into the neural network.

15. The computer system of claim 8 , wherein the crop suggestion module is further configured to execute an editing operation based on the corresponding focal region for the one or more frames.

16. One or more non-transitory computer-readable storage media storing instructions executable by a computing device to cause the computing device to perform operations comprising:

generating a mask of an object in a video segment based on an initial frame of the video segment and a first focal region associated with the object in the initial frame;

for each frame of a plurality of frames of the video segment, generating, based on the frame and the mask of the object, using a neural network, a corresponding focal region associated with the mask of the object that tracks the object through the plurality of frames; and

executing an operation based on the corresponding focal region for one or more frames of the plurality of frames.

17. The one or more non-transitory computer-readable storage media of claim 16 , the operations further comprising generating the mask using a masking network and generating the first focal region using a focal region prediction network and based on the initial frame of the video segment by extracting spatial features of the initial frame during an encoding process and decoding the spatial features.

18. The one or more non-transitory computer-readable storage media of claim 16 , the operations further comprising generating the mask by:

receiving, at a decoder block of a masking network, spatial information extracted by an encoder block of the masking network via a skip connection; and

integrating the spatial information with a representation of the first focal region.

19. The one or more non-transitory computer-readable storage media of claim 16 , the operations further comprising generating the mask by:

extracting, by an encoder block of a masking network, image features of the initial frame;

receiving the image features at a decoder branch via a first skip connection from the encoder block;

receiving a representation of the first focal region at the decoder branch via a second skip connection;

generating an integrated input that integrates the image features and the representation of the first focal region; and

generating, by the decoder branch, the mask of the object based on the integrated input.

20. The one or more non-transitory computer-readable storage media of claim 16 , wherein the operations further comprise generating the corresponding focal region for the frame based on the mask comprises generating a template image of the object by cropping the initial frame based on the mask, and feeding the template image and the frame into the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2022
From: ZHANG, LU; ZHANG, JIANMING; LIN, ZHE; MECH, RADOMIR
To: ADOBE INC.
Reel/Frame 059885/0494 →
Continuity (2)
Continuation 16900483 · Jun 12, 2020
Related Publication 20220262011A1 · Aug 18, 2022
References Cited (27)
US 9626584B2 · Lin et al. · 2017 [cited by applicant]
US 10867422B2 · Zhang et al. · 2020 [cited by applicant]
US 11367199B2 · Zhang · 2022 [cited by examiner]
US 20060215752A1 · Lee · 2006 [cited by examiner]
US 20120281127A1 · Marino et al. · 2012 [cited by applicant]
US 20160063303A1 · Cheung et al. · 2016 [cited by applicant]
US 20170324624A1 · Taine · 2017 [cited by examiner]
US 20170324785A1 · Taine · 2017 [cited by examiner]
US 20190026864A1 · Chen et al. · 2019 [cited by applicant]
US 20200143171A1 · Lee · 2020 [cited by examiner]
US 20210056661A1 · Holmes et al. · 2021 [cited by applicant]
US 20210218929A1 · Huynh Thien · 2021 [cited by examiner]
Chen, J., et al., “Automatic Image Cropping: A Computational Complexity Study”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 507-515 (2016). [cited by applicant]
Everingham, M., and Winn, J., “The Pascal Visual Object Classes Challenge”, International journal of computer vision, pp. 1-23 (2007). [cited by applicant]
Jain, S., D., et al., “FusionSeg: Learning To Combine Motion And Appearance For Fully Automatic Segmentation Of Generic Objects In Videos”, In IEEE conference on computer vision and pattern recognition (CVPR), pp. 3664-… [cited by applicant]
Jiang, L., et al., “DeepVS: A Deep Learning Based Video Saliency Prediction Approach”, In Proceedings of the european conference on computer vision (ECCV), pp. 1-16 (2018). [cited by applicant]
Jiang, M., et al., “SALICON: Saliency In Context”, In Proceedings of the IEEE conference on computer vision and pattern recognition, IEEE, pp. 1072-1080 (2015). [cited by applicant]
Li, B., et al., “High Performance Visual Tracking With Siamese Region Proposal Network”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 8971-8980 (2018). [cited by applicant]
Li, B., et al., “Siamrpn++: Evolution Of Siamese Visual Tracking With Very Deep Networks”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4282-4291 (2019). [cited by applicant]
Lu, X., et al., “See More, Know More: Unsupervised Video Object Segmentation With Co-Attention Siamese Networks”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3623-3632 (2019… [cited by applicant]
Perazzi, F., et al., “A Benchmark Dataset And Evaluation Methodology For Video Object Segmentation”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 724-732 (2016). [cited by applicant]
Tan, M., and Le, Q., V., “Efficientnet: Rethinking Model Scaling For Convolutional Neural Networks”, In 36th International Conference on Machine Learning, PMLR, pp. 1-10 (2019). [cited by applicant]
Wang, Q., et al., “Fast Online Object Tracking And Segmentation: A Unifying Approach”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1328-1338 (2019). [cited by applicant]
Wang, W., et al., “Learning Unsupervised Video Object Segmentation Through Visual Attention”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3064-3074 (2019). [cited by applicant]
Wang, W., et al., “Revisiting Video Saliency: A Large-Scale Benchmark And A New Model”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4894-4903 (2018). [cited by applicant]
Wang, W., et al., “Zero-Shot Video Object Segmentation Via Attentive Graph Neural Networks”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9236-9245 (2019). [cited by applicant]
Wei, Z., et al., “Good View Hunting: Learning Photo Composition From Dense View Pairs”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5437-5446 (2018). [cited by applicant]