IP Library › Granted Patent US 12,282,992
Granted Patent B2
US 12,282,992 · App. 17/856,362 · Granted Apr 22, 2025

Machine learning based controllable animation of still images

Inventors: Kuldeep Kulkarni (Bengaluru, IN); Aniruddha Mahapatra (Kolkata, IN)
Assignee: Adobe Inc.
G06T13/80G06T3/18G06T7/215G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,992
App. No.
17/856,362
Granted
Apr 22, 2025
Kind
B2
Abstract

Systems and methods for machine learning based controllable animation of still images is provided. In one embodiment, a still image including a fluid element is obtained. Using a flow refinement machine learning model, a refined dense optical flow is generated for the still image based on a selection mask that includes the fluid element and a dense optical flow generated from a motion hint that indicates a direction of animation. The refined dense optical flow indicates a pattern of apparent motion for the at least one fluid element. Thereafter, a plurality of video frames is generated by projecting a plurality of pixels of the still image using the refined dense optical flow.

Claims (57)

1. A system comprising:

a memory component; and

one or more processing devices coupled to the memory component, the one or more processing devices to perform operations comprising:

obtaining a still image, wherein the still image includes at least one fluid element;

obtaining a selection mask that includes the at least one fluid element;

obtaining one or more motion hints indicating motion directions of the at least one fluid element of the still image;

generating, using an optical flow prediction pipeline and the motion directions of the one or more motion hints, a sparse optical flow map;

generating, using the optical flow prediction pipeline and the sparse optical flow map, an intermediary dense optical flow for the still image;

generating, using a flow refinement machine learning model, a refined dense optical flow from the intermediary dense optical flow, the still image, and the selection mask, wherein the refined dense optical flow is generated by refining the intermediary dense optical flow by the flow refinement machine learning model using the selection mask, the still image, and the one or more motion hints, and wherein the refined dense optical flow indicates a pattern of apparent motion for the at least one fluid element; and

generating a plurality of video frames by projecting a plurality of pixels of the still image using the refined dense optical flow.

2. The system of claim 1 , the operations further comprising:

generating a plurality of flow field pairs from the refined dense optical flow, the plurality of flow field pairs each indicating a displacement of the plurality of pixels of the still image, each flow field pair comprising a forward flow field and a reverse flow field; and

generating the plurality of video frames by projecting the plurality of pixels of the still image based on the displacement indicated by the plurality of flow field pairs.

3. The system of claim 2 , wherein each flow field pair of the plurality of flow field pairs comprises either a Eularian flow field pair or a Lagrangian flow field pair.

4. The system of claim 2 , wherein for each video frame of the plurality of video frames, the operations further comprise:

projecting a displacement of pixels of the still image based on a respective flow field pair of the plurality of flow field pairs.

5. The system of claim 2 , the operations further comprising:

generating, using an encoder-decoder machine learning model, a video frame of the plurality of video frames from a selected flow field pair of the plurality of flow field pairs; and

generating an animated video based on an image frame sequence comprising the video frame generated from each of the plurality of flow field pairs.

6. The system of claim 5 , wherein the encoder-decoder machine learning model comprises a down sampling machine language encoder and an up sampling machine learning decoder.

7. The system of claim 6 , the operations further comprising:

generating, with the down sampling machine language encoder, a first plurality of down sample layers from the still image, and warping the first plurality of down sample layers based on the forward flow field of a first flow field pair to generate a plurality of forward warped down sample layers;

generating, with the down sampling machine language encoder, a second plurality of down sample layers from the still image, and warping the second plurality of down sample layers based on the reverse flow field of the first flow field pair to generate a plurality of reverse warped down sample layers;

generating, a plurality of symmetric splatting layers by applying symmetric splatting to corresponding layers of the plurality of forward warped down sample layers and the plurality of reverse warped down sample layers; and

generating the video frame, with the up sampling machine learning decoder, based on the plurality of symmetric splatting layers.

8. The system of claim 5 , wherein the encoder-decoder machine learning model is trained using a loss computed from a generated refined dense optical flow, the loss comprising at least one of a generative adversarial network (GAN) loss, a visual geometry group (VGG) loss, an L1 loss, and a discrimination feature-matching loss.

9. The system of claim 1 , wherein the one or more motion hints are based on a sparse optical flow generated from a user interaction with a display of the still image on a human machine interface.

10. The system of claim 1 , wherein the one or more motion hints further comprise an indication of a speed of animation.

11. The system of claim 1 , wherein the flow refinement machine learning model comprises a Partially ADaptivE Normalization (SPADE) neural network or a generative-adversarial network (GAN).

12. The system of claim 1 , wherein the flow refinement machine learning model is trained using a generative adversarial network (GAN) loss and a discrimination feature-matching loss computed from a generated refined dense optical flow.

13. A method comprising:

receiving a training dataset comprising a video stream, a selection mask and a dense optical flow generated from at least one motion hint, wherein the video stream includes image frames comprising one or more fluid elements and the selection mask includes the one or more fluid elements;

generating, using an optical flow prediction pipeline and motion directions of the at least one motion hint, a sparse optical flow map for the image frames;

generating, using the optical flow prediction pipeline and the sparse optical flow map, an intermediary dense optical flow for the image frames; and

training a flow refinement machine learning model, using the training dataset and the intermediary dense optical flow, to generate a refined dense optical flow for a still image, wherein the refined dense optical flow is generated by refining the intermediary dense optical flow by the flow refinement machine learning model using the selection mask, the still image, and the at least one motion hint, and wherein the refined dense optical flow indicates a pattern of apparent motion for the one or more one fluid elements.

14. The method of claim 13 , wherein the flow refinement machine learning model comprises a Partially ADaptivE Normalization (SPADE) neural network or a generative-adversarial network (GAN).

15. The method of claim 13 , the method further comprising:

training the flow refinement machine learning model using a generative adversarial network (GAN) loss and a discrimination feature-matching loss, computed from a generated refined dense optical flow.

16. The method of claim 13 , the method further comprising:

generating one or more flow field pairs from the refined dense optical flow, the one or more flow field pairs each indicating a displacement of a plurality of pixels of the still image, each flow field pair comprising a forward flow field and a reverse flow field; and

training an encoder-decoder machine learning model, using the training dataset and the one or more flow field pairs, to compute one or more video frames.

17. The method of claim 16 , wherein the encoder-decoder machine learning model is trained using a loss computed from a generated refined dense optical flow, the loss comprising at least one of a generative adversarial network (GAN) loss, a visual geometry group (VGG) loss, an L1 loss, and a discrimination feature-matching loss.

18. A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a still image;

receiving a selection mask for the still image that includes at least one fluid element of the still image and at least one motion hint that indicates a direction of animation of the at least one fluid element;

generating, using an optical flow prediction pipeline and the at least one motion hint that indicates a direction of animation of the at least one fluid element, a sparse optical flow map;

generating, using the optical flow prediction pipeline and the sparse optical flow map, an intermediary dense optical flow for the still image;

generating, using a flow refinement machine learning model, a refined dense optical flow from the intermediary dense optical flow, the still image, and the selection mask, wherein the refined dense optical flow is generated by refining the intermediary dense optical flow by the flow refinement machine learning model using the selection mask, the still image, and the at least one motion hint, and wherein the refined dense optical flow indicates a pattern of apparent motion for the at least one fluid element; and

generating, using an encoder-decoder machine learning model, a plurality of video frames by projecting a plurality of pixels of the still image based on the refined dense optical flow.

19. The non-transitory computer-readable medium of claim 18 , the operations further comprising:

generating a plurality of flow field pairs from the refined dense optical flow, the plurality of flow field pairs each indicating a displacement of the plurality of pixels of the still image, each flow field pair comprising a forward flow field and a reverse flow field; and

generating, using the encoder-decoder machine learning model, the plurality of video frames by projecting the plurality of pixels of the still image based on the displacement indicated by the plurality of flow field pairs.

20. The non-transitory computer-readable medium of claim 19 , wherein the encoder-decoder machine learning model comprises a down sampling machine language encoder and an up sampling machine learning decoder, the processing device further to perform operations comprising:

generating, with the down sampling machine language encoder, a first plurality of down sample layers from the still image, and warping the first plurality of down sample layers based on the forward flow field of a first flow field pair to generate a plurality of forward warped down sample layers;

generating, with the down sampling machine language encoder, a second plurality of down sample layers from the still image, and warping the second plurality of down sample layers based on the reverse flow field for the first flow field pair to generate a plurality of reverse warped down sample layers;

generating, a plurality of symmetric splatting layers by applying symmetric splatting to corresponding layers of the plurality of forward warped down sample layers and the plurality of reverse warped down sample layers; and

generating a video frame of the plurality of video frames, with the up sampling machine learning decoder, based on the plurality of symmetric splatting layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2022
From: KULKARNI, KULDEEP; MAHAPATRA, ANIRUDDHA
To: ADOBE INC.
Reel/Frame 061253/0066 →
Continuity (1)
Related Publication 20240005587A1 · Jan 4, 2024
References Cited (52)
US 10685472B1 · Rasheed · 2020 [cited by examiner]
US 20060244757A1 · Fang · 2006 [cited by examiner]
US 20150092856A1 · Mammou · 2015 [cited by examiner]
US 20180247418A1 · Crivelli · 2018 [cited by examiner]
US 20190164322A1 · Kong · 2019 [cited by examiner]
US 20200401835A1 · Zhao · 2020 [cited by examiner]
US 20210158539A1 · Zhang · 2021 [cited by examiner]
US 20210279840A1 · Chi · 2021 [cited by examiner]
US 20210390677A1 · Do · 2021 [cited by examiner]
US 20220262002A1 · Wang · 2022 [cited by examiner]
US 20220301184A1 · Choe · 2022 [cited by examiner]
US 20220303495A1 · Choe · 2022 [cited by examiner]
US 20220398751A1 · Parashar · 2022 [cited by examiner]
US 20230187052A1 · Chen · 2023 [cited by examiner]
US 20230281830A1 · Zhang · 2023 [cited by examiner]
US 20240249523A1 · Cole · 2024 [cited by examiner]
Blattmann, A., et al., “iPOKE: Poking a Still Image for Controlled Stochastic Video Synthesis”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14707-14717 (2021). [cited by applicant]
Blattmann, A., et al., “Understanding Object Dynamics for Interactive Image-to-Video Synthesis”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5171-5181 (2021). [cited by applicant]
Caballero, J., et al., “Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation”, In IEEE Conference on Computer Vision and Pattern Recognition, pp. 4778-4787 (2017). [cited by applicant]
Castrejon, L., et al., “Improved Conditional VRNNs for Video Prediction”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7608-7617 (2019). [cited by applicant]
Chuang, Y., Y., et al., “Animating Pictures with Stochastic Motion Textures”, In ACM SIGGRAPH 2005 Papers, pp. 853-860 (2005). [cited by applicant]
Dorkenwald, M., et al., “Stochastic Image-to-Video Synthesis using cINNs”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3742-3753 (2021). [cited by applicant]
Endo, Y., et al., “Animating Landscape: Self-Supervised Learning of Decoupled Motion and Appearance for Single-Image Video Synthesis”, ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH Asia), vol. 38, No. 6, Art… [cited by applicant]
Franceschi, J., Y., et al., “Stochastic Latent Residual Video Prediction”, In Proceedings of the 37th International Conference on Machine Learning, PMLR, pp. 3233-3246 (2020). [cited by applicant]
Goodfellow, I., J., et al., “Generative Adversarial Nets”, Advances in neural information processing systems, pp. 1-9 (Jun. 10, 2014). [cited by applicant]
Hadji, I., and Wildes, R., P., “A New Large Scale Dynamic Texture Dataset with Application to ConvNet Understanding”, In Proceedings of the European Conference on Computer Vision (ECCV), pp. 320-335 (2018). [cited by applicant]
Halperin, T., et al., “Endless Loops: Detecting and Animating Periodic Patterns in Still Images”, ACM Transactions on Graphics, vol. 40, No. 4, Article. 142, pp. 1-12 (Aug. 2021). [cited by applicant]
Hao, Z., et al., “Controllable Video Generation with Sparse Trajectories”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7854-7863 (Jun. 18-23, 2018). [cited by applicant]
Holynski, A., et al., “Animating Pictures with Eulerian Motion Fields”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2364-2373 (2021). [cited by applicant]
Holynski, A., et al., “Animating Pictures with Eulerian Motion Fields”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, pp. 5810-5819 (Jun. 2021). [cited by applicant]
Karras, T., et al., “Analyzing and Improving the Image Quality of StyleGAN”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8110-8119 (Mar. 23, 2020). [cited by applicant]
Kay, W., et al. “The Kinetics Human Action Video Dataset”, Computer Vision and Pattern Recognition, arXiv:1705.06950v1, pp. 1-22 (May 19, 2017). [cited by applicant]
Li, Y., et al., “Flow-Grounded Spatial-Temporal Video Prediction from Still Images”, In Proceedings of the European Conference on Computer Vision, ECCV, pp. 600-615 (2018). [cited by applicant]
Logacheva, E., et al., “Deeplandscape: Adversarial Modeling of Landscape Videos”, In European Conference on Computer Vision, Springer, pp. 256-272 (2020). [cited by applicant]
Mahapatra, A., and Kulkarni, K., “Controllable Animation of Fluid Elements in Still Images”, Computer Vision and Pattern Recognition, arXiv:2112.03051, pp. 1-12 (Dec. 6, 2021). [cited by applicant]
Minderer, M., et al., “Unsupervised Learning of Object Structure and Dynamics from Videos”, 33rd Conference on Neural Information Processing Systems (NeurIPS), pp. 1-20 (Nov. 12, 2019). [cited by applicant]
Niklaus, S., and Liu, F., “Softmax Splatting for Video Frame Interpolation”, In IEEE Conference on Computer Vision and Pattern Recognition, pp. 5437-5446 (2020). [cited by applicant]
Pan, J., et al., “Video Generation from Single Semantic Label Map”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3733-3742 (2019). [cited by applicant]
Park, T., et al., “Semantic Image Synthesis With Spatially-Adaptive Normalization”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2337-2346 (2019). [cited by applicant]
Reda, F., A., et al., “SDC-Net: Video prediction using spatially-displaced convolution”, In Proceedings of the European Conference on Computer Vision (ECCV), pp. 718-733 (2018). [cited by applicant]
Ronneberger, O., et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, In International Conference on Medical image computing and computer-assisted intervention, Springer, pp. 234-241 (2015). [cited by applicant]
Simonyan, K., and Zisserman, A., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, ICLR, pp. 1-13 (Dec. 13, 2014). [cited by applicant]
Szegedy, C., et al., “Rethinking the Inception Architecture for Computer Vision”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826 (2016). [cited by applicant]
Tian, Y., et al., “A Good Image Generator is What You Need for High-Resolution Video Synthesis”, In Proceedings of the International Conference on Learning Representations, ICLR, pp. 1-23 (2021). [cited by applicant]
Unterthiner, T., et al., “Towards Accurate Generative Models of Video: A New Metric & Challenges”, Computer Vision and Pattern Recognition, arXiv:1812.01717, pp. 1-17 (2018). [cited by applicant]
Villegas, R., et al., “Decomposing Motion and Content for Natural Video Sequence Prediction”, ICLR, arXiv:1706.08033, pp. 1-22 (2017). [cited by applicant]
Villegas, R., et al., “High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks”, 33rd Conference on Neural Information Processing Systems (NeurIPS), pp. 81-91 (2019). [cited by applicant]
Wang, T., C., et al., “Few-shot Video-to-Video Synthesis”, arXiv:1910.12713, pp. 1-14 (2019). [cited by applicant]
Wang, T., C., et al., “Video-to-Video Synthesis”, arXiv: 1808.06601, pp. 1-14 (2018). [cited by applicant]
Wu, Y., et al., “Future Video Synthesis with Object Motion Prediction”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5539-5548 (2020). [cited by applicant]
Xiong, W., et al., “Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2364-2373 (201… [cited by applicant]
Zhang, J., et al., “DTVNet: Dynamic Time-lapse Video Generation via Single Still Image”, In European Conference on Computer Vision, pp. 300-315 (2020). [cited by applicant]