IP Library Granted Patent US 12,530,814
Granted Patent B2
US 12,530,814 · App. 17/797,198 · Granted Jan 20, 2026

Recurrent unit for generating or processing a sequence of images

Inventors: Pauline Luc (London, GB); Aidan Clark (London, GB); Sander Etienne Lea Dieleman (London, GB); Karen Simonyan (London, GB)
Assignee: DeepMind Technologies Limited
G06T11/00G06T3/18G06T3/4046G06T7/10G06V10/764G06V10/82G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,814
App. No.
17/797,198
Granted
Jan 20, 2026
Kind
B2
Abstract

A recurrent unit is proposed which, at each of a series of time steps receives a corresponding input vector and generates an output at the time step having at least one component for each of a two-dimensional array of pixels. The recurrent unit is configured, at each of the series of time steps except the first, to receive the output of the recurrent unit at the preceding time step, and to apply to the output of the recurrent unit at the preceding time step at least one convolution which depends on the input vector at the time step. The convolution further depends upon the output of the recurrent unit at the preceding time step. This convolution generates a warped dataset which has at least one component for each pixel of the array. The output of the recurrent unit at each time step is based on the warped dataset and the input vector.

Claims (41)

1 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement:

a recurrent unit arranged, at each of a series of time steps, to receive a corresponding input vector and to generate an output having at least one respective value for each of a two -dimensional array of pixels,

the recurrent unit being configured at each of the series of time steps except a first time step: to receive the output of the recurrent unit at a preceding time step,

to apply to the output of the recurrent unit at the preceding time step at least one convolution dependent on both the input vector at the time step and the output of the recurrent unit at the preceding time step, to generate a warped dataset which has at least one component for each pixel of the array, wherein the at least one component of the warped dataset for each pixel of the array is generated by convolving the output of the recurrent unit at the preceding time step with a respective kernel that is determined based on an output of a neural network that receives the input vector and the output of the recurrent unit at the preceding time step, and

to generate the output at the time step based on the warped dataset and the input vector.

2 . The system according to claim 1 wherein the recurrent unit neural network is configured to generate the respective kernel for each pixel of the array using the input vector and the output of the recurrent unit at the preceding time step.

3 . The system according to claim 1 wherein the recurrent unit is configured to:

generate the at least one component of the warped dataset as a weighted sum of convolutions of the corresponding component of the output of the recurrent unit at the preceding time step with a respective plurality of kernels which are generated by the neural network that receives the input vector and the output of the recurrent unit at the preceding time step, the weights of the weighted sum being different for different said pixels of the array.

4 . The system according to claim 1 wherein the recurrent unit which is configured to generate the output at each time step as a sum of (i) a component-wise product of the warped dataset with a fusion vector, and (ii) a component-wise product of a vector varying inversely with the fusion vector and a refined vector generated by a rectified linear unit of the recurrent unit.

5 . The system according to claim 4 wherein the recurrent unit is configured to generate each element of the fusion vector by applying a function to: a respective component of a component-wise product of a first weight vector with a concatenation of the output of a network at the preceding time step and the input vector plus a respective first offset value.

6 . The system according to claim 4 wherein the recurrent unit is configured to generate each element of the fusion vector by applying a function to: a respective component of a component-wise product of a first weight vector with a concatenation of the warped dataset and the input vector plus a respective first offset value.

7 . The system according to claim 4 in which the rectified linear unit is configured to generate each element of the refined vector by applying a rectified linear function to: a respective component of a component-wise product of a second weight vector with a concatenation of the output of a network at the preceding time step and the input vector plus a respective second offset value.

8 . The system according to claim 4 in which the rectified linear unit is configured to generate each element of the refined vector by applying a rectified linear function to: respective components of a component-wise product of a second weight vector with a concatenation of the output of the warped dataset and the input vector plus a respective second offset value.

9 . The system according to claim 1 , wherein the instructions further cause the one or more computers to implement a generator network to generate a sequence of images representing a temporal progression and composed of values for each of a two-dimensional array of pixels, the generator network being configured to generate each of the sequence of images based on the respective output of the recurrent unit in a respective one of the time steps.

10 . The system according to claim 9 , wherein the instructions, when executed by the one or more computers, further cause the one or more computers to perform operations for jointly training the generator network and a discriminator network, the discriminator network being for distinguishing between sequences of images generated by the generator network and sequences of images which are not generated by the generator network, the operations comprising:

receiving one or more first sequences of images representing a temporal progression; and repeatedly performing iteration steps of:

generating, by the generator network, one or more second sequences of images;

generating, by the discriminator network, at least one discriminator score for one or more of the first sequences of images and for each of the second sequence of images; and

varying weights of at least one of the discriminator network and the generator network based on the at least one discriminator score.

11 . The system according to claim 10 , in which the discriminator network comprises a spatio-temporal discriminator network for discriminating based on temporal features and a spatial discriminator network for discriminating based on spatial features, the spatio-temporal discriminator network and the spatial discriminator network each comprising a multi-layer network of neurons in which each layer performs a function defined by corresponding weights;

said generation of the at least one discriminator score comprising:

(i) forming, from an input sequence, a first set of one or more images having a lower temporal resolution than the input sequence, and inputting the first set into the spatial discriminator network to determine, based on the spatial features of each image in the first set, a first discriminator score representing a probability that the input sequence has been generated by the generator network; and

(ii) forming, from the input sequence, a second set of images having a lower spatial resolution than the input sequence, and inputting the second set into the spatio-temporal discriminator network to determine, based on the temporal features of the images in the second set, a second discriminator score representing a probability that the input sequence has been generated by the generator network; and

said varying the weights of at least one of the discriminator network and the generator network comprising updating the weights based on the first discriminator score and the second discriminator score.

12 . The system according to claim 1 , wherein the instructions further cause the one or more computers to implement a segmentation network to identify within a sequence of images a portion of each image having one or more characteristics, the recurrent unit being arranged at each of series of time steps to receive an input vector comprising a corresponding one of the sequence of images, the segmentation network being configured to generate in each time step data from the output of the recurrent unit in the corresponding time step, wherein the data from the output of the recurrent unit identifies a portion of the corresponding image.

13 . The system according to claim 1 , wherein the instructions further cause the one or more computers to implement a classification network to generate data which classifies a sequence of images as being in one or more of a set of classes, the recurrent unit being arranged at each of series of time steps to receive an input vector comprising a corresponding one of the sequence of images, the classification network being configured to generate from the outputs of the recurrent unit at each of the respective series of time steps data identifying one or more of the classes.

14 . The system according to claim 1 , wherein the instructions further cause the one or more computers to implement an adaptive system for increasing the spatial and/or temporal resolution of a sequence of images, the adaptive system comprising the recurrent unit the recurrent unit being arranged at each of a series of time steps to receive an input vector comprising a one of a first sequence of images, the adaptive system being configured to generate a sequence of images having higher spatial and/or temporal resolution than images of the first sequence of images.

15 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to implement:

a recurrent unit arranged, at each of a series of time steps, to receive a corresponding input vector and to generate an output having at least one respective value for each of a two -dimensional array of pixels,

the recurrent unit being configured at each of the series of time steps except a first time step: to receive the output of the recurrent unit at a preceding time step,

to apply to the output of the recurrent unit at the preceding time step at least one convolution dependent on both the input vector at the time step and the output of the recurrent unit at the preceding time step, to generate a warped dataset which has at least one component for each pixel of the array, wherein the at least one component of the warped dataset for each pixel of the array is generated by convolving the output of the recurrent unit at the preceding time step with a respective kernel that is determined based on an output of a neural network that receives the input vector and the output of the recurrent unit at the preceding time step, and

to generate the output at the time step based on the warped dataset and the input vector.

16 . One or more non-transitory computer-readable storage media according to claim 15 wherein the neural network is configured to generate the respective kernel for each pixel of the array using the input vector and the output of the recurrent unit at the preceding time step.

17 . One or more non-transitory computer-readable storage media according to claim 15 wherein the recurrent unit is configured to:

generate the at least one component of the warped dataset as a weighted sum of convolutions of the corresponding component of the output of the recurrent unit at the preceding time step with a respective plurality of kernels which are generated by the neural network that receives the input vector and the output of the recurrent unit at the preceding time step, the weights of the weighted sum being different for different said pixels of the array.

18 . One or more non-transitory computer-readable storage media according to claim 15 wherein the recurrent unit is configured to generate the output at each time step as a sum of (i) a component-wise product of the warped dataset with a fusion vector, and (ii) a component-wise product of a vector varying inversely with the fusion vector and a refined vector generated by a rectified linear unit of the recurrent unit.

19 . A computer-implemented method, comprising:

receiving, by a recurrent unit, at each of a series of time steps, a corresponding input vector to generate an output having at least one respective value for each of a two-dimensional array of pixels, wherein, at each of the series of time steps except a first time step, the recurrent unit performs operations comprising:

receiving the output of the recurrent unit at a preceding time step;

applying, to the output of the recurrent unit at the preceding time step, at least one convolution dependent on both the input vector at the time step and the output of the recurrent unit at the preceding time step, to generate a warped dataset which has at least one component for each pixel of the array, wherein the at least one component of the warped dataset for each pixel of the array is generated by convolving the output of the recurrent unit at the preceding time step with a respective kernel that is determined based on an output of a neural network that receives the input vector and the output of the recurrent unit at the preceding time step, and

generating the output at the time step based on the warped dataset and the input vector.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2022
From: LUC, PAULINE; CLARK, AIDAN; DIELEMAN, SANDER ETIENNE LEA; SIMONYAN, KAREN
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 060823/0010 →
Continuity (2)
Provisional Application 62971639 · Feb 7, 2020
Related Publication 20230053618A1 · Feb 23, 2023
References Cited (110)
US 10176424B2 · Danihelka · 2019 [cited by examiner]
US 20170255832A1 · Jones · 2017 [cited by examiner]
US 20180288431A1 · Liu · 2018 [cited by examiner]
US 20190035113A1 · Salvi · 2019 [cited by examiner]
US 20190258938A1 · Mnih et al. · 2019 [cited by applicant]
US 20190325306A1 · Zhu · 2019 [cited by examiner]
US 20200051206A1 · Munkberg · 2020 [cited by examiner]
US 20200134804A1 · Song · 2020 [cited by examiner]
US 20210073997A1 · Vora · 2021 [cited by examiner]
AU 2018100318A4 · 2018 [cited by applicant]
CN 108460342A · 2018 [cited by applicant]
CN 109891434A · 2019 [cited by applicant]
CN 110390381A · 2019 [cited by applicant]
KR 20190125029A · 2019 [cited by applicant]
WO WO2019100065 · 2019 [cited by applicant]
Im et al., “Generating Images with Recurrent Adversarial Networks,” CoRR, submitted on Dec. 13, 2016, 20 pages (Year: 2016). [cited by examiner]
Abdolmaleki et al., “A Distributional View on Multi-Objective Policy Optimization”, CORR, Submitted on May 15, 2020, 11 pages. [cited by applicant]
Afchar et al., “Mesonet: a Compact Facial Video Forgery Detection Network,” CoRR, Submitted on Sep. 4, 2018, accepted to Workshop on Information Forensics and Security (WIFS), 2018, arXiv:1809.00888, 7 pages. [cited by applicant]
Amersfoort et al., “Transformation-Based Models of Video Sequences,” CoRR, Submitted on Jan. 29, 2017, arXiv:1701.08435, 11 pages. [cited by applicant]
Babaeizadeh et al., “Stochastic Variational Video Prediction”, CoRR, Submitted on Mar. 6, 2018, arXiv: 1710.11252v2, 15 pages. [cited by applicant]
Balaji et al., “TFGAN: Improving Conditioning for Text-to-Video Synthesis”, International Conference on Learning Representations, 2018, 18 pages. [cited by applicant]
Ballas et al., “Delving Deeper Into Convolutional Networks for Learning Video Representations”, Submitted on Nov. 19, 2015, arXiv:1511.06432, 11 pages. [cited by applicant]
Bi'nkowski et al., “High Fidelity Speech Synthesis with Adversarial Networks”, CoRR, Submitted on Sep. 25, 2019, arXiv:1909.11646, 15 pages. [cited by applicant]
Brock et al., “Large Scale GAN Training for High Fidelity Natural Image Synthesis,” CoRR, Submitted on Sep. 28, 2018, arXiv:1809.11096, 29 pages. [cited by applicant]
Burda et al., “Exploration by Random Network Distillation,” CoRR, Submitted on Oct. 30, 2018, arXiv:1810.12894, 17 pages. [cited by applicant]
Carreira et al., “A Short Note About Kinetics-600,” CoRR, Submitted on Aug. 3, 2018, arXiv:1808.01340, 6 pages. [cited by applicant]
Carreira et al., “A Short Note on the Kinetics-700 Human Action Dataset,” CoRR, Submitted on Jul. 15, 2019, arXiv:1907.06987, 6 pages. [cited by applicant]
Carreira et al., “Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset”, International Conference on Computer Vision, 2017, 10 pages. [cited by applicant]
Cho et al., “Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation,” CoRR, Submitted on Jun. 3, 2014, arXiv:1406.1078, 14 pages. [cited by applicant]
Clark et al, “Adversarial Video Generation on Complex Datasets”, ICLR 2020 Conference, Submitted on Sep. 25, 2019, 20 pages. [cited by applicant]
De Vries et al., “Modulating Early Visual Processing by Language,” Advances in Neural Information Processing Systems 30 (NIPS 2017), 2017, 11 pages. [cited by applicant]
Denton et al., “Deep Generative Image Models Using a Laplacian Pyramid of Adversarial Networks,” Advances in Neural Information Processing Systems 28 (NIPS 2015), 2015, 9 pages. [cited by applicant]
Denton et al., “Stochastic Video Generation with a Learned Prior,” International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Denton et al., “Unsupervised Learning of Disentangled Representations from Video,” Advances in Neural Information Processing Systems 30 (NIPS 2017), 2017, 10 pages. [cited by applicant]
Dumoulin et al., “A Learned Representation for Artistic Style,” International Conference on Learning Representations, 2017, 26 pages. [cited by applicant]
Ebert et al., “Self-Supervised Visual Planning with Temporal Skip Connections,” CoRR, Submitted on Oct. 15, 2017, arXiv:1710.05268, 13 pages. [cited by applicant]
Finn et al., “Deep Visual Foresight for Planning Robot Motion,” CoRR, Submitted on Oct. 3, 2016, arXiv:1610.00696, 8 pages. [cited by applicant]
Finn et al., “Unsupervised Learning for Physical Interaction Through Video Prediction,” Advances in Neural Information Processing Systems 29 (NIPS 2016), 2016, 9 pages. [cited by applicant]
Gao et al., “Disentangling Propagation and Generation for Video Prediction,” International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Gemp et al., D3C: Reducing the Price of Anarchy in Multi-Agent Learning, Submitted on Oct. 1, 2020, arXiv:2010.00575, 30 pages. [cited by applicant]
Gentine et al., “Could Machine Learning Break the Convection Parameterization Deadlock?,” Geophysical Research Letters, 2018, 45(11):5742-5751. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” Advances in Neural Information Processing Systems 27 (NIPS 2014), 2014, 9 pages. [cited by applicant]
Hao et al., “Controllable Video Generation with Sparse Trajectories,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7854-7863, 10 pages. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium,” Advances in Neural Information Processing Systems 30 (NIPS 2017), 2017, 12 pages. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, 1997, 9(8):1735-1780. [cited by applicant]
Huang et al., “Multimodal Unsupervised Image-to-Image Translation,” CoRR, Submitetd on Apr. 12, 2018, arXiv:1804.04732, 22 pages. [cited by applicant]
Im et al., “Generating Images with Recurrent Adversarial Networks,” CoRR, submitted on Dec. 13, 2016, 20 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/EP2021/052980, mailed on Aug. 18, 2022, 18 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/EP2021/052980, mailed on May 28, 2021, 21 pages. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,”, Proceedings of the 32nd International Conference on Machine Learning, 2015, 1-9. [cited by applicant]
Jaderberg et al., “Spatial Transformer Networks,” Advances in Neural Information Processing Systems 28 (NIPS 2015), 2015, 9 pages. [cited by applicant]
Jang et al., “Video Prediction with Appearance and Motion Conditions,” Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Jayaraman et al., “Time-Agnostic Prediction: Predicting Predictable Video Frames,” International Conference on Learning Presentations, 2019, 1-20. [cited by applicant]
Kalchbrenner et al., “Video Pixel Networks,” Proceedings of the 34th International Conference on Machine Learning, 2017, 9 pages. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 4401-4410. [cited by applicant]
Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation,” International Conference on Learning Representations, 2018, 1-26. [cited by applicant]
Kay et al., “The Kinetics Human Action Video Dataset,” CoRR, submitted on May 19, 2017, arXiv:1705.06950, 1-22. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization,” CoRR, Submitted on Jul. 20, 2015, arXiv:1412.6980, 13 pages. [cited by applicant]
Kosiorek et al., “Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects,” Advances in Neural Information Processing Systems 31 (NeurIPS 2018), 2018, 11 pages. [cited by applicant]
Kumar et al., “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,” Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 12 pages. [cited by applicant]
Lee et al., “Stochastic Adversarial Video Prediction,” CoRR, submitted on Apr. 4, 2018, arXiv:1804.01523, 26 pages. [cited by applicant]
Li et al., “Flow-Grounded Spatial-Temporal Video Prediction from Still Images,” CoRR, Submitted on Jul. 25, 2018, arXiv:1807.09755, 18 pages. [cited by applicant]
Li et al., “Video Generation from Text,” Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI-18), 7065-7072. [cited by applicant]
Lim et al., “Geometric GAN,” CoRR, Submitted on May 8, 2017, arXiv:1705,02894, 17 pages. [cited by applicant]
Liu et al., “Future Frame Prediction for Anomaly Detection—a New Baseline,” CORR, Submitted on Dec. 28, 2017, arXiv:1712.09867, 10 pages. [cited by applicant]
Liu et al., “Video Frame Synthesis Using Deep Voxel Fflow,” CoRR, Submitted on Feb. 8, 2017, arXiv:1702.02463, 10 pages. [cited by applicant]
Luc et al., “DVD GAN Resubmission,” Presentation Material, Dec. 6, 2019, 21 pages. [cited by applicant]
Luc et al., “Predicting Deeper Into the Future of Semantic Segmentation,” Proceedings—2017 IEEE International Conference on Computer Vision, ICVV 2017, 648-657. [cited by applicant]
Luc et al., “Predicting Future Instance Segmentation by Forecasting Convolutional Features,” European Conference on Computer Vision, 2018, 16 pages. [cited by applicant]
Luc et al., Transformation-based Adversarial Video Prediction on Large-Scale Data, CoRR, Submitted on Mar. 9, 2020 and revised on Nov. 17, 2021, arXiv:2003.04035, 22 pages. [cited by applicant]
Mathieu et al., “Deep Multi-Scale Video Prediction Beyond Mean Square Error,” CoRR, Submitted on Feb. 26, 2016, arXiv:1511.05440, 14 pages. [cited by applicant]
Miyato et al., “cGANs with Projection Discriminator,” CoRR, Submitted on Feb. 15, 2018, arXiv:1802.05637, 21 pages. [cited by applicant]
Miyato et al., “Spectral Normalization for Generative Adversarial Networks,” CoRR, Submitted on Feb. 16, 2018, arXiv:1802.05957, 26 pages. [cited by applicant]
Nguyen et al., “Deep Learning for Deepfakes Creation and Detection,” CoRR, Submitted on Sep. 25, 2019, arXiv:1909.11573, 16 pages. [cited by applicant]
Office Action in European Appln. No. 21704256.3, dated Feb. 7, 2024, 12 pages. [cited by applicant]
Oliu et al., “Folded Recurrent Neural Networks for Future Video Prediction,” European Conference on Computer Vision, 2018, 16 pages. [cited by applicant]
Patraucean et al., “Spatio-Temporal Video Autoencoder with Differentiable Memory,” CoRR, Submitted on Sep. 1, 2016, arXiv:1511.06309, 13 pages. [cited by applicant]
Perez-Pellitero et al., “Perceptual Video Super Resolution with Enhanced Temporal Consistency,” CoRR, Submited on May 2, 2019, 1-10. [cited by applicant]
Ranzato at al., “Video (Language) Modeling: A Baseline for Generative Models of Natural Videos,” CoRR, Submitted on Dec. 20, 2014, arXiv:1412.6604, 1-15. [cited by applicant]
Reda et al., “SDC-NET: Video Prediction Using Spatially-Displaced Convolution,” European Conference on Computer Vision, 2018, 16 pages. [cited by applicant]
Rolnick et al., “Trackling Climate Change with Machine Learning,” CoRR, Subumitted on Jun. 10, 2019, arXiv:1906.05433, 97 pages. [cited by applicant]
Sabir et al., “Recurrent Convolutional Strategies for Face Manipulation Detection in Videos,” CVPR Workshops, 2019, 80-87. [cited by applicant]
Saito et al., “TGANv2: Efficient Training of Large Models for Video Generation with Multiple Subsampling Layers,” CoRR, submitted on Nov. 22, 2018, arXiv:1811.09245, 1-12. [cited by applicant]
Saxe et al., “Exact Solutions to the Nonlinear Dynamics of Learning in Deep Linear Neural Networks,” CoRR, Submitted on Jan. 24, 2014, arXiv:1312.6120, 21 pages. [cited by applicant]
Shi et al., “Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting,” Advances in Neural Information Processing Systems 28 (NIPS 2015), 1-9. [cited by applicant]
Shi et al., “Deep Learning for Precipitation Nowcasting: A Benchmark and a New Model,” Advances in Neural Information Processing Systems 30 (NIPS 2017), 1-11. [cited by applicant]
Soomro et al., “UCF101: A Dataset of 101 Human Actions Classes From Videos in the Wild,” CoRR, Submitted on Dec. 3, 2012, arXiv:1212.0402, 7 pages. [cited by applicant]
Srivastava et al., “Unsupervised Learning of Video Representations Using LSTMs,” International Conference on Machine Learning, 2015, 10 pages. [cited by applicant]
Tulyakov et al., “MoCoGAN: Decomposing Motion and Content for Video Generation,” Conference on Computer Vision and Pattern Recognition, 2018, 1526-1535. [cited by applicant]
Unterthiner et al., “Towards Accurate Generative Models of Video: A New Metric & Challenges,” CoRR, Submitted on Dec. 3, 2018, arXiv:1812.01717, 16 pages. [cited by applicant]
Villegas et al, “Decomposing Motion and Content for Natural Video Sequence Prediction,” CoRR, submitted on Jun. 25, 2017, arXiv:1706.08033, 22 pages. [cited by applicant]
Villegas et al., “Learning to Generate Long-Term Future Via Hierarchical Prediction,” CoRR, Submitted on Apr. 19, 2017, arXiv:1704.05831, 20 pages. [cited by applicant]
Vondrick et al,. “Anticipating the Future by Watching Unlabeled Video,” CoRR, Submitted on Apr. 29, 2015, arXiv:1504.08023, 10 pages. [cited by applicant]
Vondrick et al., “Generating the Future with Adversarial Transformers,” Conference on Computer Vision and Pattern Recognition, 2017, 1021-1028. [cited by applicant]
Vondrick et al., “Generating Videos with Scene Dynamics,” Advances in Neural Information Processing Systems 29 (NIPS 2016), 2016, 1-9. [cited by applicant]
Wahlstroem et al., “From Pixels to Torques: Policy Learning with Deep Dynamical Models,” CoRR, Submitted on Feb. 8, 2015, arXiv:1502.02251, 10 pages. [cited by applicant]
Walker et al., “An Uncertain Future: Forecasting from Static Images Using Variational Autoencoders,” CoRR, Submitted on Jun. 25, 2016, arXiv:1606.07873, 17 pages. [cited by applicant]
Walker et al., “Dense Optical Flow Prediction from a Static Image,” International Conference on Computer Vision, 2015, 2443-2451. [cited by applicant]
Wang et al., “Image Quality Assessment: from Error Visibility to Structural Similarity,” IEEE Transactions on Image Processing, Apr. 2004, 13(4):600-612. [cited by applicant]
Watter et al., “Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images,” Advances in Neural Information Processing Systems 28 (NIPS 2015), 2015, 1-9. [cited by applicant]
Weissenborn et al., “Scaling Autoregressive Video Models,” International Conference on Learning Representations, 2020, 1-24. [cited by applicant]
Xiong et al., “Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 2364-2373. [cited by applicant]
Xu et al., “Structure Preserving Video Prediction,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 1460-1469. [cited by applicant]
Xuan et al., “On the Generalization of GAN Image Forensics,” CoRR, submitted on Feb. 27, 2019, arXiv:1902.11153, 5 pages. [cited by applicant]
Xue et al., “Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks,” Advances in Neural Information Processing Systems 29 (NIPS 2016), 2016, 1-9. [cited by applicant]
Zhang et al., “Self-Attention Generative Adversarial Networks,” Proceedings of the 36th International Conference on Machine Learning, 2019, 10 pages. [cited by applicant]
Zhang et al., “StackGAN: text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks,” Proceedings—2017 IEEE International Conference on Computer Vision, 2017, 5907-5915. [cited by applicant]
Zhu et al., “Object-Oriented Dynamics Predictor,” Advances in Neural Information Processing Systems 31 (NIPS 2018), 2018, 12 pages. [cited by applicant]
Clark et al, “Adversarial Video Generation on Complex Datasets,” ICLR 2020 Conference, Submitted on Sep. 25, 2019, 21 pages. [cited by applicant]
Office Action in Chinese Appln. No. 202180013476.7, mailed on Jun. 6, 2025, 41 pages (with English translation). [cited by applicant]