IP Library › Granted Patent US 12,248,865
Granted Patent B2
US 12,248,865 · App. 17/170,416 · Granted Mar 11, 2025

Systems and methods for modeling continuous stochastic processes with dynamic normalizing flows

Inventors: Ruizhi Deng (Coquitlam, CA); Bo Chang (Toronto, CA); Marcus Anthony Brubaker (Toronto, CA); Gregory Peter Mori (Burnaby, CA); Andreas Steffen Michael Lehrmann (Vancouver, CA)
Assignee: ROYAL BANK OF CANADA
G06N3/047G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,865
App. No.
17/170,416
Granted
Mar 11, 2025
Kind
B2
Abstract

Systems and methods for machine learning architecture for time series data prediction. The system may include a processor and a memory storing processor-executable instructions. The processor-executable instructions, when executed, may configure the processor to: obtain time series data associated with a data query; generate a predicted value based on a sampled realization of the time series data and a continuous time generative model, the continuous time generative model trained to define an invertible mapping to maximize a log-likelihood of a set of predicted values for a time range associated with the time series data; and generate a signal providing an indication of the predicted value associated with the data query.

Claims (433)

1. A system for a machine learning architecture for time series data prediction comprising:

a processor; and

a memory coupled to the processor and storing processor-executable instructions that, when executed, configure the processor to:

obtain time series data associated with a data query;

generate a predicted value by executing a machine learning application based on a sampled realization of the time series data, the machine learning application comprising a continuous time generative model trained to define an invertible mapping to maximize a log-likelihood of a set of predicted values for a time range associated with the time series data, wherein generation of the predicted value comprises:

computing the predicted value based on a joint distribution X τ =F θ (W τ ;τ), ∀τ∈[0, T], where F θ (⋅; τ): d → d is the invertible mapping parametrized by the learnable parameters θ for every τ∈[0, T], and W τ is a d-dimensional Wiener process, such that the log-likelihood

ℒ

=

log

⁢

p

x

τ

1

,

…

,

x

τ

n

(

x

τ

1

,

…

,

x

τ

n

)

is maximized, where p X (x) represents a probability density function of x;

wherein the invertible mapping is based on solving an initial value problem defined by:

d

dt

⁢

(

h

τ

(

t

)

a

τ

(

t

)

)

=

(

f

θ

(

h

τ

(

t

)

,

a

τ

(

t

)

,

t

)

g

θ

(

a

τ

(

t

)

,

t

)

)

,

(

h

τ

(

t

0

)

a

τ

(

t

0

)

)

=

(

w

τ

τ

)

,

where h τ (t)∈ d , t∈[t 0 ,t 1 ], ƒ θ : d × ×[t 0 ,t 1 ]→ d , and g θ : ×[t o ,t 1 ]→ , and the joint distribution F θ (w τ ; τ) is defined as a solution of h τ (t); and

generate a signal providing an indication of the predicted value associated with the data query for performing a downstream task.

2. The system of claim 1 , wherein the invertible mapping is parameterized by a latent variable having an isotropic Gaussian prior distribution.

3. The system of claim 1 , wherein the predicted value represents one or more observed process data points based on a sampled realization of a Weiner process and the invertible mapping to provide a time continuous observed realization of the Weiner process.

4. The system of claim 3 , wherein the obtained time series data includes an incomplete realization of a continuous stochastic process.

5. The system of claim 4 , wherein: the joint distribution is defined as the solution of h τ (t) at

F

θ

(

w

τ

;

τ

)

:

=

h

τ

(

t

1

)

=

h

τ

(

t

0

)

+

∫

t

0

t

1

f

θ

(

h

τ

(

t

)

,

a

τ

(

t

)

,

t

)

⁢

d

⁢

t

.

6. The system of claim 1 , wherein the processor-executable instructions, when executed, configure the processor to:

obtain a training dataset including irregular time series data over time associated with an incomplete realization of a continuous stochastic process; and

generate the invertible mapping associated with the continuous time generative model based on maximizing the log-likelihood of the set of predicted values.

7. The system of claim 1 , wherein the obtained time series data includes observed process data points, wherein the predicted value represents a likelihood determination of stochastic process data points based on the observed process data points and an inverse of the invertible mapping of the continuous time generative model.

8. The system of claim 1 , wherein the obtained time series data includes unobserved process data points, and wherein the predicted value represents a conditional probability density of stochastic data points based on the unobserved process data points and an inverse of the invertible mapping of the continuous time generative model.

9. The system of claim 8 , wherein the conditional probability density provides for data point interpolation associated with a stochastic process based on a Brownian bridge.

10. The system of claim 8 , wherein the conditional probability density provides for data point extrapolation of data points associated with a stochastic process based on a multivariate Gaussian conditional probability distribution.

11. The system of claim 1 , wherein the continuous time generative model is based on an augmented neural ordinary differential equation (ANODE) including a multi-layer perceptron (MLP model) having at least 4 hidden layers.

12. A method for a machine learning architecture for time series data prediction comprising:

obtaining time series data associated with a data query;

generating a predicted value by executing a machine learning application based on a sampled realization of the time series data, the machine learning application comprising a continuous time generative model trained to define an invertible mapping to maximize a log-likelihood of a set of predicted values for a time range associated with the time series data, wherein generation of the predicted value comprises:

computing the predicted value based on a joint distribution X τ =F θ (W τ ; τ), ∀τ∈[0, T], where F θ (⋅;τ): d → d is the invertible mapping parametrized by the learnable parameters θ for every τ∈[0, T], and W τ is a d-dimensional Wiener process, such that the log-likelihood

ℒ

=

log

⁢

p

x

τ

1

,

…

,

x

τ

n

(

x

τ

1

,

…

,

x

τ

n

)

is maximized, where p X (x) represents a probability density function of x;

wherein the invertible mapping is based on solving an initial value problem defined by:

d

dt

⁢

(

h

τ

(

t

)

a

τ

(

t

)

)

=

(

f

θ

(

h

τ

(

t

)

,

a

τ

(

t

)

,

t

)

g

θ

(

a

τ

(

t

)

,

t

)

)

,

(

h

τ

(

t

0

)

a

τ

(

t

0

)

)

=

(

w

τ

τ

)

,

where h τ (t)∈ d , t∈[t 0 , t 1 ], ƒ θ : d × ×[t 0 , t 1 ]→ d , and g θ : ×[t 0 , t 1 ]→ , and the joint distribution F θ (w τ ; τ) is defined as a solution of h τ (t), and

generating a signal providing an indication of the predicted value associated with the data query for performing a downstream task.

13. The method of claim 12 , wherein the invertible mapping is parameterized by a latent variable having an isotropic Gaussian prior distribution.

14. The method of claim 12 , wherein the predicted value represents one or more observed process data points based on a sampled realization of a Weiner process and the invertible mapping to provide a time continuous observed realization of the Weiner process.

15. The method of claim 14 , wherein the obtained time series data includes an incomplete realization of a continuous stochastic process.

16. The method of claim 15 , wherein: the joint distribution is defined as the solution of h τ (t) at

t

=

t

1

:

F

θ

(

w

τ

;

τ

)

:

=

h

τ

(

t

1

)

=

h

τ

(

t

0

)

+

∫

t

0

t

1

f

θ

(

h

τ

(

t

)

,

a

τ

(

t

)

,

t

)

⁢

d

⁢

t

.

17. The method of claim 12 , comprising:

obtaining a training dataset including irregular time series data over time associated with an incomplete realization of a continuous stochastic process; and

generating the invertible mapping associated with the continuous time generative model based on maximizing the log-likelihood of the set of predicted values.

18. The method of claim 12 , wherein the obtained time series data includes observed process data points, wherein the predicted value represents a likelihood determination of stochastic process data points based on the observed process data points and an inverse of the invertible mapping of the continuous time generative model.

19. The method of claim 12 , wherein the obtained time series data includes unobserved process data points, and wherein the predicted value represents a conditional probability density of stochastic data points based on the unobserved process data points and an inverse of the invertible mapping of the continuous time generative model.

20. A non-transitory computer-readable medium having stored thereon machine interpretable instructions which, when executed by a processor, cause the processor to perform a method for a machine learning architecture for time series data prediction, the method comprising:

obtaining time series data associated with a data query;

generating a predicted value by executing a machine learning application based on a sampled realization of the time series data, the machine learning application comprising a continuous time generative model trained to define an invertible mapping to maximize a log-likelihood of a set of predicted values for a time range associated with the time series data, wherein generation of the predicted value comprises;

computing the predicted value based on a joint distribution X τ =F θ (W τ ; τ), ∀τ∈[0, T], where F θ (ω;τ): d → d is the invertible mapping parametrized by the learnable parameters θ for every τ∈[0, T], and W τ is a d-dimensional Wiener process, such that the log-likelihood

ℒ

=

log

⁢

p

x

τ

1

,

…

,

x

τ

n

(

x

τ

1

,

…

,

x

τ

n

)

is maximized, where p X (x) represents a probability density function of x;

wherein the invertible mapping is based on solving an initial value problem defined by:

d

dt

⁢

(

h

τ

(

t

)

a

τ

(

t

)

)

=

(

f

θ

(

h

τ

(

t

)

,

a

τ

(

t

)

,

t

)

g

θ

(

a

τ

(

t

)

,

t

)

)

,

(

h

τ

(

t

0

)

a

τ

(

t

0

)

)

=

(

w

τ

τ

)

,

where h τ (t)∈ d , t∈[t 0 , t 1 ], ƒ θ : d ×R×[t 0 , t 1 ]→ d , and g θ : ×[t 0 , t 1 ]→ , and the joint distribution F θ (w τ ; τ) is defined as a solution of h τ (t); and

generating a signal providing an indication of the predicted value associated with the data query for performing a downstream task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2025
From: DENG, RUIZHI; CHANG, BO; BRUBAKER, MARCUS ANTHONY; MORI, GREGORY PETER; LEHRMANN, ANDREAS STEFFEN MICHAEL
To: ROYAL BANK OF CANADA
Reel/Frame 069909/0751 →
Continuity (2)
Provisional Application 62971143 · Feb 6, 2020
Related Publication 20210256358A1 · Aug 19, 2021
References Cited (53)
US 11100586B1 · Wang · 2021 [cited by examiner]
US 20180129968A1 · Osogami · 2018 [cited by examiner]
US 20190018933A1 · Oono · 2019 [cited by examiner]
Both et al., “Temporal Normalizing Flows” (Year: 2019). [cited by examiner]
De Brouwer et al., “GRU-ODE-Bayes: Continuous modeling of sporadically-observed time series” (Year: 2019). [cited by examiner]
D. Barber; Bayesian Reasoning and Machine Learning; Cambridge University Press, 2012. [cited by applicant]
L. E. Baum and T. Petrie; Statistical Inference for Probabilistic Functions of Finite State Markov Chains; The Annals of Mathematical Statistics, vol. 37, No. 6, Institute of Mathematical Statistics, 1966, pp. 1554-1563. [cited by applicant]
J. Behrmann et al.; Invertible Residual Networks; Proceedings of the 36th International Conference on Machine Learning, PMLR 97, 2019. [cited by applicant]
S.R. Bowman et al.; Generating Sentences from a Continuous Space; Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning (CONLL), 2016. [cited by applicant]
Y. Burda et al.; Importance Weighted Autoencoders; Under review as a conference paper at ICLR 2016. [cited by applicant]
O. Cappé et al.; An Overview of Existing Methods and Recent Advances in Sequential Monte Carlo; Proceedings of the IEEE, vol. 95, No. 5, pp. 899-924, May 2007. [cited by applicant]
R.T.Q. Chen et al.; Neural Ordinary Differential Equations; 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montreal, Canada. [cited by applicant]
R.T.Q. Chen et al.; Residual Flows for Invertible Generative Modeling; 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada. [cited by applicant]
K. Cho et al.; Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation; Conference on Empirical Methods in Natural Language Processing (EMNLP 2014). [cited by applicant]
J. Chung et al.; A Recurrent Latent Variable Model for Sequential Data; Advances in Neural Information Processing Systems, Jan. 2015. [cited by applicant]
L. Dinh et al.; NICE: Non-Linear Independent Components Estimation; Accepted as a workshop contribution at ICLR 2015. [cited by applicant]
L. Dinh et al.; Density Estimation Using Real NVP; Published as a conference paper at ICLR 2017. [cited by applicant]
E. Dupont et al.; Augmented Neural ODEs; 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada. [cited by applicant]
J. Durbin and S.J. Koopman; Time Series Analysis by State Space Methods; Oxford University Press, 2012. [cited by applicant]
E.B. Fox et al.; Nonparametric Bayesian Learning of Switching Linear Dynamical Systems; Advances in Neural Information Processing Systems, Jan. 2008. [cited by applicant]
M. Fraccaro et al.; Sequential Neural Models with Stochastic Layers; 30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain. [cited by applicant]
M. Garnelo et al.; Conditional Neural Processes; Proceedings of the 35th International Conference on Machine earning, Stockholm, Sweden, PMLR 80, 2018. [cited by applicant]
M. Garnelo et al.; Neural Processes; Presented at the ICML 2018 workshop on Theoretical Foundations and Applications of Deep Generative Models. [cited by applicant]
W. Grathwohl et al.; Scalable Reversible Generative Models with Free-Form Continuous Dynamics; 1st Symposium on Advances in Approximate Bayesian Inference, 2018, pp. 1-14. [cited by applicant]
J. He et al.; Lagging Inference Networks and Posterior Collapse in Variational Autoencoders; Published as a conference paper at ICLR 2019. [cited by applicant]
M.F. Hutchinson; A Stochastic Estimator of the Trace of the Influence Matrix for Laplacian Smoothing Splines; Communications in Statistics—Simulation and Computation, 19(2):433-450, 1990. [cited by applicant]
K. Ito and K. Xiong; Gaussian Filters for Nonlinear Filtering Problems; IEEE Transactions on Automatic Control, vol. 45, No. 5, May 2000, pp. 910-927. [cited by applicant]
J.M. Wang et al.; Gaussian Process Dynamical Models for Human Motion; IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, No. 2, Feb. 2008, pp. 283-297. [cited by applicant]
S.J. Julier and J.K. Uhlmann; A New Extension of the Kalman Filter to Nonlinear Systems; In Aerospace/Defense Sensing, Simulation and Controls, 1997. [cited by applicant]
R.E. Kalman; A New Approach to Linear Filtering and Prediction Problems; Transactions of the ASME—Journal of Basic Engineering, 82 (Series D): 35-45, 1960. [cited by applicant]
H. Kim et al.; Attentive Neural Processes; Published as a conference paper at ICLR 2019. [cited by applicant]
D.P. Kingma and M. Welling; Auto-Encoding Variational Bayes; In International Conference on Learning Representations, 2014. [cited by applicant]
D.P. Kingma and P. Dhariwal; Glow: Generative Flow with Invertible 1x1 Convolutions; Advances in Neural Information Processing Systems 31 (NeurIPS 2018). [cited by applicant]
D.P. Kingma et al.; Improved Variational Inference with Inverse Autoregressive Flow; 30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain. [cited by applicant]
I. Kobyzev et al.; Normalizing Flows: An Introduction and Review of Current Methods; arXiv:1908.09257, 2019. [cited by applicant]
M. Kumar et al.; VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation; Published as a conference paper at ICLR 2020, 2019. [cited by applicant]
J.-F. Le Gall; Brownian Motion, Martingales, and Stochastic Calculus; Graduate Texts in Mathematics 274, Springer, 2016. [cited by applicant]
A.M. Lehrmann et al.; Efficient Nonlinear Markov Models for Human Motion; 2014 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1314-1321. [cited by applicant]
X. Li et al.; Scalable Gradients for Stochastic Differential Equations; Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS) 2020, Palermo, Italy, PMLR: vol. 108. [cited by applicant]
R. Luo et al.; A Neural Stochastic Volatility Model; Association for the Advancement of Artificial Intelligence, 2018. [cited by applicant]
N. Mehrasa et al.; Point Process Flows; arXiv:1910.08281v3 [cs.LG] Dec. 22, 2019. [cited by applicant]
G. Papamakarios et al.; Masked Autoregressive Flow for Density Estimation; arXiv:1705.07057v4 [stat.ML] Jun. 14, 2018. [cited by applicant]
G. Papamakarios et al.; Normalizing Flows for Probabilistic Modeling and Inference; arXiv:1912.02762v1 [stat.ML] Dec. 5, 2019. [cited by applicant]
S.Qin et al.; Recurrent Attentive Neural Process for Sequential Data; arXiv:1910.09323v1 [cs.LG] Oct. 17, 2019. [cited by applicant]
C.E. Rasmussen and C.K.I. Williams; Gaussian Processes for Machine Learning; MIT (Massachusetts Institute of Technology) Press, 2006. [cited by applicant]
D. Rezende and S. Mohamed; Variational Inference with Normalizing Flows; Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 2015; JMLR: W&CP vol. 37. [cited by applicant]
Y. Rubanova et al.; Latent ODEs for Irregularly-Sampled Time Series; 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada. [cited by applicant]
O. Shchur et al.; Intensity-Free Learning of Temporal Point Processes; Published as a conference paper at ICLR 2020. [cited by applicant]
G. Singh et al.; Sequential Neural Processes; 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada. [cited by applicant]
S. Särkkä; On Unscented Kalman Filtering for State Estimation of Continuous-Time Nonlinear Systems; IEEE Transactions on Automatic Control, vol. 52, No. 9, Sep. 2007, pp. 1631-1641. [cited by applicant]
Y. Tassa et al.; DeepMind Control Suite; arXiv:1801.00690v1 [cs.Al] Jan. 2, 2018. [cited by applicant]
S. Zhang et al.; Cautionary Tales on Air-Quality Improvement in Beijing; Proc. R. Soc. A 473: 20170457 http:/dx.doi.org/10.1098/rspa.2017.0457. [cited by applicant]
O. Cappé et al.; Hidden Markov Models and Dynamical System, Chapter 4, Continuous States and Observations and Kalman Filtering, Springer, 2005. [cited by applicant]