IP Library › Granted Patent US 12,602,448
Granted Patent B2
US 12,602,448 · App. 17/304,163 · Granted Apr 14, 2026

Progressive neural ordinary differential equations

Inventors: Yi Yao (Princeton, NJ); Ajay Divakaran (Monmouth Junction, NJ); Hammad A. Ayyubi (La Jolla, CA)
Assignee: SRI International
G06F17/13A61M16/0493G06N3/08G06V10/774G06F18/213
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,448
App. No.
17/304,163
Granted
Apr 14, 2026
Kind
B2
Abstract

Techniques are described for neural networks based on Progressive Neural ODEs (PODEs). In an example, a method to progressively train a neural ordinary differential equation (NODE) model comprises processing, by a machine learning system executed by a computing system, first training data, the first training data having a first complexity, to perform training of a first layer for the NODE model; and after performing the first training, processing second training data, the second training data having a second complexity that is higher than the first complexity, to perform training of a second layer for the NODE model.

Claims (50)

1 . A computing system to progressively train a neural ordinary differential equation (NODE) model, the computing system comprising:

a machine learning system;

a memory configured to store the NODE model; and

processing circuitry coupled to the memory, the processing circuitry and the memory configured to execute the machine learning system, wherein the machine learning system is configured to:

process first training data, the first training data comprising irregularly spaced time series data derived from second training data by applying a complexity-reducing transformation that reduces frequency content of the second training data, to perform training of a first layer for the NODE model, the first layer comprising a first ordinary differential equation (ODE)-based transfer function parameterized by a first set of parameters; and

after processing the first training data to perform the training of the first layer for the NODE model, process, while retaining the first set of parameters of the first layer fixed, the second training data comprising irregularly spaced time series data to perform training of a second layer for the NODE model, the second layer different from the first layer and comprising a second ODE-based transfer function parameterized by a second set of parameters.

2 . The computing system of claim 1 , wherein applying the complexity-reducing transformation comprises applying at least one of a low-pass filter or principal component analysis to the second training data.

3 . The computing system of claim 1 , wherein the machine learning system is configured to:

process input data to perform a prediction; and

output the prediction as output data.

4 . The computing system of claim 3 , wherein the input data comprises irregularly spaced time series data having at least one of a trend and a periodicity.

5 . The computing system of claim 3 ,

wherein the input data comprises irregularly spaced data, and

wherein the output data comprises irregularly spaced data.

6 . The computing system of claim 1 , wherein the irregularly spaced time series data of the first training data and the irregularly spaced time series data of the second training data each has at least one of a trend or a periodicity.

7 . The computing system of claim 1 , wherein the memory is configured to store an encoder, the encoder configured to:

map irregularly spaced time series data to fixed length time series data comprising a fixed length embedding; and

output the fixed length time series data as input data to the NODE model.

8 . The computing system of claim 7 , wherein the machine learning system is configured to:

process the first training data to perform training of a first encoder layer of the encoder; and

after performing the training of the first layer for the NODE model, add a second encoder layer to the encoder and process the second training data to perform training of the second encoder layer of the encoder.

9 . The computing system of claim 7 , wherein the memory is configured to store a decoder configured to:

map predicted values output by the NODE model to output data comprising irregularly spaced time series data; and

output the output data.

10 . The computing system of claim 7 , wherein the memory is configured to store a decoder configured to:

map predicted values output by the NODE model to output data comprising irregularly spaced time series data; and

output the output data, and

wherein the machine learning system is configured to:

process the first training data to perform training of a first encoder layer of the encoder and training of a first decoder layer of the decoder; and

after performing the training of the first layer for the NODE model, add a second encoder layer to the encoder, add a second decoder layer to the decoder, and process the second training data to perform training of the second encoder layer of the encoder and training of the second decoder layer of the decoder.

11 . The computing system of claim 1 , wherein the NODE model comprises a set of transfer functions that are parameterized by a set of parameters, wherein the set of transfer functions include the first ODE-based transfer function and the second ODE-based transfer function.

12 . The computing system of claim 1 , wherein the machine learning system is configured to:

after performing the training of the first layer for the NODE model, add the second layer to the NODE model.

13 . The computing system of claim 12 , wherein to add the second layer to the NODE model, the machine learning system is configured to:

apply alpha blending to add the second layer to the NODE model, the alpha blending controlled by a configurable parameter.

14 . A method to progressively train a neural ordinary differential equation (NODE) model, the method comprising:

processing, by a machine learning system executed by a computing system, first training data, the first training data comprising irregularly spaced time series data derived from second training data by applying a complexity-reducing transformation that reduces frequency content of the second training data, to perform training of a first layer for the NODE model, the first layer comprising a first ordinary differential equation (ODE)-based transfer function parameterized by a first set of parameters; and

after processing the first training data to perform the training of the first layer for the NODE model, processing, while retaining the first set of parameters of the first layer fixed, the second training data comprising irregularly spaced time series data to perform training of a second layer for the NODE model, the second layer different from the first layer and comprising a second ODE-based transfer function parameterized by a second set of parameters.

15 . The method of claim 14 , further comprising:

processing irregularly spaced input data to perform a prediction; and

outputting the prediction as irregularly spaced output data.

16 . The method of claim 14 , wherein the irregularly spaced time series data of the first training data and the irregularly spaced time series data of the second training data each has at least one of a trend or a periodicity.

17 . The method of claim 14 , further comprising:

mapping, by an encoder of the machine learning system, irregularly spaced time series data to fixed length time series data comprising a fixed length embedding; and

outputting the fixed length time series data as input data to the NODE model.

18 . The method of claim 14 , further comprising:

after performing the training of the first layer for the NODE model, adding the second layer to the NODE model.

19 . Non-transitory computer readable media comprising instructions for causing processing circuitry to execute a machine learning system to progressively train a neural ordinary differential equation (NODE) model, wherein the machine learning system is configured to:

process first training data, the first training data comprising irregularly spaced time series data derived from second training data by applying a complexity-reducing transformation that reduces frequency content of the second training data, to perform training of a first layer for the NODE model, the first layer comprising a first ordinary differential equation (ODE)-based transfer function parameterized by a first set of parameters; and

after processing the first training data to perform the training of the first layer for the NODE model, process, while retaining the first set of parameters of the first layer fixed, the second training data comprising irregularly spaced time series data to perform training of a second layer for the NODE model, the second layer different from the first layer and comprising a second ODE-based transfer function parameterized by a second set of parameters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2021
From: YAO, YI; DIVAKARAN, AJAY; AYYUBI, HAMMAD A
To: SRI INTERNATIONAL
Reel/Frame 056553/0450 →
Continuity (2)
Provisional Application 63039567 · Jun 16, 2020
Related Publication 20210390400A1 · Dec 16, 2021
References Cited (28)
US 20190171936A1 · Karras · 2019 [cited by examiner]
US 20200293870A1 · Isikdogan · 2020 [cited by examiner]
Chen et al., “Neural Ordinary Differential Equations” (Year: 2018). [cited by examiner]
Du et al., “Time Series Forecasting using Sequence-to-Sequence Deep Learning Framework” (Year: 2018). [cited by examiner]
Lu et al., “On Training the Recurrent Neural Network Encoder-Decoder for Large Vocabulary End-To-End Speech Recognition” (Year: 2016). [cited by examiner]
Gunther et al., “Layer-Parallel Training of Deep Residual Neural Networks” (Year: 2019). [cited by examiner]
Kang et al., “Machine Learning: Data Pre-processing” (Year: 2018). [cited by examiner]
Elman, “Learning and development in neural networks: the importance of starting small” (Year: 1993). [cited by examiner]
Meade el al., “The Numerical Solution of Linear Ordinary Differential Equations by Feedforward Neural Networks” (Year: 1994). [cited by examiner]
Rannen et al., “Encoder Based Lifelong Learning” (Year: 2017). [cited by examiner]
Jaderberg et al., “Decoupled Neural Interfaces using Synthetic Gradients,” in Int'l Conf. Machine Learning 1627-35 (2017). (Year: 2017). [cited by examiner]
Ayyubi et al., “Progressive Growing of Neural ODEs,” arXiv.org; arXiv:2003.03695v1, Mar. 8, 2020, 6 pp. [cited by applicant]
Bengio et al., “Curriculum Learning,” in Proceedings of the 26th Annual International Conference on Machine Learning(ICML 2009) Montreal, Canada, Jun. 14-18, 2009, 8 pp. [cited by applicant]
Chen et al., “Neural Ordinary Differential Equations,” 32nd Conference on Neural Information Processing Systems(NIPS 2018) Montreal, Canada, Dec. 2-8, 2018, 13 pp. [cited by applicant]
Chen et al., “Neural Ordinary Differential Equations,” arXiv:1806.07366v5, 32nd Conference on Neural Information Processing Systems(NIPS 2018), Dec. 14, 2019, 18 pp. [cited by applicant]
Elman, “Learning and development in neural networks: The importance of starting small,” Cognition 48;1, Jul. 1993, 30 pp. [cited by applicant]
Gardner, “Exponential smoothing: The state of the art,” Journal of Forecasting, vol. 4, No. 1, Jan. 1, 1985, 38 pp. [cited by applicant]
Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation,” Nov. 3, 2017, 25 pp. [cited by applicant]
Li et al., “A scalable end-to-end Guassian process adapter for irregularly sampled time series classification,” 29th Conference on Neural Information Processing Systems, Barcelona, Spain, Oct. 28, 2016, 11 pp. [cited by applicant]
Lipton et al., “Modeling Missing Data in Clinical Time Series with RNNs” Proceedings of Machine Learning for Healthcare, vol. 56, Jun. 2016, 17 pp. [cited by applicant]
Matissen et al., “Teacher-Student Curriculum Learning” arXiv.org; arXiv:1707.00183v2, Deep Reinforcement earning Symposium (NIPS 2017), Long Beach, CA, USA, Nov. 29, 2017, 15 pp. [cited by applicant]
Mei et al., “The Neural Hawkes Process: A Neural Self-Modulating Multivariate Point Process,” arXiv.org; arXiv:1612.09328v1, Dec. 29, 2016, 16 pp. [cited by applicant]
Rehfeld et al., “Comparison of correlation analysis techniques for irregularly sampled time series,” Jun. 23, 2011, pp. 389-404. [cited by applicant]
Rubanova et al., “Latent ODEs for Irregularly-Sampled Time Series,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, Dec. 8-14, 2019, 11 pp. [cited by applicant]
Scargle, “Studies in Astronomical Time Series Analysis.II. Statistical Aspects of Spectral Anaylsis of Unevenly Spaced Data,” The Astrophysical Journal, Dec. 15, 1982, pp. 835-853. [cited by applicant]
Shukla et al., “Interpolation-Prediction Networks for Irregularly Sampled Time Series,” arXiv.org; arXiv:1909.07782v1, Sep. 13, 2019, 14 pp. [cited by applicant]
Zaremba et al., “Learning to Execute,” arXiv.org; arXiv:1410.4615v2, Dec. 21, 2014, 25 pp. [cited by applicant]
Zhang et al., “Time series forecasting using a hybrid ARIMA and neural network model,” Neurocomputing, Jan. 2003, pp. 159-175. [cited by applicant]