IP Library › Granted Patent US 12,748,953
Granted Patent B2
US 12,748,953 · App. 18/371,169 · Granted Sep 29, 2026

Multi-event time-series encoding

Inventors: Saba Zuberi (Toronto, CA); Maksims Volkovs (Toronto, CA); Aslesha Pokhrel (Toronto, CA); Alexander Jacob Labach (Toronto, CA)
Assignee: The Toronto-Dominion Bank
G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,953
App. No.
18/371,169
Granted
Sep 29, 2026
Kind
B2
Abstract

To improve processing with artificial intelligence and machine learning of multi-event time-series data, information about each event type is aggregated for a group of time bins, such that an event bin embedding represents the occurring events of that type in the time bin. The event bin embedding may be based on an aggregated event value summarizing the values of that event type in the bin and a count of those events. The event bin embeddings across event types and time bins may be combined with an embedding for static data about the data instance and a representation token for input to an encoder. The encoder may apply machine-learning blocks for an event-focused sublayer and a time-focused sublayer that attend to respective dimensions of the encoder. The model may be initially trained with self-supervised learning with time and event masking and then fine-tuned for particular applications.

Claims (32)

1 . A system for machine interpretation of multi-event time-series data, comprising:

a processor;

a non-transitory computer-readable medium having instructions executable by the processor for:

identifying an instance representation of a time-series data instance, the instance representation including a multi-dimensional representation of a plurality of event types across a plurality of time bins for the time-series data instance;

generating an encoded instance representation by applying one or more machine-learned encoding blocks to the instance representation, at least one of the machine-learned encoding blocks including:

an event-attention sublayer that applies an event-based attention layer across event types of event-sliced representations of a first sublayer input to the event-attention sublayer; and

a time-wise attention sublayer that applies a time-based attention layer across time bins of time-sliced representations of a second sublayer input to the time-attention sublayer; and

generating a decoder output by applying a machine-learned decoder to the encoded instance representation.

2 . The system of claim 1 , wherein the instructions are further executable by the processor for determining the multi-dimensional representation for each of the plurality of events at each of the plurality of time bins by applying an event embedding layer to events of an event type of the time-series data instance in the time bin.

3 . The system of claim 2 , wherein the event embedding layer is applied to an aggregated value of the events of the event type and a count of the events of the event type in an event bin.

4 . The system of claim 1 , wherein the instance representation includes, for each time bin of the plurality of time bins, a static data embedding.

5 . The system of claim 1 , wherein the instance representation includes a time bin including a learned representation token.

6 . The system of claim 1 , wherein the machine-learned encoding blocks includes a plurality of encoder blocks sequentially applied to the instance representation, the plurality of the encoder blocks including respective event-attention sublayers and time-wise attention sublayers.

7 . The system of claim 1 , wherein the instructions are further executable by the processor for training parameters of the machine-learned encoding blocks based on masked values of the time-series data instance.

8 . The system of claim 7 , wherein the instructions are further executable by the processor for fine-tuning parameters of the machine-learned encoding blocks based on labeled decoder outputs for a set of fine-tuning training data.

9 . The system of claim 1 , wherein the time-series data instance describes a sequence of health-related events.

10 . The system of claim 1 , wherein the time-series data instance describes a sequence of finance-related events.

11 . A method for machine interpretation of multi-event time-series data, the method comprising:

identifying an instance representation of a time-series data instance, the instance representation including a multi-dimensional representation of a plurality of event types across a plurality of time bins for the time-series data instance;

generating an encoded instance representation by applying one or more machine-learned encoding blocks to the instance representation, at least one of the machine-learned encoding blocks including:

an event-attention sublayer that applies an event-based attention layer across event types of event-sliced representations of a first sublayer input to the event-attention sublayer; and

a time-wise attention sublayer that applies a time-wise attention layer across time bins of time-sliced representations of a second sublayer input to the time-attention sublayer; and

generating a decoder output by applying a machine-learned decoder to the encoded instance representation.

12 . The method of claim 11 , further comprising determining the multi-dimensional representation for each of the plurality of events at each of the plurality of time bins by applying an event embedding layer to events of an event type of the time-series data instance in the time bin.

13 . The method of claim 12 , wherein the event embedding layer is applied to an aggregated value of the events of the event type and a count of the events of the event type in an event bin.

14 . The method of claim 11 , wherein the instance representation includes, for each time bin of the plurality of time bins, a static data embedding.

15 . The method of claim 11 , wherein the instance representation includes a time bin including a learned representation token.

16 . The method of claim 11 , wherein the machine-learned encoding blocks includes a plurality of encoder blocks sequentially applied to the instance representation, the plurality of the encoder blocks including respective event-attention sublayers and time-wise attention sublayers.

17 . The method of claim 11 , further comprising training parameters of the machine-learned encoding blocks based on masked values of the time-series data instance.

18 . The method of claim 17 , further comprising fine-tuning parameters of the machine-learned encoding blocks based on labeled decoder outputs for a set of fine-tuning training data.

19 . The method of claim 11 , wherein the time-series data instance describes a sequence of health-related events.

20 . The method of claim 11 , wherein the time-series data instance describes a sequence of finance-related events.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2026
From: ZUBERI, SABA; VOLKOVS, MAKSIMS; POKHREL, ASLESHA; LABACH, ALEXANDER JACOB
To: TORONTO-DOMINION BANK, THE
Reel/Frame 074238/0744 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: LAGHI, ALDO; VINT, NATHANIEL
To: ALPS SOUTH EUROPE
Reel/Frame 065187/0370 →
Continuity (2)
Provisional Application 63411932 · Sep 30, 2022
Related Publication 20240127036A1 · Apr 18, 2024
References Cited (45)
US 10614364B2 · Krumm · 2020 [cited by examiner]
US 20180164781A1 · Kubo · 2018 [cited by examiner]
US 20180285777A1 · Li · 2018 [cited by examiner]
Bak, et al., “You Can't Have AI Both Ways: Balancing Health Data Privacy and Access Fairly,” Frontiers in Genetics, vol. 13, Article 929453, Jun. 13, 2022, 7 pages; https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9234328/p… [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners,” arXiv:2005.14165v4 [cs.CL], Jul. 22, 2020, 75 pages; https://arxiv.org/pdf/2005.14165.pdf. [cited by applicant]
Cao, et al., “Data Science and AI in FinTech: An Overview,” International Journal of Data and Science Analytics, Aug. 5, 2021, 19 pages; https://arxiv.org/ftp/arxiv/papers/2007/2007.12681.pdf. [cited by applicant]
Caron, et al., “Emerging Properties in Self-Supervised Vision Transformers,” Proceedings of the IEEE/CVF international conference on computer vision, pp. 9650-9660, arXiv:2104.14294v2 [cs.CV], May 24, 2021, 21 pages; ht… [cited by applicant]
Che, et al., “Recurrent Neural Networks for Multivariate Time Series with Missing Values,” International Conference on Learning Representations, 2017, 15 pages; https://openreview.net/pdf?id=BJC8LF9ex. [cited by applicant]
Chen, et al., “Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16144-16155, Jun. 2022, 19 pages; https://ar… [cited by applicant]
Chen, et al., “Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology,” Learning Meaningful Representations of Life, arXiv:2203.00585v1 [cs.CV], Mar. 1, 2022, 11 pages; https://arxiv.org/pdf/2203.00… [cited by applicant]
Chen, et al., “XGBoost: A Scalable Tree Boosting System,” 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, arXiv:1603.02754v3 [cs.LG], Jun. 10, 2016, 13 pages; https://arxiv.org/pdf/1603.… [cited by applicant]
Cho, et al., “On the Properties of Neural Machine Translation: Encoder-Decoder Approaches,” arXiv:1409.1259v2 [cs.CL], Oct. 7, 2014, 9 pages; https:/arxiv.org/pdf/1409.1259.pdf. [cited by applicant]
Chopra, et al., “Learning a Similarity Metric Discriminatively, with Application to Face Verification,” IEEE Computer Society Conference on Computer Vision and Pattern Recognition, vol. 1, pp. 539-546, 2005, 8 pages; ht… [cited by applicant]
Darke, et al., “Benchmark Time Series Data Sets for PyTorch: The Torchtime Package,” arXiv:2207.12503v2 [cs.LG], Aug. 1, 2022, 15 pages; https://arxiv.org/pdf/2207.12503.pdf. [cited by applicant]
Devlin, et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” arXiv:1810.04805v2 [cs.CL], May 24, 2019, 16 pages; https://arxiv.org/pdf/1810.04805.pdf. [cited by applicant]
Dosovitskiy, et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” International Conference on Learning Representations, arXiv:2010.11929v2 [cs.CV], Jun. 3, 2021, 22 pages; https://arxiv.… [cited by applicant]
Gomez-Losada, et al., “Time Series Forecasting by Recommendation: An Empirical Analysis on Amazon Marketplace,” Business Information Systems, 22nd International Conference, vol. 1, May 2019, 10 pages; https://www.resear… [cited by applicant]
Grill, et al., “Bootstrap Your Own Latent a New Approach to Self-Supervised Learning,” Advances in Neural Information Processing Systems, arXiv:2006.07733v3 [cs.LG], Sep. 10, 2020, 35 pages; https://arxiv.org/pdf/2006.0… [cited by applicant]
Harutyunyan, et al., “Multitask Learning and Benchmarking with Clinical Time Series Data,” arXiv:1703.07771v1 [stat.ML], Mar. 22, 2017, 11 pages; https://www.academia.edu/34055564/Multitask_Learning_and_Benchmarking_wit… [cited by applicant]
He, et al., “Masked Autoencoders are Scalable Vision Learners,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000-16009, arXiv:2111.06377v3 [cs.CV], Dec. 19, 2021, 14 pages; https://arxiv.org/pdf… [cited by applicant]
Hochreiter, et al., “Long Short-Term Memory,” Neural Computation, vol. 9, No. 8 : 1735-1780, Dec. 1997, 32 pages; https://www.researchgate.net/publication/13853244_Long_Short-term_Memory. [cited by applicant]
Johnson, et al., “MIMIC-IV,” PhysioNet, Jun. 12, 2022, https://doi.org/10.13026/7vcr-e114. [cited by applicant]
Kidger, et al., “Neural Controlled Differential Equations for Irregular Time Series,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), arXiv:2005.08926v2 [cs.LG], Nov. 5, 2020, 25 pages; https://… [cited by applicant]
Lechner, et al., “Learning Long-Term Dependencies in Irregularly-Sampled Time Series,” arXiv:2006.04418v4 [cs.LG], Dec. 4, 2020, 19 pages; https://arxiv.org/pdf/2006.04418.pdf. [cited by applicant]
Li, et al., “Hi-BEHRT: Hierarchical Transformer-Based Model for Accurate Prediction of Clinical Events Using Multimodal Longitudinal Electronic Health Records,” IEEE journal of biomedical and health informatics, vol. 27… [cited by applicant]
Lipton, et al., “Model Missing Data in Clinical Time Series with RNNs,” Machine Learning for Healthcare, vol. 56, No. 56, arXiv:1606.04130v5 [cs.LG], Nov. 11, 2016, 17 pages; https://arxiv.org/pdf/1606.04130.pdf. [cited by applicant]
Loshchilov, et al., “Decoupled Weight Decay Regularization,” International Conference on Learning Representations 2019, arXiv:1711.05101v3 [cs.LG], Jan. 4, 2019, 19 pages; https://arxiv.org/pdf/1711.05101.pdf. [cited by applicant]
Mcdermott, et al., “A Comprehensive EHR Timeseries Pre-Training Benchmark,” Conference on Health, Inference, and Learning, pp. 257-278, Apr. 8, 2021, 23 pages; https://dl.acm.org/doi/pdf/10.1145/3450439.3451877. [cited by applicant]
Mozer, et al., Discrete-Event Continuous-Time Recurrent Nets, arXiv:1710.04110v1 [cs.NE], Oct. 11, 2017, 21 pages; https://arxiv.org/pdf/1710.04110.pdf. [cited by applicant]
Nguyen, et al., “Transformers without Tears: Improving the Normalization of Self-Attention,” arXiv:1910.05895v2 [cs.CL], Dec. 30, 2019, 11 pages; https://arxiv.org/pdf/1910.05895.pdf. [cited by applicant]
Radford, et al., “Learning Transferable Visual Models from Natural Language Supervision,” 38th International Conference on Machine Learning, arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021, 48 pages; https://arxiv.org/pdf/210… [cited by applicant]
Ren, et al., “RAPT: Pre-Training of Time-Aware Transformer for Learning Robust Healthcare Representation,” 27th ACM Sigkdd Conference on Knowledge Discovery & Data Mining, Aug. 14, 2021, 9 pages; https://www.bigscity.co… [cited by applicant]
Rubanova, et al., “Latent ODEs for Irregularly-Sampled Time Series,” Advances in Neural Information Processing Systems 32, arXiv:1907.03907v1 [cs.LG], Jul. 8, 2019, 21 pages; https://arxiv.org/pdf/1907.03907.pdf. [cited by applicant]
Shukla, et al., “Multi-Time Attention Networks for Irregularly Sampled Time Series,” International Conference on Learning Representations 2021, arXiv:2101.10318v2 [cs.LG], Jun. 7, 2021, 15 pages; https://arxiv.org/pdf/2… [cited by applicant]
Shwartz-Ziv, et al., “Tabular Data: Deep Learning is Not All You Need,” Information Fusion, 81, pp. 84-90, arXiv:2106.03253v2 [cs.LG], Nov. 23, 2021, 13 pages; https://arxiv.org/pdf/2106.03253.pdf. [cited by applicant]
Silva, et al., “Predicting Mortality of ICU Patients: The PhysioNet/Computing in Cardiology Challenge 2012,” Computing in Cardiology, pp. 245-248, 2012, 4 pages; https://lcp.mit.edu/pdf/SilvaCinC12.pdf. [cited by applicant]
Soenksen, et al., “Integrated Multimodal Artificial Intelligence Framework for Healthcare Applications,” NPJ digital medicine, vol. 5, No. 1, p. 149, Sep. 2022, 10 pages; https://www.nature.com/articles/s41746-022-00689… [cited by applicant]
Tipirneni, et al., “Self-Supervised Transformer for Sparse and Irregularly Sampled Multivariate Clinical Time-Series,” ACM Transactions on Knowledge Discovery from Data, vol. 1, No. 1, arXiv:2107.14293v2 [cs.LG], Feb. 1… [cited by applicant]
Van Der Maaten, et al., “Visualizing Data using t-SNE,” Journal of Machine Learning Research 9, 2008, 26 pages; https://www.cns.nyu.edu/events/spf/SPF_papers/JMLR_Final.pdf. [cited by applicant]
Vaswani, et al., “Attention is All you Need,” Advances in Neural Information Processing Systems 30, 2017, 11 pages; https://papers.nips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. [cited by applicant]
Wang, et al., “Linformer: Self-Attention with Linear Complexity,” arXiv:2006.04768v3 [cs.LG], Jun. 14, 2020, 12 pages; https://arxiv.org/pdf/2006.04768.pdf. [cited by applicant]
Wen, et al., “Transformers in Time Series: A Survey,” arXiv:2202.07125v5 [cs.LG], May 11, 2023, 9 pages; https://arxiv.org/pdf/2202.07125.pdf. [cited by applicant]
Xiong, et al., “On Layer Normalization in the Transformer Architecture,” 37th International Conference on Machine Learning, arXiv:2002.04745v2 [cs.LG], Jun. 29, 2020, 17 pages; https://arxiv.org/pdf/2002.04745.pdf. [cited by applicant]
Zhang, et al., “Daily-Aware Personalized Recommendation Based on Feature-Level Time Series Analysis,” 24th International Conference on World Wide Web, pp. 1373-1383, May 18, 2015, 11 pages; https://www.cs.cmu.edu/~glai1… [cited by applicant]
Zhang, et al., “Graph-Guided Network for Irregularly Sampled Multivariate Time Series,” International Conference on Learning Representations 2022, arXiv:2110.05357v2 [cs.LG], Mar. 16, 2022, 21 pages; https://arxiv.org/p… [cited by applicant]