IP Library › Granted Patent US 12,566,798
Granted Patent B2
US 12,566,798 · App. 18/895,080 · Granted Mar 3, 2026

Causal analysis with time series data

Inventors: Ajay Divakaran (Monmouth Junction, NJ); Yi Yao (Princeton, NJ); Julia Kruk (Rego Park, NY); Jesse Hostetler (Boulder, CO); Jihua Huang (Sunnyvale, CA)
Assignee: SRI International
G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,798
App. No.
18/895,080
Granted
Mar 3, 2026
Kind
B2
Abstract

In general, various aspects of the techniques are directed to causal analysis using large scale time series data. A computing system may convert large scale time series data to first time period records and second time period records according to a multi-scale time resolution. The computing system may implement a hierarchical machine learning model to generate embeddings that capture temporal characteristics of features of the large scale time series data. The computing system may generate a graph data structure indicating cause and effect correlations between features of the large scale time series data based on temporal dynamics captured in the cause and second time period records and/or the embeddings.

Claims (87)

1 . A computing system for causal analysis using time series data, the computing system comprising:

processing circuitry; and

memory comprising instructions that, when executed, cause the processing circuitry to:

generate a first time period record based on a first plurality of feature values associated with a plurality of features, wherein the first plurality of feature values include feature values for a first set of data points of time series data, the first set of data points associated with a first time period, wherein an entry of the first time period record indicates a combined first time period feature value associated with a first feature of the plurality of features;

generate a second time period record based on a second plurality of feature values associated with the plurality of features, wherein the second plurality of feature values include feature values for a second set of data points of the time series data, the second set of data points associated with a second time period following the first time period according to a time resolution, wherein an entry of the second time period record indicates a combined second time period feature value associated with the first feature of the plurality of features;

generate, based on the first time period record and the second time period record, a graph data structure indicating cause and effect correlations between features of the plurality of features; and

output an indication including the graph data structure for purposes of indicating the causal analysis of the time series data.

2 . The computing system of claim 1 , wherein to generate the first time period record, the instructions cause the processing circuitry to:

for each data point of the first set of data points:

extract, based on the first time period and timestamps associated with the first set of data points, one or more first time period feature values of the first plurality of feature values associated the plurality of features; and

generate the entry of the first time period record by combining extracted first time period feature values associated with the first feature for each data point of the first set of data points.

3 . The computing system of claim 1 , wherein to generate the second time period record, the instructions cause the processing circuitry to:

for each data point of the second set of data points:

extract, based on the second time period and timestamps associated with the second set of data points, one or more second time period feature values of the second plurality of feature values associated the plurality of features; and

generate the entry of the second time period record by combining extracted second time period feature values associated with the first feature for each data point of the second set of data points.

4 . The computing system of claim 1 , wherein to generate the graph data structure, the instructions cause the processing circuitry to:

determine a target variable associated with a second feature of the plurality of features;

assign, based on the target variable, flags to one or more data points of the second set of data points; and

generate the graph data structure further based on the flags and external events associated with the target variable.

5 . The computing system of claim 1 , wherein to generate the graph data structure, the instructions cause the processing circuitry to:

generate an initial graph that fully connects each feature of the plurality of features; and

prune, based on temporal dynamics associated with the first time period record and the second time period record, the initial graph to generate the graph data structure.

6 . The computing system of claim 1 , wherein to generate the graph data structure, the instructions cause the processing circuitry to:

generate a third time period record based on a third plurality of feature values associated with the plurality of features, wherein the third plurality of feature values include feature values for a third set of data points of the time series data, the third set of data points associated with a third time period following the second time period according to the time resolution, wherein an entry of the third time period record indicates a combined second time period feature value associated with the first feature of the plurality of features; and

generate the graph data structure, further based on the third time period record.

7 . Non-transitory computer-readable storage media comprising machine readable instructions for configuring processing circuitry to:

generate a first time period record based on a first plurality of feature values associated with a plurality of features, wherein the first plurality of feature values include feature values for a first set of data points of time series data, the first set of data points associated with a first time period, wherein an entry of the first time period record indicates a combined first time period feature value associated with a first feature of the plurality of features;

generate a second time period record based on a second plurality of feature values associated with the plurality of features, wherein the second plurality of feature values include feature values for a second set of data points of the time series data, the second set of data points associated with a second time period following the first time period according to a time resolution, wherein an entry of the second time period record indicates a combined second time period feature value associated with the first feature of the plurality of features;

generate, based on the first time period record and the second time period record, a graph data structure indicating cause and effect correlations between features of the plurality of features; and

output an indication including the graph data structure for purposes of indicating causal analysis of the time series data.

8 . A method comprising:

generating, by processing circuitry, a first time period record based on a first plurality of feature values associated with a plurality of features, wherein the first plurality of feature values include feature values for a first set of data points of time series data, the first set of data points associated with a first time period, wherein an entry of the first time period record indicates a combined first time period feature value associated with a first feature of the plurality of features;

generating, by the processing circuitry, a second time period record based on a second plurality of feature values associated with the plurality of features, wherein the second plurality of feature values include feature values for a second set of data points of the time series data, the second set of data points associated with a second time period following the first time period according to a time resolution, wherein an entry of the second time period record indicates a combined second time period feature value associated with the first feature of the plurality of features;

generating, by the processing circuitry and based on the first time period record and the second time period record, a graph data structure indicating cause and effect correlations between features of the plurality of features; and

outputting, by the processing circuitry, an indication including the graph data structure for purposes of indicating causal analysis of the time series data.

9 . A computing system for causal analysis using time series data, the computing system comprising:

processing circuitry; and

memory comprising instructions that, when executed, cause the processing circuitry to:

generate a first embedding that characterizes a first plurality of feature values associated with a plurality of features for a first set of data points included in time series data, the first set of data points associated with a first time period;

generate a second embedding that characterizes a second plurality of feature values associated with the plurality of features for a second set of data points included in the time series data, the second set of data points associated with a second time period that follows the first time period;

generate, based on the first embedding and the second embedding, a sequence embedding that characterizes one or more temporal trends associated with the plurality of features from the first time period to the second time period;

generate, based on the sequence embedding, the first embedding, and the second embedding, a graph data structure indicating cause and effect correlations between the plurality of features; and

output an indication including the graph data structure for purposes of indicating the causal analysis of the time series data.

10 . The computing system of claim 9 , wherein to generate the first embedding, the instructions cause the processing circuitry to:

extract the first plurality of feature values from the first set of data points;

provide the first plurality of feature values to a machine learning model; and

generate, using the machine learning model and based on the first plurality of feature values, the first embedding.

11 . The computing system of claim 9 , wherein to generate the second embedding, the instructions cause the processing circuitry to:

extract the second plurality of feature values from the second set of data points;

provide the second plurality of feature values to a machine learning model; and

generate, using the machine learning model and based on the second plurality of feature values, the second embedding.

12 . The computing system of claim 9 , wherein to generate the sequence embedding, the instructions cause the processing circuitry to compare the first plurality of feature values as characterized in the first embedding to the second plurality of feature values as characterized in the second embedding.

13 . The computing system of claim 9 , wherein the instructions further cause the processing circuitry to:

determine a counterfactual hypothesis as a different feature value for a feature value of the first plurality of feature values;

generate, based on the first plurality of feature values including the different feature value, a third embedding that characterizes the first plurality of features including the different feature value;

generate, based on the third embedding and the second embedding, a second sequence embedding; and

predict, based on the second sequence embedding, a defect rate for counterfactual analysis.

14 . The computing system of claim 9 , wherein the instructions further cause the processing circuitry to interpolate one or more feature values of the first plurality of feature values.

15 . The computing system of claim 9 , wherein the sequence embedding is a first sequence embedding, and wherein to generate the graph data structure the instructions cause the processing circuitry to:

generate a third embedding that characterizes a third plurality of feature values associated with the plurality of features for a third set of data points included in time series data, the third set of data points associated with a third time period that follows the second time period;

generate a fourth embedding that characterizes a fourth plurality of feature values associated with the plurality of features for a fourth set of data points included in time series data, the fourth set of data points associated with a fourth time period that follows the third time period;

generate, based on the third embedding and the fourth embedding, a second sequence embedding that characterizes one or more temporal trends associated with the plurality of features from the third time period to the fourth time period;

generate, based on the first sequence embedding and the second sequence embedding, a third sequence embedding that characterizes one or more temporal trends associated with the plurality of features from the first time period to the fourth time period; and

generate the graph data structure based on the first sequence embedding, the second sequence embedding, the third sequence embedding, the first embedding, the second embedding, the third embedding, and the fourth embedding.

16 . The computing system of claim 9 , wherein to generate the graph data structure, the instructions cause the processing circuitry to:

generate an initial graph data structure that fully connects features of the plurality of features;

determine a gradient attribution vector based on the first embedding, the second embedding, and the sequence embedding; and

generate the graph data structure based on the initial graph data structure and the gradient attribution vector.

17 . Non-transitory computer-readable storage media comprising machine readable instructions for configuring processing circuitry to:

generate a first embedding that characterizes a first plurality of feature values associated with a plurality of features for a first set of data points included in time series data, the first set of data points associated with a first time period;

generate a second embedding that characterizes a second plurality of feature values associated with the plurality of features for a second set of data points included in the time series data, the second set of data points associated with a second time period that follows the first time period;

generate, based on the first embedding and the second embedding, a sequence embedding that characterizes one or more temporal trends associated with the plurality of features from the first time period to the second time period;

generate, based on the sequence embedding, the first embedding, and the second embedding, a graph data structure indicating cause and effect correlations between the plurality of features; and

output an indication including the graph data structure for purposes of indicating causal analysis of the time series data.

18 . The non-transitory computer-readable storage media of claim 17 , wherein to generate the sequence embedding, the machine readable instructions configure the processing circuitry to compare the first plurality of feature values as characterized in the first embedding to the second plurality of feature values as characterized in the second embedding.

19 . A method comprising:

generating, by processing circuitry, a first embedding that characterizes a first plurality of feature values associated with a plurality of features for a first set of data points included in time series data, the first set of data points associated with a first time period;

generating, by the processing circuitry, a second embedding that characterizes a second plurality of feature values associated with the plurality of features for a second set of data points included in the time series data, the second set of data points associated with a second time period that follows the first time period;

generating, by the processing circuitry and based on the first embedding and the second embedding, a sequence embedding that characterizes one or more temporal trends associated with the plurality of features from the first time period to the second time period;

generating, by the processing circuitry and based on the sequence embedding, a graph data structure indicating cause and effect correlations between the plurality of features; and

outputting, by the processing circuitry, data for a graphical user interface to include the graph data structure.

20 . The method of claim 19 , wherein the sequence embedding is a first sequence embedding, and wherein generating the graph data structure comprises:

generating, by the processing circuitry, a third embedding that characterizes a third plurality of feature values associated with the plurality of features for a third set of data points included in time series data, the third set of data points associated with a third time period that follows the second time period;

generating, by the processing circuitry, a fourth embedding that characterizes a fourth plurality of feature values associated with the plurality of features for a fourth set of data points included in time series data, the fourth set of data points associated with a fourth time period that follows the third time period;

generating, by the processing circuitry and based on the third embedding and the fourth embedding, a second sequence embedding that characterizes one or more temporal trends associated with the plurality of features from the third time period to the fourth time period;

generating, by the processing circuitry and based on the first sequence embedding and the second sequence embedding, a third sequence embedding that characterizes one or more temporal trends associated with the plurality of features from the first time period to the fourth time period; and

generating, by the processing circuitry, the graph data structure based on the first sequence embedding, the second sequence embedding, the third sequence embedding, the first embedding, the second embedding, the third embedding, and the fourth embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2024
From: DIVAKARAN, AJAY; YAO, YI; KRUK, JULIA; HOSTETLER, JESSE; HUANG, JIHUA
To: SRI INTERNATIONAL
Reel/Frame 068799/0038 →
Continuity (2)
Provisional Application 63542249 · Oct 3, 2023
Related Publication 20250110989A1 · Apr 3, 2025
References Cited (26)
US 7221987B2 · Bett · 2007 [cited by examiner]
US 7454311B2 · Maeda · 2008 [cited by examiner]
US 7711734B2 · Leonard · 2010 [cited by examiner]
US 9589031B2 · Lin · 2017 [cited by examiner]
US 10565171B1 · Bruno · 2020 [cited by examiner]
US 11210315B2 · Willson · 2021 [cited by examiner]
US 20140172867A1 · Lin · 2014 [cited by examiner]
US 20160217384A1 · Leonard · 2016 [cited by examiner]
US 20170031742A1 · Jilani · 2017 [cited by examiner]
US 20190026351A1 · Maor · 2019 [cited by examiner]
US 20200082284A1 · Moghtaderi · 2020 [cited by examiner]
US 20200301972A1 · Wang · 2020 [cited by examiner]
US 20220035354A1 · Mostafavi · 2022 [cited by examiner]
US 20220335064A1 · Gottemukkala · 2022 [cited by examiner]
US 20230072173A1 · Gupta · 2023 [cited by examiner]
US 20230186174A1 · Nitzken · 2023 [cited by examiner]
“Covid Data Tracker”, U.S Centers for Disease Control and Prevention, Accessed from: https://covid.cdc.gov/covid-data-tracker/#datatracker-home, Retrieved on: Jun. 11, 2024, 2 pp. [cited by applicant]
“Our Governance—The Organisation”, Transparency International, Accessed from: https://www.transparency.org/en/the-organisation/our-governance, Retrieved on: Jun. 11, 2024, 1 pp. [cited by applicant]
“Package edu.cmu.tetrad.search”, Accessed from: https://www.phil.cmu.edu/tetrad-javadocs/7.6.6/edu/cmu/tetrad/search/package-summary.html, Retrieved on: Jun. 11, 2024, 6 pp. [cited by applicant]
“Survey on Coverage, Operational Reach, and Effectiveness (SCORE)”, Humanitarian Outcomes, Accessed from: https://humanitarianoutcomes.org/projects/core, Retrieved on: Jun. 11, 2024, 7 pp. [cited by applicant]
“Tetrad User Manual”, Center for Causal Discovery, Accessed from: https://htmlpreview.github.io/?https:///github.com/cmu-phil/tetrad/blob/development/tetrad-lib/src/main/resources/docs/manual/index.html, May 2020, 109 p… [cited by applicant]
Asad et al., “Mexico-U.S. Migration in Time: From Economic to Social Mechanisms”, The Annals of the American Academy of Political and Social Science, vol. 684, No. 1, Jul. 2019, pp. 60-84. [cited by applicant]
Costa et al., “How Deep Learning Sees the World: A Survey on Adversarial Attacks & Defenses”, arXiv:2305.10862v1, May 18, 2023, 20 pp. [cited by applicant]
Eberhardt, “Introduction to the foundations of causal discovery”, International Journal of Data Science and Analytics, vol. 3, Dec. 28, 2016, pp. 81-91. [cited by applicant]
Feng et al., “TwiBot-20: A Comprehensive Twitter Bot Detection Benchmark”, arXiv:2106.13088v4, Aug. 27, 2021, 10 pp. [cited by applicant]
Sayyadiharikandeh et al., “Detection of Novel Social Bots by Ensembles of Specialized Classifiers”, arXiv:2006.06867v2, Aug. 14, 2020, 8 pp. [cited by applicant]