IP Library Granted Patent US 12,412,665
Granted Patent B2
US 12,412,665 · App. 18/339,726 · Granted Sep 9, 2025

Methods for evaluating the effect of the start date for cancer treatment with a cancer medication using propensity scoring

Inventors: Ashraf Hafez (Woodinville, WA); Caroline Epstein (Chicago, IL)
Assignee: Tempus AI, Inc.
G16H50/30A61B5/4848G06F3/011G16H20/40G16H50/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,665
App. No.
18/339,726
Granted
Sep 9, 2025
Kind
B2
Abstract

An evaluation of a cancer treatment start date for a cancer medication identifies a first plurality of subjects and, for each, the treatment start date. A subject is selected for a second plurality of subjects by applying features for the subject to a model at a propensity value threshold. One subset of the features is associated with a time period and another subset is static. The applying obtains anchor point predictions, each associated with a time in the time period and including a probability that the time is the start date. The time of the anchor point prediction having the greatest probability is assigned the anchor point of the subject. The start date is evaluated with a survival objective based on the start date for each subject in the first plurality and the assigned anchor point for each subject in the second plurality.

Claims (80)

1. A computer-implemented method of quantifying a survival duration caused by a cancer treatment, for a cancer, performed at a computer system comprising one or more processors, the method comprising:

(A) identifying, via the one or more processors, a first plurality of subjects from a base population, wherein each subject in the base population has the cancer and wherein each respective subject in the first plurality of subjects incurred the cancer treatment for the cancer on a corresponding treatment start date;

(B) determining, via the one or more processors, a second plurality of subjects from the base population, wherein the second plurality of subjects are other than the first plurality of subjects, and wherein each subject in the second plurality of subjects has not undergone the cancer treatment;

(C) training, via the one or more processors, a propensity machine learning model, using a classification algorithm with the survival duration caused by the cancer treatment as an objective response variable, to generate a trained propensity machine learning model;

(D) determining, via the one or more processors, for each respective subject in the second plurality of subjects, for each respective time period in a corresponding plurality of time periods for the respective subject, a corresponding propensity score that the respective subject incurred cancer treatment in the corresponding time period by applying, a corresponding value for each feature in a plurality of features for the respective subject to the trained propensity machine learning model, wherein the trained propensity machine learning model uses both (i) a binary indication of whether or not the respective subject incurred cancer treatment in the corresponding time period and (ii) the corresponding value for each feature in the plurality of features for the respective subject, to determine the corresponding propensity score, wherein a first subset of the corresponding plurality of features for the respective subject for which data is acquired for the respective subject is associated with the respective time period and a second subset of the corresponding plurality of features for which data was acquired for the respective subject is static, thereby determining a corresponding propensity score, in a plurality of propensity scores, for each respective subject in the second plurality of subjects;

(E) assigning, via the one or more processors, for each respective subject in the second plurality of subjects a corresponding anchor point for the respective subject to be the respective time period in the corresponding plurality of time periods for the respective subject having the highest propensity score;

(F) creating, via the one or more processors, a first survival curve over a plurality of time intervals, for the first plurality of subjects by computing a Kaplan-Meier estimate at each time interval in the plurality of time intervals using (i) the corresponding treatment start date for each respective subject in the first plurality of subjects, (ii) the corresponding propensity scores computed for each subject in the first plurality of subjects and (iii) a number of deaths incurred in each time interval in the plurality of time intervals in the first plurality of subjects, wherein the first survival curve provides a respective survival function estimate for the first plurality of subjects at each respective time interval in the plurality of time intervals after the respective treatment start date of each subject in the first plurality of subjects;

(G) creating, via the one or more processors, a second survival curve over the plurality of time intervals, for the second plurality of subjects by computing a Kaplan-Meier estimate at each time interval in the plurality of time intervals using (i) the corresponding anchor point for each respective subject in the second plurality of subjects, (ii) the corresponding propensity scores computed for each subject in the second plurality of subjects and (iii) a number of deaths incurred in each time interval in the plurality of time intervals in the second plurality of subjects, wherein the second survival curve provides a respective survival function estimate for the second plurality of subjects at each respective time interval in the plurality of time intervals after the respective anchor point of each subject in the second plurality of subjects;

(H) computing, via the one or more processors, a first kernel density estimation plot by computing a first kernel density estimation through application of a kernel function to the corresponding plurality of propensity scores for each respective subject in the first plurality of subjects;

(I) computing, via the one or more processors, a second kernel density estimation plot by computing a second kernel density estimation through application of a kernel function to the corresponding plurality of propensity scores for each respective subject in the second plurality of subjects; and

(J) quantifying, via the one or more processors, the survival duration caused by the cancer treatment as a difference in the first survival curve and the second survival curve when the kernel density estimation plot overlaps the second kernel density estimation plot.

2. The computer-implemented method of claim 1 , the method further comprising:

filtering, via the one or more processors, the first plurality of subjects and the second plurality of subjects by a procedure comprising:

(i) displaying on a user interface a first propensity histogram for the first plurality of subjects at treatment start date and a second propensity histogram for the second plurality of subjects at anchor point,

(ii) obtaining a propensity value threshold from the user interface, and

(iii) removing from the first plurality of subjects and the second plurality of subjects those subjects whose anchor points fail to satisfy the propensity value threshold; and wherein

the first propensity histogram and the second propensity histogram are superimposed on each other.

3. The computer-implemented method of claim 1 , wherein

the first survival curve further provides a respective confidence interval for the first plurality of subjects at each respective time interval in the plurality of time intervals, and

the second survival curve further provides a respective confidence interval for the second plurality of subjects at each respective time interval in the plurality of time intervals.

4. The computer-implemented method of claim 1 , wherein the method further comprises:

filtering, via the one or more processors, the first plurality of subjects and the second plurality of subjects by a procedure comprising:

(i) displaying on a user interface a first propensity histogram for the first plurality of subjects at treatment start date and a second propensity histogram for the second plurality of subjects at anchor point,

(ii) obtaining a propensity value threshold from the user interface, and

(iii) removing from the first plurality of subjects and the second plurality of subjects those subjects whose anchor points fail to satisfy the propensity value threshold; and wherein

the first propensity histogram and the second propensity histogram are superimposed on each other,

the first survival curve further provides a respective confidence interval for the first plurality of subjects at each respective time interval in the plurality of time intervals, and

the second survival curve further provides a respective confidence interval for the second plurality of subjects at each respective time interval in the plurality of time intervals.

5. The computer-implemented method of claim 1 , wherein the method further comprises: filtering, via the one or more processors, the first plurality of subjects and the second plurality of subjects by a procedure comprising: (i) displaying on a user interface a first propensity histogram for the first plurality of subjects at treatment start date and a second propensity histogram for the second plurality of subjects at anchor point, (ii) obtaining a propensity value threshold from the user interface, and (iii) removing from the first plurality of subjects and the second plurality of subjects those subjects whose anchor points fail to satisfy the propensity value threshold; and wherein the obtaining a propensity value threshold from the user interface comprises providing a slidable lower bound threshold and a slideable upper bound threshold and the computer-implemented method further comprises repeating the filtering, creating (F), and creating (G) without further human intervention each time the slidable lower bound threshold or the slideable upper bound threshold is moved on the user interface.

6. The computer-implemented method of claim 1 , wherein the method further comprises:

filtering, via the one or more processors, the first plurality of subjects and the second plurality of subjects by a procedure comprising:

(i) displaying on a user interface a first propensity histogram for the first plurality of subjects at treatment start date and a second propensity histogram for the second plurality of subjects at anchor point,

(ii) obtaining a propensity value threshold from the user interface, and

(iii) removing from the first plurality of subjects and the second plurality of subjects those subjects whose anchor points fail to satisfy the propensity value threshold; and wherein

the first propensity histogram and the second propensity histogram are superimposed on each other,

the first survival curve further provides a respective confidence interval for the first plurality of subjects at each respective time interval in the plurality of time intervals,

the second survival curve further provides a respective confidence interval for the second plurality of subjects at each respective time interval in the plurality of time intervals, and

the obtaining a propensity value threshold from the user interface comprises providing a slidable lower bound threshold and a slideable upper bound threshold and the computer-implemented method further comprises repeating the filtering, creating (F), and creating (G) without further human intervention each time the slidable lower bound threshold or the slideable upper bound threshold is moved on the user interface.

7. The computer-implemented method of claim 1 , wherein the propensity machine learning model is a random forest model, a gradient boosting model, a neural network model, a boosted Classification and Regressions Trees (CART) model, a generalized boosted model, or a greedy nearest neighbor model.

8. The computer-implemented method of claim 1 , wherein the treatment for cancer comprises application of a medication to a subject or a medical procedure performed on a subject.

9. The computer-implemented method of claim 1 , wherein the treatment is a surgical procedure or a radiation treatment.

10. The computer-implemented method of claim 1 , wherein the survival function estimate is a time until death, a time until progression of the cancer, or a time until an adverse event associated with the cancer is incurred.

11. The computer-implemented method of claim 1 , wherein the corresponding plurality of features for the respective subject comprises a corresponding plurality of demographic features for the respective subject and a plurality of clinical temporal data for the respective subject.

12. The computer-implemented method of claim 11 , wherein the corresponding plurality of features further comprises a corresponding plurality of genomic features for the respective subject.

13. The computer-implemented method of claim 1 , wherein the method further comprises receiving an indication of the cancer and the treatment based on user input received via the user interface prior to the identifying (A).

14. The computer-implemented method of claim 1 , wherein the propensity value threshold is a propensity value range.

15. The method of claim 14 , wherein the propensity value range is between 0 and 1.

16. The computer-implemented method of claim 1 , wherein the cancer is breast cancer, colon cancer, lung cancer, ovary cancer, or prostate cancer.

17. The computer-implemented method of claim 1 , wherein the corresponding plurality of time periods spans a period of days, months or years.

18. The computer-implemented method of claim 1 , wherein a feature in the second subset of the corresponding plurality of features is gender, race, or year of birth, family history, body weight, size, or body mass index.

19. The computer-implemented method of claim 1 , wherein a feature in the first subset of the corresponding plurality of features is months since birth, smoking status, menopausal status, time since menopause, time since last smoked, primary cancer site observed, metastasis site observed, cancer recurrence site observed, tumor characterization, medical procedure performed, medication type administered, radiotherapy treatment administered, time since primary diagnosis, time since predefined cancer stage diagnosed, time since metastasis, time since last recurrence of cancer, time since medical procedure performed, time since predefined medication taken, time since radiotherapy treatment administered, imaging procedure performed, change in tumor characteristic, rate of change in tumor characteristic, or predetermined response observed.

20. The computer-implemented method of claim 1 , wherein a first feature in the corresponding plurality of features is obtained from a biological sample of the respective subject and corresponds to a DNA for a predetermined human gene.

21. The computer-implemented method of claim 20 , wherein the first feature is a count of somatic mutations observed for the DNA in the biological sample of the respective subject.

22. The computer-implemented method of claim 1 , wherein a first feature in the corresponding plurality of features is a number of somatic mutations on a predetermined chromosome as determined by sequencing DNA from a biological sample obtained from the respective subject.

23. The computer-implemented method of claim 1 , wherein a first feature in the corresponding plurality of features is a number of genes with mutations on a predetermined chromosome as determined by sequencing DNA from a biological sample obtained from the respective subject.

24. The computer-implemented method of claim 1 , wherein a first feature in the corresponding plurality of features is a mutation density of a predetermined chromosome as determined by sequencing DNA from a biological sample obtained from the respective subject.

25. The computer-implemented method of claim 1 , wherein a first feature in the corresponding plurality of features is a number of mutations of a defined mutational class of a predetermined chromosome as determined by sequencing DNA from a biological sample obtained from the respective subject.

26. The computer-implemented method of claim 25 , wherein the defined mutational class is single nucleotide polymorphism (SNP), multiple nucleotide polymorphism (MNP), insertions (INS), deletion (DEL), or translocation.

27. A computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors, the one or more programs comprising instructions for performing a method of quantifying a survival duration caused by a cancer treatment, for a cancer, the method comprising:

(A) identifying, via the one or more processors, a first plurality of subjects from a base population, wherein each subject in the base population has the cancer and wherein each respective subject in the first plurality of subjects incurred the cancer treatment for the cancer on a corresponding treatment start date;

(B) determining, via the one or more processors, a second plurality of subjects from the base population, wherein the second plurality of subjects are other than the first plurality of subjects, and wherein each subject in the second plurality of subjects has not undergone the cancer treatment;

(C) training, via the one or more processors, a propensity machine learning model, using a classification algorithm with the survival duration caused by the cancer treatment as an objective response variable, to generate a trained propensity machine learning model;

(D) determining, via the one or more processors, for each respective subject in the second plurality of subjects, for each respective time period in a corresponding plurality of time periods for the respective subject, a corresponding propensity score that the respective subject incurred cancer treatment in the corresponding time period by applying, a corresponding value for each feature in a plurality of features for the respective subject to the trained propensity machine learning model, wherein the trained propensity machine learning model uses both (i) a binary indication of whether or not the respective subject incurred cancer treatment in the corresponding time period and (ii) the corresponding value for each feature in the plurality of features for the respective subject, to determine the corresponding propensity score, wherein a first subset of the corresponding plurality of features for which data was acquired for the respective subject is associated with the respective time period and a second subset of the corresponding plurality of features for which data was acquired for the respective subject is static, thereby determining a corresponding propensity score, in a plurality of propensity scores, for each respective subject in the second plurality of subjects;

(E) assigning, via the one or more processors, for each respective subject in the second plurality of subjects, a corresponding anchor point for the respective subject to be the respective time period in the corresponding plurality of time periods for the respective subject having the highest propensity score;

(F) creating, via the one or more processors, a first survival curve over a plurality of time intervals, for the first plurality of subjects by computing a Kaplan-Meier estimate at each time interval in the plurality of time intervals using (i) the corresponding treatment start date for each respective subject in the first plurality of subjects, (ii) the corresponding propensity scores computed for each subject in the first plurality of subjects and (iii) a number of deaths incurred in each time interval in the plurality of time intervals in the first plurality of subjects, wherein the first survival curve provides a respective survival function estimate for the first plurality of subjects at each respective time interval in the plurality of time intervals after the respective treatment start date of each subject in the first plurality of subjects;

(G) creating, via the one or more processors, a second survival curve over the plurality of time intervals, for the second plurality of subjects by computing a Kaplan-Meier estimate at each time interval in the plurality of time intervals using (i) the corresponding anchor point for each respective subject in the second plurality of subjects, (ii) the corresponding propensity scores computed for each subject in the second plurality of subjects and (iii) a number of deaths incurred in each time interval in the plurality of time intervals in the second plurality of subjects, wherein the second survival curve provides a respective survival function estimate for the second plurality of subjects at each respective time interval in the plurality of time intervals after the respective anchor point of each subject in the second plurality of subjects;

(H) computing, via the one or more processors, a first kernel density estimation plot by computing a first kernel density estimation through application of a kernel function to the corresponding plurality of propensity scores for each respective subject in the first plurality of subjects;

(I) computing, via the one or more processors, a second kernel density estimation plot by computing a second kernel density estimation through application of a kernel function to the corresponding plurality of propensity scores for each respective subject in the second plurality of subjects; and

(J) quantifying, via the one or more processors, the survival duration caused by the cancer treatment as a difference in the first survival curve and the second survival curve when the kernel density estimation plot overlaps the second kernel density estimation plot.

28. A non-transitory computer readable storage medium, wherein the non-transitory computer readable storage medium stores instructions, which when executed by a computer system having one or more processors, cause the computer system to perform a method of quantifying a survival duration caused by a cancer treatment, for a caner, the method comprising:

(A) identifying, via the one or more processors, a first plurality of subjects from a base population, wherein each subject in the base population has the cancer and wherein each respective subject in the first plurality of subjects incurred a cancer treatment for the cancer on a corresponding treatment start date;

(B) determining, via the one or more processors, a second plurality of subjects from the base population, wherein the second plurality of subjects are other than the first plurality of subjects, and wherein each subject in the second plurality of subjects has not undergone the cancer treatment;

(C) training, via the one or more processors, a propensity machine learning model, using a classification algorithm with the survival duration caused by the cancer treatment as an objective response variable, to generate a trained propensity machine learning model;

(D) determining, via the one or more processors, for each respective subject in the second plurality of subjects, for each respective time period in a corresponding plurality of time periods for the respective subject, a corresponding propensity score that the respective subject incurred cancer treatment in the corresponding time period by applying, a corresponding value for each feature in the plurality of features for the respective subject to the trained propensity machine learning model, wherein the trained propensity machine learning model uses both (i) a binary indication of whether or not the respective subject incurred cancer treatment in the corresponding time period and (ii) the corresponding value for each feature in a plurality of features for the respective subject, to determine the corresponding propensity score, wherein a first subset of the corresponding plurality of features for which data was acquired for the respective subject is associated with the respective time period and a second subset of the corresponding plurality of features for which data was acquired for the respective subject is static, thereby determining a corresponding propensity score, in a plurality of propensity scores, for each respective subject in the second plurality of subjects;

(E) assigning, via the one or more processors, for each respective subject in the second plurality of subjects a corresponding anchor point for the respective subject to be the respective time period in the corresponding plurality of time periods for the respective subject having the highest propensity score;

(F) creating, via the one or more processors, a first survival curve over a plurality of time intervals, for the first plurality of subjects by computing a Kaplan-Meier estimate at each time interval in the plurality of time intervals using (i) the corresponding treatment start date for each respective subject in the first plurality of subjects, (ii) the corresponding propensity scores computed for each subject in the first plurality of subjects and (iii) a number of deaths incurred in each time interval in the plurality of time intervals in the first plurality of subjects, wherein the first survival curve provides a respective survival function estimate for the first plurality of subjects at each respective time interval in the plurality of time intervals after the respective treatment start date of each subject in the first plurality of subjects;

(G) creating, via the one or more processors, a second survival curve over the plurality of time intervals, for the second plurality of subjects by computing a Kaplan-Meier estimate at each time interval in the plurality of time intervals using (i) the corresponding anchor point for each respective subject in the second plurality of subjects, (ii) the corresponding propensity scores computed for each subject in the second plurality of subjects and (iii) a number of deaths incurred in each time interval in the plurality of time intervals in the second plurality of subjects, wherein the second survival curve provides a respective survival function estimate for the second plurality of subjects at each respective time interval in the plurality of time intervals after the respective anchor point of each subject in the second plurality of subjects;

(H) computing, via one or more processors, a first kernel density estimation plot by computing a first kernel density estimation through application of a kernel function to the corresponding plurality of propensity scores for each respective subject in the first plurality of subjects;

(I) computing, via one or more processors, a second kernel density estimation plot by computing a second kernel density estimation through application of a kernel function to the corresponding plurality of propensity scores for each respective subject in the second plurality of subjects; and

(J) quantifying, via one or more processors, the survival duration caused by the cancer treatment as a difference in the first survival curve and the second survival curve when the kernel density estimation plot overlaps the second kernel density estimation plot.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded May 14, 2026
From: ARES CAPITAL CORPORATION, AS COLLATERAL AGENT
To: TEMPUS AI, INC. (F/K/A TEMPUS LABS, INC.)
Reel/Frame 075577/0513 →
SECURITY INTEREST Recorded Jun 2, 2025
From: TEMPUS AI, INC.
To: ARES CAPITAL CORPORATION, AS COLLATERAL AGENT
Reel/Frame 071468/0107 →
CHANGE OF NAME Recorded Feb 29, 2024
From: TEMPUS LABS, INC.
To: TEMPUS AI, INC.
Reel/Frame 066707/0382 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2023
From: HAFEZ, ASHRAF; EPSTEIN, CAROLINE
To: TEMPUS LABS, INC.
Reel/Frame 064886/0811 →
Continuity (2)
Continuation 16679054 · Nov 8, 2019
Related Publication 20240423536A1 · Dec 26, 2024
References Cited (22)
US 7912734B2 · Kil · 2011 [cited by applicant]
US 20130144636A1 · Pouliot et al. · 2013 [cited by applicant]
US 20140343959A1 · Hasegawa · 2014 [cited by examiner]
US 20170083682A1 · McNutt · 2017 [cited by examiner]
US 20170107577A1 · Al-Ejeh · 2017 [cited by examiner]
US 20210090694A1 · Colley · 2021 [cited by examiner]
US 20220028534A1 · Sugiyama · 2022 [cited by examiner]
US 20230106057A1 · Wolf · 2023 [cited by examiner]
WO WO2014183023A1 · 2014 [cited by applicant]
Ross, R. et al.(2021). Veridical causal inference using propensity score methods for comparative effectiveness research with medical claims. Health Services & Outcomes Research Methodology, 21(2), 206-228. doi:http://dx… [cited by examiner]
Mahalanobis distance | Wikipedia, the free encyclopedia, 2004. [Online; accessed Jul. 19, 2018], en.wikipedia.org/wiki/Mahalanobis_distance, pp. 1-5. [cited by applicant]
Austin, Peter C., “The use of propensity score methods with survival or time-to-event outcomes: reporting measures of effect similar to those used in randomized experiments”, Statistics in Medicine, 33, Sep. 3, 2013, pp… [cited by applicant]
Austin, Peter C., et al. “Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies”, Statistics i… [cited by applicant]
Cole, Stephen R., et al. “Adjusted survival curves with inverse probability weights”, Computer Methods and Programs in Biomedicine, Elsevier, 2004, 75, pp. 45-49. [cited by applicant]
Hong, et al. “Feasibility Study Using Propensity Score Matching Methods for the Pseudo-Common Person Equating Requirement”, OTJR: Occupation, Participation and Health, 2019, vol. 39(1), pp. 32-40. [cited by applicant]
Jackson, et al. “Propensity Scores in Pharmacoepidemiology: Beyond the Horizon”, Curr Epidemiol Rep., Dec. 2017, 4(4), pp. 271-280. [cited by applicant]
Reinisch, June M., et al. “In Utero Exposure to Phenobarbital and Intelligence Deficits in Adult Men”, Journal of the American Medical Association, Nov. 15, 1995, vol. 274, No. 19, pp. 1518-1525. [cited by applicant]
Rosenbaum, Paul R., et al. “The Central Role of the Propensity Score in Observational Studies for Causal Effects”, Biometrika, vol. 70, No. 1, Apr. 1983, pp. 41-55. [cited by applicant]
Rosenbaum, Paul R., et al. “Constructing a Control Group Using Multivariate Matched Sampling Methods that Incorporate the Propensity Score”, American Statistical Association, Feb. 1985, vol. 39, No. 1, pp. 33-38. [cited by applicant]
Ross, R. et al. “Veridical causal inference using propensity score methods for comparative effectiveness research with medical claims”, Health Services & Outcomes Research Methodology, (2021), 21(2), pp. 206-228, doi:ht… [cited by applicant]
Shi, et al. “Evaluation of the benefit of post-mastectomy radiotherapy in patients with earlystage breast cancer: A propensity score matching study”, Oncology Letters, 17: 2019, pp. 4851-4858. [cited by applicant]
Stuart, Elizabeth A. “Matching methods for causal inference: a review and a look Forward”, Statistical Science, 25(1), Feb. 1, 2010, pp. 1-29. [cited by applicant]