IP Library › Granted Patent US 10,579,930
Granted Patent B2
US 10,579,930 · App. 15/217,820 · Granted Mar 3, 2020

Discovery and scoring of causal assertions in scientific publications

Inventors: Adam Grossman (Los Angeles, CA); Lauren Caston (Los Angeles, CA); Ryan Irvine (Los Angeles, CA); David Loughran (Los Angeles, CA); Robert Thomas Reville (Los Angeles, CA)
Assignee: Praedicat, Inc.
G06N5/041G06F16/93G06F17/11G06N5/04G06N7/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,579,930
App. No.
15/217,820
Granted
Mar 3, 2020
Kind
B2
Abstract

Examples of the disclosure are directed toward generating a causation score with respect to an agent and an outcome, and projecting a future causation score distribution. For example, a causation score may be determined with respect to a hypothesis that a given agent causes a given outcome, and the score may indicate the acceptance of that hypothesis in the scientific community, as described by scientific literature. A future causation score distribution, then, may indicate a probability distribution over possible future causation scores, thereby predicting the scientific acceptance of the hypothesis at some specific date in the future. A future causation score distribution can be projected by first generating one or more future publication datasets, and then determining causation scores for each of the one or more future publication datasets.

Claims (28)

1. A computer-implemented method of computing a future causation score for a hypothesis that an agent causes an outcome, the method comprising:

determining a plurality of distributions based on a publication dataset, wherein each distribution is a subset of the publication dataset and at least one of the plurality of distributions is limited to publications that are relevant to the agent and at least one of the plurality of distributions includes publications that are not relevant to the agent;

generating a plurality of future causation scores via a Monte Carlo simulation, including, for each iteration of the Monte Carlo simulation:

generating a future publication dataset based on multiple random samples from a weighted mixture of the plurality of distributions, and

after the future publication dataset is generated for a respective iteration of the Monte Carlo simulation, computing a future causation score based on articles from the future publication dataset that are relevant to the hypothesis that the agent causes the outcome and not based on articles from the future publication dataset that are not relevant to the hypothesis that the agent causes the outcome; and

generating a visual representation of the plurality of future causation scores for display on a display device.

2. The method of claim 1 , further comprising:

generating a plurality of additional future publication datasets based on the weighted mixture of the plurality of distributions; and

computing a causation score distribution based on the plurality of additional future publication datasets.

3. The method of claim 1 , wherein the publication dataset consists of publication metadata.

4. The method of claim 1 , further comprising:

determining an additional plurality of distributions based on the future publication dataset, wherein each of the additional plurality of distributions is a subset of the future publication dataset; and

generating an additional future publication dataset based on a weighted mixture of the additional plurality of distributions.

5. The method of claim 1 , wherein generating a future publication dataset includes sampling the weighted mixture of the plurality of distributions a number of times, wherein the number of times is a sampling count determined based on a random walk from a publication rate of the publication dataset.

6. A non-transitory computer readable storage medium storing instructions executable to perform a method of computing a future causation score for a hypothesis that an agent causes an outcome, the method comprising:

determining a plurality of distributions based on a publication dataset, wherein each distribution is a subset of the publication dataset and at least one of the plurality of distributions is limited to publications that are relevant to the agent and at least one of the plurality of distributions includes publications that are not relevant to the agent;

generating a plurality of future causation scores via a Monte Carlo simulation, including, for each iteration of the Monte Carlo simulation:

generating a future publication dataset based on multiple random samples from a weighted mixture of the plurality of distributions, and

after the future publication dataset is generated for a respective iteration of the Monte Carlo simulation, computing a future causation score based on articles from the future publication dataset that are relevant to the hypothesis that the agent causes the outcome and not based on articles from the future publication dataset that are not relevant to the hypothesis that the agent causes the outcome; and

generating a visual representation of the plurality of future causation scores for display on a display device.

7. The non-transitory computer readable storage medium of claim 6 , the method further comprising:

generating a plurality of additional future publication datasets based on the weighted mixture of the plurality of distributions; and

computing a causation score distribution based on the plurality of additional future publication datasets.

8. The non-transitory computer readable storage medium of claim 6 , wherein the publication dataset consists of publication metadata.

9. The non-transitory computer readable storage medium of claim 6 , the method further comprising:

determining an additional plurality of distributions based on the future publication dataset, wherein each of the additional plurality of distributions is a subset of the future publication dataset; and

generating an additional future publication dataset based on a weighted mixture of the additional plurality of distributions.

10. The non-transitory computer readable storage medium of claim 6 , wherein generating a future publication dataset includes sampling the weighted mixture of the plurality of distributions a number of times, wherein the number of times is a sampling count determined based on a random walk from a publication rate of the publication dataset.

Continuity (2)
Continuation 14135436 · Dec 19, 2013
Related Publication 20160328652A1 · Nov 10, 2016
Cited By (1)
US 12,619,829