IP Library Granted Patent US 7,516,050
Granted Patent B2
US 7,516,050 · App. 11/077,373 · Granted Apr 7, 2009

Defining the semantics of data through observation

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,516,050
App. No.
11/077,373
Granted
Apr 7, 2009
Kind
B2
Abstract

Techniques for estimating the structure and meaning of data using probability are described. The techniques include retrieving data as data strings from a data source, producing a dataset from the retrieved data strings and building a statistical model of parent-child relationships from data strings in the dataset. Building the statistical model includes determining incidence values for the data strings in the dataset and concatenating the incident values with the data strings to provide child variables. The techniques include analyzing the child variables and the parent variables to produce statistical relationships between the child variables and a parent variable, determining probabilities values based on the determined parent child relationships and building an ontological representation of the data based on subsequent conditional probabilities values.

Claims (59)

1. A method executed in a computer system for estimating a probability associated with a parent variable that an event corresponding to the parent variable will occur, the method comprises:

retrieving data as data strings from a data source;

producing a dataset from the retrieved data strings;

building a statistical model of parent-child relationships from data strings in the dataset by:

determining incidence values for the data strings in the dataset; and

concatenating the incident values with the data strings to provide child variables;

analyzing the child variables and the parent variable to produce statistical relationships between the child variables and the parent variable;

determining probabilities values for the event based on the determined parent child relationships; and

building an ontological representation of the data based on subsequent conditional probabilities values.

2. The method of claim 1 wherein determining probabilities values uses conditional probabilities.

3. The method of claim 1 wherein determining probabilities values uses basic probabilities.

4. The method of claim 1 wherein the parent variable represents an outcome and the child variables represent prior knowledge relevant to the probability of the outcome.

5. The method of claim 4 wherein prior knowledge data that is not in the parent variable.

6. The method of claim 1 wherein analyzing the child variables and the parent variables to produce statistical relationships uses a Bayesian probability algorithm engine.

7. The method of claim 1 wherein multiple routines determine conditional probability by measuring condition probability of each child variable based on the relevance of each child variable to the parent variable.

8. The method of claim 7 further comprising

aggregating the conditional probabilities; and

comparing the aggregated conditional probabilities to parent.

9. The method of claim 1 further comprising

performing a value of information analysis to determine which child variable is more relevant to estimating a probability associated with a parent variable than other child variables.

10. The method of claim 1 wherein building the ontological representation is used to determine a structure of child variables as those child variables relate to the parent variable.

11. The method of claim 1 further comprising:

predicting a value for the parent variable based on the ontological representation.

12. The method of claim 1 wherein the text strings represent any alphanumeric text data.

13. The method of claim 1 further comprising:

filtering noise from the data retrieved from the data source to provide the data strings.

14. The method of claim 1 further comprising:

filtering context-specific noise from data in the data set.

15. A computer program product residing on a computer readable medium for estimating a probability associated with a parent variable that an event corresponding to the parent variable will occur, the computer program product comprising instructions for causing a computer to:

retrieve data as data strings from a data source;

produce a dataset from the retrieved data strings;

build a statistical model of parent-child relationships from data strings in the dataset by:

determine incidence values for the data strings in the dataset; and

concatenate the incident values with the data strings to provide child variables;

analyze the child variables and the parent variable to produce statistical relationships between the child variables and the parent variable;

determine probabilities values for the event based on the determined parent child relationships; and

build an ontological representation of the data based on subsequent conditional probabilities values.

16. The computer program product of claim 15 wherein the parent variable represents an outcome and the child variables represent prior knowledge relevant to the probability of the outcome.

17. The computer program product claim 15 wherein instructions to filter the child variables and the parent variables to produce statistical relationships uses a Bayesian probability algorithm.

18. The computer program product of claim 15 wherein multiple routines determine conditional probability by instructions to measure condition probability of each child variable based on the relevance of each child variable to the parent variable.

19. The computer program product of claim 15 further comprising instructions to:

aggregate the conditional probabilities; and

compare the aggregated conditional probabilities to parent.

20. The computer program product of claim 15 further comprising instructions to:

perform a value of information analysis to determine which child variable is more relevant to estimating a probability associated with a parent variable than other child variables.

21. The computer program product of claim 15 further comprising instructions to predict a value for the parent variable based on the ontological representation.

22. The computer program product of claim 15 further comprising instructions to filter noise from the data retrieved from the data source to provide the data strings.

23. An apparatus comprising:

a processor; and

a computer readable medium storing a computer program product for estimating a probability associated with a parent variable that an event corresponding to the parent variable will occur, the computer program product comprising instructions for causing the processor to:

retrieve data as data strings from a data source;

produce a dataset from the retrieved data strings;

build a statistical model of parent-child relationships from data strings in the dataset by:

determine incidence values for the data strings in the dataset; and

concatenate the incident values with the data strings to provide child variables;

analyze the child variables and the parent variables to produce statistical relationships between the child variables and the parent variable;

determine probabilities values for the event based on the determined parent child relationships; and

build an ontological representation of the data based on subsequent conditional probabilities values.

24. The apparatus of claim 23 wherein the computer program product further comprises instructions to predict a value for the parent variable based on the ontological representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2025
From: POULIN, CHRISTIAN D.
To: POULIN HOLDINGS, LLC
Reel/Frame 072149/0251 →
Continuity (1)
Related Publication 20060206293A1 · Sep 14, 2006