IP Library Patent Application 12031671
Patent Application
App. No. 12/031,671

Efficiently Compiling Co-associating Attributes

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
12/031,671
Abstract

A method, software, database and system are presented in which attribute profiles of query-attribute-positive individuals and query-attribute-negative individuals are compared, and combinations of attributes that occur at a higher frequency in the group of query-attribute-positive individuals are identified and stored to generate a compilation of attribute combinations that co-associate with the query attribute (i.e., an attribute of interest). Several computationally efficient approaches for identifying the attribute combinations are incorporated.

Claims (117)

1 . A method for generating a compilation containing combinations of attributes associated with a query attribute, comprising:

a) receiving a query attribute;

b) accessing a query-attribute-positive set of attribute profiles associated with a group of query-attribute-positive individuals and a query-attribute-negative set of attribute profiles associated with a group of query-attribute-negative individuals; and

c) storing one or more attribute combinations having higher frequencies of occurrence in the query-attribute-positive set of attribute profiles than in the query-attribute-negative set of attribute profiles, to generate a compilation containing attribute combinations associated with the query attribute.

2 . The method of claim 1 , further comprising:

d) storing, in the compilation, corresponding statistical results that are computed based on the frequencies of occurrence to indicate the strength of association of each of the attribute combinations with the query attribute.

3 . The method of claim 1 , wherein each of the attribute profiles comprises pangenetic attributes and non-pangenetic attributes, and wherein each of the attribute combinations contains at least one pangenetic attribute, at least one physical attribute, at least one behavioral attribute, and at least one situational attribute.

4 . The method of claim 1 , wherein step (c) further comprises:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a first set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of candidate attributes that occurs within both attribute profiles of a corresponding pairing of the first attribute profile with another attribute profile from the query-attribute-positive set;

iv) identifying, as a primary attribute combination, the largest attribute combination contained in the first set of attribute combinations;

v) generating a second set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of attributes that occurs within both attribute combinations of a corresponding pairing of the primary attribute combination with another attribute combination from the first set of attribute combinations;

vi) partitioning the query-attribute-positive subset of attribute profiles into at least: 1) a first subset comprising the two attribute profiles corresponding to the primary attribute combination as well as attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are equal to or larger than a predetermined fraction of the size of the primary attribute combination, and 2) a second subset comprising attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are smaller than a predetermined fraction of the size of the primary attribute combination;

vii) storing, as a candidate attribute combination in a set of candidate attribute combinations, the largest combination of attributes that occurs within every one of the attribute profiles of the first subset;

viii) redesignating the second subset as the query-attribute-positive subset of attribute profiles; and

ix) repeating steps (i)-(viii) until step (viii) yields an empty set for the query-attribute-positive subset of attribute profiles.

5 . The method of claim 4 , further comprising:

x) generating subcombinations of candidate attributes, wherein each subcombination comprises candidate attributes selected from one of the candidate attribute combinations, the number of candidate attributes in each subcombination being fewer than that of the candidate attribute combination from which the subcombination is generated;

xi) determining frequencies of occurrence, in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, of each of the candidate attribute combinations and each of the subcombinations;

xii) identifying, based on the frequencies of occurrence, each subcombination that has a lower strength of association with the query attribute than the candidate attribute combination from which it was generated;

xiii) identifying each candidate attribute that differs between each identified subcombination and the candidate attribute combination from which the identified subcombination was generated; and

xiv) storing combinations of the candidate attributes identified in step (xiii), each combination comprising candidate attributes identified from the same candidate attribute combination, to generate a compilation containing attribute combinations associated with the query attribute.

6 . The method of claim 5 , wherein in step (x) the number of attributes in a subcombination is n−1, where n is the number of attributes in the candidate attribute combination from which the subcombination is generated.

7 . The method of claim 1 , wherein step (c) further comprises:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a set containing n-tuple combinations of candidate attributes, where n is a specified positive integer value designating the number of candidate attributes in each n-tuple combination;

iv) storing each of the n-tuple combinations that is statistically associated with the query attribute, based on frequencies of occurrence of the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, to generate a compilation of attribute combinations associated with the query attribute;

v) generating (n+1)-tuple combinations of candidate attributes, wherein each (n+1)-tuple combination is generated by combining one of the stored n-tuple combinations with a candidate attribute that does not occur in that stored n-tuple combination;

vi) storing within the compilation of attribute combinations associated with the query attribute, based on frequencies of occurrence of the (n+1)-tuple combinations and the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, each (n+1)-tuple combination that has a stronger statistical association with the query attribute than the n-tuple combination from which it was generated; and

vii) incrementing the value of n by one and repeating steps (v)-(vi), if at least one (n+1)-tuple combination was stored in step (vi).

8 . The method of claim 7 , wherein if no n-tuple combinations are determined to be statistically associated with the query attribute in step (iv), the value of n is incremented by one and execution of the method is redirected to step (iii).

9 . A program storage device readable by a machine and containing a set of instructions which, when read by the machine, causes execution of a method for generating a compilation containing combinations of attributes associated with a query attribute, comprising:

a) receiving a query attribute;

b) accessing a query-attribute-positive set of attribute profiles associated with a group of query-attribute-positive individuals and a query-attribute-negative set of attribute profiles associated with a group of query-attribute-negative individuals; and

c) storing one or more attribute combinations having higher frequencies of occurrence in the query-attribute-positive set of attribute profiles than in the query-attribute-negative set of attribute profiles, to generate a compilation containing attribute combinations associated with the query attribute.

10 . The program storage device of claim 9 , wherein step (c) further comprises:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a first set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of candidate attributes that occurs within both attribute profiles of a corresponding pairing of the first attribute profile with another attribute profile from the query-attribute-positive set;

iv) identifying, as a primary attribute combination, the largest attribute combination contained in the first set of attribute combinations;

v) generating a second set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of attributes that occurs within both attribute combinations of a corresponding pairing of the primary attribute combination with another attribute combination from the first set of attribute combinations;

vi) partitioning the query-attribute-positive subset of attribute profiles into at least: 1) a first subset comprising the two attribute profiles corresponding to the primary attribute combination as well as attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are equal to or larger than a predetermined fraction of the size of the primary attribute combination, and 2) a second subset comprising attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are smaller than a predetermined fraction of the size of the primary attribute combination;

vii) storing, as a candidate attribute combination in a set of candidate attribute combinations, the largest combination of attributes that occurs within every one of the attribute profiles of the first subset;

viii) redesignating the second subset as the query-attribute-positive subset of attribute profiles; and

ix) repeating steps (i)-(viii) until step (viii) yields an empty set for the query-attribute-positive subset of attribute profiles.

11 . The program storage device of claim 10 , further comprising:

x) generating subcombinations of candidate attributes, wherein each subcombination comprises candidate attributes selected from one of the candidate attribute combinations, the number of candidate attributes in each subcombination being fewer than that of the candidate attribute combination from which the subcombination is generated;

xi) determining frequencies of occurrence, in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, of each of the candidate attribute combinations and each of the subcombinations;

xii) identifying, based on the frequencies of occurrence, each subcombination that has a lower strength of association with the query attribute than the candidate attribute combination from which it was generated;

xiii) identifying each candidate attribute that differs between each identified subcombination and the candidate attribute combination from which the identified subcombination was generated; and

xiv) storing combinations of the candidate attributes identified in step (xiii), each combination comprising candidate attributes identified from the same candidate attribute combination, to generate a compilation containing attribute combinations associated with the query attribute.

12 . The program storage device of claim 9 , wherein step (c) further comprises:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a set containing n-tuple combinations of candidate attributes, where n is a specified positive integer value designating the number of candidate attributes in each n-tuple combination;

iv) storing each of the n-tuple combinations that is statistically associated with the query attribute, based on frequencies of occurrence of the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, to generate a compilation of attribute combinations associated with the query attribute;

v) generating (n+1)-tuple combinations of candidate attributes, wherein each (n+1)-tuple combination is generated by combining one of the stored n-tuple combinations with a candidate attribute that does not occur in that stored n-tuple combination;

vi) storing within the compilation of attribute combinations associated with the query attribute, based on frequencies of occurrence of the (n+1)-tuple combinations and the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, each (n+1)-tuple combination that has a stronger statistical association with the query attribute than the n-tuple combination from which it was generated; and

vii) incrementing the value of n by one and repeating steps (v)-(vi), if at least one (n+1)-tuple combination was stored in step (vi).

13 . A database system for generating a compilation containing combinations of attributes associated with a query attribute, comprising a memory containing a first data structure containing a query-attribute-positive set of attribute profiles associated with a group of query-attribute-positive individuals and a query-attribute-negative set of attribute profiles associated with a group of query-attribute-negative individuals, and a processor for:

a) receiving a query attribute;

b) accessing the first data structure; and

c) storing, in a second data structure in the memory, one or more attribute combinations having higher frequencies of occurrence in the query-attribute-positive set of attribute profiles than in the query-attribute-negative set of attribute profiles, to generate a compilation containing attribute combinations associated with the query attribute.

14 . The database system of claim 13 , wherein part (c) further comprises:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a first set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of candidate attributes that occurs within both attribute profiles of a corresponding pairing of the first attribute profile with another attribute profile from the query-attribute-positive set;

iv) identifying, as a primary attribute combination, the largest attribute combination contained in the first set of attribute combinations;

v) generating a second set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of attributes that occurs within both attribute combinations of a corresponding pairing of the primary attribute combination with another attribute combination from the first set of attribute combinations;

vi) partitioning the query-attribute-positive subset of attribute profiles into at least: 1) a first subset comprising the two attribute profiles corresponding to the primary attribute combination as well as attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are equal to or larger than a predetermined fraction of the size of the primary attribute combination, and 2) a second subset comprising attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are smaller than a predetermined fraction of the size of the primary attribute combination;

vii) storing, as a candidate attribute combination in a set of candidate attribute combinations, the largest combination of attributes that occurs within every one of the attribute profiles of the first subset;

viii) redesignating the second subset as the query-attribute-positive subset of attribute profiles; and

ix) repeating steps (i)-(viii) until step (viii) yields an empty set for the query-attribute-positive subset of attribute profiles.

15 . The database system of claim 14 , further comprising:

x) generating subcombinations of candidate attributes, wherein each subcombination comprises candidate attributes selected from one of the candidate attribute combinations, the number of candidate attributes in each subcombination being fewer than that of the candidate attribute combination from which the subcombination is generated;

xi) determining frequencies of occurrence, in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, of each of the candidate attribute combinations and each of the subcombinations;

xii) identifying, based on the frequencies of occurrence, each subcombination that has a lower strength of association with the query attribute than the candidate attribute combination from which it was generated;

xiii) identifying each candidate attribute that differs between each identified subcombination and the candidate attribute combination from which the identified subcombination was generated; and

xiv) storing combinations of the candidate attributes identified in step (xiii), each combination comprising candidate attributes identified from the same candidate attribute combination, to generate a compilation containing attribute combinations associated with the query attribute.

16 . The method of claim 13 , wherein part (c) further comprises:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a set containing n-tuple combinations of candidate attributes, where n is a specified positive integer value designating the number of candidate attributes in each n-tuple combination;

iv) storing each of the n-tuple combinations that is statistically associated with the query attribute, based on frequencies of occurrence of the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, to generate a compilation of attribute combinations associated with the query attribute;

v) generating (n+1)-tuple combinations of candidate attributes, wherein each (n+1)-tuple combination is generated by combining one of the stored n-tuple combinations with a candidate attribute that does not occur in that stored n-tuple combination;

vi) storing within the compilation of attribute combinations associated with the query attribute, based on frequencies of occurrence of the (n+1)-tuple combinations and the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, each (n+1)-tuple combination that has a stronger statistical association with the query attribute than the n-tuple combination from which it was generated; and

vii) incrementing the value of n by one and repeating steps (v)-(vi), if at least one (n+1)-tuple combination was stored in step (vi).

17 . A computer-based system for generating a compilation containing combinations of attributes associated with a query attribute, comprising:

a) a data communications subsystem for receiving a query attribute;

b) a data accessing subsystem for accessing a query-attribute-positive set of attribute profiles associated with a group of query-attribute-positive individuals and a query-attribute-negative set of attribute profiles associated with a group of query-attribute-negative individuals;

c) a data processing subsystem for storing one or more attribute combinations having higher frequencies of occurrence in the query-attribute-positive set of attribute profiles than in the query-attribute-negative set of attribute profiles, to generate a compilation containing attribute combinations associated with the query attribute.

18 . The computer-based system of claim 17 , wherein the data processing system is also for:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a first set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of candidate attributes that occurs within both attribute profiles of a corresponding pairing of the first attribute profile with another attribute profile from the query-attribute-positive set;

iv) identifying, as a primary attribute combination, the largest attribute combination contained in the first set of attribute combinations;

v) generating a second set of attribute combinations, wherein each of the attribute combinations comprises the largest combination of attributes that occurs within both attribute combinations of a corresponding pairing of the primary attribute combination with another attribute combination from the first set of attribute combinations;

vi) partitioning the query-attribute-positive subset of attribute profiles into at least: 1) a first subset comprising the two attribute profiles corresponding to the primary attribute combination as well as attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are equal to or larger than a predetermined fraction of the size of the primary attribute combination, and 2) a second subset comprising attribute profiles corresponding to attribute combinations in the second set of attribute combinations that are smaller than a predetermined fraction of the size of the primary attribute combination;

vii) storing, as a candidate attribute combination in a set of candidate attribute combinations, the largest combination of attributes that occurs within every one of the attribute profiles of the first subset;

viii) redesignating the second subset as the query-attribute-positive subset of attribute profiles; and

ix) repeating steps (i)-(viii) until step (viii) yields an empty set for the query-attribute-positive subset of attribute profiles.

19 . The computer-based system of claim 18 , further comprising:

x) generating subcombinations of candidate attributes, wherein each subcombination comprises candidate attributes selected from one of the candidate attribute combinations, the number of candidate attributes in each subcombination being fewer than that of the candidate attribute combination from which the subcombination is generated;

xi) determining frequencies of occurrence, in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, of each of the candidate attribute combinations and each of the subcombinations;

xii) identifying, based on the frequencies of occurrence, each subcombination that has a lower strength of association with the query attribute than the candidate attribute combination from which it was generated;

xiii) identifying each candidate attribute that differs between each identified subcombination and the candidate attribute combination from which the identified subcombination was generated; and

xiv) storing combinations of the candidate attributes identified in step (xiii), each combination comprising candidate attributes identified from the same candidate attribute combination, to generate a compilation containing attribute combinations associated with the query attribute.

20 . The computer-based system of claim 17 , wherein the data processing system is also for:

i) selecting a first attribute profile from the query-attribute-positive set of attribute profiles;

ii) identifying, as candidate attributes, attributes from the first attribute profile that do not occur in a portion of the attribute profiles of the query-attribute-negative set;

iii) generating a set containing n-tuple combinations of candidate attributes, where n is a specified positive integer value designating the number of candidate attributes in each n-tuple combination;

iv) storing each of the n-tuple combinations that is statistically associated with the query attribute, based on frequencies of occurrence of the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, to generate a compilation of attribute combinations associated with the query attribute;

v) generating (n+1)-tuple combinations of candidate attributes, wherein each (n+1)-tuple combination is generated by combining one of the stored n-tuple combinations with a candidate attribute that does not occur in that stored n-tuple combination;

vi) storing within the compilation of attribute combinations associated with the query attribute, based on frequencies of occurrence of the (n+1)-tuple combinations and the n-tuple combinations in the query-attribute-positive set of attribute profiles and in the query-attribute-negative set of attribute profiles, each (n+1)-tuple combination that has a stronger statistical association with the query attribute than the n-tuple combination from which it was generated; and

vii) incrementing the value of n by one and repeating steps (v)-(vi), if at least one (n+1)-tuple combination was stored in step (vi).

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2022
From: EXPANSE BIOINFORMATICS, INC.
To: 23ANDME, INC.
Reel/Frame 058650/0006 →
CHANGE OF NAME Recorded Sep 23, 2013
From: EXPANSE NETWORKS, INC.
To: EXPANSE BIOINFORMATICS, INC.
Reel/Frame 031289/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2008
From: KENEDY, ANDREW A.
To: EXPANSE NETWORKS, INC.
Reel/Frame 020724/0084 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2008
From: ELDERING, CHARLES A.
To: EXPANSE NETWORKS, INC.
Reel/Frame 020724/0086 →