IP Library Patent Application 17968285
Patent Application
App. No. 17/968,285

OPTIMIZED BURDEN TEST BASED ON NESTED T-TESTS THAT MAXIMIZE SEPARATION BETWEEN CARRIERS AND NON-CARRIERS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/968,285
Abstract

A computer-implemented method of performing an optimized burden test for a particular gene, in which an optimal combination of a maximum allele count and a minimum pathogenicity score threshold that maximize significance of burden testing for rare deleterious variants are determined using a grid search protocol. Each combination of maximum allele count and minimum pathogenicity score threshold is tested with a t-test to obtain effect size and p-value. The combination of allele count and pathogenicity score threshold with the most significant p-value is selected as the optimal parameters for the rare deleterious variant burden test for a particular gene.

Claims (41)

1 . A computer-implemented method of performing an optimized burden test, including:

determining an optimal combination of a maximum allele count and a minimum pathogenicity score threshold that maximizes significance of burden testing effects of rare pathogenic variants in a particular gene on a particular phenotype, including:

grid searching a plurality of allele counts and a plurality of pathogenicity score thresholds, including:

generating a plurality of combinations of allele counts and pathogenicity score thresholds from the plurality of allele counts and the plurality of pathogenicity score thresholds;

identifying a plurality of groups of rare pathogenic variants corresponding to the plurality of combinations of allele counts and pathogenicity score thresholds; and

burden testing the plurality of groups of rare pathogenic variants in dependence upon a carrier status that separates carriers of a particular group of rare pathogenic variants in a cohort of individuals from non-carriers of the particular group of rare pathogenic variants in the cohort of individuals, and determining a plurality of effect sizes and p-values corresponding to the plurality of combinations of allele counts and pathogenicity score thresholds;

selecting, from the plurality of combinations of allele counts and pathogenicity score thresholds, a particular combination of an allele count and a pathogenicity score threshold that has a most significant p-value; and

using the particular combination as the optimal combination.

2 . The computer-implemented method of claim 1 , wherein allele counts in the plurality of allele counts correspond to groups of rare pathogenic variants observed in the particular gene across the cohort of individuals.

3 . The computer-implemented method of claim 2 , wherein pathogenicity score thresholds in the plurality of pathogenicity score thresholds correspond to pathogenicity score quantiles of pathogenicity scores determined for rare pathogenic variants in the groups of rare pathogenic variants.

4 . The computer-implemented method of claim 3 , wherein the pathogenicity scores are generated by a convolutional neural network.

5 . The computer-implemented method of claim 1 , wherein the particular phenotype is a quantitative biomarker phenotype.

6 . The computer-implemented method of claim 5 , wherein the quantitative biomarker phenotype is burden tested using a two-tailed t-test.

7 . The computer-implemented method of claim 6 , wherein the two-tailed t-test is executed a*p times in a nested fashion, where a is a number of allele counts in the plurality of allele counts, and where p is a number of the pathogenicity score thresholds in the plurality of pathogenicity score thresholds.

8 . The computer-implemented method of claim 7 , wherein intermediate statistics are shared between each execution of the two-tailed t-test.

9 . The computer-implemented method of claim 1 , wherein the particular phenotype is a categorical clinical diagnosis phenotype.

10 . The computer-implemented method of claim 9 , wherein the categorical clinical diagnosis phenotype is burden tested using logistic regression.

11 . The computer-implemented method of claim 10 , wherein the logistic regression is executed a*p times in a nested fashion.

12 . The computer-implemented method of claim 11 , wherein intermediate statistics are shared between each execution of the logistic regression.

13 . The computer-implemented method of claim 1 , wherein the grid searching further includes performing a first grid search through the plurality of allele counts.

14 . The computer-implemented method of claim 1 , wherein the grid searching further includes performing a second grid search through the plurality of pathogenicity score thresholds.

15 . The computer-implemented method of claim 1 , further including correcting the most significant p-value using an adaptive permutation false discovery rate.

16 . The computer-implemented method of claim 1 , further including correcting the most significant p-value using a Benjamini-Hochberg false discovery rate.

17 . A system including one or more processors coupled to memory, the memory loaded with computer instructions to perform an optimized burden test, the instructions, when executed on the processors, implement actions comprising:

determining an optimal combination of a maximum allele count and a minimum pathogenicity score threshold that maximizes significance of burden testing effects of rare pathogenic variants in a particular gene on a particular phenotype, including:

grid searching a plurality of allele counts and a plurality of pathogenicity score thresholds, including:

generating a plurality of combinations of allele counts and pathogenicity score thresholds from the plurality of allele counts and the plurality of pathogenicity score thresholds;

identifying a plurality of groups of rare pathogenic variants corresponding to the plurality of combinations of allele counts and pathogenicity score thresholds; and

burden testing the plurality of groups of rare pathogenic variants in dependence upon a carrier status that separates carriers of a particular group of rare pathogenic variants in a cohort of individuals from non-carriers of the particular group of rare pathogenic variants in the cohort of individuals, and determining a plurality of effect sizes and p-values corresponding to the plurality of combinations of allele counts and pathogenicity score thresholds;

selecting, from the plurality of combinations of allele counts and pathogenicity score thresholds, a particular combination of an allele count and a pathogenicity score threshold that has a most significant p-value; and

using the particular combination as the optimal combination.

18 . The system of claim 17 , wherein allele counts in the plurality of allele counts correspond to groups of rare pathogenic variants observed in the particular gene across the cohort of individuals.

19 . The system of claim 18 , wherein pathogenicity score thresholds in the plurality of pathogenicity score thresholds correspond to pathogenicity score quantiles of pathogenicity scores determined for rare pathogenic variants in the groups of rare pathogenic variants.

20 . A non-transitory computer readable storage medium impressed with computer program instructions perform an optimized burden test, the instructions, when executed on a processor, implement a method comprising:

determining an optimal combination of a maximum allele count and a minimum pathogenicity score threshold that maximizes significance of burden testing effects of rare pathogenic variants in a particular gene on a particular phenotype, including:

grid searching a plurality of allele counts and a plurality of pathogenicity score thresholds, including:

generating a plurality of combinations of allele counts and pathogenicity score thresholds from the plurality of allele counts and the plurality of pathogenicity score thresholds;

identifying a plurality of groups of rare pathogenic variants corresponding to the plurality of combinations of allele counts and pathogenicity score thresholds; and

burden testing the plurality of groups of rare pathogenic variants in dependence upon a carrier status that separates carriers of a particular group of rare pathogenic variants in a cohort of individuals from non-carriers of the particular group of rare pathogenic variants in the cohort of individuals, and determining a plurality of effect sizes and p-values corresponding to the plurality of combinations of allele counts and pathogenicity score thresholds;

selecting, from the plurality of combinations of allele counts and pathogenicity score thresholds, a particular combination of an allele count and a pathogenicity score threshold that has a most significant p-value; and

using the particular combination as the optimal combination.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: FIZIEV, PETKO PLAMENOV; MCRAE, JEREMY FRANCIS; FARH, KAI-HOW
To: ILLUMINA, INC.
Reel/Frame 061582/0828 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: FIZIEV, PETKO PLAMENOV; MCRAE, JEREMY FRANCIS; FARH, KAI-HOW
To: ILLUMINA, INC.
Reel/Frame 061583/0252 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: FIZIEV, PETKO PLAMENOV; MCRAE, JEREMY FRANCIS; FARH, KAI-HOW
To: ILLUMINA, INC.
Reel/Frame 061583/0546 →