IP Library › Granted Patent US 12,633,377
Granted Patent B2
US 12,633,377 · App. 17/251,293 · Granted May 19, 2026

Methods for detecting variants in next-generation sequencing genomic data

Inventors: Zhenyu Xu (Saint-Sulpice, CH); Lin Song (Saint-Sulpice, CH)
Assignee: Sophia Genetics S.A.
G16B30/10G16B5/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,377
App. No.
17/251,293
Granted
May 19, 2026
Kind
B2
Abstract

A genomic data analyzer may be configured to detect and characterize, with a variant calling module, genomic variants from next generation sequencing reads out of a pool of enriched genomic patient samples without suffering from next generation sequencing workflow biases such as those introduced by sequencing errors in particular in repeat patterns regions of the human genome such as homopolymers or heteropolymers. The variant calling module may estimate the probability distribution of the length of the repeat pattern for each patient sample and cross-analyze it against other samples in a single experimental pool to identify best-fit variant models for each pair of samples. The variant calling module may further group samples according to their matching best-fit variant models and identify which group of patient samples carries the wild type reference without the need for control data in the pool. The variant calling module may subsequently characterize the homozygous or heterozygous repeat patterns variants for each patient sample with improved specificity and accuracy even in the presence of next generation sequencing biases.

Claims (778)

1 . A method for detecting and reporting, with a processor, a variant as a repeat of at least two nucleotide patterns in a genomic sequence of a patient sample, the method comprising:

(a) identifying a reference repeat pattern P ref =N*l as the repeat of l (l>=2) genomic patterns N in a genomic region of a human genome reference sequence;

(b) obtaining, with a next generation sequencer, n patient sets of next generation sequencing data reads covering the reference repeat pattern genomic region S={S 1 , S 2 , . . . , S i , . . . , S n } from a pool of n enriched genomic patient samples, each set S i being associated with a patient sample, the number n of enriched genomic patient samples being at least 4, wherein the pool of n enriched genomic patient samples is obtained by aligning and re-aligning the next generation sequencing data reads, via a realignment algorithm executed by a genomic data analyzer, to provide refined data variant calling data comprising the enriched genomic patient samples, wherein aligning and re-aligning the next generation sequencing reads comprises the steps of:

receiving a next generation sequencing analysis request;

identifying, via the genomic data analyzer, a first set of characteristics associated with the next generation sequencing analysis request;

executing, via the genomic data analyzer, a first data alignment; and

identifying, from the first data alignment, via the genomic data analyzer, a second set of characteristics including a data alignment pattern identifier,

wherein the genomic data analyzer automatically identifies one or more reads with soft clipping regions from the first data alignment, and

wherein the genomic data analyzer configures a data alignment module to execute a data re-alignment on the one or more reads with soft clipping regions;

(c) for each patient sample i in the set S of patient samples, measuring the distribution P i of the length of the repeat pattern in the set of next generation sequencing reads S i ,

(d) for a possible pair of patient samples i and j, j>i, estimating a best-fit model

[

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

j

2

]

of the two allele variants for sample i relative to sample j, with a confidence level L ij ,

(e) for each possible triplet of patient samples i, j>i, k>j, comparing their respective best-fit models

[

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

j

2

]

,

[

V

j

⁢

k

1

⁢

❘

"\[LeftBracketingBar]"

V

j

⁢

k

2

]

,

[

V

i

⁢

k

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

k

2

]

,

grouping matching best-fit models into groups of best-fit variant models with an increased confidence level, and iterating the comparison until stable groups of best-fit variant models are formed;

(f) identifying the most likely group carrying the wild type variant from the stable groups of best-fit variant models;

(g) for each patient sample in the most likely group carrying the wild type variant, reporting the sample variant as the wild type reference repeat pattern P ref =N*l;

(h) for each patient sample out of the most likely group carrying the wild type variant, unbiasing the best-fit variant model of the group comprising this sample as a function of the best-fit variant model for the identified wild type group and reporting the sample variant as the unbiased best-fit model variant to determine whether a patient corresponding with each patient sample out of the most likely group carrying the wild type variant has either a hereditary disease or a somatic disease.

2 . The method of claim 1 , wherein estimating a best-fit model

[

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

j

2

]

of the two allele variants for sample i relative to sample j comprises:

(d1) estimating for sample j, under the assumption that sample i carries the wild type human genome reference homopolymer pattern P ref N*l for each allele, a best-fit model

[

V

j

|

i

1

⁢

❘

"\[LeftBracketingBar]"

V

j

|

i

2

]

of the two allele variants for sample j, with a confidence level L j|i , as well as the smallest distance D j|i between the measured distribution P j for sample j and a predicted unimodal or bimodal distribution for the best-fit variant model

[

V

j

|

i

1

⁢

❘

"\[LeftBracketingBar]"

V

j

|

i

2

]

;

(d2) estimating for sample i, under the assumption that sample j carries the wild type human genome reference homopolymer pattern P ref =N*l for each allele, a best-fit model

[

V

i

|

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

|

j

2

]

of the two allele variants for sample i, with a confidence level L i|j , as well as the smallest distance D i|j between the measured distribution P i for sample i and the predicted unimodal or bimodal distribution for the best fit variant model

[

V

i

|

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

|

j

2

]

;

(d3) if D i|j ≥D j|i , selecting for the pair of samples (i,j) the best-fit model

[

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

j

2

]

=

[

-

V

j

|

i

1

⁢

❘

"\[LeftBracketingBar]"

-

V

j

|

i

2

]

as the best-fit variant model of the two allele variants and the confidence level L ij =L j|i as the confidence level value for this best-fit match with sample i as the reference sample of pair (i,j);

(d4) else if D i|j <D j|i selecting for the pair of samples (i,j) the model

[

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

j

2

]

=

[

V

i

|

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

|

j

2

]

as the best-fit variant model of the two allele variants and the confidence level L ij =L i|j as the confidence level value for this best-fit match with sample j as the reference sample of pair (i,j).

3 . The method of claim 2 , further comprising estimating a secondary best-fit model

[

V

j

|

i

1

⁢

′

⁢

❘

"\[LeftBracketingBar]"

V

j

|

i

2

⁢

′

]

of the two allele variants for sample j, under the assumption that sample i carries the wild type human genome reference homopolymer pattern P ref =N*l for each allele, as well as the second smallest distance D j|i ′ between the measured distribution P j for sample j and the predicted unimodal or bimodal distribution for the secondary best-fit variant model

[

V

j

|

i

1

⁢

′

⁢

❘

"\[LeftBracketingBar]"

V

j

|

i

2

⁢

′

]

,

estimating a secondary best-fit model

[

V

i

|

j

1

⁢

′

⁢

❘

"\[LeftBracketingBar]"

V

i

|

j

2

⁢

′

]

of the two allele variants for sample i, under the assumption that sample j carries the wild type human genome reference homopolymer pattern P ref N*l for each allele, as well as the second smallest distance D i|j ′ between the measured distribution P i for sample i and the predicted unimodal or bimodal distribution for the secondary best-fit variant model

[

V

j

|

i

1

⁢

′

⁢

❘

"\[LeftBracketingBar]"

V

j

|

i

2

⁢

′

]

,

and calculating the confidence level L ij of the estimation

[

V

i

⁢

j

1

|

V

i

⁢

j

2

]

as:

L

i

⁢

j

=

{

1

-

D

i

|

j

/

D

i

|

j

′

,

if

⁢

D

i

|

j

>

D

j

|

i

1

-

D

j

|

i

/

D

j

|

i

′

,

if

⁢

D

i

|

j

≤

D

j

|

i

.

4 . The method of claim 1 , further comprising grouping together into q different groups of samples (1≤q≤n−1) based on

[

V

ij

.

k

1

|

V

ij

.

k

2

]

values such that within each group of samples G r (1≤r≤q), all results agree with each other, and calculating the overall confidence level for this group as

L

ij

.

Gr

=

1

-

∏

k

∈

G

r

⁢

(

1

-

L

ij

.

k

)

.

5 . The method of claim 4 , wherein the best-fit models

[

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

j

2

]

corresponding to different types of heterozygous mutations are excluded from the grouping of matching best-fit models.

6 . The method of claim 4 , comprising selecting the group G h with the highest confidence level L ij.Gh , setting the value

[

V

ij

.

Gh

1

⁢

❘

"\[LeftBracketingBar]"

V

ij

.

Gh

2

]

of all samples in this group and calculating the new confidence level for pair i,j as

L

ij

.

new

=

max

(

0

,

1

-

(

1

-

L

ij

.

Gh

)

*

∏

1

≤

r

≤

q

,

r

≠

h

⁢

(

1

-

L

ij

.

Gr

)

-

1

.

7 . The method of claim 4 , further comprising iterating grouping together groups of samples until the results are stable.

8 . The method of claim 1 , comprising for each possible triplet of patient samples i, j>i, k>j, comparing their respective best-fit models

[

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

j

2

]

,

[

V

j

⁢

k

1

⁢

❘

"\[LeftBracketingBar]"

V

j

⁢

k

2

]

,

[

V

i

⁢

k

1

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

k

2

]

and if all three best-fit models for the triplet of patient samples match each other, increasing their confidence levels L ij , L jk , L ik ; else the three best-fit models do not match each other, replacing the best-fit model with the lowest confidence level out of the subset by a best-fit model calculated from the other two samples of the subset, and decreasing the confidence levels L ij , L jk , L ik of all the best-fit models for the triplet of patient samples, and repeating the comparison for all possible triplets until the results are no longer varying.

9 . The method of claim 8 wherein the confidence levels for each pair in a triplet subset i, j, k of matching best-fit models are increased as L ij ′=1−(1−L ij )(1−L jk *L ik ), L jk ′=1−(1−L jk )(1−L ij *L ik ) and L ik ′=1−(1−L ik )(1−L ij *L jk ).

10 . The method of claim 8 wherein the lowest initial confidence level is L ik for the pair j,k within the triplet, the confidence levels for each pair in a triplet subset i, j, k of non-matching best-fit models are decreased as L ij ′=L ij (1−L jk )*L ik , L jk ′=L jk (1−L ij )*L ik and L ik ′=max(0, L ij *L jk −L ik ) and the best-fit model for the pair j,k with the lowest confidence level out of the subset is replaced by

[

V

i

⁢

k

1

′

⁢

❘

"\[LeftBracketingBar]"

V

i

⁢

k

2

′

]

=

[

V

i

⁢

j

1

+

V

i

⁢

j

1

⁢

❘

"\[LeftBracketingBar]"

V

j

⁢

k

2

+

V

j

⁢

k

2

]

.

11 . The method of claim 10 , wherein identifying the subset of one or more samples corresponding to the wild type reference in the pool of patient samples consists of selecting as the wild type the homozygous best-fit variant model group [V G |V G ] with which the largest number of samples i, j, . . . have been associated out of cross-analyzing the pool of samples.

12 . The method of claim 10 , wherein identifying the subset of one or more samples corresponding to the wild type reference in the pool of patient samples comprises verifying for each group G of homozygous best-fit variant model

[

V

G

1

⁢

❘

"\[LeftBracketingBar]"

V

G

2

]

,

if

[

V

G

⁢

′

1

⁢

❘

"\[LeftBracketingBar]"

V

G

⁢

′

2

]

is a possible variant for each other group G′ of best-fit variant model

[

V

G

⁢

′

1

⁢

❘

"\[LeftBracketingBar]"

V

G

⁢

′

2

]

,

under the hypothesis that the homozygous best-fit variant model

[

V

G

1

⁢

❘

"\[LeftBracketingBar]"

V

G

2

]

is the wild type, and if that is not the case, excluding group G as carrying the wild type pattern.

13 . The method of claim 12 , wherein

V

G

⁢

′

1

<

V

G

1

⁢

and

⁢

V

G

⁢

′

2

<

V

G

2

,

comprising verifying that the length of the repeat pattern with the

[

V

G

⁢

′

1

⁢

❘

"\[LeftBracketingBar]"

V

G

⁢

′

2

]

best-fit variant model is long enough to be detected as a plausible deletion variant.

14 . The method of claim 12 , wherein

V

G

⁢

′

1

>

V

G

1

⁢

and

/

or

⁢

V

G

⁢

′

2

>

V

G

2

,

comprising verifying that the length of the repeat pattern with the

[

V

G

⁢

′

1

⁢

❘

"\[LeftBracketingBar]"

V

G

⁢

′

2

]

best-fit variant model is short enough to be detected as a plausible insertion variant.

15 . The method of claim 10 , further comprising estimating an error rate based on the average homopolymer length h and the standard deviation SD for each group of homozygous best-fit variant model

[

V

G

1

⁢

❘

"\[LeftBracketingBar]"

V

G

2

]

,

and if h is within a pre-defined threshold threshold_h to the nearest integer [ h ], that is if abs ( h −[ h ])<threshold_h, and if SD is small enough to be under a predefined threshold threshold_sd, that is if SD<threshold_sd, selecting the homozygous best-fit variant model

[

V

G

1

⁢

❘

"\[LeftBracketingBar]"

V

G

2

]

as the wild type reference with the lowest error rate.

16 . The method of claim 15 , wherein threshold_h is chosen in the range of 0 to 0.1.

17 . The method of claim 15 , wherein threshold_sd is chosen in the range of 0 to 0.1.

18 . The method of claim 1 , wherein aligning and re-aligning the next generation sequencing reads further comprises the steps of:

refining the aligned sequencing data in accordance with at least one characteristic of the first set of characteristics and at least one characteristic of the second set of characteristics;

identifying a third set of characteristics associated with the refined alignment sequencing data;

detecting a first set of genomic variants associated with the refined alignment sequencing data in accordance with at least one characteristic of the first set of characteristics, at least one characteristic of the second set of characteristics, and at least one characteristic of the third set of characteristics;

identifying a fourth set of characteristics associated with the detected first set of genomic variants;

detecting variants associated with the refined alignment data in accordance with at least one characteristic of the first set of characteristics, at least one characteristic of the second set of characteristics, at least one characteristic of the third set of characteristics and at least one characteristic of the fourth set of characteristics, wherein the pool n of enriched genomic patient samples comprises the variants;

wherein the first set of characteristics includes at least one of:

a target enrichment technology identifier,

a sequencing technology identifier, and

a genomic context identifier;

wherein, during the first data alignment, the genomic data analyzer removes one or more assay specific adapters from the next generation sequencing data reads; and

wherein, the data alignment pattern comprises one or more soft clip patterns.

19 . The method of claim 1 , wherein the n patient sets of next generation sequencing data reads are sourced from a plurality of next generation sequencing laboratories.

20 . The method of claim 1 , wherein the next generation sequencer processes n patient sets of next generation sequencing data reads in parallel and generates one or more corresponding sequence reads in a computer-readable file.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2026
From: SONG, LIN; XU, ZHENYU
To: SOPHIA GENETIS S.A.
Reel/Frame 074194/0400 →
SECURITY AGREEMENT Recorded May 3, 2024
From: SOPHIA GENETICS SA
To: PERCEPTIVE CREDIT HOLDINGS IV, LP
Reel/Frame 067307/0266 →
Priority Claims (1)
EP 18177876 · Jun 14, 2018 · regional
Continuity (1)
Related Publication 20210125689A1 · Apr 29, 2021
References Cited (12)
US 20090026082A1 · Rothberg et al. · 2009 [cited by applicant]
US 20140052381A1 · Utiramerur · 2014 [cited by examiner]
WO WO2016009224A1 · 2016 [cited by examiner]
WO 2018104466A1 · 2018 [cited by applicant]
Heard, Nicholas A. “Iterative Reclassification in Agglomerative Clustering.” Journal of Computational and Graphical Statistics, vol. 20, No. 4, 2011, pp. 920-936. (Year: 2011). [cited by examiner]
Min Zhao et al. Computational tools for copy number variation (CNV) detection using next-generation sequencing data: features and perspectives. BMC Bioinformatics 2013, 14(Suppl 11):S1 (Year: 2013). [cited by examiner]
Treangen, Todd & Salzberg, Steven. (2011). Repetitive DNA and next-generation sequencing: Computational challenges and solutions. Nature reviews. Genetics. 13. 36-46. (Year: 2012). [cited by examiner]
Feng, Weixing er al. (2016). Improving alignment accuracy on homopolymer regions for semiconductor-based sequencing technologies. BMC Genomics. 17. (Year: 2016). [cited by examiner]
Duan J, Deng HW, Wang YP. Common copy number variation detection from multiple sequenced samples. IEEE Trans Biomed Eng. Mar. 2014;61(3):928-37. (Year: 2014). [cited by examiner]
Poulet et al., “Improved Efficiency and Reliability of NGS Amplicon Sequencing Data Analysis for Genetic Diagnostic Procedures Using AGSA Software”, Biomed Research International, vol. 2016, XP055620816, ISSN: 2314-6133… [cited by applicant]
International Search Report, mailed Oct. 1, 2019 by the European Patent Office 9EPO), in International Application No. PCT/EP2019/065777. [cited by applicant]
Written Opinion of the International Searching Authority, mailed Oct. 1, 2019 by the European Patent Office (EPO), in International Application No. PCT/EP2019/065777. [cited by applicant]