IP Library › Granted Patent US 12,731,658
Granted Patent B2
US 12,731,658 · App. 15/777,091 · Granted Sep 8, 2026

Methods for detecting copy-number variations in next-generation sequencing

Inventors: Dmitri Ivanov (Pully, CH); Zhenyu Xu (Nyon, CH)
Assignee: Sophia Genetics, S.A.
G16B20/10G16B20/00G16B30/00G16B30/10G16B20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,658
App. No.
15/777,091
Granted
Sep 8, 2026
Kind
B2
Abstract

Copy Number Variants (CNV) detection methods described herein may efficiently integrate CNV detection into the workflow for a next generation sequencer (NGS) data processing, in parallel with SNP and INDEL variant calling. CNV detection methods as described herein may be performed by analyzing the coverage pattern across a suitable set of genomic regions or amplicons and across a batch of samples from different patients. The proposed methods do not require the use of specifically chosen reference samples as inputs to the workflow, but rather automatically select a set of reference samples from the same batch, for each sample being tested. The CNV detection methods may reliably detect CNVs in a set of samples without prior assumptions about the CNV status of any of those samples. Embodiments described herein may also apply the CNV detection scheme iteratively to further improve the detection performance, especially in the case of more frequent CNV occurrence. Since the knowledge on the CNVs in reference samples may improve their comparison with the sample being tested, the proposed methods may further comprise the step of iteratively feeding back the information about the CNVs found in the samples from any detection step into the next iteration step. The proposed methods may also further use additional information available from the NGS workflow about the samples, such as information on SNP fractions, as input to the NGS CNV detection.

Claims (35)

1 . A method for detecting copy-number values (CNV), wherein the detection of CNVs is integrated into a single, targeted high-throughput sequencing experiment, the method comprising the steps of:

(A) enriching a pool of DNA samples with a target enrichment technology, each enriched DNA sample being associated with a library of pooled fragments from a set of amplicons/regions;

(B) sequencing each amplicon/region with a high-throughput sequencer to generate raw sequencing data; and

(C) analyzing the raw sequencing data with a genomic data analyzer to determine the CNVs for each sample and each amplicon/region, comprising:

(i) cleaning the sequencing data to remove low-quality bases and adapter sequences, and aligning the cleaned reads to a reference genome;

(ii) generating from the aligned cleaned reads, with a data processing unit, a coverage count for each sample and for each amplicon/region; and

(iii) repeating, with a data processing unit, over a plurality of iterations, the following steps to estimate the copy-number values for each sample and for each amplicon/region:

(a) normalizing, with the data processing unit, the coverage count associated with each sample based on a prior estimate of the copy number values for each sample,

wherein, if a first iteration of a plurality of iterations, the prior estimate of the copy number values is an initial estimate of the copy-number values, and

wherein, if not the first iteration of a plurality of iterations, the prior estimate of the copy number values is the copy-number values calculated in the course of a previous iteration;

(b) selecting, automatically, with the data processing unit, for each sample, a set of reference samples as the samples with the closest normalized coverage count to the normalized coverage count of said sample, the number of reference samples in each subset of reference samples being a function of the total number of samples,

wherein the selecting of the reference samples does not require the use of specifically chosen, dedicated control samples, and

wherein the closest coverage pattern is selected by calculating for each sample a distance from a current sample and sorting said samples in order of increasing distance, and choosing the set of reference samples from the top of said order having the smallest distances;

(c) normalizing the coverage count associated with each amplicon/region based on the normalized coverage count of the current sample and the normalized coverage count of said set of reference samples, and

(d) for each sample, estimating, using a Hidden Markov Model (HMM), the copy-number values in said sample as a function of at least the coverage counts in said sample and of at least the coverage counts in the selected set of reference samples for said sample and utilizing the estimate of the copy-number values calculated over previous iterations; and

(e) stopping the iteration and outputting the inferred copy-number values if the estimates of the copy-number values converge over iterations, if the estimates of the copy-number values reaches a cycle over multiple iterations, or if the number of iterations reaches a pre-defined limit.

2 . The method of claim 1 , wherein the number of reference samples NR in each set of reference samples is given by NR=[0.25*N]+2, where N is the total number of samples.

3 . The method of claim 1 , wherein the estimate of the copy-number values is calculated using information on the single-nucleotide polymorphisms (SNP) fractions and coverage, and wherein a percentage of SNP fractions is indicative of a duplication.

4 . The method of claim 1 , further comprising: applying a principal-component filter to the coverage count generated for each sample and for each amplicon/region.

5 . The method of claim 1 , wherein normalizing the coverage count associated with each sample in one iteration differs from normalizing the coverage count associated with each sample in subsequent iterations.

6 . The method of claim 1 , wherein selecting the reference samples in one iteration differs from selecting the reference samples in subsequent iterations.

7 . The method of claim 5 , wherein the number of reference samples in one iteration differs from the number of reference samples in subsequent iterations.

8 . The method of claim 7 , wherein in one iteration the number of reference samples in a set equals the total number of samples N, and wherein in subsequent iterations the number of reference samples in a set is different from the total number of samples N.

9 . The method of claim 1 , wherein the step of estimating, using a Hidden Markov Model (HMM), the copy number values in the said samples as a function of at least the coverage counts in the selected set of reference samples for said sample and utilizing the estimate of the copy-number values calculated over the plurality of iterations, further comprises the step of:

estimating, via the HMM, a likelihood for each of the copy-number values and a confidence level for each amplicon/region.

10 . The method of claim 9 , wherein the step of estimating, via the HMM, a likelihood for each of the copy-number values and a confidence level for each amplicon/region, further comprises the steps of:

determining the likelihood for each copy-number value as a log-likelihood, the log-likelihood being defined as

L a (r)=min(( c a,s /( r C a )−1) 2 /(2 δ a,s 2 ), L max )

wherein C a,s , C a , δ a,s are, respectively, a coverage level, a reference normalized coverage level, and a noise level for the current sample s and amplicon/region a,

determining an HMM score defined as

S HMM ({ r a })=Σ a ( L a ( r a )+ P nb ( r a )+ P sw ( r a ,r a+1 )

wherein the HMM score, S HMM ({r a }), is a function of a set of assumed copy numbers r a for every amplicon/region in the current sample, L a (r a ) are the log-likelihoods calculated at the previous step, and p nb (r a ) and p sw (r a r a+1 ) are additional penalties associated with a non- normal copy number and with a transition between different copy numbers among neighboring amplicons/regions, denoted as α and α+1, and

finding a set of CNV states {r a }, via a forward-backward algorithm, which minimize the HMM score and the set of confidence values, each confidence value at a position a defined as the minimal possible increase of the HMM score with the state r a differing from its optimal value.

11 . The method of claim 10 , further comprising the step of excluding each of the copy-number values with the confidence level for each amplicon/region below a threshold.

12 . The method of claim 1 , further comprising the step of excluding each of the copy-number values with a confidence level for each amplicon/region below a threshold.

Assignments (2)
SECURITY AGREEMENT Recorded May 3, 2024
From: SOPHIA GENETICS SA
To: PERCEPTIVE CREDIT HOLDINGS IV, LP
Reel/Frame 067307/0266 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2018
From: IVANOV, DMITRI; XU, ZHENYU
To: SOPHIA GENETICS S.A.
Reel/Frame 045894/0159 →
Continuity (2)
Provisional Application 62256748 · Nov 18, 2015
Related Publication 20180330046A1 · Nov 15, 2018
References Cited (35)
US 20140296094A1 · Domanus · 2014 [cited by applicant]
US 20160017412A1 · Srinivasan · 2016 [cited by examiner]
US 20160275240A1 · Huelga · 2016 [cited by examiner]
US 20160300013A1 · Ashutosh · 2016 [cited by examiner]
US 20160340722A1 · Platt · 2016 [cited by applicant]
US 20180330046A1 · Ivanov · 2018 [cited by examiner]
US 20220101944A1 · Ivanov et al. · 2022 [cited by applicant]
US 20220130488A1 · Ivanov et al. · 2022 [cited by applicant]
CN 104221022A · 2014 [cited by applicant]
CN 104830993A · 2015 [cited by applicant]
WO 2014083147 · 2014 [cited by applicant]
WO 2014151511 · 2014 [cited by applicant]
WO 2015112619 · 2016 [cited by applicant]
Backenroth, d. et al. (2014) CANOES: detecting rare copy number variants from whole exome sequencing data. Nucleic Acids Research 42:12 e97, 9 pages. (Year: 2014). [cited by examiner]
Amarasinghe et al. CoNVEX: copy number variation estimation in exome sequencing data using HMM. BMC bioinformatics 2013 14 (suppl 2) S2, 9 pages. (Year: 2013). [cited by examiner]
Jiang et al. CODEX: a normalization and copy number variation detection method for whole exome sequencing. Nucleic Acid Research (2015) 43:6 e39 12 pages. (Year: 2015). [cited by examiner]
Plagnol, V. (2012) A robust model for read count data in exome sequencing experiments and implications for copy number variant calling. Bioinformatics vol. 28, No. 21, p. 2747-2754 (and supplemental information). (Year:… [cited by examiner]
Wang, Weibo et al. (Apr. 16, 2015) Allele-specific copy-number discovery from whole-genome and whole-exome sequencing. Nucleic Acids Research, vol. 43, No. 14, e90, 18 pages, plus some supplemental information. (Year: 2… [cited by examiner]
Eddy (2004) “What is a hidden Markov model?” Nature Biotechnology, vol. 22, No. 10 p. 1315-1316. (Year: 2004). [cited by examiner]
Kadalayil et al (Aug. 2014) Exome sequence read depth methods for identifying copy number changes. Briefings in Bioinformatics, vol. 16, No. 3, p. 380-392 (Year: 2014). [cited by examiner]
Illumina (2015) Sequencing power for every scale. 16 pages. (Year: 2015). [cited by examiner]
Illumina (2015) MiSeq® system specification sheet. 4 pages. (Year: 2015). [cited by examiner]
Illumina (2015) Nextera® Rapid Capture Enrichment Protocol Guide. 18 pages. (Year: 2015). [cited by examiner]
Illumina (2016) Nextera® Rapid Capture Enrichment Reference Guide, 60 pages. (Year: 2016). [cited by examiner]
Illumina (2012-2014) Nextera® Rapid Capture Exomes data sheet, 4 pages. (Year: 2014). [cited by examiner]
PCT Search Report and IPER in PCT/EP2016/078113. [cited by applicant]
Agilent Technologies (2012) SureSelect RNA Target Enrichment for Illumina Paired-End Sequencing Protocol, version 2.2.1, Feb. 2012. 76 pages. (Year: 2012). [cited by applicant]
Chen, C. et al. (2014) Software for pre-processing Illumina next-generation short read sequences. Source Code for Biology and Medicine, vol. 9:8, 11 pages. (Year: 2014). [cited by applicant]
Final Office Action of the United States Patent and Trademark Office in related U.S. Appl. No. 17/505,934, dated Oct. 8, 2025, 40 pages. [cited by applicant]
Final Office Action of the United States Patent and Trademark Office in related U.S. Appl. No. 17/505,943, dated Oct. 14, 2025, 26 pages. [cited by applicant]
Fromer, M. (Apr. 24, 2014) Using XHMM software to detect copy number variation in whole exome sequencing data. Current Protocols in Human Genetics, vol. 81: 7.23.1-7.23.21. (Year: 2014). [cited by applicant]
Illumina (2016) Nextera® Rapid Capture Enrichment Reference Guide, 60 pages. Revision history illustrates capabilities added between 2013 and 2016. The disclosure includes quantifying enriched targets in the Appendix. (… [cited by applicant]
Marioni J. C, et al. (2007) Breaking the waves: improved detection of copy number variation from microarray based comparative genomic hybridization. Genome Biology, vol. 8: R228. (Year: 2007). [cited by applicant]
TAO Dan, “Study on the Genetic Variation Profile of Lung Squamous Cell Carcinoma in the Chinese Population”, China Doctoral Dissertations Full-text Database (CDFD), Medical and Health Sciences Collection, vol. 2015, Iss… [cited by applicant]
Third Office Action of the China Patent Office in related Chinese Appl. No. 201680067423.2 , dated Apr. 30, 2026, 12 pages. [cited by applicant]