IP Library Granted Patent US 12,374,462
Granted Patent B2
US 12,374,462 · App. 18/171,360 · Granted Jul 29, 2025

Method for diagnosing cancer and predicting cancer type by using terminal sequence motif frequency and size of cell-free nucleic acid fragment

Inventors: Eun Hae Cho (Gyeonggi-do, KR); Tae-Rim Lee (Gyeonggi-do, KR); Sook Ryun Park (Seoul, KR)
Assignee: GC GENOME CORPORATION
G16H50/20G16B30/10G16B40/00G01N2800/7028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,462
App. No.
18/171,360
Granted
Jul 29, 2025
Kind
B2
Abstract

Disclosed is a method for diagnosing cancer and predicting a cancer type using fragment end motif frequencies and sizes of cell-free nucleic acid, and more preferably, to a method for diagnosing cancer and predicting a cancer type by extracting nucleic acids from a biological sample to obtain sequence information, acquiring fragment end motif frequencies and sizes of nucleic acids based on the aligned reads, converting the fragment end motif frequencies and sizes of nucleic acids into vectorized data, inputting the vectorized data to a trained artificial intelligence model and analyzing a resulting calculated value. The method includes generating vectorized data and analyzing the same using an AI algorithm and thus is useful due to high sensitivity and accuracy thereof even in the case of low read coverage.

Claims (162)

1. A method for measuring fragment end motif frequency and size of nucleic acids, the method comprising:

a) extracting nucleic acids from a biological sample, the extracted nucleic acids comprising cell free DNA (cfDNA);

b) sequencing the nucleic acids to a depth of at least 1 million reads to obtain sequence reads;

c) aligning the sequence reads with a reference genome database, wherein the length of the reads is from 5 to 5,000 bp;

d) acquiring end motif frequencies and sizes of nucleic acid fragments based on the aligned sequence reads, wherein the end motif frequencies are selected from a total of 256 4-mer motifs, and wherein the nucleic acid fragments comprise fragment sizes ranging from 90 to 250 base pairs;

e) generating vectorized data using the end motif frequencies and sizes of nucleic acid fragments, wherein the vectorized data is expressed by a type of an end motif of the nucleic acid fragments and the size of the nucleic acid fragments plotted on a single or multidimensional axis;

f) inputting the generated vectorized data to a trained artificial intelligence model, wherein the artificial intelligence model transforms the vectorized data of step (e) into an output value, wherein the artificial intelligence model has been trained to adjust the output value to about 1 if there is cancer or to adjust the output value to about 0 if there is no cancer; and

g) determining whether or not cancer develops, wherein the determining comprises comparing the output value of step (f) with a cut-off value to determine whether or not cancer develops, wherein if the output value is greater than a cut-off value, it is determined that there is cancer, and if the output value is less than the cut-off value, it is determined that there is no cancer,

wherein the artificial intelligence model in step (f) has been trained to distinguish between vectorized data of a healthy subject and vectorized data of a cancer patient.

2. The method according to claim 1 , wherein step (a) comprises:

a-i. obtaining nucleic acids from the blood, semen, vaginal cells, hair, saliva, urine, oral cells, amniotic fluid containing placental cells or fetal cells, tissue cells, or a mixture thereof;

a-ii. removing proteins, fats, and other residues from the collected nucleic acids using a salting-out method, a column chromatography method, or a bead method to obtain purified nucleic acids;

a-iii. producing a single-end sequencing or paired-end sequencing library for the purified nucleic acids or nucleic acids randomly fragmented by an enzymatic digestion, pulverization, or Hydroshear method;

a-iv. reacting the produced library with a next-generation sequencer; and

a-v. obtaining sequence reads of the nucleic acids in the next-generation sequencer.

3. The method according to claim 1 , wherein the end motif of each nucleic acid fragment in step (d) has a sequence pattern of 2 to 30 bases at both ends of the nucleic acid fragment.

4. The method according to claim 1 , wherein the frequency of end motifs of the nucleic acid fragments in step (d) corresponds to the number of motifs detected in all the nucleic acid fragments.

5. The method according to claim 1 , wherein the size of each nucleic acid fragment in step (d) corresponds to the number of bases from the 5′ end to the 3′ end of the nucleic acid fragment.

6. The method according to claim 1 , wherein the vectorized data in step (d) is expressed by the type of the end motif of the nucleic acid fragment plotted on an X-axis and the size of the nucleic acid fragment plotted on a Y-axis.

7. The method according to claim 6 , wherein the vectorized data further comprises a sum of frequencies for end motifs of nucleic acid fragments and a sum of frequencies for sizes of nucleic acid fragments.

8. The method according to claim 1 , wherein the artificial intelligence model is a convolutional neural network (CNN).

9. The method according to claim 8 , wherein, a loss function for performing binary classification is represented by Equation 1 below and a loss function for performing multi-class classification is represented by Equation 2 below:

Equation

1

:

Binary

classification

loss

(

model

(

x

)

,

y

)

=

-

1

n

[

i

=

1

n

(

y

i

log

(

model

(

x

i

)

)

+

(

1

-

y

y

)

log

(

1

-

model

(

x

t

)

)

)

]

Model (x i )=Artificial intelligence model output in response to i th input

y=Actual label value

n=Number of input data

Equation

2

:

Multi

-

class

classification

loss

(

model

(

x

)

,

y

)

=

-

1

n

i

=

1

n

(

j

=

1

c

(

y

ij

log

(

model

(

x

i

)

)

j

)

Model (x i ) j =j th artificial intelligence model output in response to i th input

y=Actual label value

n=Number of input data

c=Number of classes.

10. The method according to claim 1 , wherein the output value resulting from analysis of the input vectorized data by the artificial intelligence model in step (f) is a deep probability index (DPI).

11. The method according to claim 1 , wherein the cut-off value of step (e) is 0.5 and a determination is made that cancer has developed when the output value is 0.5 or more.

12. The method according to claim 1 , wherein the artificial intelligence model is a deep neural network (DNN).

13. The method according to claim 1 , wherein the artificial intelligence model is a recurrent neural network (RNN).

14. A non-transitory computer-readable storage medium comprising an instruction configured to be executed by a processor for diagnosing cancer through the following steps comprising:

(a) extracting nucleic acids from a biological sample to obtain sequence reads, the extracted nucleic acids comprising cell free DNA (cfDNA);

(b) aligning the sequence reads with a reference genome database;

(c) acquiring end motif frequencies and sizes of nucleic acid fragments based on the aligned sequence reads, wherein the end motif frequencies are selected from a total of 256 4-mer motifs, and wherein the nucleic acid fragments comprise fragment sizes ranging from 90 to 250 base pairs;

(d) generating vectorized data using the motif frequencies and sizes of nucleic acid fragments;

(e) inputting the generated vectorized data into a trained artificial intelligence model, wherein the artificial intelligence model transforms the vectorized data of step (d) into an output value, wherein the artificial intelligence model has been trained to adjust the output value to about 1 if there is cancer or to adjust the output value to about 0 if there is no cancer; and

(f) determining whether or not cancer develops, wherein the determining comprises comparing the output value of step (e) with a cut-off value to determine whether or not cancer develops, wherein if the output value is greater than a cut-off value, it is determined that there is cancer, and if the output value is less than the cut-off value, it is determined that there is no cancer,

wherein the artificial intelligence model in step (f) has been trained to distinguish between vectorized data of a healthy subject and vectorized data of a cancer patient.

15. The non-transitory computer-readable storage medium of claim 14 , further comprising predicting a cancer type, wherein the output value in step (e) comprises a deep probability index (DPI), wherein predicting a cancer type comprises comparing the DPI calculated for multiple cancer types, wherein the cancer type showing the highest DPI is predicted to be the cancer type of the sample.

16. A method for measuring fragment end motif frequency and size of nucleic acids, the method comprising:

a) extracting nucleic acids from a biological sample, the extracted nucleic acids comprising cell free DNA (cfDNA);

b) sequencing the nucleic acids to a depth of at least 1 million reads to obtain sequence reads;

c) aligning the sequence reads with a reference genome database, wherein the length of the reads is from 5 to 5,000 bp;

d) acquiring end motif frequencies and sizes of nucleic acid fragments based on the aligned sequence reads, wherein the end motif frequencies are selected from a total of 256 4-mer motifs, and wherein the nucleic acid fragments comprise fragment sizes ranging from 90 to 250 base pairs;

e) generating vectorized data using the end motif frequencies and sizes of nucleic acid fragments, wherein the vectorized data is expressed by a type of an end motif of the nucleic acid fragments and the size of the nucleic acid fragments plotted on a single or multidimensional axis;

f) inputting the generated vectorized data to a trained artificial intelligence model, wherein the artificial intelligence model transforms the vectorized data of step (e) into an output value, wherein the artificial intelligence model has been trained to adjust the output value to about 1 if there is cancer or to adjust the output value to about 0 if there is no cancer;

g) determining whether or not cancer develops, wherein the determining comprises comparing the output value of step (f) with a cut-off value to determine whether or not cancer develops, wherein if the output value is greater than a cut-off value, it is determined that there is cancer, and if the output value is less than the cut-off value, it is determined that there is no cancer; and

h) predicting a cancer type, wherein the output value in step (f) comprises a deep probability index (DPI), wherein predicting a cancer type comprises comparing the DPI calculated for multiple cancer types, wherein the cancer type showing the highest DPI is predicted to be the cancer type of the sample,

wherein the artificial intelligence model in step (f) has been trained to distinguish between vectorized data of a healthy subject and vectorized data of a cancer patient.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2023
From: CHO, EUN HAE; LEE, TAE-RIM; PARK, SOOK RYUN
To: GC GENOME CORPORATION
Reel/Frame 062743/0591 →
Priority Claims (1)
KR 10-2021-0068891 · May 28, 2021 · national
Continuity (2)
Continuation PCTKR2022007651 · May 30, 2022
Related Publication 20230260655A1 · Aug 17, 2023
References Cited (69)
US 7244567B2 · Chen et al. · 2007 [cited by applicant]
US 10975431B2 · Velculescu · 2021 [cited by examiner]
US 20050088236A1 · Matsumoto et al. · 2005 [cited by applicant]
US 20060246497A1 · Huang et al. · 2006 [cited by applicant]
US 20060275779A1 · Li et al. · 2006 [cited by applicant]
US 20070087362A1 · Church et al. · 2007 [cited by applicant]
US 20070194225A1 · Zorn · 2007 [cited by applicant]
US 20190189242A1 · Angiuoli et al. · 2019 [cited by applicant]
KR 1020190036494A · 2019 [cited by applicant]
KR 101984611B1 · 2019 [cited by applicant]
KR 102084683B1 · 2020 [cited by applicant]
KR 1020200011471A · 2020 [cited by applicant]
KR 1020200087427A · 2020 [cited by applicant]
KR 1020200101106A · 2020 [cited by applicant]
KR 1020200108938A · 2020 [cited by applicant]
WO 2020125709A1 · 2020 [cited by applicant]
Alzubaidi et al. (Journal of Big Data, 2021 vol. 8:pp. 1-74). [cited by examiner]
Ma et al. (Plos One, 2017. vol. 12:e0169231, pp. 1-18). [cited by examiner]
Huang et al. (Cancer (2019) vol. 11:e-pp. 1-15) (Year: 2019). [cited by examiner]
Jiang et al. (Cancer Discov 2020; 10:664-73) (Year: 2020). [cited by examiner]
Trends Cancer. Apr. 2021 ; 7(4): 283-292 (Year: 2021). [cited by examiner]
“Applied Biosystems, Carlsbad, California, USA”, Applied Biosystems, 2009, Corona Lite Introduction 27 pages. [cited by applicant]
Branton, D., et al., “The potential and challenges of nanopore sequencing”, Nature Biotechnology, 2008, pp. 1146-1153, vol. 26, No. 10, Publisher: Nature Publishing Group. [cited by applicant]
Butler, J., et al., “ALLPATHS De novo assembly of whole-genome shotgun microreads”, Genome Research, 2008, pp. 810-820, vol. 18, Publisher: Cold Spring Harbor Laboratory Press. [cited by applicant]
Campagna, D., et al., “PASS a program to align short sequences”, Bioinformatics, 2009, pp. 967-968; doi:10.1093/bioinformatics/btp087, vol. 25, No. 7, Publisher: Oxford University Press. [cited by applicant]
Chen Y., et al., “PerM: efficient mapping of short sequencing reads with periodic full sensitive spaced seeds”, Bioinformatics, 2009, pp. 2514-2521; doi: 10.1093/bioinformatics/btp486, vol. 25, No. 19. [cited by applicant]
Clement, N.L., et al., “The GNUMAP algorithim unbiased probablistic mapping of oligonucleotides from next-generation sequencing”, Bioinformatics, 2010, pp. 38-45, vol. 26, No. 1, Publisher: Oxford University Press. [cited by applicant]
De Bona, F., et al., “Optimal spliced alignments of short sequence reads”, Bioinformatics, 2008, pp. i174-i180, vol. 24, No. 16, Publisher: Oxford University Press. [cited by applicant]
Eaves, H.L., et al., “MOM: maximum oligonucleotide mapping”, Bioinformatics, 2009, pp. 969-970; doi. 10.1093/bioinformatics/btp092, vol. 25, No. 7, Publisher: Oxford University Press. [cited by applicant]
Edwards, J.R., et al., “Mass-spectrometry DNA sequencing”, Mutation Research, 2005, pp. 3-12, vol. 573, Publisher: Elsevier. [cited by applicant]
Fahlgren, N., “Computational and analytical framework for small RNA profiling by high-throughput sequencing”, RNA, 2009, pp. 992-1002, vol. 15, Publisher: ResearchGate. [cited by applicant]
Gnirke, A., et al., “Solution hybrid selection with ultra-long oligonucleotides for massively parallel targeted sequencing”, Nature Biotechnology, 2009, pp. 182-189, vol. 27, No. 2, Publisher: Nature America, Inc. [cited by applicant]
Hanna, G.J., et al., “Comparison of Sequencing by Hybridization and Cycle Sequencing for Genotyping of Human Immunodeficiency Virus Type 1 Reverse Transcriptase”, Journal of Clinical Microbiology, 2000, pp. 2715-2721, v… [cited by applicant]
Hinton, G., et al., “Deep Neural Networks for Acoustic Modeling in Speech Recognition”, IEEE Signal Processing Magazine, 2012, pp. 82-97, vol. 29, No. 6. [cited by applicant]
Homer, N., et al., “BFAST An Alignment Tool for Large Scale Genome Resequencing”, PLoS One, 2009, e7767, vol. 4, No. 11, Publisher: www.plosone.org, 12 pages. [cited by applicant]
Jiang, H., et al., “SeqMap: mapping massive amount of oligonucleotides to the genome”, Bioinformatics, 2008, pp. 2395-2396; doi: 10.1093/bioinformatics/btn429, vol. 24, No. 20, Publisher: Oxford University Press. [cited by applicant]
Jiang, P., et al., “Plasma DNA End-Motif Profiling as a Fragmentomic Marker in Cancer, Pregnancy, and Transplantation”, Cancer Discovery, 2020, pp. 664-673, vol. 10. [cited by applicant]
Kent, W.J., “BLAT—The Blast-Like Alignment Tool”, Genome Research, 2002, vol. 12, No. 656-664, Publisher: Cold Spring Laboratory Press. [cited by applicant]
Kim, Y.J., et al., “ProbeMatch rapid alignment of oligonucleotides to genome allowing both gaps and mismatches”, Bioinformatics, 2009, pp. 1424-1425; doi:10.1093/bioinformatics/btp178, vol. 25, No. 11, Publisher: Oxford… [cited by applicant]
Krishnakumar, S., et al., “A comprehensive assay for targeted multiplex amplification of human DNA sequences”, PNAS, 2008, pp. 9296-9310; www.pnas.org/cgi/doi/10.1073/pnas.0803240105, vol. 105, No. 27. [cited by applicant]
Langmead, B., et al., “Ultrafast and memory-efficient alignment of short DNA sequenes to the human genome”, Genome Biology, 2009, doi:10.1186/GB-2009-10-3-r25, vol. 10, No. R25, 10 pages. [cited by applicant]
Lasken, R.S., “Single-cell genomic sequening using Multiple Displacement Amplification”, Current Opinion in Microbiology, 2007, pp. 510-516, vol. 10, Publisher: Elsevier. [cited by applicant]
Li, H., et al., “Mapping short DNA sequencing reads and calling variants using mapping quality scores”, Genome Research, 2008, pp. 1851-1858, vol. 18, Publisher: CSH Press. [cited by applicant]
Li, R., et al., “SOAP: short oligonucleotide alignment program”, Bioinformatics, 2008, pp. 713-714; doi:10.1093/bioinformatics/btn025, vol. 24, No. 5, Publisher: Oxford University Press. [cited by applicant]
Li, H., et al., “Fast and accurate short read alignment with Burrows-Wheeler transform”, Bioinformatics, 2009, pp. 1754-1760; doi:10.1093/bioinformatics/btp324, vol. 25, No. 14. [cited by applicant]
Li, R., et al., “SOAP2 an improved ultrafast tool for short read alignment”, Bioinformatics, 2009, pp. 1966-1967; doi:10.1093/bioinformatics/btp336, vol. 25, No. 15, Publisher: Oxford University Press. [cited by applicant]
Li, H., et al., “Fast and accurate long-read alignment with Burrows-Wheeler transform”, Bioinformatics, 2010, pp. 589-595; doi:10.1093/bioinformatics/btp698, vol. 26, No. 5. [cited by applicant]
Lunter, G, et al., “Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads”, Genome Research, 2011, pp. 936-939, vol. 21, Publisher: Cold Spring Harbor Laboratory Press. [cited by applicant]
Malhis, N., et al., “Slider-maximum use of probability information for alignment of short sequence reads and SNP detection”, Bioinformatics, 2009, pp. 6-13; doi:10.1093/bioinformatics/btn565, vol. 25, No. 1. [cited by applicant]
Metzker, M.L., “Sequencing technologies—the next generation”, Nature Reviews Genetics, 2010, pp. 31-46, vol. 11, Publisher: Macmillan Publishers Limited. [cited by applicant]
Muller, T., et al., “Non-symmetric score matrices and the detection of homologous transmembrane proteins”, Bioinformatics, 2001, pp. S182-S189, vol. 17, No. 1, Publisher: Oxford University Press. [cited by applicant]
Ning, Z., et al., “SSAHA: A Fast Search Method for Large DNA Databases”, Genome Research, 2001, pp. 1725-1729, vol. 11, Publisher: Cold Harbor Laboratory Press. [cited by applicant]
Ondov, B.D., et al., “Efficient mapping of Applied Biosystems SOLiD sequence data to a reference genome for functional genomic applications”, Bioinformatics, 2008, pp. 2776-2777; doi:10.1093/bioinformatics/btn512, vol. … [cited by applicant]
Porreca, G.J., et al., “Multiplex amplification of large sets of human exons”, Nature Methods, 2007, pp. 931-936, vol. 4, No. 11, Publisher: Nature Publishing Group. [cited by applicant]
Prufer, K., et al., “PatMaN rapid alignment of short sequences to large databases”, Bioinformatics, 2008, pp. 1530-1532; doi:10.1093/bioinformatics/btn223, vol. 24, No. 13. [cited by applicant]
Rumble, S.M., et al., “SHRiMP: Accurate Mapping of Short Color-spade Reads”, PLoS Computational Biology, 2009, e1000386, vol. 5, No. 5, 11 pages. [cited by applicant]
Salmela, L., “Correction of sequencing errors in a mixed set of reads”, Bioinformatics, 2010, pp. 1284-1290; doi:10.1093/bioinformatics/btq151, vol. 26, No. 10, Publisher: Oxford University Press. [cited by applicant]
Schatz, M. C., “CloudBurst: highly sensitive read mapping with MapReduce”, Bioinformatics, 2009, pp. 1363-1369; doi.10.1093/bioinformatics/btp236, vol. 25, No. 11. [cited by applicant]
Shi, H, et al., “A Parallel Algorithim for Error Correction in High-Throughput Short-Read Data on CUDA-Enabled Graphics Hardware”, Journal of Computational Biology, 2010, pp. 603-615; DOI:10.1089/cmb.2009.0062, vol. 17,… [cited by applicant]
Smith, A.D., et al., “Updates to the RMAP short-read mapping software”, Bioinformatics, 2009, pp. 2841-2842; doi:10.1093/bioinformatics/btp533, vol. 25, No. 21, Publisher: Oxford University Press. [cited by applicant]
Tewhey, R., et al., “Microdroplet-based PCR enrichment for large-scale targeted sequencing”, Nature Biotechnology, 2009, pp. 1025-1031, vol. 27, No. 11, Publisher: Nature Publishing Group. [cited by applicant]
Trapnell, C., “How to map billions of short reads onto genomes”, Nature Biotechnology, 2009, pp. 455-457, vol. 27. [cited by applicant]
Turner, E.H., et al., “Massively parallel exon capture and library-free resequencing across 16 genomes”, Nature Methods, 2009, pp. 315-316, vol. 6, No. 5, Publisher: Nature America, Inc. [cited by applicant]
Warren, R.L., et al., “Assembling millions of short DNA sequences using SSAKE”, Bioinformatics, 2007, pp. 500-501; doi: 10.1093/bioinformatics/btl629, vol. 23, No. 4. [cited by applicant]
Weese, D., et al., “RazerS-fast read mapping with sensitivity control”, Genome Research, 2009, pp. 1646-1654, vol. 19, Publisher: Cold Spring Harbor Laboratory Press. [cited by applicant]
Wu, T.D., et al., “GMAP a genmic mapping and alignment program for mRNA and EST sequences”, Bioinfomatics, 2005, pp. 1859-1875, vol. 21, No. 9. [cited by applicant]
Wu, T.D., et al., “Fast and SNP-tolerant detection of complex variants and splicing in short reads”, Bioinformatics, 2010, pp. 873-881, vol. 26, No. 7, Publisher: Oxford University Press. [cited by applicant]
Zerbino, D.R., et al., “Velvet: Algorithms for de novo read assembly using de Bruijin graphs”, Genome Research, 2008, pp. 821-829, vol. 18, Publisher: Cold Spring Harbor Laboratory Press. [cited by applicant]
Zhou, X., et al., “CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing”, bioRxiv, 2020, http://doi.org/10.1101/2020.07.16.201350, 34 pages. [cited by applicant]