IP Library Granted Patent US 12,236,346
Granted Patent B2
US 12,236,346 · App. 17/489,458 · Granted Feb 25, 2025

Systems and methods for using a convolutional neural network to detect contamination

Inventors: Christopher-James A. V. Yakym (Mountain View, CA); Onur Sakarya (Redwood City, CA)
Assignee: Grail, Inc.
G06N3/08G06F18/2148G16B20/20G16B30/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,346
App. No.
17/489,458
Granted
Feb 25, 2025
Kind
B2
Abstract

A method for training a convolutional neural net for contamination analysis is provided. A training dataset is obtained comprising, for each respective training subject in a plurality of subjects, a variant allele frequency of each respective single nucleotide variant in a respective plurality of single nucleotide variants, and a respective contamination indication. First and second subsets of the plurality of training subjects have first and second contamination indication values, respectively. A corresponding first channel comprising a first plurality of parameters that include a respective parameter for a single nucleotide variant allele frequency of each respective single nucleotide variant in a set of single nucleotide variants in a reference genome is constructed for each respective training subject. An untrained or partially trained convolutional neural net is trained using, for each respective training subject, at least the corresponding first channel of the respective training subject as input against the respective contamination indication.

Claims (34)

1. A method of determining a contamination status of a test biological sample obtained from a test subject, comprising:

(a) obtaining, in electronic format, one or more training subject datasets, each training subject dataset comprising a corresponding training variant allele frequency of each respective training single nucleotide variant in a plurality of training single nucleotide variants;

(b) training a computational neural network based on the one or more training subject datasets, wherein the computational neural network comprises a pre-trained convolutional neural network and an untrained classifier;

(c) obtaining, in electronic format, a test subject dataset comprising a corresponding test variant allele frequency of each respective test single nucleotide variant in a plurality of test single nucleotide variants; and

(d) determining the contamination status for the test biological sample based on the trained computational neural network and the test subject dataset.

2. The method of claim 1 , wherein the corresponding training variant allele frequency is determined by sequencing of one or more nucleic acids in a respective training biological sample obtained from a respective training subject.

3. The method of claim 2 , wherein the respective training biological sample is substantially cell-free sample of blood plasma or blood serum obtained from the respective training subject.

4. The method of claim 1 , wherein the corresponding training variant allele frequency is between 0 and 1.

5. The method of claim 1 , wherein the training dataset subject dataset further comprises one or more contamination indications.

6. The method of claim 5 , wherein each of the one or more contamination indications is at least 0%, 0.1%, 0.3%, 0.5%, 1%, 5%, 10%, 15%, 20%, or 25%.

7. The method of claim 5 , wherein the training comprises constructing one or more channels comprising one or more parameters, wherein the one or more parameters comprise at least one parameter associated with the one or more contamination indications.

8. The method of claim 1 , wherein the training comprises constructing one or more channels comprising one or more parameters, wherein the one or more parameters comprise at least one parameter associated with the corresponding variant allele frequency.

9. The method of claim 1 , wherein the plurality of training single nucleotide variants comprises 100 or more single nucleotide variants, 200 or more single nucleotide variants, 500 or more single nucleotide variants, 1000 or more single nucleotide variants, 2000 or more single nucleotide variants, 2000 or more single nucleotide variants, 4000 or more single nucleotide variants, or 10000 or more single nucleotide variants.

10. The method of claim 1 , wherein the pre-trained convolutional neural network comprises LeNet, AlexNet, VGGNet 16, GoogLeNet, or ResNet.

11. The method of claim 1 , wherein the corresponding test variant allele frequency is determined by sequencing of one or more nucleic acids in the test biological sample obtained from the test subject.

12. The method of claim 1 , wherein the test biological sample is substantially cell-free sample of blood plasma or blood serum obtained from the test subject.

13. The method of claim 1 , wherein the contamination status comprises an estimation of contamination percentage for the test biological sample of the test subject.

14. A system for determining a contamination status of a test biological sample obtained from a test subject:

a storage device that stores instructions; and

at least one processor that executes the instructions in order for:

(a) obtaining, in electronic format, one or more training subject datasets, each training subject dataset comprising a corresponding training variant allele frequency of each respective training single nucleotide variant in a plurality of training single nucleotide variants;

(b) training a computational neural network based on the one or more training subject datasets, wherein the computational neural network comprises a pre-trained convolutional neural network and an untrained classifier;

(c) obtaining, in electronic format, a test subject dataset comprising a corresponding test variant allele frequency of each respective test single nucleotide variant in a plurality of test single nucleotide variants; and

(d) determining the contamination status for the test biological sample based on the trained computational neural network and the test subject dataset.

15. The system of claim 14 , wherein the corresponding training variant allele frequency is determined by sequencing of one or more nucleic acids in a respective training biological sample obtained from a respective training subject.

16. The system of claim 15 , wherein the respective training biological sample is substantially cell-free sample of blood plasma or blood serum obtained from the respective training subject.

17. The system of claim 14 , wherein the corresponding training variant allele frequency is between 0 and 1.

18. The system of claim 14 , wherein the training dataset subject dataset further comprises one or more contamination indications.

19. The system of claim 18 , wherein each of the one or more contamination indications is at least 0%, 0.1%, 0.3%, 0.5%, 1%, 5%, 10%, 15%, 20%, or 25%.

20. A non-transitory computer-readable medium storing instructions for determining a contamination status of a test biological sample obtained from a test subject comprising:

(a) obtaining, in electronic format, one or more training subject datasets, each training subject dataset comprising a corresponding training variant allele frequency of each respective training single nucleotide variant in a plurality of training single nucleotide variants;

(b) training a computational neural network based on the one or more training subject datasets, wherein the computational neural network comprises a pre-trained convolutional neural network and an untrained classifier;

(c) obtaining, in electronic format, a test subject dataset comprising a corresponding test variant allele frequency of each respective test single nucleotide variant in a plurality of test single nucleotide variants; and

(d) determining the contamination status for the test biological sample based on the trained computational neural network and the test subject dataset.

Assignments (2)
CHANGE OF NAME Recorded Jan 14, 2025
From: GRAIL, LLC
To: GRAIL, INC.
Reel/Frame 069897/0799 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2022
From: YAKYM, CHRISTOPHER-JAMES A.V.; SAKARYA, ONUR
To: GRAIL, LLC
Reel/Frame 059621/0442 →
Continuity (2)
Provisional Application 63085369 · Sep 30, 2020
Related Publication 20220101135A1 · Mar 31, 2022
References Cited (34)
US 10496884B1 · Nguyen · 2019 [cited by examiner]
US 20140332994A1 · Danes · 2014 [cited by examiner]
US 20180237838A1 · Sakarya · 2018 [cited by examiner]
US 20180237863A1 · Namsaraev · 2018 [cited by examiner]
US 20180373832A1 · Sakarya · 2018 [cited by examiner]
US 20190287649A1 · Filippova et al. · 2019 [cited by applicant]
US 20190287652A1 · Gross et al. · 2019 [cited by applicant]
US 20200003440A1 · Kim · 2020 [cited by examiner]
US 20200013024A1 · Armstrong · 2020 [cited by examiner]
US 20200134461A1 · Chai · 2020 [cited by examiner]
US 20200279368A1 · Tada · 2020 [cited by examiner]
US 20200303078A1 · Mayhew · 2020 [cited by examiner]
US 20200340064A1 · Gross · 2020 [cited by examiner]
US 20200385813A1 · Venn · 2020 [cited by applicant]
US 20210104297A1 · Venn et al. · 2021 [cited by applicant]
US 20210158308A1 · Armstrong · 2021 [cited by examiner]
US 20210187732A1 · Chae · 2021 [cited by examiner]
US 20210292845A1 · Melton et al. · 2021 [cited by applicant]
US 20210299706A1 · Filler · 2021 [cited by examiner]
US 20210327534A1 · Nicula · 2021 [cited by examiner]
US 20220053005A1 · Liu · 2022 [cited by examiner]
US 20220101135A1 · Yakym · 2022 [cited by examiner]
US 20230093535A1 · Karlík · 2023 [cited by examiner]
US 20230296313A1 · Melhem · 2023 [cited by examiner]
US 20240312561A1 · Calef · 2024 [cited by examiner]
US 20240312564A1 · Nohzadeh-Malakshah · 2024 [cited by examiner]
US 20240410807A1 · Mareuge · 2024 [cited by examiner]
US 20240412821A1 · Sakarya · 2024 [cited by examiner]
WO WO2018081130A1 · 2018 [cited by applicant]
WO WO2019178289A1 · 2019 [cited by applicant]
WO WO2020132148A1 · 2020 [cited by applicant]
Klein, E. et al., “Development of a comprehensive cell-free DNA (cfDNA) assay for early detection of multiple tumor types: The Circulating Cell-free Genome Atlas (CCGA) study” 2018 ASCO Annual Meeting, Meeting Abstract,… [cited by applicant]
Liu, M.C. et al., “Genome-wide cell-free DNA (cfDNA) methylation signatures and effect on tissue of origin (TOO) performance” 2019 ASCO Annual Meeting, Meeting Abstract, Jun. 2019, pp. 1. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US2021/052709, Jan. 28, 2022, 20 pages. [cited by applicant]