IP Library Granted Patent US 9,218,352
Granted Patent B2
US 9,218,352 · App. 14/604,165 · Granted Dec 22, 2015

Methods and systems for storing sequence read data

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,218,352
App. No.
14/604,165
Granted
Dec 22, 2015
Kind
B2
Abstract

The present invention generally relates to storing sequence read data. The invention can involve obtaining a plurality of sequence reads from a sample, identifying one or more sets of duplicative sequence reads within the plurality of sequence reads, and storing only one of the sequence reads from each set of duplicative sequence reads in a text file using nucleotide characters.

Claims (27)

1. A method for storing sequence read data, the method comprising:

sequencing a nucleic acid from a sample to generate a plurality of sequence reads;

identifying one or more sets of duplicative sequence reads within the plurality of sequence reads;

storing only one sequence read from each of the one or more sets of duplicative sequence in a master read file;

collecting meta information for each of the plurality of sequence reads, appending the meta information into a compressed file, and matching the meta information to a single read in the master read file; and

later retrieving the plurality of sequence reads from the compressed file and the master read file.

2. A method for storing sequence read data, the method comprising:

obtaining a plurality of sequence reads from a sample by sequencing a nucleic acid from the sample;

identifying one or more sets of duplicative sequence reads within the plurality of sequence reads;

storing in a master read file only one sequence read from each of the one or more sets of duplicative sequence reads; and

separately retrieving the plurality of sequence reads from the master read file.

3. A method for using stored sequence read data, the method comprising:

using a computer system comprising a memory coupled to a processor for:

obtaining a master sequence read file that includes only one sequence read from each of one or more sets of duplicative sequence reads obtained from a sample and a compressed file that includes lines of metadata for the sequence reads obtained from the sample;

for each line of metadata in the compressed file, retrieving an associated read from the master sequence read file and appending that line of metadata and the associated read to an output sequence read file, wherein the output sequence read file contains the sequence reads as originally obtained from the sample.

4. The method of claim 3 , wherein the master sequence read file, the compressed file, and the output sequence read file each have a format that is human-readable.

5. The method of claim 4 , wherein sequence reads are stored in the master sequence read file and the output sequence read file using IUPAC characters.

6. The method of claim 4 , wherein the format of the output sequence read file is one selected from the group consisting of: a FASTA file; a FASTQ file; and a VCF file.

7. The method of claim 4 , wherein the format of the master sequence read file is a text file.

8. The method of claim 4 , wherein the computer system comprises a text editor program capable of opening the files having the format.

9. The method of claim 8 , wherein the text editor program is capable of displaying the files on a computer screen showing the metadata and the sequence reads in a human-readable format.

10. The method of claim 3 , wherein each line of metadata comprises a sequence read ID.

11. The method of claim 3 , wherein the output sequence read file contains a perfectly lossless retrieval of the sequence reads obtained from the sample.

12. The method of claim 3 , further comprising obtaining the master sequence read file and the compressed file by:

sequencing a nucleic acid from the sample to generate the sequence reads;

identifying the one or more sets of duplicative sequence reads within the sequence reads; and

storing the only one sequence read from each of the one or more sets of duplicative sequence reads in the master read file.

Assignments (9)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2024
From: INVITAE CORPORATION
To: LABORATORY CORPORATION OF AMERICA HOLDINGS
Reel/Frame 068822/0025 →
SECURITY INTEREST Recorded Mar 13, 2023
From: INVITAE CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 063787/0148 →
RELEASE OF SECURITY INTEREST Recorded Mar 6, 2023
From: PERCEPTIVE CREDIT HOLDINGS III, LP
To: INVITAE CORPORATION; GOOD START GENETICS, INC.; SINGULAR BIO, INC.; YOUSCRIPT, LLC
Reel/Frame 063282/0538 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE SCHEDULE A OF THE CONFIRMATORY ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 056756 FRAME: 0884. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 11, 2021
From: GOOD START GENETICS, INC.
To: INVITAE CORPORATION
Reel/Frame 057772/0828 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2021
From: GOOD START GENETICS, INC.
To: INVITAE CORPORATION
Reel/Frame 056756/0884 →
PATENT SECURITY AGREEMENT Recorded Oct 2, 2020
From: INVITAE CORPORATION; GOOD START GENETICS, INC.; SINGULAR BIO, INC.; YOUSCRIPT, LLC
To: PERCEPTIVE CREDIT HOLDINGS III, LP
Reel/Frame 054234/0872 →
RELEASE OF SECURITY INTEREST Recorded Sep 11, 2019
From: INN SA LLC
To: INVITAE CORPORATION; GOOD START GENETICS, INC.; COMBIMATRIX CORPORATION
Reel/Frame 050454/0559 →
SECURITY INTEREST Recorded Nov 6, 2018
From: INVITAE CORPORATION; GOOD START GENETICS, INC.; COMBIMATRIX CORPORATION
To: INN SA LLC
Reel/Frame 047889/0836 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2015
From: KENNEDY, CALEB J.; CHENNAGIRI, NIRU
To: GOOD START GENETICS, INC.
Reel/Frame 035085/0094 →