IP Library Granted Patent US 12,205,679
Granted Patent B2
US 12,205,679 · App. 16/998,236 · Granted Jan 21, 2025

Systems and methods for sequence encoding, storage, and compression

Inventor: Vladimir Semenyuk (Monterey, CA)
Assignee: SEVEN BRIDGES GENOMICS INC.
G16B50/50G06F3/0608G06F16/9024H03M7/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,679
App. No.
16/998,236
Granted
Jan 21, 2025
Kind
B2
Abstract

Genomic data is written to disk in a compact format by dividing the data into segments and encoding each segment with the smallest number of bits per character necessary for whatever alphabet of characters appears in that segment. A computer system dynamically chooses the segment boundaries for maximum space savings. A first one of the segments may use a different number of bits per character than a second one of the segments. In one embodiment, dividing the data into segments comprises scanning the data and keeping track of a number of unique characters, noting positions in the sequence where the number increases to a power of two, calculating a compression that would be obtained by dividing the genomic data into one of the plurality of segments at ones of the noted positions, and dividing the genomic data into the plurality of segments at the positions that yield the best compression.

Claims (15)

1. A method for encoding genomic data, the method comprising:

dividing genomic data into a plurality of segments, wherein dividing the genomic data into the plurality of segments comprises:

scanning the genomic data and keeping track of a number of unique characters scanned;

noting positions in the sequence where the number increases to a power of two; and

dividing the genomic data into the plurality of segments based on the positions;

encoding characters within each of the plurality of segments with a smallest number of bits per character needed to represent all unique characters present in each segment, wherein an identity and number of each character is encoded, wherein encoding the characters within each of the plurality of segments comprises creating for each segment a segment header and a corresponding data portion, the header comprising information about an alphabet of characters that is encoded by the corresponding data portion; and

storing the encoded characters in a computer-readable medium.

2. The method of claim 1 , wherein a first one of the segments uses a different number of bits per character than a second one of the segments.

3. The method of claim 1 , wherein, within a segment, every character is encoded using the same number of bits.

4. The method of claim 1 , wherein encoding the characters within each of the plurality of segments comprises:

determining a number N of unique characters in the segment; and

encoding the segment using X bits per character, where 2{circumflex over ( )}(x−1)<N<2{circumflex over ( )}x.

5. The method of claim 1 , wherein the segment header comprises at least two bits that indicate a length of the segment header and a plurality of flag bits, wherein the flag bits identify the alphabet of characters that is encoded by the corresponding data portion.

6. The method of claim 5 , wherein the flag bits identify the alphabet using a bitmap to identify included characters.

7. The method of claim 1 , wherein the genomic data comprises a directed graph representing at least a portion of a plurality of genomes, and wherein encoding the characters within each of the plurality of segments transforms the directed graph into a sequence of bit, the method further comprising transferring the sequence of bits serially from a computer memory device a computer system to a second computer system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: SEMENYUK, VLADIMIR
To: SEVEN BRIDGES GENOMICS INC.
Reel/Frame 066077/0598 →
Continuity (6)
Continuation 15597464 · May 17, 2017
Provisional Application 62400123 · Sep 27, 2016
Provisional Application 62400411 · Sep 27, 2016
Provisional Application 62400388 · Sep 27, 2016
Provisional Application 62338706 · May 19, 2016
Related Publication 20210050074A1 · Feb 18, 2021
References Cited (11)
US 5298895A · Van Maren · 1994 [cited by examiner]
US 5410671A · Elgamal · 1995 [cited by examiner]
US 5608396A · Cheng · 1997 [cited by examiner]
US 5870036A · Franaszek · 1999 [cited by examiner]
US 20040153255A1 · Ahn · 2004 [cited by examiner]
US 20060273934A1 · Delfs · 2006 [cited by examiner]
US 20090292699A1 · Knowles · 2009 [cited by examiner]
US 20090298702A1 · Su · 2009 [cited by examiner]
US 20130031092A1 · Bhola · 2013 [cited by examiner]
WO WO9901940A1 · 1999 [cited by examiner]
Raphael, B. J.; Zhi, D.; Tang, H.; Pevzner, P. A. A Novel Method for Multiple Alignment of Sequences with Repeated and Shuffled Elements. Genome Research 2004, 14 (11), 2336-2346. [cited by examiner]