IP Library › Granted Patent US 12,287,764
Granted Patent B2
US 12,287,764 · App. 17/990,361 · Granted Apr 29, 2025

FastQ/FastA compression systems and methods

Inventors: Foad Nazari (Malvern, PA); Sneh Patel (Malvern, PA); Emma K. Murray (Malvern, PA); Giana J. Schena (Malvern, PA)
Assignee: Rajant Health Incorporated
G06F16/1744G06F16/162
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,287,764
App. No.
17/990,361
Granted
Apr 29, 2025
Kind
B2
Abstract

Systems and methods to analyze and significantly compress FastQ and/or FastA datasets are disclosed. The methodology includes algorithms to compress sequences, quality scores and identifiers of read files. The method relies on reducing the dimension and redundancy in genomic data in a unique and optimal way and in the binary format. The methodology also includes the decoding protocols to decompress the compressed data with zero loss.

Claims (42)

1. A method for data compression of genomic data comprising:

receiving a data file comprised of sequence bases, quality scores, and identifiers, the sequence bases comprising regular bases and irregular bases, the regular bases comprising adenine (A), cytosine (C), guanine (G), and thymine (T);

applying an optimization algorithm to the quality scores, sequence k-mers from the sequence bases, and the identifiers;

splitting long reads of the sequence bases and the quality scores into smaller segments;

deleting duplicated and semi-duplicated reads for the sequence bases and the quality scores;

performing a dimensionality reduction on the sequence bases;

storing a template of the identifier that is consistent across the data file;

detecting and storing the location and type of each irregular base;

encoding the data file in a binary format such that each regular base is represented by one of the four two-digit binary numbers and all irregular bases are represented by one of the four two-digit binary numbers; and

compressing the encoded data file.

2. The method of claim 1 , wherein the identifiers are comprised of sequencing run data and cluster data.

3. The method of claim 1 , wherein the quality scores comprise the sequence of quality values for the sequence bases.

4. The method of claim 1 , wherein:

the irregular bases include N and a plurality of additional irregular bases; and

storing the location and type of each irregular base comprises storing the location of each irregular base and the type of each additional irregular base.

5. The method of claim 4 , wherein dimensionality reduction is performed using binary labels from a k-mers dictionary and the base sequences.

6. The method of claim 1 , wherein the data file is a FastQ file or a FastA file.

7. The method of claim 1 , wherein the optimization algorithm determines an optimal value for encoding algorithm hyper-parameters, the sequence k-mers from the bases, and the identifiers.

8. The method of claim 1 , wherein the optimization algorithm is a function of protocol hyperparameters and distribution indices of the data file.

9. The method of claim 1 , wherein the compression is reference-free.

10. The method of claim 1 , wherein the compression is tuned based on characteristics of the data file.

11. A non-transitory computer storage medium that stores a program thereon that causes a computer to execute a process comprising:

receiving a data file comprised of sequence bases, quality scores, and identifiers, the sequence bases comprising regular bases and irregular bases, the regular bases comprising adenine (A), cytosine (C), guanine (G), and thymine (T);

applying an optimization algorithm to the quality scores, sequence k-mers from the sequence bases, and the identifiers;

splitting long reads of the sequence bases and the quality scores into smaller segments;

deleting duplicated and semi-duplicated reads for the sequence bases and the quality scores;

performing a dimensionality reduction on the sequence bases;

storing a template of the identifier that is consistent across the data file;

detecting and storing the location and type of each irregular base;

encoding the data file in a binary format such that each regular base is represented by one of the four two-digit binary numbers and all irregular bases are represented by one of the four two-digit binary numbers; and

compressing the data file, wherein the compression is lossless.

12. The non-transitory computer medium of claim 11 , wherein the identifiers are comprised of sequencing run data and cluster data.

13. The non-transitory computer medium of claim 11 , wherein the quality scores comprise the sequence of quality values for the sequence bases.

14. The non-transitory computer medium of claim 11 , wherein:

the irregular bases include N and a plurality of non-N irregular bases; and

storing the location and type of each irregular base comprises storing the location of each irregular base and the type of each non-N irregular base.

15. The non-transitory computer medium of claim 14 , wherein dimensionality reduction is performed using binary labels from a k-mers dictionary and the regular bases.

16. The non-transitory computer medium of claim 11 , wherein the data file is a FastQ file or a FastA file.

17. The non-transitory computer medium of claim 11 , wherein the optimization algorithm determines the optimal hyperparameters values for each of the quality scores, the sequence k-mers from the bases, and the identifiers.

18. The non-transitory computer medium of claim 11 , wherein the optimization algorithm is a function of protocol hyperparameters and distribution indices of the data file.

19. The non-transitory computer medium of claim 11 , wherein the compression is reference-free.

20. The non-transitory computer medium of claim 11 , wherein the compression is tuned based on characteristics of the data file.

Assignments (2)
SECURITY INTEREST Recorded Nov 26, 2025
From: RAJANT HEALTH INCORPORATED; RAJANT CORPORATION
To: MERIDIAN BANK
Reel/Frame 073042/0178 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2023
From: NAZARI, FOAD; PATEL, SNEH; MURRAY, EMMA K; SCHENA, GIANA J
To: RAJANT HEALTH, INC.
Reel/Frame 062679/0316 →
Continuity (3)
Provisional Application 63409993 · Sep 26, 2022
Provisional Application 63280721 · Nov 18, 2021
Related Publication 20240134825A1 · Apr 25, 2024
References Cited (23)
US 8856089B1 · Briggs · 2014 [cited by examiner]
US 9292327B1 · von Thenen · 2016 [cited by examiner]
US 9424185B1 · Botelho · 2016 [cited by examiner]
US 10554220B1 · Constantinescu · 2020 [cited by examiner]
US 10972742B2 · Gisquet · 2021 [cited by examiner]
US 11068444B2 · Sharangpani · 2021 [cited by examiner]
US 11081208B2 · Jaffe · 2021 [cited by examiner]
US 20050025316A1 · Pelly · 2005 [cited by examiner]
US 20050028192A1 · Hooper · 2005 [cited by examiner]
US 20080091698A1 · Cook · 2008 [cited by examiner]
US 20100172543A1 · Winkler · 2010 [cited by examiner]
US 20120254333A1 · Chandramouli · 2012 [cited by examiner]
US 20130338934A1 · Asadi · 2013 [cited by examiner]
US 20170060896A1 · Ito · 2017 [cited by examiner]
US 20170237445A1 · Cox et al. · 2017 [cited by applicant]
US 20180075262A1 · Auh · 2018 [cited by examiner]
US 20180152535A1 · Sade et al. · 2018 [cited by applicant]
US 20180364949A1 · Aston · 2018 [cited by examiner]
US 20190205542A1 · Kao · 2019 [cited by examiner]
US 20190287655A1 · Wesselman · 2019 [cited by examiner]
US 20210366576A1 · Bartov · 2021 [cited by examiner]
Notification of Transmittal of The International Search Report and the Written Opinion of the International Searching Authority, or the Declaration, International Search Report, and Written Opinion of the International … [cited by applicant]
Behjati et al., “What is next generation sequencing?” Arch Dis Child Educ Pract Ed 2013; 98:236-238. [cited by applicant]