IP Library Granted Patent US 11,811,428
Granted Patent B2
US 11,811,428 · App. 18/305,305 · Granted Nov 7, 2023

System and method for data compression using genomic encryption techniques

Inventors: Joshua Cooper (Columbia, SC); Aliasghar Riahi (Orinda, CA); Mojgan Haddad (Orinda, CA); Ryan Kourosh Riahi (Orinda, CA); Razmin Riahi (Orinda, CA); Charles Yeomans (Orinda, CA)
Assignee: ATOMBEAM TECHNOLOGIES INC.
H03M7/3059G06N20/00H03M7/6005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,811,428
App. No.
18/305,305
Filed
Apr 21, 2023
Granted
Nov 7, 2023
Kind
B2
Art Unit
2845
USPC
707/693
Abstract

A system and method for data compression with genomic encryption, which uses frequency analysis on data blocks within an input data stream to produce a prefix table, representing a first layer of transformation, and which applies a Burrow's-Wheeler transform (BWT) to the data inside the prefix table, representing a second layer of transformation, and which compresses the transformed data. In some implementations, the system and method may further include applying the BWT to a conditioned stream of genomic data, wherein the conditioned stream of genomic data is accompanied by an error stream comprising the differences between the original data and the encrypted data.

Claims (45)

1. A system for genomic data compression with encryption, comprising:

a computing device comprising a processor and a memory;

a stream analyzer comprising a first plurality of programming instructions stored in the memory which, when operating on the processor, causes the computing device to:

receive an input data stream comprising genomic data;

analyze the frequency distribution of a plurality of data blocks within the input data stream to determine whether the input data stream meets a configured threshold for data conditioning;

if the input data stream meets or exceeds the configured threshold, send the input data stream to a stream conditioner; and

the stream conditioner comprising a second plurality of programming instructions stored in the memory which, when operating on the processor, causes the computing device to:

receive the input data stream from the stream analyzer;

produce a conditioned data stream and an error stream by, for each of a plurality of data blocks within the data stream:

analyzing the data block within the data stream to compare the data block's real frequency within the data stream against an ideal frequency;

if the difference between the data block's real frequency and ideal frequency exceeds a configured conditioning threshold, applying a conditioning rule to the data block;

applying a logical XOR operation to the data block;

appending the output of the logical XOR operation to the error stream;

send the conditioned data stream and the error stream as output.

2. The system of claim 1 , wherein the stream analyzer is further configured to:

analyze the frequency distribution of the plurality of data blocks within the input data stream to determine the most frequent bytes or strings of bytes that occur at the beginning of each of the plurality of data blocks;

designate the most-frequent bytes or strings of bytes as prefixes;

compile a prefix table based on the results of the frequency distribution, the prefix table comprising the designated prefixes and the block lengths;

send the prefix table to a data transformer.

3. The system of claim 2 , further comprising the data transformer comprising a third plurality of programming instructions stored in the memory which, when operating on the processor, causes the computing device to:

receive the prefix table;

transform each of the plurality of prefixes within the prefix table by applying a Burrow's-Wheeler transform (BWT) and generating as output a plurality BWT-prefixes; and

send the plurality of BWT-prefixes as output.

4. The system of claim 2 , wherein each of the plurality of data blocks are k-mers of the genomic data and the prefixes are one or more base pairs that occur at the beginning of each k-mer.

5. A method for genomic data compression with encryption, comprising the steps of:

receiving an input data stream comprising genomic data;

analyzing the frequency distribution of a plurality of data blocks within the input data stream to determine whether the input data stream meets a configured threshold for data conditioning;

if the input data stream meets or exceeds the configured threshold, sending the input data stream to a stream conditioner;

receiving the input data stream from the stream analyzer;

producing a conditioned data stream and an error stream by, for each of a plurality of data blocks within the data stream:

analyzing the data block within the data stream to compare the data block's real frequency within the data stream against an ideal frequency;

if the difference between the data block's real frequency and ideal frequency exceeds a configured conditioning threshold, applying a conditioning rule to the data block;

applying a logical XOR operation to the data block;

appending the output of the logical XOR operation to the error stream; and

sending the conditioned data stream and the error stream as output.

6. The method of claim 5 , further comprising the steps of:

analyzing the frequency distribution of the plurality of data blocks within the input data stream to determine the most frequent bytes or strings of bytes that occur at the beginning of each of the plurality of data blocks;

designating the most-frequent bytes or strings of bytes as prefixes;

compiling a prefix table based on the results of the frequency distribution, the prefix table comprising the designated prefixes and the block lengths; and

sending the prefix table to a data transformer.

7. The method of claim 6 , further comprising the steps of:

receiving the prefix table using the data transformer;

transforming each of the plurality of prefixes within the prefix table by applying a Burrow's-Wheeler transform (BWT) and generating as output a plurality BWT-prefixes using the data transformer; and

sending the plurality of BWT-prefixes as output using the data transformer.

8. The method of claim 6 , wherein each of the plurality of data blocks are k-mers of the genomic data and the prefixes are one or more base pairs that occur at the beginning of each k-mer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2023
From: COOPER, JOSHUA; RIAHI, ALIASGHAR; HADDAD, MOJGAN; RIAHI, RYAN KOUROSH; RIAHI, RAZMIN; YEOMANS, CHARLES
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 064333/0044 →
Continuity (20)
Continuation In Part 18190044 · Mar 24, 2023
Continuation In Part 17875201 · Jul 27, 2022
Continuation In Part 17727913 · Apr 25, 2022
Continuation 17514913 · Oct 29, 2021
Continuation 17458747 · Aug 27, 2021
Continuation 17727913
Continuation 17404699 · Aug 17, 2021
Continuation In Part 17404699 · Aug 17, 2021
Continuation In Part 16923039 · Jul 7, 2020
Continuation In Part 16716098 · Dec 16, 2019
Continuation In Part 16455655 · Jun 27, 2019
Continuation In Part 16455655 · Jun 27, 2019
Continuation In Part 16200466 · Nov 26, 2018
Continuation In Part 15975741 · May 9, 2018
Provisional Application 63485518 · Feb 16, 2023
Provisional Application 63388411 · Jul 12, 2022
Provisional Application 63027166 · May 19, 2020
Provisional Application 62926723 · Oct 28, 2019
Provisional Application 62578824 · Oct 30, 2017
Related Publication 20230253981A1 · Aug 10, 2023