IP Library › Granted Patent US 12,216,623
Granted Patent B2
US 12,216,623 · App. 18/822,208 · Granted Feb 4, 2025

System and method for random-access manipulation of compacted data files

Inventors: Joshua Cooper (Columbia, SC); Charles Yeomans (Orinda, CA); Brian Galvin (Silverdale, WA)
Assignee: ATOMBEAM TECHNOLOGIES INC
G06F16/1752G06F3/0608G06F3/0641G06F3/067
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,216,623
App. No.
18/822,208
Filed
Sep 1, 2024
Granted
Feb 4, 2025
Kind
B2
Art Unit
2136
USPC
707/692
Abstract

A system and method for random-access manipulation of compacted data files, utilizing a reference codebook, a random-access engine, a data deconstruction engine, and a data deconstruction engine. The system may receive a data query pertaining to a data read or data write request, wherein the data file to be read from or written to is a compacted data file. A random-access engine may facilitate data manipulation processes by transforming the codebook into a hierarchical representation and then traversing the representation scanning for specific codewords associated with a data query request. In an embodiment, an estimator module is present and configured to utilize cardinality estimation to determine a starting codeword to begin searching the compacted data file for the data associated with the data query. The random-access engine may encode the data to be written, insert the encoded data into a compacted data file, and update the codebook as needed.

Claims (55)

1. A system for random-access manipulation of a compacted data file, comprising:

a computing device comprising a memory, a processor, and a non-volatile data storage device;

a random access engine comprising a first plurality of programming instructions that, when operating on the processor, cause the computing device to:

receive a data search query directed to the compacted data file, wherein the compacted file is compacted using a reference codebook;

organize the reference codebook into a hierarchical representation;

traverse the hierarchical representation to identify a start codeword corresponding to the beginning of the data search query; and

send the start codeword and a plurality of immediately following codewords from the compacted data file to a decoder; and

an estimator module comprising a second plurality of programming instructions that, when operating on the processor, cause the computing device to:

estimate a first starting bit location in the compacted data file;

refine the first starting bit location by:

determining a plurality of codeword boundaries by performing distinct value estimation on the compacted data file, wherein the distinct value estimates correspond to a codeword boundary; and

determining whether a bit sequence starting at the first starting bit location corresponds to a codeword boundary of the plurality of codeword boundaries and, if not, traversing the hierarchical representation until a codeword boundary is located at a new starting bit.

2. The system of claim 1 , further comprising a dyadic distribution compression and encryption subsystem comprising a third plurality of programming instructions that, when operating on the processor, cause the computing device to:

analyze input data to determine its properties;

create a transformation matrix based on the properties of the input data;

transform the input data into a dyadic distribution;

generate a main data stream of transformed data and a secondary data stream of transformation information; and

compress the main data stream.

3. The system of claim 2 , further comprising a large codeword model with a latent transformer core comprising a fourth plurality of programming instructions that, when operating on the processor, cause the computing device to:

receive the compressed data as input vectors;

generate latent space vectors by processing the input vectors through a variational autoencoder's encoder;

process the latent space vectors through a transformer to learn relationships between the vectors, wherein the transformer does not include an embedding layer and a positional encoding layer; and

decode output latent space vectors through a variational autoencoder's decoder to produce final output data.

4. The system of claim 1 , wherein the random access engine uses the refined starting bit location as an initial point to begin traversing the hierarchical representation.

5. The system of claim 4 , further comprising an end to end training subsystem configured to jointly train the random access engine, the dyadic distribution compression and encryption module, the large codeword model, and the estimator module.

6. The system of claim 5 , wherein the end to end training subsystem is configured to optimize performance across all components from initial compression and encryption to final random access capabilities.

7. The system of claim 2 , wherein the dyadic distribution compression and encryption subsystem uses trainable parameters for transforming the input data and compressing the main data stream.

8. The system of claim 3 , wherein the large codeword model is configured to adapt its processing based on the output of the dyadic distribution compression and encryption module.

9. The system of claim 4 , wherein the estimator module uses learnable parameters to estimate the starting bit location and refine the estimate.

10. A method for random-access manipulation of a compacted data file, comprising the steps of:

receiving a data search query directed to the compacted data file, wherein the compacted file is compacted using a reference codebook;

organizing the reference codebook into a hierarchical representation;

traversing the hierarchical representation to identify a start codeword corresponding to the beginning of the data search query;

sending the start codeword and a plurality of immediately following codewords from the compacted data file to a decoder;

estimating, using an estimator module, a first starting bit location in the compact data file; and

refining the first starting bit location by:

determining a plurality of codeword boundaries by performing distinct value estimation on the compacted data file, wherein the distinct value estimates correspond to a codeword boundary; and

determining whether a bit sequence starting at the first starting bit location corresponds to a codeword boundary of the plurality of codeword boundaries and, if not, traversing the hierarchical representation until a codeword boundary is located at a new starting bit.

11. The method of claim 10 , further comprising:

analyzing input data to determine its properties;

creating a transformation matrix based on the properties of the input data;

transforming the input data into a dyadic distribution;

generating a main data stream of transformed data and a secondary data stream of transformation information; and

compressing the main data stream.

12. The method of claim 10 , further comprising:

receiving the compressed data as input vectors;

generating latent space vectors by processing the input vectors through a variational autoencoder's encoder;

processing the latent space vectors through a transformer to learn relationships between the vectors, wherein the transformer does not include an embedding layer and a positional encoding layer; and

decoding output latent space vectors through a variational autoencoder's decoder to produce final output data.

13. The method of claim 10 , wherein traversing the hierarchical representation uses the refined starting bit location as an initial point.

14. The method of claim 11 , further comprising jointly training, in an end-to-end manner, the steps of organizing, traversing, and sending along with the steps of analyzing, creating, transforming, generating, compressing, and estimating.

15. The method of claim 14 , wherein the joint training optimizes performance across all steps from initial compression and encryption to final random access capabilities.

16. The method of claim 12 , wherein transforming the input data and compressing the main data stream use trainable parameters.

17. The method of claim 12 , wherein processing the latent space vectors adapts based on the output of the compression step.

18. The method of claim 13 , wherein estimating the first starting bit location and refining the estimate use learnable parameters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2024
From: COOPER, JOSHUA; YEOMANS, CHARLES; GALVIN, BRIAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 069044/0699 →
Continuity (14)
Continuation In Part 18412439 · Jan 12, 2024
Continuation In Part 18078909 · Dec 9, 2022
Continuation 17734052 · Apr 30, 2022
Continuation 17180439 · Feb 19, 2021
Continuation In Part 16923039 · Jul 7, 2020
Continuation In Part 16716098 · Dec 16, 2019
Continuation 16455655 · Jun 27, 2019
Continuation In Part 16200466 · Nov 26, 2018
Continuation In Part 15975741 · May 9, 2018
Provisional Application 63140111 · Jan 21, 2021
Provisional Application 63027166 · May 19, 2020
Provisional Application 62926723 · Oct 28, 2019
Provisional Application 62578824 · Oct 30, 2017
Related Publication 20240427739A1 · Dec 26, 2024
References Cited (9)
US 5408234A · Chu · 1995 [cited by applicant]
US 7154416B1 · Savage · 2006 [cited by applicant]
US 8453040B2 · Henderson, Jr. et al. · 2013 [cited by applicant]
US 9294589B2 · Crosta et al. · 2016 [cited by applicant]
US 9524392B2 · Naehrig et al. · 2016 [cited by applicant]
US 9727255B2 · Matsushita · 2017 [cited by applicant]
US 10255315B2 · Kalevo et al. · 2019 [cited by applicant]
US 20180196609A1 · Niesen · 2018 [cited by applicant]
US 20200395955A1 · Choi et al. · 2020 [cited by applicant]