IP Library › Granted Patent US 11,916,576
Granted Patent B2
US 11,916,576 · App. 17/767,070 · Granted Feb 27, 2024

System and method for effective compression, representation and decompression of diverse tabulated data

Inventors: Shubham Chandak (Menlo Park, CA); Yee Him Cheung (Boston, MA)
Assignee: Koninklijke Philips N.V.
H03M7/6082G06F16/13G06F21/604H03M7/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,916,576
App. No.
17/767,070
Granted
Feb 27, 2024
Kind
B2
Abstract

A method for controlling compression of data includes accessing genomic annotation data in one of a plurality of first file formats, extracting attributes from the genomic annotation data, dividing the genomic annotation data into multiple chunks, and processing the extracted attributes and chunks into correlated information. The method also includes selecting different compressors for the attributes and chunks identified in the correlated information and generating a file in a second file format that includes the correlated information and information indicative of the different compressors for the chunks and attributes indicated in the correlated information. The information indicative of the different compressors is processed into the second file format to allow selective decompression of the attributes and chunks indicated in correlated information.

Claims (43)

1. A method for compression of data, comprising:

accessing tabulated data in at least two of a plurality of first file formats wherein the at least two of a plurality of first file formats are incompatible with one another;

extracting attributes from the tabulated data;

dividing the tabulated data into chunks;

processing the extracted attributes and chunks into correlated information;

selecting different compressors and dependency attributes for the attributes and chunks identified in the correlated information;

storing global data for compressors shared across chunks;

generating a file in a second file format that includes the correlated information, and further includes information indicative of the different compressors being used to compress the chunks and attributes, wherein the second file format is different from the first file formats; and

generating access control policy information for the correlated information, and integrating the access control policy information into the file of the second file format, wherein the access control policy information includes first information indicating a first level of access for a first portion of the correlated information and second information indicating a second level of access for a second portion of the correlated information, wherein the second level of access is different from a first level of access.

2. The method of claim 1 , wherein the correlated information includes at least one table including:

first information indicative of one or more of the attributes, and second information corresponding to chunks associated with the one or more attributes indicated in the first information.

3. The method of claim 2 , wherein the first information includes a two-dimensional array of cells and wherein each cell identifies one or more corresponding attributes included in the chunks corresponding to the two-dimensional array of cells.

4. The method of claim 3 , wherein the first information includes at least one one-dimensional table of dimension-specific attributes relating to the chunks.

5. The method of claim 1 , wherein code of or an executable of at least one of the different compressors is embedded in the file.

6. The method of claim 1 , wherein at least one of the different compressors encodes one or more attributes of the tabulated data as a sparse array.

7. The method of claim 1 , wherein the tabulated data is divided into chunks of different sizes.

8. The method of claim 1 , further comprising:

generating one or more attribute-specific indexes; and

incorporating the one or more attribute-specific indexes into the file of the second format, wherein the attribute-specific indexes enable identification of chunks containing a value or range of a specific attribute of the tabulated data.

9. The method of claim 1 , further comprising:

integrating the file in the second file format into an MPEG-G file.

10. The method of claim 1 , wherein code of or an executable of at least one of the different compressors is embedded in the file.

11. The method of claim 1 , wherein at least one of the different compressors encodes one or more attributes of the tabulated data as a sparse array.

12. An apparatus for compression of data, comprising:

a memory configured to store instructions; and

at least one processor configured to execute the instructions to perform operations including:

accessing tabulated data in at least two of a plurality of first file formats, wherein the at least two of a plurality of first file formats are incompatible with one another;

extracting attributes from the tabulated data;

dividing the tabulated data into chunks;

processing the extracted attributes and chunks into correlated information;

selecting different compressors and dependency attributes for the attributes and chunks identified in the correlated information;

storing global data for compressors shared across chunks; and

generating a file in a second file format that includes the correlated information and further including information indicative of the different compressors used for compressing the chunks and attributes, wherein the second file format is different from the first file formats; and

generating access control policy information for the correlated information, and integrate the access control policy information into the file of the second file format, wherein the access control policy information includes first information indicating a first level of access for a first portion of the correlated information and second information indicating a second level of access for a second portion of the correlated information, and wherein the second level of access is different from a first level of access.

13. The apparatus of claim 12 , wherein the correlated information includes at least one table including:

first information indicative of one or more of the attributes, and second information corresponding to chunks associated with the one or more attributes indicated in the first information.

14. The apparatus of claim 13 , wherein the first information includes a two-dimensional array of cells and wherein each cell identifies one or more corresponding attributes included in the chunks corresponding to the two-dimensional array of cells.

15. The apparatus of claim 14 , wherein the first information includes at least one one-dimensional table of dimension-specific attributes relating to the chunks.

16. The apparatus of claim 12 , wherein the tabulated data is divided into chunks of different sizes.

17. The apparatus of claim 12 , wherein the at least one processor is configured to execute the instructions to:

generate one or more attribute-specific indexes; and

incorporate the one or more attribute-specific indexes into the file of the second format, wherein the attribute-specific indexes enable identification of chunks containing a value or range of a specific attribute of the tabulated data.

18. The apparatus of claim 12 , wherein the at least one processor is configured to execute the instructions to integrate the file in the second file format into an MPEG-G file.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2022
From: CHANDAK, SHUBHAM; CHEUNG, YEE HIM
To: KONINKLIJKE PHILIPS N.V.
Reel/Frame 059526/0342 →
Continuity (3)
Provisional Application 62956952 · Jan 3, 2020
Provisional Application 62923141 · Oct 18, 2019
Related Publication 20220368347A1 · Nov 17, 2022