IP Library Granted Patent US 10,043,138
Granted Patent B2
US 10,043,138 · App. 14/869,074 · Granted Aug 7, 2018

Metadata representation and storage

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,043,138
App. No.
14/869,074
Granted
Aug 7, 2018
Kind
B2
Abstract

At least one original data set is obtained. Header type metadata is extracted from the original data set and the extracted header type metadata is stored in a document-oriented database. Content type metadata is extracted from the original data set and the extracted content type metadata is stored in a table-structured database. The original data set is stored in a data store. The document-oriented database comprises one or more links to access the content type metadata in the table-structured database and the original data set in the data store. By way of example only, the data storage techniques may be used for bioinformatics applications.

Claims (42)

1. A method comprising:

obtaining at least one original data set;

extracting header type metadata from the original data set and storing the extracted header type metadata in a document-oriented database implemented on a first processing element, wherein the header type metadata comprises identifying data associated with the original data set;

extracting content type metadata from the original data set and storing the extracted content type metadata in a table-structured database implemented on a second processing element, wherein the content type metadata comprises one or more sets of data generated by analysis of the original data set, and wherein the first processing element is operatively coupled to the second processing element via a communication network;

storing the original data set in a data store implemented on a third processing device, wherein the first processing element and the second processing element are operatively coupled to the third processing element via the communication network;

wherein the document-oriented database comprises one or more links to access the content type metadata in the table-structured database and the original data set in the data store via the communication network operatively coupling the first processing element, the second processing element, and the third processing element;

receiving a query to access one or more of at least a portion of the original data set, at least a portion of the header type metadata, and at least a portion of the content type metadata; and

processing the query through the document-oriented database and accessing the content type metadata in the table-structured database and the original data set in the data store using the one or more links from the document-oriented database.

2. The method of claim 1 , wherein one or more of the extracting steps comprises a parser configured to automatically determine a schema associated with the extracted metadata.

3. The method of claim 2 , wherein the schema is automatically determined by the parser by importing a schema.

4. The method of claim 2 , wherein the schema is automatically determined by the parser by applying a machine learning algorithm.

5. The method of claim 2 , wherein the parser is configured to read data formats associated with the original data set, decompress metadata associated with the original data set, and separate the metadata into the header type metadata and the content type metadata.

6. The method of claim 1 , wherein the header type metadata is stored in the document-oriented database as one or more objects.

7. The method of claim 1 , wherein the content type metadata is searchable via structured query language type queries.

8. The method of claim 1 , wherein the original data set is not parsed before storing in the data store.

9. The method of claim 1 , wherein the one or more links in the document-based database used to access the content type metadata in the table-structured database and the original data set in the data store comprise one or more identifiers that correspond to data stored in the table-structured database and the data store.

10. The method of claim 1 , wherein the original data set comprises biological data.

11. The method of claim 10 , wherein the extracted header type metadata comprises one or more of: data identifying the particular mechanism used to obtain the biological data; data identifying the one or more biological species from which the biological data was taken; a file name of a data file in which the sequence of biological data is stored; and at least one of a creation data and a modification date of the data file.

12. The method of claim 10 , wherein the extracted header type metadata comprises header information from a sequence alignment map file.

13. The method of claim 10 , wherein the extracted content type metadata comprises one or more of: reads generated from the biological data; and variants generated from the biological data.

14. The method of claim 10 , wherein the content type metadata comprises reference-aligned reads from a sequence alignment map file.

15. An article of manufacture comprising a processor-readable storage medium having encoded therein executable code of one or more software programs, wherein the one or more software programs when executed by one or more processing devices implement steps of:

obtaining at least one original data set;

extracting header type metadata from the original data set and storing the extracted header type metadata in a document-oriented database implemented on a first processing element, wherein the header type metadata comprises identifying data associated with the original data set;

extracting content type metadata from the original data set and storing the extracted content type metadata in a table-structured database implemented on a second processing element, wherein the content type metadata comprises one or more sets of data generated by analysis of the original data set, and wherein the first processing element is operatively coupled to the second processing element via a communication network;

storing the original data set in a data store implemented on a third processing device, wherein the first processing element and the second processing element are operatively coupled to the third processing element via the communication network;

wherein the document-oriented database comprises one or more links to access the content type metadata in the table-structured database and the original data set in the data store via the communication network operatively coupling the first processing element, the second processing element, and the third processing element;

receiving a query to access one or more of at least a portion of the original data set, at least a portion of the header type metadata, and at least a portion of the content type metadata; and

processing the query through the document-oriented database and accessing the content type metadata in the table-structured database and the original data set in the data store using the one or more links from the document-oriented database.

16. A data storage system comprising:

one or more processors operatively coupled to one or more memories and configured to:

obtain at least one original data set;

extract header type metadata from the original data set and storing the extracted header type metadata in a document-oriented database implemented on a first processing element, wherein the header type metadata comprises identifying data associated with the original data set;

extract content type metadata from the original data set and storing the extracted content type metadata in a table-structured database implemented on a second processing element, wherein the content type metadata comprises one or more sets of data generated by analysis of the original data set, and wherein the first processing element is operatively coupled to the second processing element via a communication network;

store the original data set in a data store implemented on a third processing device, wherein the first processing element and the second processing element are operatively coupled to the third processing element via the communication network;

wherein the document-oriented database comprises one or more links to access the content type metadata in the table-structured database and the original data set in the data store via the communication network operatively coupling the first processing element, the second processing element, and the third processing element;

receive a query to access one or more of at least a portion of the original data set, at least a portion of the header type metadata, and at least a portion of the content type metadata; and

process the query through the document-oriented database and accessing the content type metadata in the table-structured database and the original data set in the data store using the one or more links from the document-oriented database.

17. The system of claim 16 , wherein the one or more processors are further configured to perform one or more of the extracting steps using a parser configured to automatically determine a schema associated with the extracted metadata.

18. The system of claim 17 , wherein the schema is automatically determined by the parser by at least one of: importing a schema; and applying a machine learning algorithm.

19. The system of claim 17 , wherein the parser is configured to read data formats associated with the original data set, decompress metadata associated with the original data set, and separate the metadata into the header type metadata and the content type metadata.

20. The system of claim 16 , wherein the original data set comprises biological data.

Assignments (5)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2017
From: EMC CORPORATION
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 041872/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2015
From: SUVOROV, VLADIMIR ALEXANDROVICH
To: EMC CORPORATION
Reel/Frame 036682/0508 →