IP Library Granted Patent US 10,901,961
Granted Patent B2
US 10,901,961 · App. 15/716,447 · Granted Jan 26, 2021

Systems and methods for generating schemas that represent multiple data sources

Inventors: Rick Morrison (Palo Alto, CA); Matthew Saffer (Palo Alto, CA); Jud Gardner (Palo Alto, CA)
Assignee: SAAMA TECHNOLOGIES, INC.
G06F16/211G06F16/25G06F16/2468G06F16/256G06F16/835G06F16/951
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,901,961
App. No.
15/716,447
Granted
Jan 26, 2021
Kind
B2
Abstract

Systems and methods generating schemas that represent multiple data sources are provided herein. According to some embodiments, methods may include determining a schema for each of the multiple data sources via a computing device communicatively couplable with each of the multiple data sources, each of the multiple data sources including one or more data structures that define how data is stored in the data source, generating a negotiated schema by comparing the schemas of the multiple data sources to one another and interrelating data points of the multiple data sources based upon the schemas, interrelating the negotiated schema with the schema for each of the multiple data sources based upon the interrelation of the data points, and storing the negotiated schema in a storage media by way of the computing device.

Claims (55)

1. A method, comprising:

extracting data blobs from a plurality of compressed data sources, wherein the data blobs have a blob data structure;

evaluating the data blobs for a data structure for each of the plurality of data sources;

identifying a schema template for each of the plurality of data sources by executing schema template matching based on at least one machine learning heuristic using the data structure of each of the plurality of data sources;

generating a negotiated schema from the matching schema templates, the negotiated schema directly linking data points between the matching schema templates and indirectly linking data points by applying a transform to the indirectly linked data points;

receiving a query;

executing the query against the plurality of data sources using the negotiated schema; and

assembling a response to the query from data obtained from the plurality of data sources using the directly and indirectly linking of the negotiated schema.

2. The method according to claim 1 , wherein a schema template comprises a representation of a data structure of a data source.

3. The method according to claim 2 , further comprising determining a fuzzy template match between the data structure and one or more of a plurality of schema templates.

4. The method according to claim 3 , further comprising selecting a schema template representing the data structure that is a fuzzy match between the data structure and a schema template of one or more of the plurality of schema templates.

5. The method according to claim 4 , further comprising:

comparing the plurality of matching schema templates to one another, the plurality of matching schema templates including the schema template representing that is the fuzzy template match; and

interrelating data points of the plurality of data sources that correspond to one another using the plurality of matching schema templates; and

storing the negotiated schema in a storage media by way of a computing device.

6. The method according to claim 5 , wherein the step of interrelating data points of the plurality of data sources includes establishing at least one of a fuzzy or a concrete relationship between data points.

7. The method according to claim 5 , wherein the step of interrelating data points of the plurality of data sources includes associating related data points with metadata that describe an interrelationship between the data points.

8. The method according to claim 7 , wherein the plurality of data sources comprise metadata, the metadata may include any of a data attribute, a schema information for each data source, and a confidence level for interrelated sets of data points.

9. The method according to claim 8 , further comprising receiving verification from an end user that the interrelationship between data points is correct.

10. The method according to claim 1 , automatically updating the negotiated schema when the data structure changes.

11. The method according to claim 1 , further comprising at least one of:

selecting one or more alternative data sources when one or more of the plurality of data sources are unavailable; and

marking metadata in the response appropriately if no alternative data source is available.

12. The method according to claim 1 , further comprising decompressing and extracting the data blobs from a compressed file.

13. The method according to claim 1 , wherein at least a portion of the plurality of data sources is proprietary.

14. A system comprising:

means for extracting data blobs from a plurality of compressed data sources, wherein the data blobs have a blob data structure;

means for evaluating the data blobs for a data structure for each of the plurality of data sources;

means for identifying a schema template for each of the plurality of data sources by executing schema template matching based on at least one machine learning heuristic using the data structure of each of the plurality of data sources;

means for generating a negotiated schema from the matching schema templates, the negotiated schema directly linking data points between the matching schema templates and indirectly linking data points by applying a transform to the indirectly linked data points; and

means for:

receiving a query;

executing the query against the plurality of data sources using the negotiated schema; and

assembling a response to the query from data obtained from the plurality of data sources using the linking of the negotiated schema.

15. The system according to claim 14 , wherein a schema template comprises a representation of a data structure of a data source.

16. The system according to claim 15 , further comprising:

means for determining a fuzzy template match between the data structure and one or more of a plurality of schema templates; and

means for selecting a schema template representing the data structure that is a fuzzy match between the data structure and a schema template of one or more of the plurality of schema templates.

17. A system comprising:

a processor; and

a memory for storing instructions, the instructions comprising:

an interrogation module configured to:

extract data blobs from a plurality of compressed data sources, wherein the data blobs have a blob data structure;

evaluate the data blobs for a data structure for each of the plurality of data sources; and

identify a schema template by performing schema template matching, for each of the plurality of data sources, based on at least one machine learning heuristic using the data structure of each of the plurality of data sources;

a schema generator configured to:

generate a negotiated schema from the matching schema templates, the negotiated schema directly linking data points between the matching schema templates and indirectly linking data points by applying a transform to the indirectly linked data points; and

a query engine configured to:

receive a query;

execute the query against the plurality of data sources using the negotiated schema; and

assemble a response to the query from data obtained from the plurality of data sources using the linking of the negotiated schema.

18. The system according to claim 17 , wherein a schema template comprises a representation of a data structure of a data source.

19. The system according to claim 18 , wherein the interrogation module is further configured to determine a fuzzy template match between the data structure and one or more of a plurality of schema templates.

20. The system according to claim 19 , wherein the interrogation module is further configured to select a schema template representing the data structure that is a fuzzy match between the data structure and a schema template of one or more of the plurality of schema templates.

21. The system according to claim 20 , wherein the schema generator is further configured to store the negotiated schema in a storage media by way of a computing device.

Assignments (5)
SECURITY INTEREST Recorded Jun 30, 2023
From: SAAMA TECHNOLOGIES, LLC
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 064127/0314 →
ENTITY CONVERSION Recorded Jun 29, 2023
From: SAAMA TECHNOLOGIES, INC.
To: SAAMA TECHNOLOGIES, LLC
Reel/Frame 064165/0578 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PROPERTY NUMBER PREVIOUSLY RECORDED AT REEL: 50117 FRAME: 017. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 22, 2019
From: COMPREHEND SYSTEMS, INC.
To: SAAMA TECHNOLOGIES, INC.
Reel/Frame 050139/0612 →
MERGER Recorded Aug 21, 2019
From: COMPREHEND SYSTEMS, INC.
To: SAAMA TECHNOLOGIES, INC.
Reel/Frame 050117/0017 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2017
From: MORRISON, RICK; GARDNER, JUD; SAFFER, MATTHEW
To: COMPREHEND SYSTEMS, INC.
Reel/Frame 043813/0183 →
Continuity (3)
Continuation 14667272 · Mar 24, 2015
Continuation 13251149 · Sep 30, 2011
Related Publication 20180018353A1 · Jan 18, 2018