IP Library › Granted Patent US 10,866,994
Granted Patent B2
US 10,866,994 · App. 15/188,769 · Granted Dec 15, 2020

Systems and methods for instant crawling, curation of data sources, and enabling ad-hoc search

Inventor: Ramesh Panuganty (Cupertino, CA)
Assignee: SPLUNK INC.
G06F16/951
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,866,994
App. No.
15/188,769
Filed
Jun 21, 2016
Granted
Dec 15, 2020
Kind
B2
Art Unit
2167
USPC
707/710
Abstract

Improved crawling and curation of data and metadata from diverse data sources is described. In some embodiments, improvements are achieved by interpreting the context, vocabulary and relationships of data element, to enable relational data search capability for users. The user querying process is improved by systematic identification of the data objects, context, and relationships across data objects and elements, aggregation methods and operators on the data objects and data elements as identified in the curation process. User query suggestions and recommendations can be adjusted based on the context, relationships between the data elements, user profile, and the data sources. When the user query is executed, the query text is translated into an equivalent of one or more query statements, such as SQL or PostGre statements, and the query is performed on the identified data sources. Results are assembled to present the answer in a meaningful visualization for the user query.

Claims (48)

1. A computer-implemented method comprising:

analyzing one or more data sources to provide a context for information contained in the data sources;

grouping data included in the data sources into sets of related data entities;

analyzing each set of related data entities to attribute a characteristic to the respective set of related data entities;

analyzing attributed characteristics among the sets of related data entities to identify logical relationships between the sets of related data entities;

generating a relationship between two or more of the sets of related data entities by interpreting a first attribute of the two or more of the sets based on a first probability that indicates a likelihood that the first attribute represents a first type of characteristic and a second probability that indicates a likelihood that the first attribute represents a second type of characteristic; and

in accordance with at least the identified logical relationships and the relationship, enabling a natural language search to be conducted using the sets of data entities.

2. A method as described in claim 1 , wherein at least one data source comprises a public data source.

3. A method as described in claim 1 , wherein at least one data source comprises a data source other than a public data source.

4. A method as described in claim 1 , wherein said analyzing one or more data sources comprises analyzing the data sources based on a data source name, a sub-name, frequency of use, access restrictions, or data format.

5. A method as described in claim 1 , wherein the sets of related data entities comprise columns.

6. A method as described in claim 1 , wherein the characteristic includes one or more of names or a characteristic associated with a numeric value.

7. A method as described in claim 1 , wherein the sets of related data entities comprise columns, and wherein interpreting the first attribute comprises interpreting attributes of adjacent columns.

8. A method as described in claim 1 further comprising analyzing randomness of data in each set of data entities to classify the data as finite or infinite.

9. A method as described in claim 1 , wherein the sets of related data entities comprise columns, and wherein interpreting the first attribute comprises interpreting attributes of adjacent columns, and further comprising analyzing randomness of data in each set of data entities to classify the data as finite or infinite.

10. A method as described in claim 1 , wherein generating the relationship between two or more of the sets of related data entities by interpreting a first attribute of the two or more of the sets comprises:

identifying a first candidate relationship between the two or more data entities and determining the first probability for the first candidate relationship;

identifying a second candidate relationship between the two or more data entities and determining the second probability for the second candidate relationship;

adjusting the first probability and the second probability in accordance with searches and acceptance of search results by users; and

in accordance with a higher of the first probability and the second probability, generating as the relationship between the two or more data entities one of the first candidate relationship and the second candidate relationship.

11. One or more non-transitory computer-readable storage media storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising:

analyzing one or more data sources to provide a context for information contained in the data sources;

grouping data included in the data sources into sets of related data entities;

analyzing each set of related data entities to attribute a characteristic to the respective set of related data entities;

analyzing attributed characteristics among the sets of related data entities to identify logical relationships between the sets of related data entities;

generating a relationship between two or more of the sets of related data entities by interpreting a first attribute of the two or more of the sets based on a first probability that indicates a likelihood that the first attribute represents a first type of characteristic and a second probability that indicates a likelihood that the first attribute represents a second type of characteristic; and

in accordance with at least the identified logical relationships and the relationship, enabling a natural language search to be conducted using the sets of data entities.

12. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein at least one data source comprises a data source other than a public data source.

13. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein said analyzing one or more data sources comprises analyzing the data sources based on a data source name, a sub-name, frequency of use, access restrictions, or data format.

14. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein the sets of related data entities comprise columns.

15. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein the characteristic includes one or more of names or a characteristic associated with a numeric value.

16. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein the sets of related data entities comprise columns, and wherein interpreting the first attribute comprises interpreting attributes of adjacent columns.

17. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein the operations further comprise analyzing randomness of data in each set of data entities to classify the data as finite or infinite.

18. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein the sets of related data entities comprise columns, and wherein interpreting the first attribute comprises interpreting attributes of adjacent columns, and wherein the operations further comprise analyzing randomness of data in each set of data entities to classify the data as finite or infinite.

19. The one or more non-transitory computer-readable storage media as described in claim 11 , wherein generating the relationship between two or more of the sets of related data entities by interpreting a first attribute of the two or more of the sets comprises:

identifying a first candidate relationship between the two or more data entities and determining the first probability for the first candidate relationship;

identifying a second candidate relationship between the two or more data entities and determining the second probability for the second candidate relationship;

adjusting the first probability and the second probability in accordance with searches and acceptance of search results by users; and

in accordance with a higher of the first probability and the second probability, generating as the relationship between the two or more data entities one of the first candidate relationship and the second candidate relationship.

20. A system, comprising:

a memory including instructions; and

a processor that is coupled to the memory and, when executing the instructions, is configured to perform the steps of:

analyzing one or more data sources to provide a context for information contained in the data sources;

grouping data included in the data sources into sets of related data entities;

analyzing each set of related data entities to attribute a characteristic to the respective set of related data entities;

analyzing attributed characteristics among the sets of related data entities to identify logical relationships between the sets of related data entities;

generating a relationship between two or more of the sets of related data entities by interpreting a first attribute of the two or more of the sets based on a first probability that indicates a likelihood that the first attribute represents a first type of characteristic and a second probability that indicates a likelihood that the first attribute represents a second type of characteristic; and

in accordance with at least the identified logical relationships and the relationship, enabling a natural language search to be conducted using the sets of data entities.

Assignments (4)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2017
From: DRASTIN, INC.
To: SPLUNK INC.
Reel/Frame 042878/0925 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2016
From: PANUGANTY, RAMESH
To: DRASTIN, INC.
Reel/Frame 039020/0501 →
Continuity (2)
Provisional Application 62183194 · Jun 23, 2015
Related Publication 20160378867A1 · Dec 29, 2016