IP Library Granted Patent US 9,842,301
Granted Patent B2
US 9,842,301 · App. 14/747,922 · Granted Dec 12, 2017

Systems and methods for improved knowledge mining

Inventor: Abhishek Gunjan (Gaya, IN)
Assignee: WIPRO LIMITED
G06N99/005G06F17/2785G06F17/3071G06F17/30734G06N5/02G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,842,301
App. No.
14/747,922
Granted
Dec 12, 2017
Kind
B2
Abstract

This disclosure relates to systems and methods for improved knowledge mining. In one embodiment, a method is disclosed, which comprises filtering aggregated data encoded according to multiple data formats, using a combination of sliding-window and boundary-based filtration techniques. Machine learning and natural language processing are applied to the filtered data to generate a business ontology. Also, using a prediction analysis, one or more recommended classification techniques are automatically identified. The filtered data is clustered into an automatically determined number of categories based on the automatically recommended one or more classification techniques. The one or more classification techniques may utilize iterative feedback between a supervised learning technique and an unsupervised learning technique. Furthermore, the method includes generating automatically correlations between the business ontology and the automatically determined number of categories, and generating a knowledge base using the correlations between the business ontology and the automatically determined number of categories.

Claims (58)

1. A processor-implemented automated knowledge mining method, comprising:

aggregating, via one or more hardware processors, data encoded according to a plurality of data formats;

filtering, via the one or more hardware processors, the aggregated data using a combination of sliding-window and boundary-based filtration techniques to obtain filtered data;

applying, via the one or more hardware processors, machine learning and natural language processing to the filtered data to generate a business ontology;

identifying automatically, via the one or more hardware processors, using a prediction analysis, one or more recommended classification techniques to apply to the filtered data;

clustering, via the one or more hardware processors, the filtered data into an automatically determined number of categories based on the automatically recommended one or more classification techniques;

wherein the one or more classification techniques utilize iterative feedback between a supervised learning technique and an unsupervised learning technique;

generating automatically, via the one or more hardware processors, correlations between the business ontology and the automatically determined number of categories; and

generating, via the one or more hardware processors, a knowledge base using the correlations between the business ontology and the automatically determined number of categories.

2. The method of claim 1 , further comprising:

generating, via the one or more hardware processors, a hierarchical relationship between the categories and clustered data that is clustered within the categories.

3. The method of claim 1 , further comprising:

detecting, via the one or more hardware processors, one or more key terms using a natural language processing technique; and

generating an observation regarding the aggregated data using the detected one or more key terms.

4. The method of claim 3 , further comprising:

detecting, via the one or more hardware processors, an anomaly in the aggregated data based on the generated observation.

5. The method of claim 1 , wherein filtering the aggregated data includes performing a combined time-frequency traffic analysis of the aggregated data.

6. The method of claim 1 , wherein clustering the filtered data includes testing accuracy of the automatically recommended one or more classification techniques.

7. The method of claim 1 , wherein clustering is performed without use of any training related to the automatically recommended one or more classification techniques.

8. The method of claim 1 , wherein a number of iterations for the iterative feedback between the supervised learning technique and the unsupervised learning technique is based on a precision and a recall value associated with clustered data that is clustered within the categories.

9. An automated knowledge mining system, comprising:

one or more hardware processors; and

one or more memory units storing instructions executable by the one or more hardware processors for:

aggregating data encoded according to a plurality of data formats;

filtering the aggregated data using a combination of sliding-window and boundary-based filtration techniques to obtain filtered data;

applying machine learning and natural language processing to the filtered data to generate a business ontology;

identifying automatically, using a prediction analysis, one or more recommended classification techniques to apply to the filtered data;

clustering the filtered data into an automatically determined number of categories based on the automatically recommended one or more classification techniques;

wherein the one or more classification techniques utilize iterative feedback between a supervised learning technique and an unsupervised learning technique;

generating automatically correlations between the business ontology and the automatically determined number of categories; and

generating a knowledge base using the correlations between the business ontology and the automatically determined number of categories.

10. The system of claim 9 , further storing instructions for:

generating a hierarchical relationship between the categories and clustered data that is clustered within the categories.

11. The system of claim 9 , further storing instructions for:

detecting one or more key terms using a natural language processing technique; and

generating an observation regarding the aggregated data using the detected one or more key terms.

12. The system of claim 11 , further storing instructions for:

detecting an anomaly in the aggregated data based on the generated observation.

13. The system of claim 9 , wherein filtering the aggregated data includes performing a combined time-frequency traffic analysis of the aggregated data.

14. The system of claim 9 , wherein clustering the filtered data includes testing accuracy of the automatically recommended one or more classification techniques.

15. The system of claim 9 , wherein clustering is performed without use of any training related to the automatically recommended one or more classification techniques.

16. The system of claim 9 , wherein a number of iterations for the iterative feedback between the supervised learning technique and the unsupervised learning technique is based on a precision and a recall value associated with clustered data that is clustered within the categories.

17. A non-transitory computer-readable medium storing computer-executable automated knowledge mining instructions comprising instructions for:

aggregating data encoded according to a plurality of data formats;

filtering the aggregated data using a combination of sliding-window and boundary-based filtration techniques to obtain filtered data;

applying machine learning and natural language processing to the filtered data to generate a business ontology;

identifying automatically, using a prediction analysis, one or more recommended classification techniques to apply to the filtered data;

clustering the filtered data into an automatically determined number of categories based on the automatically recommended one or more classification techniques;

wherein the one or more classification techniques utilize iterative feedback between a supervised learning technique and an unsupervised learning technique;

generating automatically correlations between the business ontology and the automatically determined number of categories; and

generating a knowledge base using the correlations between the business ontology and the automatically determined number of categories.

18. The medium of claim 17 , further storing instructions for:

generating a hierarchical relationship between the categories and clustered data that is clustered within the categories.

19. The medium of claim 17 , further storing instructions for:

detecting one or more key terms using a natural language processing technique; and

generating an observation regarding the aggregated data using the detected one or more key terms.

20. The medium of claim 19 , further storing instructions for:

detecting an anomaly in the aggregated data based on the generated observation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2023
From: WIPRO LIMITED
To: WORKDAY, INC.
Reel/Frame 062290/0745 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2015
From: GUNJAN, ABHISHEK
To: WIPRO LIMITED
Reel/Frame 035915/0427 →
Priority Claims (1)
IN 1424/CHE/2015 · Mar 20, 2015 · national
Continuity (1)
Related Publication 20160275152A1 · Sep 22, 2016