IP Library Granted Patent US 11,586,970
Granted Patent B2
US 11,586,970 · App. 15/922,983 · Granted Feb 21, 2023

Systems and methods for initial learning of an adaptive deterministic classifier for data extraction

Inventor: Samrat Saha (Bangalore, IN)
Assignee: Wipro Limited
G06N20/00G06F16/35G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,586,970
App. No.
15/922,983
Granted
Feb 21, 2023
Kind
B2
Abstract

This disclosure relates to initial learning of a classifier for automating extraction of structured data from unstructured or semi-structured data. In one embodiment, a method is disclosed, comprising: identifying at least one expected relation class associated with at least one expected relation data; populating at least one expected name entity data from the at least one identified expected relation class; generating training data by tagging the at least one expected relation data and the at least one identified expected relation class with unstructured or semi-structured data; generating feedback data for a relation data and relation class, using a convergence technique on the tagged training data; retuning a NE classifier cluster and a relation classifier cluster by continuously tagging new training data or generating new cascaded expression for a deterministic classifier and a statistical classifier; and extracting the structured data when the NE classifier cluster and the relation classifier cluster converge.

Claims (33)

1. A processing system for data extraction, comprising:

one or more hardware processors;

a memory communicatively coupled to the one or more hardware processors, wherein the memory stores instructions, which, when executed, cause the one or more hardware processors to:

identify at least one expected relation class associated with at least one expected relation data;

populate at least one expected name entity data from the at least one identified expected relation class;

generate training data by tagging the at least one expected relation data and the at least one identified expected relation class with unstructured or semi-structured data;

generate feedback data for a relation data and relation class, using a convergence technique on the tagged training data, wherein the convergence technique for generating the feedback data for the relation data and the relation class uses at least a conditional random field classifier;

retune a NE classifier cluster and a relation classifier cluster based on the feedback data by continuously tagging new training data or generating new cascaded expression for a deterministic classifier and a statistical classifier; and

complete extraction of the structured data when the NE classifier cluster and the relation classifier cluster undergoes convergence through the retuning, wherein the convergence occurs when a composite F-score exceeds a minimum threshold score.

2. The processing system of claim 1 , wherein the identified expected relation class takes as input a triplet of variables comprising a subject, a predicate, and an object.

3. The processing system of claim 1 , wherein at least one statistical relation classifier trainer is used to automate generation of the training data.

4. The processing system of claim 1 , wherein the conditional random field classifier is trained by automatically tagging the training data.

5. The processing system of claim 1 , wherein the convergence is per relation predicate.

6. A hardware processor-implemented method for data extraction, comprising:

identifying, via one or more hardware processors, at least one expected relation class associated with at least one expected relation data;

populating, via the one or more hardware processors, at least one expected name entity data from the at least one identified expected relation class;

generating, via the one or more hardware processors, training data by tagging the at least one expected relation data and the at least one identified expected relation class with unstructured or semi-structured data;

generating, via the one or more hardware processors, feedback data for a relation data and relation class, using a convergence technique on the tagged training data, wherein the convergence technique for generating the feedback data for the relation data and the relation class uses at least a conditional random field classifier;

retuning, via the one or more hardware processors, a NE classifier cluster and a relation classifier cluster based on the feedback data by continuously tagging new training data or generating new cascaded expression for a deterministic classifier and a statistical classifier; and

completing extraction, via the one or more hardware processors, of the structured data when the NE classifier cluster and the relation classifier cluster undergoes convergence through the retuning, wherein the convergence occurs when a composite F-score exceeds a minimum threshold score.

7. The method of claim 6 , wherein the identified expected relation class takes as input a triplet of variables comprising a subject, a predicate, and an object.

8. The method of claim 6 , wherein at least one statistical relation classifier trainer is used to automate generation of the training data.

9. The method of claim 6 , wherein the conditional random field classifier is trained by automatically tagging the training data.

10. The method of claim 6 , wherein the convergence is per relation predicate.

11. A non-transitory, computer-readable medium storing data extraction instructions that, when executed by a hardware processor, cause the hardware processor to:

identify at least one expected relation class associated with at least one expected relation data;

populate at least one expected name entity data from the at least one identified expected relation class;

generate training data by tagging the at least one expected relation data and the at least one identified expected relation class with unstructured or semi-structured data;

generate feedback data for a relation data and relation class, using a convergence technique on the tagged training data, wherein the convergence technique for generating the feedback data for the relation data and the relation class uses at least a conditional random field classifier;

retune a NE classifier cluster and a relation classifier cluster based on the feedback data by continuously tagging new training data or generating new cascaded expression for a deterministic classifier and a statistical classifier; and

complete extraction of the structured data when the NE classifier cluster and the relation classifier cluster undergoes convergence through the retuning, wherein the convergence occurs when a composite F-score exceeds a minimum threshold score.

12. The medium of claim 11 , wherein the identified expected relation class takes as input a triplet of variables comprising a subject, a predicate, and an object.

13. The medium of claim 11 , wherein at least one statistical relation classifier trainer is used to automate generation of the training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2018
From: SAHA, SAMRAT
To: WIPRO LIMITED
Reel/Frame 045245/0160 →
Priority Claims (1)
IN 201841003537 · Jan 30, 2018 · national
Continuity (1)
Related Publication 20190236492A1 · Aug 1, 2019