IP Library › Granted Patent US 11,720,649
Granted Patent B2
US 11,720,649 · App. 16/458,520 · Granted Aug 8, 2023

System and method for classification of data in a machine learning system

Inventors: Niraj Kunnumma (Bangalore, IN); Rajeshwari Ganesan (Palo Alto, CA); Bhavana Bhasker (San Jose, CA)
Assignee: EDGEVERVE SYSTEMS LIMITED
G06F18/217G06F17/18G06F18/24G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,720,649
App. No.
16/458,520
Granted
Aug 8, 2023
Kind
B2
Abstract

Disclosed are a system, method and apparatus for classification of data in a machine learning system. In one aspect, a method for classification of data in a machine learning system through one or more computer processors is disclosed. Further, generating, through one or more computer processors, a data classifier using a first dataset and determining an accuracy value of the data classifier to achieve a predefined model accuracy threshold. Still further, iterating, through one or more computer processors, calibration of the first dataset based on a set of parameters until the accuracy value matches or exceeds the predefined model accuracy threshold value. Further, the calibration comprises a user input to indicate a correctness of a presented subset of data from a second dataset and using the above to generate an enhanced data classifier for the classification of data.

Claims (42)

1. A method for data classification in a machine learning system, the method comprising:

generating, by a computing device, using training data a data classifier for a first dataset;

and

iterating, by the computing device, calibration of the data classifier based on a set of one or more parameters until an accuracy value for the data classifier matches or exceeds a predefined model accuracy threshold value, wherein the calibration comprises:

receiving a user input comprising an annotation of a presented subset of the first dataset to thereby disambiguate the presented subset of data, wherein the annotation comprises an indication of correctness;

generating using the disambiguated subset of data and the training data an enhanced version of the data classifier; and

determining the accuracy value based on an application of the enhanced version of the data classifier on at least another subset of the first dataset.

2. The method of claim 1 , wherein the set of parameters further comprises a certainty score, a click utilization value, a certainty threshold value, a model accuracy value, an average machine learning time, or an average annotation time.

3. The method of the claim 1 , wherein the set of parameters comprises at least a certainty threshold value and a number of candidates included in the presented subset of data is determined based on the certainty threshold value.

4. The method of claim 2 , wherein the user input received over the presented subset of data is used to generate the click utilization value or the average annotation time.

5. The method of claim 1 , further comprising generating, by the computing device, the presented subset of data based on a plurality of clusters created based on computed candidate vector distances.

6. The method of claim 1 , wherein the first dataset comprises a plurality of candidates, the set of parameters comprises a certainty threshold value, and the method further comprises:

generating, by the computing device, a certainty score for each of the plurality of candidates in a current iteration of the calibration;

comparing, by the computing device, the certainty score to the certainty threshold value; and

determining, by the computing device, based on the comparison a number of the plurality of candidates included in the presented subset of data in a subsequent iteration of the calibration.

7. The method of claim 6 , further comprising:

generating, by the computing device, a click utilization value or an average annotation time based on the user input, wherein the click utilization value comprises a percentage of the plurality of candidates that are annotated by the user in the current iteration of the calibration; and

adjusting, by the computing device, for the subsequent iteration of the calibration the certainty threshold value based on the click utilization value or the average annotation time.

8. A data classification system comprising:

a processor; and

a memory coupled to the processor and comprising programmed instructions stored thereon that, when executed by the processor, are configured to cause the processor to:

generate using training data a data classifier for a first dataset;

and

iterate calibration of the data classifier based on a set of one or more parameters until an accuracy value for the data classifier matches or exceeds a predefined model accuracy threshold value, wherein the calibration comprises:

receiving a user input comprising an annotation of a presented subset of the first dataset to thereby disambiguate the presented subset of data, wherein the annotation comprises an indication of correctness;

generating using the disambiguated subset of data and the training data an enhanced version of the data classifier; and

determining the accuracy value based on an application of the enhanced version of the data classifier on at least another subset of the first dataset.

9. The system of claim 8 , wherein the set of parameters further comprises a certainty score, a click utilization value, a certainty threshold value, a model accuracy value, an average machine learning time, or an average annotation time.

10. The system of the claim 8 , wherein the set of parameters comprises at least a certainty threshold value and a number of candidates included in the presented subset of data is determined based on the certainty threshold value.

11. The system of claim 9 , wherein the programmed instructions, when executed by the processor, are further configured to cause the processor to generate the click utilization value or the average annotation time based on the user input received over the presented subset of data.

12. The system of claim 8 wherein the programmed instructions, when executed by the processor, are further configured to cause the processor to generate the presented subset of data based on a plurality of clusters created based on computed candidate vector distances.

13. A non-transitory computer readable medium including instruction for data classification stored thereon that when executed by at least one processor cause the at least one processor to:

generate using training data a data classifier for a first dataset;

and

iterate calibration of the data classifier based on a set of one or more parameters until an accuracy value for the data classifier matches or exceeds a predefined model accuracy threshold value, wherein the calibration comprises:

receiving a user input comprising an annotation of a presented subset of the first dataset to thereby disambiguate the presented subset of data, wherein the annotation comprises an indication of correctness;

generating using the disambiguated subset of data and the training data an enhanced version of the data classifier; and

determining the accuracy value based on an application of the enhanced version of the data classifier on at least another subset of the first dataset.

14. The non-transitory computer readable medium of claim 13 , wherein the set of parameters further comprises a certainty score, a click utilization value, a certainty threshold value, a model accuracy value, an average machine learning time, or an average annotation time.

15. The non-transitory computer readable medium of the claim 13 , wherein the set of parameters comprises at least a certainty threshold value and a number of candidates included in the presented subset of data is determined based on the certainty threshold value.

16. The non-transitory computer readable medium of claim 14 , wherein the instructions, when executed by the processor, further cause the processor to generate the click utilization value or the average annotation time based on the user input received over the presented subset of data.

17. The non-transitory computer readable medium of claim 13 wherein the instructions, when executed by the processor, further cause the processor to generate the presented subset of data based on a plurality of clusters created based on computed candidate vector distances.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2019
From: KUNNUMMA, NIRAJ; GANESAN, RAJESHWARI; BHASKER, BHAVANA
To: EDGEVERVE SYSTEMS LIMITED
Reel/Frame 050206/0215 →
Priority Claims (1)
IN 201941013239 · Apr 2, 2019 · national
Continuity (1)
Related Publication 20200320430A1 · Oct 8, 2020