IP Library Granted Patent US 9,418,148
Granted Patent B2
US 9,418,148 · App. 13/731,651 · Granted Aug 16, 2016

System and method to label unlabeled data

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,418,148
App. No.
13/731,651
Granted
Aug 16, 2016
Kind
B2
Abstract

An embodiment of the invention provides a technique for permitting a machine to discover classes and topics that data contains and to annotate data objects with those identified classes. The technique enables machines to group and annotate data objects in ways that are meaningful and intuitive for a user of the data objects. An interactive method uses clustering, along with feedback from a user on the clustering output, to discover a set of classes. The feedback from the user is used to guide the clustering process in the later stages, which results in better and better discovery of classes and annotations with more and more human feedback. A method can be used to produce labeled data that involves discovering classes and annotating a given dataset with the discovered class labels. This is advantageous for building a classifier that has wide applications, such as call routing and intent discovery.

Claims (34)

1. A device for labeling unlabeled data, the device comprising:

a memory with computer code instructions stored thereon, the memory with one or more processors, and the computer code instructions being configured to cause the device to implement:

a clustering processor configured to group the unlabeled data to produce at least one data group;

a feedback processor configured to:

enable a user to provide feedback on the at least one data group, the feedback including at least one of: (i) feedback on membership of a data object in a data group, and (ii) feedback on a current labeling of a data group, and

discover a representative set of data objects from the at least one data group on which to seek feedback from the user, the representative set of data objects comprising at least two data objects and being smaller than the at least one data group, each data object of the representative set of data objects being similar to a respective sub-group of the at least one data group, each sub-group comprising a plurality of data objects from the at least one data group;

the clustering processor further configured to regroup the at least one data group using at least one constraint based on the feedback provided by the user; and

the feedback processor further configured to apply a topic label to a data group of the at least one data group to produce at least one labeled data group, the topic label indicating contents of the labeled data group and the topic label based on the feedback provided by the user after the grouping or on feedback provided by the user after at least one regrouping by the clustering processor.

2. The device according to claim 1 , wherein the feedback processor is further configured to enable the user to provide feedback based on the discovered representative set of data objects.

3. The device according to claim 1 , wherein the clustering processor comprises a K-means clustering processor.

4. The device according to claim 1 , wherein the feedback processor is further configured to receive feedback provided by the user including at least one of the following: a new name of a data group; an acceptance of membership of a data object in a data group; a rejection of membership of a data object in a data group; a splitting of a data group; a merging of a data group; an acceptance of a member of a cluster centroid of a data group; and a rejection of a member of a cluster centroid of a data group.

5. The device according to claim 1 , wherein the feedback processor is further configured to display at least one data object of the at least one labeled data group to the user.

6. The device according to claim 1 , wherein the clustering processor is further configured to determine, for a data group of the at least one data group: (i) at least one data object belonging to the data group; and (ii) a context vector for the data group.

7. The device according to claim 1 , wherein the clustering processor is further configured to determine, for a data group of the at least one data group, a scoring measure for the data group.

8. The device according to claim 1 , wherein the clustering processor is further configured to group the data based on at least one initial constraint provided by the user.

9. A method for labeling unlabeled data, the method comprising:

grouping the unlabeled data using clustering to produce at least one data group;

discovering a representative set of data objects from the at least one data group on which to seek feedback from the user, the representative set of data objects comprising at least two data objects and being smaller than the at least one data group, each data object of the representative set of data objects being similar to a respective sub-group of the at least one data group, each sub-group comprising a plurality of data objects from the at least one data group;

enabling a user to provide feedback on the at least one data group, the feedback including at least one of the following: (i) feedback on membership of a data object in a data group, and (ii) feedback on a current labeling of a data group;

regrouping the at least one data group using further clustering including at least one constraint based on the feedback provided by the user; and

applying a topic label to a data group of the at least one data group to produce at least one labeled data group, the topic label indicating contents of the labeled data group and the topic label based on the feedback provided by the user after the grouping or on feedback provided by the user after at least one regrouping.

10. The method according to claim 9 , wherein the enabling the user to provide feedback is performed based on the discovered representative set of data objects.

11. The method according to claim 9 , wherein the grouping the data using clustering comprises performing K-means clustering on the data.

12. The method according to claim 9 , wherein the feedback provided by the user comprises at least one of the following: a new name of a data group; an acceptance of membership of a data object in a data group; a rejection of membership of a data object in a data group; a splitting of a data group; a merging of a data group; an acceptance of a member of a cluster centroid of a data group; and a rejection of a member of a cluster centroid of a data group.

13. The method according to claim 9 , further comprising displaying at least one data object of the at least one labeled data group to the user.

14. The method according to claim 9 , wherein grouping the data comprises, for a data group of the at least one data group, determining: (i) at least one data object belonging to the data group; and (ii) a context vector for the data group.

15. The method according to claim 9 , wherein grouping the data comprises, for a data group of the at least one data group, determining a scoring measure for the data group.

16. The method according to claim 9 , wherein the grouping the data using clustering is performed based on at least one initial constraint provided by the user.

17. A non-transitory computer-readable storage medium having computer-readable code stored thereon, which, when executed by a computer processor, causes the computer processor to label unlabeled data, by causing the processor to:

group the unlabeled data using clustering to produce at least one data group;

discover a representative set of data objects from the at least one data group on which to seek feedback from the user, the representative set of data objects comprising at least two data objects and being smaller than the at least one data group, each data object of the representative set of data objects being similar to a respective sub-group of the at least one data group, each sub-group comprising a plurality of data objects from the at least one data group;

enable a user to provide feedback on the at least one data group, the feedback including at least one of the following: (i) feedback on membership of a data object in a data group, and (ii) feedback on a current labeling of a data group;

regroup the at least one data group using further clustering including at least one constraint based on the feedback provided by the user; and

apply a topic label to a data group of the at least one data group to produce at least one labeled data group, the topic label indicating contents of the labeled data group and the topic label based on the feedback provided by the user after the grouping or on feedback provided by the user after at least one regrouping.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2013
From: JOSHI, SACHINDRA; BODBOLE, SHATANU RAVINDRA; VERMA, ASHISH
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 029945/0430 →