IP Library › Granted Patent US 11,748,393
Granted Patent B2
US 11,748,393 · App. 16/203,000 · Granted Sep 5, 2023

Creating compact example sets for intent classification

Inventors: Abhishek Shah (Jersey City, NJ); Tin Kam Ho (Millburn, NJ)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/35G06F18/217G06F18/2148G06F18/23213G06V30/274G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,748,393
App. No.
16/203,000
Granted
Sep 5, 2023
Kind
B2
Abstract

Embodiments for creating compact example subsets for intent classification in a conversational system are provided. A set of content used for training an intent classifier is received from a conversational corpus. Entries within the set of content are separated into a first subset and a second subset, and a cross-validation operation is performed on the first and second subsets to identify a correctly labeled portion and an incorrectly labeled portion of the set of content. A reduced content used for performing a final training of the intent classifier is formed by combining a first number of the entries from the correctly labeled portion and a second number of the entries from the incorrectly labeled portion of the set of content.

Claims (52)

1. A method for creating compact example subsets for intent classification, by a processor, comprising:

receiving a set of content used for training an intent classifier;

separating entries within the set of content into a first subset and a second subset;

performing a cross-validation operation on the first and second subsets to identify a correctly labeled portion and an incorrectly labeled portion of the set of content, wherein the cross-validation operation further comprises:

performing an initial training of the intent classifier utilizing the first subset to form a first subset trained classifier;

utilizing the first subset trained classifier against the second subset to identify a correctly labeled subset and an incorrectly labeled subset of the second subset;

performing a secondary training of the intent classifier, subsequent to the initial training, to form a second subset trained classifier; and

utilizing the second subset trained classifier against the first subset to identify a correctly labeled subset and an incorrectly labeled subset of the first subset; and

forming a reduced content used for performing a final training of the intent classifier by combining a first number of the entries from the correctly labeled portion and a second number of the entries from the incorrectly labeled portion of the set of content, wherein an anti-clustering procedure, performed separately and independently on each of the first subset and the second subset, is utilized to select members of the first number of the entries and members of the second number of the entries by:

computing a vector representation for each of the entries from the correctly labeled portion and the entries from the incorrectly labeled portion of the set of content;

clustering each of a set of vectors representing the correctly labeled portion and the incorrectly labeled portion into k clusters, wherein k equals a desired size of the reduced content;

selecting a longest entry in each of the k clusters as a cluster representative to yield maximally-spread samples among all of the k clusters, wherein each other in-cluster members of each of the k clusters are ignored; and

using the selected longest entry in each of the k clusters aggregately as the first number of the entries and the second number of the entries comprising the reduced content.

2. The method of claim 1 , further comprising organizing the correctly labeled subset of the first subset and the correctly labeled subset of the second subset into the correctly labeled portion; and

organizing the incorrectly labeled subset of the first subset and the incorrectly labeled subset of the second subset into the incorrectly labeled portion.

3. The method of claim 1 , wherein, commensurate with the combining, the first number of entries is larger than the second number of entries.

4. The method of claim 1 , wherein the content comprises utterances received from a conversational corpus.

5. A system for creating compact example subsets for intent classification, comprising:

a processor executing instructions stored in a memory device; wherein the processor:

receives a set of content used for training an intent classifier;

separates entries within the set of content into a first subset and a second subset;

performs a cross-validation operation on the first and second subsets to identify a correctly labeled portion and an incorrectly labeled portion of the set of content, wherein the cross-validation operation further comprises:

performing an initial training of the intent classifier utilizing the first subset to form a first subset trained classifier;

utilizing the first subset trained classifier against the second subset to identify a correctly labeled subset and an incorrectly labeled subset of the second subset;

performing a secondary training of the intent classifier, subsequent to the initial training, to form a second subset trained classifier; and

utilizing the second subset trained classifier against the first subset to identify a correctly labeled subset and an incorrectly labeled subset of the first subset; and

forms a reduced content used for performing a final training of the intent classifier by combining a first number of the entries from the correctly labeled portion and a second number of the entries from the incorrectly labeled portion of the set of content, wherein an anti-clustering procedure, performed separately and independently on each of the first subset and the second subset, is utilized to select members of the first number of the entries and members of the second number of the entries by:

computing a vector representation for each of the entries from the correctly labeled portion and the entries from the incorrectly labeled portion of the set of content;

clustering each of a set of vectors representing the correctly labeled portion and the incorrectly labeled portion into k clusters, wherein k equals a desired size of the reduced content;

selecting a longest entry in each of the k clusters as a cluster representative to yield maximally-spread samples among all of the k clusters, wherein each other in-cluster members of each of the k clusters are ignored; and

using the selected longest entry in each of the k clusters aggregately as the first number of the entries and the second number of the entries comprising the reduced content.

6. The system of claim 5 , wherein the processor organizes the correctly labeled subset of the first subset and the correctly labeled subset of the second subset into the correctly labeled portion; and

organizes the incorrectly labeled subset of the first subset and the incorrectly labeled subset of the second subset into the incorrectly labeled portion.

7. The system of claim 5 , wherein, commensurate with the combining, the first number of entries is larger than the second number of entries.

8. The system of claim 5 , wherein the content comprises utterances received from a conversational corpus.

9. A computer program product for creating compact example subsets for intent classification, by a processor, the computer program product embodied on a non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions comprising:

an executable portion that receives a set of content used for training an intent classifier;

an executable portion that separates entries within the set of content into a first subset and a second subset;

an executable portion that performs a cross-validation operation on the first and second subsets to identify a correctly labeled portion and an incorrectly labeled portion of the set of content, wherein the cross-validation operation further comprises:

performing an initial training of the intent classifier utilizing the first subset to form a first subset trained classifier;

utilizing the first subset trained classifier against the second subset to identify a correctly labeled subset and an incorrectly labeled subset of the second subset;

performing a secondary training of the intent classifier, subsequent to the initial training, to form a second subset trained classifier; and

utilizing the second subset trained classifier against the first subset to identify a correctly labeled subset and an incorrectly labeled subset of the first subset; and

an executable portion that forms a reduced content used for performing a final training of the intent classifier by combining a first number of the entries from the correctly labeled portion and a second number of the entries from the incorrectly labeled portion of the set of content, wherein an anti-clustering procedure, performed separately and independently on each of the first subset and the second subset, is utilized to select members of the first number of the entries and members of the second number of the entries by:

computing a vector representation for each of the entries from the correctly labeled portion and the entries from the incorrectly labeled portion of the set of content;

clustering each of a set of vectors representing the correctly labeled portion and the incorrectly labeled portion into k clusters, wherein k equals a desired size of the reduced content;

selecting a longest entry in each of the k clusters as a cluster representative to yield maximally-spread samples among all of the k clusters, wherein each other in-cluster members of each of the k clusters are ignored; and

using the selected longest entry in each of the k clusters aggregately as the first number of the entries and the second number of the entries comprising the reduced content.

10. The computer program product of claim 9 , further comprising an executable portion that organizes the correctly labeled subset of the first subset and the correctly labeled subset of the second subset into the correctly labeled portion; and

organizes the incorrectly labeled subset of the first subset and the incorrectly labeled subset of the second subset into the incorrectly labeled portion.

11. The computer program product of claim 9 , wherein, commensurate with the combining, the first number of entries is larger than the second number of entries.

12. The computer program product of claim 9 , wherein the content comprises utterances received from a conversational corpus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2018
From: SHAH, ABHISHEK; HO, TIN KAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047610/0641 →
Continuity (1)
Related Publication 20200167604A1 · May 28, 2020