IP Library Granted Patent US 10,460,257
Granted Patent B2
US 10,460,257 · App. 15/259,541 · Granted Oct 29, 2019

Method and system for training a target domain classifier to label text segments

Inventors: Raksha Sharma (Gwalior, IN); Sandipan Dandapat (Kolkata, IN); Himanshu Sharad Bhatt (Nagar, IN)
Assignee: CONDUENT BUSINESS SERVICES, LLC
G06N20/00G06F16/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,460,257
App. No.
15/259,541
Granted
Oct 29, 2019
Kind
B2
Abstract

The disclosed embodiments illustrate methods of data processing for training a target domain classifier to label text segments. The method includes identifying a set of common keywords with same label from a set of source keywords and a set of target keywords. The method includes training a first classifier, based on the set of common keywords, to label a first set of target text segments. The method includes training a second classifier based on at least a subset of the labeled first set of target text segments. The method includes training a third classifier, based on the first classifier and the second classifier, to label a second set of target text segments, wherein a subset of the labeled second set of target text segments is utilized for re-training the second classifier. The method further includes determining labels of another plurality of target text segments based on the re-trained second classifier.

Claims (42)

1. A method of data processing for training a target domain classifier to label text segments, the method comprising:

identifying, by one or more processors in a computing device, a set of common keywords with same label from a set of source keywords and a set of target keywords, wherein each keyword in the set of source keywords and the set of target keywords is associated with a label;

training, by the one or more processors in the computing device, a first classifier, based on the set of common keywords, to label a first set of target text segments from a plurality of target text segments, wherein each of the labeled first set of target text segments is associated with a first confidence score;

training, by the one or more processors in the computing device, a second classifier based on at least a subset of the labeled first set of target text segments for which the first confidence score exceeds a confidence threshold value;

training, by the one or more processors in the computing device, a third classifier, based on the trained first classifier and the trained second classifier, to label a second set of target text segments from the plurality of target text segments,

wherein each of the labeled second set of target text segments is associated with a second confidence score,

wherein a subset of the labeled second set of target text segments, for which the second confidence score exceeds the confidence threshold value, is utilized for re-training the second classifier; and

determining, by the one or more processors in the computing device, labels of another plurality of target text segments based on the re-trained second classifier that corresponds to the target domain classifier.

2. The method of claim 1 , further comprising receiving, by one or more transceivers in the computing device, from another computing device, a plurality of source text segments associated with a source domain and the plurality of target text segments associated with a target domain.

3. The method of claim 2 , wherein the plurality of source text segments comprises a plurality of source keywords and the plurality of target text segments comprises a plurality of target keywords.

4. The method of claim 3 , wherein the set of source keywords is identified from the plurality of source keywords, when a first significance score associated with a source keyword exceeds a first significance threshold value.

5. The method of claim 3 , wherein the set of target keywords is identified from the plurality of target keywords, when a second significance score associated with a target keyword is less than a second significance threshold value.

6. The method of claim 2 , wherein each of the plurality of source text segments is associated with a label and each of the plurality of target text segments is independent of a label.

7. The method of claim 1 , wherein each keyword in the set of source keywords is labeled based on a first score, wherein the first score of a source keyword in the set of source keywords is determined based on one or more statistical techniques.

8. The method of claim 1 , wherein each keyword in the set of target keywords is labeled based on a second score, wherein the second score for a target keyword in the set of target keywords is determined based on one or more similarity measures between the target keyword and a set of pre-defined keywords.

9. The method of claim 1 , wherein the re-training of the second classifier is based on a count of target text segments in the second set of target text segments.

10. The method of claim 1 , wherein the third classifier is a weighted combination of the trained first classifier and the trained second classifier, wherein the weights of the trained first classifier and the trained second classifier are determined based on an accuracy parameter associated with the trained first classifier and the trained second classifier.

11. A system of data processing for training a target domain classifier to label text segments, the system comprises:

one or more processors in a computing device configured to:

identify a set of common keywords with same label from a set of source keywords and a set of target keywords, wherein each keyword in the set of source keywords and the set of target keywords is associated with a label;

train a first classifier, based on the set of common keywords, to label a first set of target text segments from a plurality of target text segments, wherein each of the labeled first set of target text segments is associated with a first confidence score;

train a second classifier based on at least a subset of the labeled first set of target text segments for which the first confidence score exceeds a confidence threshold value;

train a third classifier, based on the trained first classifier and the trained second classifier, to label a second set of target text segments from the plurality of target text segments,

wherein each of the labeled second set of target text segments is associated with a second confidence score,

wherein a subset of the labeled second set of target text segments, for which the second confidence score exceeds the confidence threshold value is utilized for re-training the second classifier; and

determine labels of another plurality of target text segments based on the re-trained second classifier that corresponds to the target domain classifier.

12. The system of claim 11 , wherein one or more transceivers in the computing device are configured to receive, from another computing device, a plurality of source text segments associated with a source domain and the plurality of target text segments associated with a target domain.

13. The system of claim 12 , wherein the plurality of source text segments comprises a plurality of source keywords and the plurality of target text segments comprises a plurality of target keywords.

14. The system of claim 13 , wherein the set of source keywords is identified from the plurality of source keywords, when a first significance score associated with a source keyword exceeds a first significance threshold value, wherein the set of target keywords is identified from the plurality of target keywords, when a second significance score associated with a target keyword is less than a second significance threshold value.

15. The system of claim 12 , wherein each of the plurality of source text segments is associated with a label and each of the plurality of target text segments is independent of a label.

16. The system of claim 11 , wherein each keyword in the set of source keywords is labeled based on a first score, wherein the first score of a source keyword in the set of source keywords is determined based on one or more statistical techniques.

17. The system of claim 11 , wherein each keyword in the set of target keywords is labeled based on a second score, wherein the second score for a target keyword in the set of target keywords is determined based on one or more similarity measures between the target keyword and a set of pre-defined keywords.

18. The system of claim 11 , wherein the re-training of the second classifier is based on a count of target text segments in the second set of target text segments.

19. The system of claim 11 , wherein the third classifier is a weighted combination of the trained first classifier and the trained second classifier, wherein the weights of the trained first classifier and the trained second classifier are determined based on an accuracy parameter associated with the trained first classifier and the trained second classifier.

20. A computer program product for use with a computer, the computer program product comprising a non-transitory computer readable medium, wherein the non-transitory computer readable medium stores a computer program code of data processing for training a target domain classifier to label text segments, wherein the computer program code is executable by one or more processors in a computing device to:

identify a set of common keywords with same label from a set of source keywords and a set of target keywords, wherein each keyword in the set of source keywords and the set of target keywords is associated with a label;

train a first classifier, based on the set of common keywords, to label a first set of target text segments from a plurality of target text segments, wherein each of the labeled first set of target text segments is associated with a first confidence score;

train a second classifier based on at least a subset of the labeled first set of target text segments for which the first confidence score exceeds a confidence threshold value;

train a third classifier, based on the trained first classifier and the trained second classifier, to label a second set of target text segments from the plurality of target text segments,

wherein each of the labeled second set of target text segments is associated with a second confidence score,

wherein a subset of the labeled second set of target text segments, for which the second confidence score exceeds the confidence threshold value is utilized for re-training the second classifier; and

determine labels of another plurality of target text segments based on the re-trained second classifier that corresponds to the target domain classifier.

Assignments (6)
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: U.S. BANK, NATIONAL ASSOCIATION
Reel/Frame 057969/0445 →
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 057970/0001 →
RELEASE OF SECURITY INTEREST Recorded Oct 18, 2021
From: JPMORGAN CHASE BANK, N.A.
To: CONDUENT BUSINESS SERVICES, LLC; CONDUENT STATE & LOCAL SOLUTIONS, INC.; CONDUENT TRANSPORT SOLUTIONS, INC.; ADVECTIS, INC.; CONDUENT COMMERCIAL SOLUTIONS, LLC; CONDUENT BUSINESS SOLUTIONS, LLC; CONDUENT CASUALTY CLAIMS SOLUTIONS, LLC; CONDUENT HEALTH ASSESSMENTS, LLC
Reel/Frame 057969/0180 →
SECURITY AGREEMENT Recorded Mar 19, 2020
From: CONDUENT BUSINESS SERVICES, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 052189/0698 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 28, 2017
From: XEROX CORPORATION
To: CONDUENT BUSINESS SERVICES, LLC
Reel/Frame 041542/0022 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2016
From: SHARMA, RAKSHA , ,; DANDAPAT, SANDIPAN , ,; BHATT, HIMANSHU SHARAD, ,
To: XEROX CORPORATION
Reel/Frame 039680/0960 →
Continuity (1)
Related Publication 20180068231A1 · Mar 8, 2018