IP Library › Granted Patent US 11,321,527
Granted Patent B1
US 11,321,527 · App. 17/154,060 · Granted May 3, 2022

Effective classification of data based on curated features

Inventors: Maithreyi Gopalarao (Bangalore, IN); Manveer Singh Sandhu (Amritsar, IN); Rohit Athradi Shetty (Bangalore, IN); Amit Meel (Didwana, IN)
Assignee: International Business Machines Corporation
G06F40/279G06F40/166G06N20/00G06V30/413
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,321,527
App. No.
17/154,060
Granted
May 3, 2022
Kind
B1
Abstract

Techniques for machine learning using curated features are provided. A plurality of key terms is identified for a first document type of a plurality of document types. A document associated with the first document type is received, and the document is modified by inserting one or more of the plurality of key terms. A vector is generated for the modified document, and a machine learning model is trained to categorize input into the plurality of document types based on the modified document.

Claims (78)

1. A method, comprising:

identifying a predefined set of key terms for a first document type of a plurality of document types;

receiving a first document of the first document type;

in response to identifying a first instance of a first key term, of the predefined set of key terms, in text of the first document, modifying the first document by inserting a second instance of the first key term into the text of the first document;

generating a first document vector for the modified first document including at least the first instance and the second instance of the first key term;

training a machine learning model based on the first document vector for the modified document comprising:

associating the first key term with an increased weight in response to the second instance of the key term inserted into the first document, as compared to a weight of at least one term that is not included in the predefined set of key terms; and

categorizing the first document into at least one of the plurality of document types based on the first document vector for the modified first document.

2. The method of claim 1 , wherein inserting the second instance of the first key term into the text of the first document comprises appending the first key term at end of the first document.

3. The method of claim 2 , wherein modifying the first document further comprises:

upon failing to locate a second key term, of the predefined set of key terms, in the text of the first document, refraining from inserting the second key term into the text of the first document.

4. The method of claim 2 , wherein modifying the first document comprises:

for each respective instance of the first key term located in the text of the first document, inserting a respective new instance of the first key term into the text of the first document.

5. The method of claim 1 , the method further comprising:

receiving a second document for classification using the trained machine learning model;

upon locating a third instance of the first key term, of the predefined set of key terms, in text of the second document, modifying the second document by inserting a fourth instance of the first key term into the text of the second document;

generating a second document vector for the modified second document; and

classifying the second document by processing the second document vector using the trained machine learning model.

6. The method of claim 1 , the method further comprising:

identifying a new key term that was not used when training the machine learning model;

receiving a second document for classification using the trained machine learning model;

upon locating a first instance of the new key term, of the plurality predefined set of key terms, in text of the second document, modifying the second document by inserting a second instance of the new key term into the text of the second document;

generating a second document vector for the modified second document; and

classifying the second document by processing the second document vector using the trained machine learning model.

7. The method of claim 6 , wherein identifying the new key term is performed upon determining that the second document was misclassified by the trained machine learning model.

8. One or more non-transitory computer-readable storage medium collectively containing computer program code that, when executed by operation of one or more computer processors, performs an operation comprising:

identifying a predefined set of key terms for a first document type of a plurality of document types;

receiving a first document of the first document type;

in response to identifying a first instance of a first key term, of the predefined set of key terms, in text of the first document, modifying the first document by inserting a second instance of the first key term into the text of the first document;

generating a first document vector for the modified first document including at least the first and second instances of the first key term; and

training a machine learning model based on the first document vector for the modified document comprising:

associating the first key term with an increased weight in response to the second instance of the first key term inserted into the first document, as compared to a weight of at least one term that is not included in the predefined set of key terms; and

categorizing the first document into at least one of the plurality of document types based on the first document vector for the modified first document.

9. The computer-readable storage medium of claim 8 , wherein modifying the first document comprises:

inserting the second instance of the first key term into the text of the first document comprises appending the first key term at end of the first document.

10. The computer-readable storage medium of claim 9 , wherein modifying the first document further comprises:

upon failing to identify a second key term, of the predefined set of key terms, in the text of the first document, refraining from inserting the second key term into the text of the first document.

11. The computer-readable storage medium of claim 9 , wherein modifying the first document comprises:

for each respective instance of the first key term located in the text of the first document, inserting a respective new instance of the first key term into the text of the first document.

12. The computer-readable storage medium of claim 8 , the operation further comprising:

receiving a second document for classification using the trained machine learning model;

in response to identifying a third instance of the first key term, of the predefined set of key terms, in text of the second document, modifying the second document by inserting a fourth instance of the first key term into the text of the second document;

generating a second document vector for the modified second document; and

classifying the second document by processing the second document vector using the trained machine learning model.

13. The computer-readable storage medium of claim 8 , the operation further comprising:

identifying a new key term that was not used when training the machine learning model;

receiving a second document for classification using the trained machine learning model;

in response to identifying a first instance of the new key term, of the predefined set of key terms, in text of the second document, modifying the second document by inserting a second instance of the new key term into the text of the second document;

generating a second document vector for the modified second document; and

classifying the second document by processing the second document vector using the trained machine learning model.

14. The computer-readable storage medium of claim 13 , wherein identifying the new key term is performed upon determining that the second document as misclassified by the trained machine learning model.

15. A system comprising:

One or more computer processors; and

One or more memories collectively containing one or more programs which when executed by the one or more computer processors performs an operation, the operation comprising:

identifying a predefined set of key terms for a first document type of a plurality of document types;

receiving a first document of the first document type;

in response to identifying a first instance of a first key term, of the predefined set of key terms, in text of the first document, modifying the first document by inserting a second instance of the first key term into the text of the first document;

generating a first document vector for the modified first document including at least the first instance and the second instance of the first key term;

training a machine learning model based on the first document vector for the modified document comprising:

associating the first key term with an increased weight in response to the second instance of the key term inserted into the first document, as compared to a weight of at least one term that is not included in the predefined set of key terms; and

categorizing the first document into at least one of the plurality of document types based on the first document vector for the modified first document.

16. The system of claim 15 , wherein

inserting the second instance of the first key term into the text of the first document comprises appending the first term at end of the first document.

17. The system of claim 16 , wherein modifying the first document further comprises:

upon failing to locate a second key term, of the predefined set key terms, in the text of the first document, refraining from inserting the second key term into the text of the first document.

18. The system of claim 16 , wherein modifying the first document comprises:

for each respective instance of the first key term located in the text of the first document, inserting a respective new instance of the first key term into the text of the first document.

19. The system of claim 15 , the operation further comprising:

receiving a second document for classification using the trained machine learning model;

upon locating a third instance of the first key term, of the predefined set of key terms, in text of the second document, modifying the second document by inserting a fourth instance of the first key term into the text of the second document;

generating a second document vector for the modified second document; and

classifying the second document by processing the second document vector using the trained machine learning model.

20. The system of claim 15 , the operation further comprising:

identifying a new key term that was not used when training the machine learning model;

receiving a second document for classification using the trained machine learning model;

upon locating a first instance of the new key term, of the predefined set of key terms, in text of the second document, modifying the second document by inserting a second instance of the new key term into the text of the second document;

generating a second document vector for the modified second document; and

classifying the second document by processing the second document vector using the trained machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2021
From: GOPALARAO, MAITHREYI; SANDHU, MANVEER SINGH; SHETTY, ROHIT ATHRADI; MEEL, AMIT
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054981/0816 →
Cited By (3)
US 12,321,428 US 12,554,754 US 12,682,217