IP Library Granted Patent US 12,299,545
Granted Patent B2
US 12,299,545 · App. 17/125,120 · Granted May 13, 2025

Systems and methods for automatic extraction of classification training data

Inventors: Igal Mazor (Tel-Aviv, IL); Yaron Ismah-Moshe (Tel-Aviv, IL)
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,545
App. No.
17/125,120
Granted
May 13, 2025
Kind
B2
Abstract

A method for training a multi-class classification model includes receiving training data corresponding to a plurality of classes. For each class in the plurality of classes, the method includes training a binary classification model configured to determine whether or not an observation of training data belongs to the class and for each observation of training data identified as belonging to the class, extracting one or more class identification features from the observation of training data based on activations of an intermediate attention layer in the binary classification model. A multi-class classification model is trained using the class identification features extracted for each of the plurality of classes.

Claims (33)

1. A computer-implemented method for training a multi-class classification model, comprising:

receiving, by a processor, training data corresponding to a plurality of classes, wherein the training data comprises a first corpus and a second corpus, wherein the first corpus is constructed such that each observation of training data is relevant to a respective class of the plurality of classes, and wherein the second corpus includes collected real world data;

for each class in the plurality of classes:

training, by the processor, a binary classification model configured to determine whether or not an observation of training data of the first corpus belongs to a class based on activations of an intermediate attention layer, wherein the binary classification model comprises a plurality of layers including an input layer and the intermediate attention layer, the input layer configured to divide an input text document into a plurality of sub-features and to evaluate a relative importance of each of a plurality of sub-features in the input text document, and the intermediate attention layer including a plurality of nodes corresponding to the plurality of sub-features of the input text document, wherein the binary classification model uses the relative importance to identify whether the input text document belongs to the class; and

extracting one or more class identification features from the observation of training data of the second corpus based on activations of the intermediate attention layer in the binary classification model; and

training, by the processor, the multi-class classification model using the class identification features extracted for each of the plurality of classes.

2. The method of claim 1 , wherein the multi-class classification model is a text classification model, and the training data comprises a plurality of text documents.

3. The method of claim 2 , wherein each class in the plurality of classes corresponds to an intended purpose of a text document.

4. The method of claim 2 , wherein the class identification features correspond to sentences extracted from the plurality of text documents.

5. The method of claim 2 , wherein extracting the class identification features comprises dividing each of the training data text documents into a plurality of sentences.

6. The method of claim 5 , wherein extracting the class identification features comprises evaluating a metric for each sentence in the plurality of text documents.

7. The method of claim 1 , wherein each class identification feature is extracted based on whether or not an activation weight of the intermediate attention layer is higher than a predefined threshold.

8. The method of claim 1 , wherein extracting the class identification features includes validating each feature using the corresponding binary classification model.

9. The method of claim 1 , wherein the binary classification model comprises the intermediate attention layer followed by a fully connected layer.

10. The method of claim 1 , further comprising labelling one or more observation of training data which do not belong to any of the plurality of classes.

11. A processing apparatus comprising a processor configured to execute a method comprising the steps of:

receiving, by the processor, training data corresponding to a plurality of classes, wherein the training data comprises a first corpus and a second corpus, wherein the first corpus is constructed such that each observation of training data is relevant to a respective class of the plurality of classes, and wherein the second corpus includes collected real world data; and, for each class in the plurality of classes:

training, by the processor, a binary classification model configured to determine whether or not an observation of training data of the first corpus belongs to a class based on activations of an intermediate attention layer, wherein the binary classification model comprises a plurality of layers including an input layer and the intermediate attention layer, the input layer configured to divide an input text document into a plurality of sub-features and to evaluate a relative importance of each of a plurality of sub-features in the input text document, and the intermediate attention layer including a plurality of nodes corresponding to the plurality of sub-features of the input text document, wherein the binary classification model uses the relative importance to identify whether the input text document belongs to the class; and

extracting one or more class identification features from the observation of training data of the second corpus based on activations of the intermediate attention layer in the binary classification model; and

training, by the processor, a multi-class classification model using the class identification features extracted for each of the plurality of classes.

12. The processing apparatus of claim 11 , wherein the multi-class classification model is a text classification model, and the training data comprises a plurality of text documents.

13. The processing apparatus of claim 12 , wherein each class in the plurality of classes corresponds to an intended purpose of a text document.

14. The processing apparatus of claim 12 , wherein the class identification features correspond to sentences extracted from the plurality of training text documents.

15. The processing apparatus of claim 12 , wherein extracting the class identification features comprises dividing each of the training data text documents into a plurality of sentences, and wherein extracting the class identification features comprises evaluating a metric for each sentence in the plurality of text documents.

16. A non-transitory, computer-readable medium configured to store instructions which, when executed by a processor, causes the processor to execute a method comprising the steps of:

receiving, by the processor, training data corresponding to a plurality of classes, wherein the training data comprises a first corpus and a second corpus, wherein the first corpus is constructed such that each observation of training data is relevant to a respective class of the plurality of classes, and wherein the second corpus includes collected real world data; and, for each class in the plurality of classes:

training, by the processor, a binary classification model configured to determine whether or not an observation of training data of the first corpus belongs to a class based on activations of an intermediate attention layer, wherein the binary classification model comprises a plurality of layers including an input layer and the intermediate attention layer, the input layer configured to divide an input text document into a plurality of sub-features and to evaluate a relative importance of each of a plurality of sub-features in the input text document, and the intermediate attention layer including a plurality of nodes corresponding to the plurality of sub-features of the input text document, wherein the binary classification model uses the relative importance to identify whether the input text document belongs to the class; and

extracting one or more class identification features from the observation of training data of the second corpus based on activations of the intermediate attention layer in the binary classification model; and

training, by the processor, a multi-class classification model using the class identification features extracted for each of the plurality of classes.

17. The non-transitory, computer-readable medium of claim 16 , wherein each class identification feature is extracted based on whether or not an activation weight of the intermediate attention layer is higher than a predefined threshold.

18. The non-transitory, computer-readable medium of claim 16 , wherein extracting the class identification features includes validating each feature using the corresponding binary classification model.

19. The non-transitory, computer-readable medium of claim 16 , wherein the binary classification model comprises the intermediate attention layer followed by a fully connected layer.

20. The non-transitory, computer-readable medium of claim 16 , wherein the method further comprising labelling one or more observation of training data which do not belong to any of the plurality of classes.

Assignments (4)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 064367/0879 Recorded Feb 4, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070098/0287 →
SECURITY AGREEMENT Recorded Jul 24, 2023
From: GENESYS CLOUD SERVICES, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 064367/0879 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: MAZOR, IGAL; ISMAH-MOSHE, YARON
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 063827/0307 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: EXCEED.AI LTD.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 059199/0159 →
Continuity (1)
Related Publication 20220198316A1 · Jun 23, 2022
References Cited (21)
US 20130024407A1 · Thompson · 2013 [cited by examiner]
US 20150310862A1 · Dauphin · 2015 [cited by examiner]
US 20190370398A1 · He · 2019 [cited by examiner]
US 20200142999A1 · Pedersen · 2020 [cited by examiner]
US 20200285702A1 · Padhi · 2020 [cited by examiner]
US 20200344194A1 · Hosseinisianaki · 2020 [cited by examiner]
US 20200401844A1 · Han · 2020 [cited by examiner]
US 20210034988A1 · Adel-Vu · 2021 [cited by examiner]
CN 111061881A · 2020 [cited by examiner]
CN 111475648A · 2020 [cited by examiner]
CN 111930939A · 2020 [cited by examiner]
CN 112214595A · 2021 [cited by examiner]
Jiang, “A Hierarchical Model with Recurrent Convolutional Neural Networks for Sequential Sentence Classification”, NLPCC 2019, LNAI 11839, pp. 78-89, 2019. (Previously provided). (Year: 2019). [cited by examiner]
Qing, “A Novel Neural Network-Based Method for Medical Text Classification”, Future Internet 2019. (Previously provided). (Year: 2019). [cited by examiner]
Parwez, “Multi-Label Classification of Microblogging Texts Using Convolution Neural Network”, IEEE Access, vol. 7, 2019. (Previously provided). (Year: 2019). [cited by examiner]
Li, “A Survey on Text Classification: From Shallow to Deep Learning”, IEEE Transactions on Neural Networks and Learning Systems, vol. 31, No. 11, Oct. 2020. (Previously provided). (Year: 2020). [cited by examiner]
International Search Report and Written Opinion for co-pending PCT application PCT/US2021/063970 mailed May 9, 2022. [cited by applicant]
Hiitlee: “SALNet Semi-Supervised Few-Shot Text Classification with Attention-Based Lexicon”, GitHub, Dec. 12, 2020, XP055912048, Retrieved from the internet—URL: https://github.com/HiitLee/SALNet/commit/c49cefd94cbeb15a… [cited by applicant]
Hiitlee: “SALNet: Semi-Supervised Few-Shot Text Classification with Attention-based Lexicon Construction”, Dec. 12, 2020, XP055912091, Retrieved from the internet—URL: https://github.com/HiitLee/SALNet. [cited by applicant]
Lee Ju-Hyoung et al.: “SALNet: Semi-Supervised Few-Shot Text Classification with Attention-Based Lexicon Construction”, Proceedings of the AAAI Conference on Artificial Intelligence, Feb. 9, 2021, pp. 13189-13197, XP055… [cited by applicant]
Matthew Tang et al., “Progress Notes Classification and Keyword Extraction using Attention-based Deep Learning Models with BERT”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 148… [cited by applicant]