IP Library Granted Patent US 11,281,858
Granted Patent B1
US 11,281,858 · App. 17/373,810 · Granted Mar 22, 2022

Systems and methods for data classification

Inventors: Igal Mazor (Tel-Aviv, IL); Yaron Ismah-Moshe (Tel-Aviv, IL)
G06F40/284G06F16/3344G06F40/166G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,281,858
App. No.
17/373,810
Granted
Mar 22, 2022
Kind
B1
Abstract

A computer implemented method of document classification includes receiving a text document. A first classification is generated for the document, and a text corpus is searched for one or more terms from the document. Searched terms having an incidence in the text corpus lower than a threshold incidence are flagged, and at least one classification is generated after removing at least one flagged term from the document. An output is generated if the further classification is different from the first classification.

Claims (26)

1. A computer implemented method of document classification, comprising:

receiving a text document;

generating, using a text classification model, a first classification for the document;

searching for one or more terms from the document within a text corpus, wherein the corpus comprises each word in each training document used to train the model;

flagging one or more of the searched terms having an incidence in the text corpus lower than a threshold incidence;

removing at least one flagged term from the document and generating at least one further classification; and

generating an output, if the further classification is different from the first classification.

2. The method of claim 1 , wherein searching for one or more terms includes: identifying one or more significant terms which contribute to the first classification of the document, and searching for the identified significant terms within the text corpus.

3. The method of claim 2 , wherein the significant terms are based on an attention layer of the text classification model.

4. The method of claim 1 , wherein searching further includes, in response to identifying a searched term having an incidence in the text corpus lower than the threshold incidence:

generating one or more words that are similar to the identified low incidence term,

searching for the one or more similar words in the text corpus; and

flagging the identified low incidence term if a combined incidence of the low incidence term and the one or more similar words is below the threshold incidence.

5. The method of claim 1 , wherein the threshold incidence is zero.

6. The method of claim 1 , wherein generating the further classification comprises generating a plurality of further classifications for each a plurality of flagged terms.

7. The method of claim 1 , wherein generating the further classification includes replacing a removed term with an entity token corresponding to the removed term.

8. The method of claim 1 , wherein generating the further classification includes:

identifying a word that is similar to a removed term with a higher incidence in the text corpus, and

replacing the removed term with the identified similar word.

9. The method of claim 1 , wherein the output includes outputting the further classification as a classification result.

10. The method of claim 1 , wherein the output includes an indication for a user to manually check a classification result.

11. The method of claim 1 , wherein the output includes updating the text corpus to include the text document with the flagged term removed.

12. The method of claim 1 , wherein the output includes replacing a removed term with an entity token corresponding to the removed term and updating the text corpus to include the text document with the replaced term.

13. The method of claim 1 , wherein the searching includes searching the text corpus using a pre-trained model, or fine-tuning a pre-trained model using the text corpus and searching the text corpus using the fine-tuned model, or training a dedicated word embedding model using the text corpus and searching the text corpus using the dedicated word embedding model.

14. A data processing apparatus comprising a processor configured to execute the method of claim 1 .

15. A computer-readable medium configured to store instructions which, when executed by a processor, cause the processor to perform the method of claim 1 .

Assignments (4)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 064367/0879 Recorded Feb 4, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070098/0287 →
SECURITY AGREEMENT Recorded Jul 24, 2023
From: GENESYS CLOUD SERVICES, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 064367/0879 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2022
From: EXCEED.AI LTD.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 059012/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2021
From: MAZOR, IGAL; ISMAH-MOSHE, YARON
To: EXCEED AI LTD
Reel/Frame 056845/0213 →