IP Library › Granted Patent US 12,572,746
Granted Patent B2
US 12,572,746 · App. 18/121,360 · Granted Mar 10, 2026

Deep learning systems and methods to disambiguate false positives in natural language processing analytics

Inventors: Oleksandr Iakovenko (Odesa, UA); William Kelly (Jamaica Plain, MA); Stephen Land Stewart (Jenkintown, PA)
Assignee: Nuix Limited
G06F40/295G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,746
App. No.
18/121,360
Granted
Mar 10, 2026
Kind
B2
Abstract

Embodiments improve a document classification system by adjusting the class vector of each class based on embedded vectors of documents known to have been misclassified by the system. Some embodiments use a set of misclassified documents to adjust the class vector of the class into which a document has been misclassified. Some embodiments use a set of misclassified documents to adjust the class vector of the pre-assigned class of each misclassified document. Embodiments of adjusting the class vector may be described as machine learning.

Claims (60)

1 . A computer-implemented method comprising:

training a natural language processing-based classification system by:

providing a plurality of pre-specified class definitions, wherein each pre-specified class definition of the plurality of pre-specified class definitions is defined pursuant to a Neural Network Model, and each class definition of the plurality of pre-specified class definitions is uniquely associated with a corresponding document class selected from a plurality of document classes;

providing a test set comprising a plurality of test documents, each test document in the test set having a pre-specified class assignment to a class selected from the plurality of document classes;

automatically classifying each document of the plurality of test documents in the test set into a class selected from a plurality of document classes, such that each class has a plurality of correctly classified test documents and a set of misclassified test documents;

for each document class of the plurality of document classes:

identifying a set of test documents from the test set that have been misclassified into said document class by the act of automatically classifying, by:

for each test document classified into said document class by the act of automatically classifying, comparing the document's pre-specified class assignment to the test document's resulting class, and

determining which of said test documents have a resulting class different from its pre-specified class, each document having a resulting class different from its pre-specified class being a misclassified document;

identifying, from the set of misclassified documents, a corresponding subset of misclassified documents;

modifying the class definition of the class to move the class definition in semantic space away from the misclassified documents by retraining a neural network based on content of the misclassified documents; and

removing the misclassified documents from the test set to produce a reduced test set.

2 . The computer-implemented method of claim 1 , the method further comprising:

identifying each misclassified document from the test set;

for each such identified misclassified document, identifying the misclassified document's pre-specified class assignment, to produce a set of target classes; and

for each target class of the set of target classes,

modifying the class definition of said target class to move the class definition in semantic space toward said misclassified document.

3 . The computer-implemented method of claim 1 , further comprising, subsequent to producing the reduced test set:

classifying each document of the plurality of test documents in the reduced test set into a class selected from a plurality of document classes, such that each class has a new plurality of correctly classified test documents and a new set of misclassified test documents; and subsequently for each document class of the plurality of document classes:

identifying a corresponding subset of misclassified documents;

modifying the class definition of the class using (a) positively weighted correctly classified test documents and (b) negatively weighted misclassified test documents; and

removing the new set of misclassified documents from the reduced test set to produce a further reduced test set.

4 . A system comprising:

a computer processor coupled to a set of non-transitory computer memories, the non-transitory computer memories storing computer-executable instructions that, when executed by the computer processor, cause the computer processor to perform a process comprising:

training a natural language processing-based classification system by:

providing a plurality of pre-specified class definitions, wherein each pre-specified class definition of the plurality of pre-specified class definitions is defined pursuant to a Neural Network Model, and each class definition of the plurality of pre-specified class definitions is uniquely associated with a corresponding document class selected from a plurality of document classes;

providing a test set comprising a plurality of test documents, each test document in the test set having a pre-specified class assignment to a class selected from the plurality of document classes;

automatically classifying each document of the plurality of test documents in the test set into a class selected from a plurality of document classes, such that each class has a plurality of correctly classified test documents and a set of misclassified test documents;

for each document class of the plurality of document classes:

identifying a set of test documents from the test set that have been misclassified into said document class by the act of automatically classifying, by:

for each test document classified into said document class by the act of automatically classifying, comparing the document's pre-specified class assignment to the test document's resulting class, and

determining which of said test documents have a resulting class different from its pre-specified class, each document having a resulting class different from its pre-specified class being a misclassified document;

identifying, from the set of misclassified documents, a corresponding subset of misclassified documents;

modifying the class definition of the class to move the class definition in semantic space away from the misclassified documents by retraining a neural network based on content of the misclassified documents; and

removing the misclassified documents from the test set to produce a reduced test set.

5 . The system of claim 4 , the process further comprising:

identifying each misclassified document from the test set;

for each such identified misclassified document, identifying the misclassified document's pre-specified class assignment, to produce a set of target classes; and

for each target class of the set of target classes, modifying the class definition of said target class to move the class definition in semantic space toward said misclassified document.

6 . The system of claim 4 , wherein the process further includes, subsequent to producing the reduced test set:

classifying each document of the plurality of test documents in the reduced test set into a class selected from a plurality of document classes, such that each class has a new plurality of correctly classified test documents and a new set of misclassified test documents; and subsequently for each document class of the plurality of document classes:

identifying a corresponding subset of misclassified documents;

modifying the class definition of the class using (a) positively weighted correctly classified test documents and (b) negatively weighted misclassified test documents; and

removing the new set of misclassified documents from the reduced test set to produce a further reduced test set.

7 . A non-transitory computer readable medium storing computer-executable code, the computer-executable code comprising:

code for training a natural language processing-based classification system by:

code for providing a plurality of pre-specified class definitions, wherein each pre-specified class definition of the plurality of pre-specified class definitions is defined pursuant to a Neural Network Model, and each class definition of the plurality of pre-specified class definitions is uniquely associated with a corresponding document class selected from a plurality of document classes;

code for providing a test set comprising a plurality of test documents, each test document in the test set having a pre-specified class assignment to a class selected from the plurality of document classes;

code for automatically classifying each document of the plurality of test documents in the test set into a class selected from a plurality of document classes, such that each class has a plurality of correctly classified test documents and a set of misclassified test documents;

code for, for each document class of the plurality of document classes:

identifying a set of test documents from the test set that have been misclassified into said document class by the act of automatically classifying, by:

for each test document classified into said document class by the act of automatically classifying, comparing the document's pre-specified class assignment to the test document's resulting class, and

determining which of said test documents have a resulting class different from its pre-specified class, each document having a resulting class different from its pre-specified class being a misclassified document;

identifying, from the set of misclassified documents, a corresponding subset of misclassified documents;

code for modifying the class definition of the class to move the class definition in semantic space away from the misclassified documents by retraining a neural network based on content of the misclassified documents; and

code for removing the misclassified documents from the test set to produce a reduced test set.

8 . The non-transitory computer readable medium of claim 7 , the computer-executable code further comprising:

code for identifying each misclassified document from the test set;

code for, for each such identified misclassified document, identifying the misclassified document's pre-specified class assignment, to produce a set of target classes; and

code for, for each target class of the set of target classes, modifying the class definition of said target class to move the class definition in semantic space toward said misclassified document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2023
From: IAKOVENKO, OLEKSANDR; KELLY, WILLIAM; STEWART, STEPHEN LAND
To: NUIX LIMITED
Reel/Frame 063109/0294 →
Continuity (2)
Continuation In Part 17694243 · Mar 14, 2022
Related Publication 20230289531A1 · Sep 14, 2023
References Cited (25)
US 8024413B1 · Kolcz · 2011 [cited by examiner]
US 8509396B2 · Jan · 2013 [cited by examiner]
US 8620842B1 · Cormack · 2013 [cited by applicant]
US 8713023B1 · Cormack et al. · 2014 [cited by applicant]
US 8838606B1 · Cormack et al. · 2014 [cited by applicant]
US 9122681B2 · Cormack et al. · 2015 [cited by applicant]
US 9678957B2 · Cormack et al. · 2017 [cited by applicant]
US 9928234B2 · Kolotienko · 2018 [cited by examiner]
US 10229117B2 · Cormack et al. · 2019 [cited by applicant]
US 10242001B2 · Cormack et al. · 2019 [cited by applicant]
US 10353961B2 · Cormack et al. · 2019 [cited by applicant]
US 10445374B2 · Cormack et al. · 2019 [cited by applicant]
US 10671675B2 · Cormack et al. · 2020 [cited by applicant]
US 10817781B2 · Skiles · 2020 [cited by examiner]
US 11080340B2 · Cormack et al. · 2021 [cited by applicant]
US 11157829B2 · Kurata · 2021 [cited by examiner]
US 20120030157A1 · Tsuchida · 2012 [cited by examiner]
US 20150019211A1 · Simard · 2015 [cited by examiner]
US 20150324451A1 · Cormack · 2015 [cited by examiner]
US 20180052818A1 · Bethard · 2018 [cited by examiner]
US 20180349388A1 · Skiles et al. · 2018 [cited by applicant]
US 20180357531A1 · Giridhari · 2018 [cited by examiner]
US 20220198316A1 · Mazor · 2022 [cited by examiner]
International Search Report and Written Opinion for International Application No. PCT/US2023/015192, mailed Jun. 1, 2023 (11 pages). [cited by applicant]
Le, Q., et al.—Distributed Representations of Sentences and Documents, dated May 22, 2014, 9 pages. [cited by applicant]