IP Library Granted Patent US 12,242,710
Granted Patent B2
US 12,242,710 · App. 18/541,901 · Granted Mar 4, 2025

Natural language processing system and method for documents

Inventors: Nicholas E. Vandivere (Spring, TX); Michael B. Kuykendall (Spring, TX)
Assignee: Thomson Reuters Enterprise Centre GmbH
G06F3/0482G06F3/0484G06F16/93G06F40/106G06F40/30G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,710
App. No.
18/541,901
Granted
Mar 4, 2025
Kind
B2
Abstract

In various embodiments, the disclosed systems and methods may receive documents, analyze the documents, categorize portions of the analyzed documents, and present the images of the documents and at least a portion of the categories. The analysis may include identification of categories and the presentation may include indicia of the portion of the image of the document related to the category. The systems and methods disclosed may allow querying and/or reporting of a plurality of documents to facilitate processing.

Claims (39)

1. A method for categorizing electronic documents, the method comprising:

receiving, by a processor, a plurality of electronic documents;

associating, by a plurality of trained machine learning models comprising a paragraph model trained to identify one or more categories associated with paragraphs of text and a sentence model trained to identify one or more subcategories of the one or more categories associated with sentences of text, a category and a subcategory for each of the plurality of electronic documents, the one or more categories corresponding to conceptual context of a content of the text;

identifying, by the processor, a conflict between a category and a subcategory associated with a first document of the plurality of electronic documents and a category and a subcategory associated with a second document of the plurality of documents;

removing, based on the identified conflict, an association of the category and the subcategory from the first document of the plurality of electronic documents; and

generating a graphical user interface comprising a navigable document image of the first document and the second document and a list of the associated category and the subcategory within the image of the first document and the second document.

2. The method of claim 1 , further comprising:

ordering, by a second trained machine learning model, the plurality of electronic documents chronologically.

3. The method of claim 2 wherein the second trained machine learning model: identifies a portion of the first document and a portion of the second document, each portion associated with a temporal component of the first document and the second document; and sorts the first document and the second document based on the temporal components of each portion.

4. The method of claim 3 wherein the second trained machine learning model identifies a text portion of the first document including a date as the temporal component.

5. The method of claim 3 wherein the second trained machine learning model identifies a text portion of the first document including a semantic identifier of a date as the temporal component.

6. The method of claim 1 wherein the second document of the plurality of electronic documents is more recent than the first document of the plurality of electronic documents, the removing of the association of the category and the subcategory from the first document of the plurality of electronic documents further based on the second document being more recent than the first document.

7. The method of claim 1 wherein identifying the conflict between the category and the subcategory associated with the first document and the category and a subcategory associated with the second document comprises comparing a vectorized string of text of the first document to a vectorized string of text of the second document.

8. The method of claim 1 , wherein the paragraph model is fed a paragraph of text and outputs whether the text conforms to the associated category.

9. The method of claim 8 , wherein the sentence model is fed a sentence of text and outputs whether the text conforms to the associated subcategory.

10. The method of claim 1 , wherein the electronic document is received as an image file and converted to a text format using optical character recognition software.

11. A system for categorizing electronic documents, the system comprising:

a processor; and

a memory storing instructions that, when executed, cause the processor to perform operations comprising:

receiving an electronic document;

associating, by a plurality of trained machine learning models comprising a paragraph model trained to identify one or more categories associated with paragraphs of text and a sentence model trained to identify one or more subcategories of the one or more categories associated with sentences of text, a category and a subcategory for each of the plurality of electronic documents, the one or more categories corresponding to conceptual context of a content of the text;

identifying, by the processor, a conflict between a category and a subcategory associated with a first document of the plurality of electronic documents and a category and a subcategory associated with a second document of the plurality of documents;

removing, based on the identified conflict, an association of the category and the subcategory from the first document of the plurality of electronic documents; and

generating a graphical user interface comprising a navigable document image of the first document and the second document and a list of the associated category and the subcategory within the image of the first document and the second document.

12. The system of claim 11 wherein the instructions further cause the processor to perform the operation of:

ordering, by a second trained machine learning model, the plurality of electronic documents chronologically.

13. The system of claim 12 wherein the second trained machine learning model: identifies a portion of the first document and a portion of the second document, each portion associated with a temporal component of the first document and the second document; and sorts the first document and the second document based on the temporal components of each portion.

14. The system of claim 13 wherein the second trained machine learning model identifies a text portion of the first document including a date as the temporal component.

15. The system of claim 13 wherein the second trained machine learning model identifies a text portion of the first document including a semantic identifier of a date as the temporal component.

16. The system of claim 11 wherein the second document of the plurality of electronic documents is more recent than the first document of the plurality of electronic documents, the removing of the association of the category and the subcategory from the first document of the plurality of electronic documents further based on the second document being more recent than the first document.

17. The system of claim 11 wherein identifying the conflict between the category and the subcategory associated with the first document and the category and a subcategory associated with the second document comprises comparing a vectorized string of text of the first document to a vectorized string of text of the second document.

18. The system of claim 11 , wherein the paragraph model is fed a paragraph of text and outputs whether the text conforms to the associated category.

19. The system of claim 18 , wherein the sentence model is fed a sentence of text and outputs whether the text conforms to the associated subcategory.

20. A non-transitory computer readable medium containing instructions which, when executed by a computer, cause the computer to perform the operations of:

receiving a plurality of electronic documents;

associating, by a plurality of trained machine learning models comprising a paragraph model trained to identify one or more categories associated with paragraphs of text and a sentence model trained to identify one or more subcategories of the one or more categories associated with sentences of text, a category and a subcategory for each of the plurality of electronic documents, the one or more categories corresponding to conceptual context of a content of the text;

identifying a conflict between a category and a subcategory associated with a first document of the plurality of electronic documents and a category and a subcategory associated with a second document of the plurality of documents;

removing, based on the identified conflict, an association of the category and the subcategory from the first document of the plurality of electronic documents; and

generating a graphical user interface comprising a navigable document image of the first document and the second document and a list of the associated category and the subcategory within the image of the first document and the second document.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2024
From: VANDIVERE, NICHOLAS E.; KUYKENDALL, MICHAEL B.
To: AGILE UPSTREAM GROUP, INC.
Reel/Frame 067804/0440 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2024
From: THOUGHTTRACE, INC.
To: WEST PUBLISHING CORPORATION
Reel/Frame 067804/0873 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2024
From: WEST PUBLISHING CORPORATION
To: THOMSON REUTERS ENTERPRISE CENTRE GMBH
Reel/Frame 067804/0904 →
CHANGE OF NAME Recorded Jun 21, 2024
From: AGILE UPSTREAM GROUP, INC.
To: THOUGHTTRACE, INC.
Reel/Frame 067807/0168 →
Continuity (6)
Continuation 17545662 · Dec 8, 2021
Continuation 15887689 · Feb 2, 2018
Provisional Application 62584527 · Nov 10, 2017
Provisional Application 62573542 · Oct 17, 2017
Provisional Application 62454648 · Feb 3, 2017
Related Publication 20240111396A1 · Apr 4, 2024
References Cited (14)
US 11226720B1 · Vandivere · 2022 [cited by examiner]
US 11861143B1 · Vandivere · 2024 [cited by examiner]
US 20040163034A1 · Colbath · 2004 [cited by examiner]
US 20060282442A1 · Lennon · 2006 [cited by examiner]
US 20090092374A1 · Kulas · 2009 [cited by examiner]
US 20100280989A1 · Mehra et al. · 2010 [cited by applicant]
US 20150372955A1 · Janakiraman · 2015 [cited by examiner]
US 20160162476A1 · Munro et al. · 2016 [cited by applicant]
US 20160210551A1 · Lee et al. · 2016 [cited by applicant]
US 20170228659A1 · Lin et al. · 2017 [cited by applicant]
US 20170357852A1 · Cai et al. · 2017 [cited by applicant]
US 20180165708A1 · Bajaj et al. · 2018 [cited by applicant]
US 20200160199A1 · Al Hasan et al. · 2020 [cited by applicant]
US 20200184016A1 · Roller · 2020 [cited by applicant]