IP Library Granted Patent US 11,954,098
Granted Patent B1
US 11,954,098 · App. 16/455,465 · Granted Apr 9, 2024

Natural language processing system and method for documents

Inventors: Joel M. Hron, II (The Woodlands, TX); Nicholas E. Vandivere (Spring, TX); Michael B. Kuykendall (Spring, TX)
Assignee: Thomson Reuters Enterprise Centre GmbH
G06F16/243G06F16/285G06F16/29G06F40/30G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,954,098
App. No.
16/455,465
Granted
Apr 9, 2024
Kind
B1
Abstract

In various embodiments, the disclosed systems and methods may receive documents, analyze the documents, categorize portions of the analyzed documents, and present the images of the documents and at least a portion of the categories. The analysis may include identification of categories and the presentation may include indicia of the portion of the image of the document related to the category. The systems and methods disclosed may allow querying and/or reporting of a plurality of documents to facilitate processing.

Claims (40)

1. A method for linking similar documents, the method comprising:

generating, by a processor and based on a first document, a first feature in the first document associated with an ontology, the first feature including a portion of a first vector containing predetermined values;

generating, by the processor and based on a second document, a second feature in the second document associated with the ontology, the second feature including a portion of a second vector containing the predetermined values; and

linking, by the processor, the first document and the second document by the first feature and the second feature, the linking being based on a measure of similarity of the first vector to the second vector being within a defined threshold range;

wherein the first vector, the second vector, and the predetermined values are identified by a plurality of trained machine learning models executed on the first document and the second document, the plurality of trained machine learning models comprising a learned paragraph model trained to identify an ontological category and a learned sentence model trained to identify an ontological sub-category of the ontological category that, when executed sequentially, generate the first feature and the second feature, wherein an output of the learned paragraph model is an input to the learned sentence model.

2. The method of claim 1 , wherein the first feature includes one of a first feature type, a first geographic location, or a semantic reference to the first geographic location and the second feature includes one of a second feature type, a second geographic location, or a semantic reference to the second geographic location.

3. The method of claim 2 , wherein the semantic reference to the first geographic location and the semantic reference to the second geographic location each include one of a latitude and longitude, a state, a country, a land grid, or a deed.

4. The method of claim 2 , further comprising identifying, by the processor, one or more of a shared geographic location including the first geographic location and the second geographic location or a shared entity identification between the first feature type and the second feature type.

5. The method of claim 4 , further comprising generating, by the processor, a map including a visual marking based on one of the first geographic location, the second geographic location, or the shared geographic location.

6. The method of claim 1 , further comprising:

receiving, by a processor, a selection of a portion of the first document, the portion including the first feature; and

retrieving, by the processor, one of the second document or a portion of the second document based on the linking of the first document and the second document.

7. The method of claim 1 , wherein a linkage between the first document and the second document, when linked, comprises a first value of a dimension, the first value associated with the first document, and a second value of the dimension, the second value associated with the second document, the first value and the second value within a threshold range of each other.

8. The method of claim 7 , wherein the linkage is determined at a time of search.

9. A system for linking similar documents, the system comprising:

a processor; and

a memory comprising instructions that, when executed, cause the processor to:

generate, based on a first document, a first feature in the first document associated with an ontology, the first feature including a portion of a first vector containing certain predetermined values;

generate, based on a second document, a second feature in the second document associated with the ontology, the second feature including a portion of a second vector containing certain predetermined values; and

link the first document and the second document by the first feature and the second feature, the first document and the second document being linked based on a similarity of the first vector to the second vector exceeding a defined threshold;

wherein the first vector, the second vector, and the certain predetermined values are identified by a plurality of trained machine learning models executed on the first document and the second document, the plurality of trained machine learning models comprising a learned paragraph model trained to identify an ontological category and a learned sentence model trained to identify an ontological sub-category of the ontological category that, when executed sequentially, generate the first feature and the second feature, wherein an output of the learned paragraph model is an input to the learned sentence model.

10. The system of claim 9 , wherein the first feature includes one of a first feature type, a first geographic location, or a semantic reference to the first geographic location and the second feature includes one of a second feature type, a second geographic location, or a semantic reference to the second geographic location.

11. The system of claim 10 , wherein the semantic reference to the first geographic location and the semantic reference to the second geographic location each include one of a latitude and longitude, a state, a country, a land grid, or a land deed.

12. The system of claim 10 , wherein the memory further comprises instructions that, when executed, cause the processor to identify one or more of a shared geographic location including the first geographic location and the second geographic location or a shared entity identification between the first feature type and the second feature type.

13. The system of claim 12 , wherein the memory further comprises instructions that, when executed, cause the processor to generate a map including a visual marking based on one of the first geographic location, the second geographic location, or the shared geographic location.

14. The system of claim 9 , wherein the memory further comprises instructions that, when executed, cause the processor to:

receive a selection of a portion of the first document, the portion including the first feature; and

retrieve the second document based on a linkage between the first document and the second document.

15. The system of claim 9 , wherein a linkage between the first document and the second document, when linked, comprises a first value of a dimension, the first value associated with the first document, and a second value of the dimension, the second value associated with the second document, the first value and the second value within a threshold range of each other.

16. The system of claim 15 , wherein the linkage is determined at a time of search.

17. A non-transitory computer readable medium which, when executed by one or more processors, causes the one or more processors to:

generate, based on a first document, a first feature in the first document associated with an ontology, the first feature including a portion of a first vector containing certain predetermined values;

generate, based on a second document, a second feature in the second document associated with the ontology, the second feature including a portion of a second vector containing certain predetermined values; and

link the first document and the second document by the first feature and the second feature, the first document and the second document being linked based on a similarity of the first vector to the second vector exceeding a defined threshold and comprising a first value of a dimension, the first value associated with the first document, and a second value of the dimension, the second value associated with the second document, the first value and the second value within a range of each other based on the defined threshold;

wherein the first vector, the second vector, and the certain predetermined values are identified by a plurality of trained machine learning models executed on the first document and the second document, the plurality of trained machine learning models comprising a learned paragraph model trained to identify an ontological category and a learned sentence model trained to identify an ontological sub-category of the ontological category that, when executed sequentially, generate the first feature and the second feature, wherein an output of the learned paragraph model is an input to the learned sentence model.

18. The non-transitory computer readable medium of claim 17 , wherein the first feature includes one of a first feature type, a first geographic location, or a semantic reference to the first geographic location and the second feature includes one of a second feature type, a second geographic location, or a semantic reference to the second geographic location, and further causing the one or more processors to identify one or more of a shared geographic location including the first geographic location and the second geographic location or a shared entity identification between the first feature type and the second feature type.

19. The non-transitory computer readable medium of claim 17 , further causing the one or more processors to:

receive a selection of a portion of the first document, the portion including the first feature; and

retrieve the second document based on a linkage between the first document and the second document.

20. The non-transitory computer readable medium of claim 19 , wherein the linkage is determined at a time of search.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2023
From: WEST PUBLISHING CORPORATION
To: THOMSON REUTERS ENTERPRISE CENTRE GMBH
Reel/Frame 063730/0782 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2023
From: THOUGHTTRACE, INC.
To: WEST PUBLISHING CORPORATION
Reel/Frame 063686/0329 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2021
From: HRON, JOEL M., II; VANDIVERE, NICHOLAS E.; KUYKENDALL, MICHAEL B.
To: THOUGHTTRACE, INC.
Reel/Frame 057565/0718 →
Continuity (5)
Continuation In Part 15887689 · Feb 2, 2018
Provisional Application 62690759 · Jun 27, 2018
Provisional Application 62584527 · Nov 10, 2017
Provisional Application 62573542 · Oct 17, 2017
Provisional Application 62454648 · Feb 3, 2017
Cited By (1)
US 12,314,328