IP Library Granted Patent US 12,499,702
Granted Patent B2
US 12,499,702 · App. 18/130,647 · Granted Dec 16, 2025

System and method for unsupervised document ontology generation of corpus of documents and partitioning the documents based on partitioning parameter

Inventor: Joel M. Hron, II (The Woodlands, TX)
Assignee: Thomson Reuters Enterprise Centre GmbH
G06V30/413G06V10/762G06V30/416G06T2207/30144G06V30/1444G06V30/19173
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,702
App. No.
18/130,647
Granted
Dec 16, 2025
Kind
B2
Abstract

Aspects of the present disclosure involve an automated, machine-learning technique for generating a representation of an ontology of a corpus of documents. This unsupervised generation of the ontology of the content of the documents may describe, based on the semantics of the language in the corpus and on the structure and format of the documents in that corpus, potentially key differentiable topics and sub-topics within the documents and the potential relationship between the topics and sub-topics. The unsupervised, or automated, generation of the ontology may provide a foundation of potential topics and sub-topics of a corpus of documents from which a complete ontology for the corpus of documents may be created. This ontology may be both pertinent in defining a structure through which an end user may interpret the data identified from a document or set of documents and/or to inform a machine-learning model to extract document information and classification.

Claims (52)

1 . A method for generating an ontology for a corpus of documents, the method comprising:

accessing, by a processor and from a database, a plurality of electronic documents;

partitioning, based on a partitioning parameter, each of the plurality of electronic documents into a plurality of partitions;

computing, by the processor, a word sequence embedding vector for each of the plurality of partitions;

clustering, based on a clustering parameter, the word sequence embedding vectors into one or more clusters of corresponding vectors by:

comparing the one or more clusters of corresponding vectors to a clustering criteria value comprising at least one of a number of corresponding vectors in the one or more clusters of corresponding vectors, an indication of each word sequence embedding vector is clustered with another word sequence embedding vector, or an indication of each of the plurality of electronic documents corresponds to at least one of a clustered vector; and

adjusting, based on the comparison to the clustering criteria value, the clustering parameter; and

assigning a subset of the plurality of partitions corresponding to a cluster of corresponding vectors to an ontology topic tier for the plurality of electronic documents.

2 . The method of claim 1 further comprising:

partitioning, based on a second partitioning parameter, each of the plurality of partitions into a plurality of sub-partitions;

computing, by the processor, a word sequence embedding vector for each of the plurality of sub-partitions; and

clustering, based on a second clustering parameter, the word sequence embedding vectors for each of the plurality of sub-partitions into one or more clusters of corresponding vectors for the plurality of sub-partitions.

3 . The method of claim 2 further comprising:

assigning the plurality of sub-partitions to an ontology sub-topic tier for the plurality of electronic documents, the ontology sub-topic tier dependent on the ontology topic tier.

4 . The method of claim 2 further comprising:

recursively partitioning the plurality of electronic documents and clustering the partitions until a stopping criteria value is obtained.

5 . The method of claim 1 wherein the partitioning parameter comprises at least one of partitioning based on a paragraph indicator, partitioning based on a section indicator, or partitioning based on a sentence indicator of the plurality of electronic documents.

6 . The method of claim 1 wherein the partitioning parameter comprises a span of x words of the plurality of electronic documents.

7 . The method of claim 1 wherein computing the word sequence embedding vector comprises executing a continuous bag-of-word model or a continuous skip-gram model for each of the plurality of partitions.

8 . A system for aggregating related documents, the system comprising:

a processor; and

a memory comprising instructions that, when executed, cause the processor to:

partition, based on a partitioning parameter, each of a plurality of electronic documents into a plurality of partitions;

compute, by the processor, a word sequence embedding vector for each of the plurality of partitions;

cluster, based on a clustering parameter, the word sequence embedding vectors into one or more clusters of corresponding vectors;

compare the one or more clusters of corresponding vectors to a clustering criteria value comprising at least one of a number of corresponding vectors in the one or more clusters of corresponding vectors, an indication of each word sequence embedding vector is clustered with another word sequence embedding vector, or an indication of each of the plurality of electronic documents corresponds to at least one of a clustered vector;

adjust, based on the comparison to the clustering criteria value, the clustering parameter;

associate a subset of the plurality of partitions corresponding to a cluster of corresponding vectors to an ontology topic tier for the plurality of electronic documents; and

generate a graphical user interface including a first portion displaying a visual representation of the plurality of partitions.

9 . The system of claim 8 wherein the processor is further caused to:

partition, based on a second partitioning parameter, each of the plurality of partitions into a plurality of sub-partitions;

compute a word sequence embedding vector for each of the plurality of sub-partitions; and

cluster, based on a second clustering parameter, the word sequence embedding vectors for each of the plurality of sub-partitions into one or more clusters of corresponding vectors for the plurality of sub-partitions.

10 . The system of claim 9 wherein the processor is further caused to:

assign the plurality of sub-partitions to an ontology sub-topic tier for the plurality of electronic documents, the ontology sub-topic tier dependent on the topic tier.

11 . The system of claim 9 wherein the processor is further caused to:

recursively partition the plurality of electronic documents and clustering the partitions until a stopping criteria value is obtained.

12 . One or more non-transitory computer-readable storage media storing computer-executable instructions for performing a computer process on a computing system, the computer process comprising:

accessing, by a processor and from a database, a plurality of electronic documents;

partitioning, based on a partitioning parameter, each of the plurality of electronic documents into a plurality of partitions;

computing, by the processor, a word sequence embedding vector for each of the plurality of partitions;

clustering, based on a clustering parameter, the word sequence embedding vectors into one or more clusters of corresponding vectors by:

comparing the one or more clusters of corresponding vectors to a clustering criteria value comprising at least one of a number of corresponding vectors in the one or more clusters of corresponding vectors, an indication of each word sequence embedding vector is clustered with another word sequence embedding vector, or an indication of each of the plurality of electronic documents corresponds to at least one of a clustered vector; and

adjusting, based on the comparison to the clustering criteria value, the clustering parameter; and

assigning a subset of the plurality of partitions corresponding to a cluster of corresponding vectors to an ontology topic tier for the plurality of electronic documents.

13 . The one or more non-transitory computer-readable storage media of claim 12 , the computer process further comprising:

partitioning, based on a second partitioning parameter, each of the plurality of partitions into a plurality of sub-partitions;

computing, by the processor, a word sequence embedding vector for each of the plurality of sub-partitions; and

clustering, based on a second clustering parameter, the word sequence embedding vectors for each of the plurality of sub-partitions into one or more clusters of corresponding vectors for the plurality of sub-partitions.

14 . The one or more non-transitory computer-readable storage media of claim 13 , the computer process further comprising:

assigning the plurality of sub-partitions to an ontology sub-topic tier for the plurality of electronic documents, the ontology sub-topic tier dependent on the ontology topic tier.

15 . The one or more non-transitory computer-readable storage media of claim 12 wherein the partitioning parameter comprises at least one of partitioning based on a paragraph indicator, partitioning based on a section indicator, or partitioning based on a sentence indicator of the plurality of electronic documents.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2023
From: THOUGHTTRACE, INC.
To: WEST PUBLISHING CORPORATION
Reel/Frame 064186/0751 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2023
From: WEST PUBLISHING CORPORATION
To: THOMSON REUTERS ENTERPRISE CENTRE GMBH
Reel/Frame 064186/0882 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2023
From: HRON, JOEL M., II
To: THOUGHTTRACE, INC.
Reel/Frame 063241/0480 →
Continuity (2)
Provisional Application 63329172 · Apr 8, 2022
Related Publication 20230326222A1 · Oct 12, 2023
References Cited (11)
US 10339375B2 · Castillo · 2019 [cited by examiner]
US 20060242180A1 · Graf et al. · 2006 [cited by applicant]
US 20060248053A1 · Sanfilippo et al. · 2006 [cited by applicant]
US 20130232147A1 · Mehra et al. · 2013 [cited by applicant]
US 20150161248A1 · Majkowska · 2015 [cited by applicant]
US 20150269431A1 · Haji · 2015 [cited by examiner]
US 20170004208A1 · Podder et al. · 2017 [cited by applicant]
US 20200311113A1 · Gao · 2020 [cited by examiner]
US 20200327172A1 · Coquard · 2020 [cited by examiner]
US 20210303838A1 · Sickert · 2021 [cited by examiner]
PCT App. No. PCT/US2023/017430, International Search Report and Written Opinion, Jun. 14, 2023, 9 pages. [cited by applicant]