IP Library Granted Patent US 12,380,271
Granted Patent B2
US 12,380,271 · App. 17/990,025 · Granted Aug 5, 2025

Federated system and method for analyzing language coherency, conformance, and anomaly detection

Inventor: Joel M. Hron, II (The Woodlands, TX)
Assignee: Thomson Reuters Enterprise Centre GmbH
G06F40/237G06F16/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,271
App. No.
17/990,025
Granted
Aug 5, 2025
Kind
B2
Abstract

Aspects of the present disclosure involve systems and methods for evaluating a piece of text or document against many corpuses of text or documents located on sources which may be the same and/or different from the text of interest in a tensorized manner and aggregating the coherence/anomaly score against some or all of the entire corpus. This joining of multiple data sources for evaluating the given piece or text may be a “federated” system as disparate data sources, each of which may contain confidential or otherwise private information, may be considered as a single repository of texts or documents. The systems and methods provide for a coherency and/or anomaly check of a piece of text of a document against similar pieces of text to determine a similarity of the piece of text to a large corpus of documents stored in disparate locations.

Claims (45)

1. A system for processing an electronic document, the system comprising:

a processor; and

a memory comprising instructions that, when executed, cause the processor to:

transmit an initial tensor generated from a text portion of an electronic document and a relative location of the text portion to a plurality of computing environments, each of the plurality of computing environments hosting a language coherency system remotely located separate from the processor and configured to calculate a comparison score based on a comparison of the initial tensor to a corpus of local electronic documents of a respective language coherency system;

generate a distribution of comparison scores received from each of the language coherency systems;

execute a risk scoring model to generate an assigned risk score of the initial tensor based on the comparison of the initial tensor to the corpus of local electronic documents, the assigned risk score associated with the comparison scores received from each of the language coherency systems;

display, on a display device, the distribution of comparison scores for the initial tensor for comparison of the text portion to a plurality of corresponding tensors from the corpus of local electronic documents and the assigned risk score; and

identify, based on the distribution of comparison scores, a similarity of the text portion of the electronic document to the corpus of local electronic documents of the respective language coherency systems while maintaining inaccessibility of the corpus of local electronic documents by the processor.

2. The system of claim 1 wherein the instructions further cause the processor to:

execute a conversion algorithm to convert the text portion of the electronic document into the initial tensor.

3. The system of claim 2 wherein the conversion algorithm is one of a hashing algorithm, a term frequency-inverse document frequency (tf-idf) algorithm, or a trained machine learning-based embedding model.

4. The system of claim 1 wherein the instructions further cause the processor to:

receive the initial tensor from a computing device in communication with the processor.

5. The system of claim 1 wherein the language coherency system is further configured to:

identify a portion of the corpus of local electronic documents of a same type as the text portion of the electronic document; and

generate the corresponding tensors based on the identified portion of the corpus of local electronic documents.

6. The system of claim 5 wherein calculating the comparison score comprises:

executing a distance-based scoring algorithm to determine a plurality of distance calculations in a dimensional ontological space, each of the plurality of distance calculations corresponding to a similarity of the text portion of the electronic document associated with the initial tensor to the identified portion of the corpus of local electronic documents.

7. The system of claim 1 wherein the plurality of computing environments is one of a public cloud computing environment, a private cloud computing environment, or a private tenant network.

8. The system of claim 1 wherein the distribution of comparison scores indicates a similarity of the text portion of the electronic document to the corpus of local electronic documents.

9. A method for processing a portion of an electronic document, the method comprising:

executing, via a processing device, a conversion algorithm to convert a text portion of the electronic document into an initial tensor;

transmitting the initial tensor and a relative location of the text portion to a plurality of computing environments different than the processing device, each of the plurality of computing environments hosting a language coherency system to calculate a similarity score through a comparison of the initial tensor to a corpus of local electronic documents of a respective language coherency system;

generating, via the processing device, a distribution of similarity scores generated by each of the language coherency systems based on the comparison of the initial tensor to a plurality of comparison tensors of local electronic documents; and

executing a risk scoring model to generate an assigned risk score of the initial tensor based on the comparison of the initial tensor to the corpus of local electronic documents, the assigned risk score associated with the comparison scores received from each of the language coherency systems.

10. The method of claim 9 further comprising:

executing a conversion algorithm to convert the text portion of the electronic document into the initial tensor.

11. The method of claim 10 wherein the conversion algorithm is one of a hashing algorithm, a term frequency-inverse document frequency (tf-idf) algorithm, or a trained machine learning-based embedding model.

12. The method of claim 9 further comprising:

identifying a portion of the corpus of local electronic documents of a same type as the text portion of the electronic document; and

generating the comparison tensors based on the identified portion of the corpus of local electronic documents.

13. The method of claim 12 wherein calculating the similarity score comprises:

executing a distance-based scoring algorithm to determine a plurality of distance calculations in a dimensional ontological space, each of the plurality of distance calculations corresponding to a similarity of the text portion of the electronic document associated with the initial tensor to the identified portion of the corpus of local electronic documents.

14. The method of claim 9 wherein the plurality of computing environments is one of a public cloud computing environment, a private cloud computing environment, or a private tenant network.

15. The method of claim 9 wherein the distribution of comparison scores indicates a similarity of the text portion of the electronic document to the corpus of local electronic documents.

16. One or more non-transitory computer-readable storage media storing computer-executable instructions for performing a computer process on a computing system, the computer process comprising:

executing, via a processing device, a conversion algorithm to convert a text portion of a received electronic document into an initial tensor;

transmitting the initial tensor and a relative location of the text portion to a plurality of computing environments, each of the plurality of computing environments hosting a language coherency system to calculate a similarity score through a comparison of the initial tensor to a corpus of local electronic documents of a respective language coherency system;

generating, via the processing device, a distribution of similarity scores generated by each of the language coherency systems based on the comparison of the initial tensor to a plurality of comparison tensors of local electronic documents; and

executing a risk scoring model to generate an assigned risk score of the initial tensor based on the comparison of the initial tensor to the corpus of local electronic documents, the assigned risk score associated with the comparison scores received from each of the language coherency systems.

17. The one or more non-transitory computer-readable storage media of claim 16 storing computer-executable instructions for performing the computer process on the computing system, the computer process further comprising:

executing a conversion algorithm to convert the text portion of the electronic document into the initial tensor, the conversion algorithm is one of a hashing algorithm, a term frequency-inverse document frequency (tf-idf) algorithm, or a trained machine learning-based embedding model.

18. The one or more non-transitory computer-readable storage media of claim 16 storing computer-executable instructions for performing the computer process on the computing system, the computer process further comprising:

identifying a portion of the corpus of local electronic documents of a same type as the text portion of the electronic document; and

generating the comparison tensors based on the identified portion of the corpus of local electronic documents, wherein calculating the similarity score comprises executing a distance-based scoring algorithm to determine a plurality of distance calculations in a dimensional ontological space, each of the plurality of distance calculations corresponding to a similarity of the text portion of the electronic document associated with the initial tensor to the identified portion of the corpus of local electronic documents.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2023
From: THOUGHTTRACE, INC.
To: WEST PUBLISHING CORPORATION
Reel/Frame 064186/0751 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2023
From: WEST PUBLISHING CORPORATION
To: THOMSON REUTERS ENTERPRISE CENTRE GMBH
Reel/Frame 064186/0882 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: HRON, JOEL M., II
To: THOUGHTTRACE, INC.
Reel/Frame 064046/0873 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2023
From: HRON, JOEL M., II
To: THOUGHTTRACE, INC.
Reel/Frame 063965/0184 →
Continuity (2)
Provisional Application 63283049 · Nov 24, 2021
Related Publication 20230161958A1 · May 25, 2023
References Cited (12)
US 7529719B2 · Liu · 2009 [cited by examiner]
US 9183203B1 · Tuchman et al. · 2015 [cited by applicant]
US 11226720B1 · Vandivere et al. · 2022 [cited by applicant]
US 20050010819A1 · Williams et al. · 2005 [cited by applicant]
US 20180349477A1 · Jaech · 2018 [cited by examiner]
US 20200394197A1 · Gutfreund · 2020 [cited by examiner]
US 20210216928A1 · O'Toole · 2021 [cited by examiner]
US 20220300505A1 · Wang · 2022 [cited by examiner]
PCT App. No. PCT/US2022/050368, International Search Report and Written Opinion, Mar. 6, 2023, 9 pages. [cited by applicant]
Miltsakaki et al., “Evaluation of text coherence for electronic essay scoring systems,” Natural Languange Engineering 10(1): 25-55, 2004. [cited by applicant]
Bader et al., “MATLAB Tensor Classes for Fast Algorithm Prototyping.” ACM Transactions on Mathematical Softward(TOMS) 32.4 (2006): 635-653. [cited by applicant]
Ansah et al., “Leveraging burst in twitter network communities for event detection,” Mar. 4, 2020, Retrieved Jan. 24, 2023, https://link.springer.com/article/10.1007/s11280-020-0786-y , 27 pages. [cited by applicant]