IP Library › Granted Patent US 12,361,742
Granted Patent B2
US 12,361,742 · App. 18/314,618 · Granted Jul 15, 2025

Method and system for assessing similarity of documents

Inventors: Jeroen Mattijs van Rotterdam (Fort Lauderdale, FL); Michael T Mohen (Columbia, MD); Chao Chen (Shanghai, CN); Kun Zhao (Shanghai, CN)
Assignee: OPEN TEXT CORPORATION
G06V30/418G06F16/335G06F16/93G06V30/40G06V30/416G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,742
App. No.
18/314,618
Granted
Jul 15, 2025
Kind
B2
Abstract

Systems and methods for assessing similarity of documents are provided. Embodiments of the systems and methods include extracting a reference document text from a reference document, extracting an archived document text from an archived document, and quantifying the reference document and the archived document. The systems and methods may also include determining a document similarity value of the quantified reference document and the archived document. Determining the document similarity value includes calculating a set of vector similarity values for a set of combinations of a reference document text vector and an archived document text vector, and calculating the document similarity value, including a sum of the plurality of vector similarity values.

Claims (88)

1. A system for assessing similarity of documents, comprising:

a processor;

a document repository storing archived documents, wherein each of the archived documents comprises:

a plurality of archived document sentences; and

archived document metadata;

a non-transitory computer readable medium, storing instructions for:

tokenizing the sentences of each archived document;

vectorizing the tokens of the sentences of each archived document, thereby obtaining an archived document sentence vector for each of the archived documents;

tokenizing the metadata of each archived document;

vectorizing the tokens of the metadata of each archived document, thereby obtaining an archived document metadata vector for each of the archived documents;

obtaining reference document data associated with a reference document, the reference document data comprising:

a plurality of reference document sentences; and

reference document metadata;

tokenizing the sentences of the reference document; and

vectorizing the tokens of the sentences of the reference document;

tokenizing the reference document metadata for the reference document;

vectorizing the tokens of the reference document metadata;

in response to quantifying the reference document by tokenizing and vectorizing tokens of the reference document metadata for the reference document, performing a similarity analysis between the reference document and each of the one or more archived document by:

determining a degree of similarity between the reference document and the respective archived document based on:

comparing the reference document sentence vector to the respective archived document sentence vector; and

comparing the reference document metadata vector to the respective archived document metadata vector;

and

in response to performing the similarity analysis, identifying a number of the one or more archived documents based on the degree of similarity of the each of the one or more archived documents to the reference document.

2. The system of claim 1 , wherein the reference document data comprises text content or non-text content and each archived document comprises text content or non-text content.

3. The system of claim 2 , wherein the instructions further comprise instructions for obtaining an archived text vector for a portion of text of each of the archived documents, the archived document text vector for the portion of the archived document created by tokenizing the portion of text of the archived document and vectorizing the tokenized portion to obtain the archived document text vector for the portion for that archived document,

tokenizing a portion of the reference document data comprising text content and vectorizing the tokenized portion of the text content of the reference document data to obtain a reference document text vector for the portion of the text content of the reference document data, wherein the similarity analysis between the reference document and the archived document is based on the reference document text vector and the archived document text vector.

4. The system of claim 3 , wherein the portion of the reference document data comprising text content comprises the entire text content of the reference document data, a word, a linguistic unit, or a paragraph.

5. The system of claim 4 , wherein the instructions further comprise instructions for:

generating a reference document path for the portion of the reference document metadata or the portion of the text content of the reference document data;

for each archived document, generating an archived document path for the portion of the archived document metadata or the portion of the text of the archived document, wherein the similarity analysis between the reference document and the archived document is based on the reference document path and the archived document path.

6. The system of claim 1 , wherein identifying the number of the one or more archived documents comprises ranking the one or more archived documents based on the degree of similarity determined for each of the one or more archived documents and identifying the number of top ranked archived documents or identifying the number of the one or more archived documents having a degree of similarity with the reference document over a threshold.

7. A method for assessing similarity of documents, comprising:

accessing a document repository storing archived documents, wherein each of the archived documents comprises:

a plurality of archived document sentences; and

archived document metadata;

tokenizing the sentences of each archived document;

vectorizing the tokens of the sentences of each archived document, thereby obtaining an archived document sentence vector for each of the archived documents;

tokenizing the metadata of each archived document;

vectorizing the tokens of the metadata of each archived document, thereby obtaining an archived document metadata vector for each of the archived documents;

obtaining reference document data associated with a reference document, the reference document data comprising:

a plurality of reference document sentences; and

reference document metadata;

tokenizing the sentences of the reference document; and

vectorizing the tokens of the sentences of the reference document;

tokenizing the reference document metadata for the reference document;

vectorizing the tokens of the reference document metadata;

in response to quantifying the reference document by tokenizing and vectorizing tokens of the reference document metadata for the reference document, performing a similarity analysis between the reference document and each archived document by:

determining a degree of similarity between the reference document and the respective archived document based on;

comparing the reference document sentence vector to the respective archived document sentence vector; and

comparing the reference document metadata vector to the respective archived document metadata vector;

and

in response to performing the similarity analysis, identifying a number of the one or more archived documents based on the degree of similarity of the each of the one or more archived documents to the reference document.

8. The method of claim 7 , wherein the reference document data comprises text content or non-text content and each archived document comprises text content or non-text content.

9. The method of claim 8 , further comprising:

obtaining an archived text vector for a portion of text of each of the archived documents, the archived document text vector for the portion of the archived document created by tokenizing the portion of text of the archived document and vectorizing the tokenized portion to obtain the archived document text vector for the portion for that archived document,

tokenizing a portion of the reference document data comprising text content and vectorizing the tokenized portion of the text content of the reference document data to obtain a reference document text vector for the portion of the text content of the reference document data, wherein the similarity analysis between the reference document and the archived document is based on the reference document text vector and the archived document text vector.

10. The method of claim 9 , wherein the portion of the reference document data comprising text content comprises the entire text content of the reference document data, a word, a linguistic unit, a sentence or a paragraph.

11. The method of claim 10 , further comprising:

generating a reference document path for the portion of the reference document metadata or the portion of the text content of the reference document data;

for each archived document, generating an archived document path for the portion of the archived document metadata or the portion of the text of the archived document, wherein the similarity analysis between the reference document and the archived document is based on the reference document path and the archived document path.

12. The method of claim 7 , wherein identifying the number of the one or more archived documents comprises ranking the one or more archived documents based on the degree of similarity determined for each of the one or more archived documents and identifying the number of top ranked archived documents or identifying the number of the one or more archived documents having a degree of similarity with the reference document over a threshold.

13. A non-transitory computer readable medium, storing instructions for:

accessing a document tokenizing the sentences of each archived document;

vectorizing the tokens of the sentences of each archived document, thereby obtaining an archived document sentence vector for each of the archived documents;

tokenizing the metadata of each archived document;

vectorizing the tokens of the metadata of each archived document, thereby obtaining an archived document metadata vector for each of the archived documents;

obtaining reference document data associated with a reference document, the reference document data comprising:

a plurality of reference document sentences; and

reference document metadata;

tokenizing the sentences of the reference document; and

vectorizing the tokens of the sentences of the reference document;

tokenizing the reference document metadata for the reference document;

vectorizing the tokens of the reference document metadata;

in response to quantifying the reference document by tokenizing and vectorizing tokens of the reference document metadata for the reference document, performing a similarity analysis between the reference document and each archived document by:

determining a degree of similarity between the reference document and the respective archived document based on;

comparing the reference document sentence vector to the respective archived document sentence vector; and

comparing the reference document metadata vector to the respective archived document metadata vector;

and

in response to performing the similarity analysis, identifying a number of the one or more archived documents based on the degree of similarity of the each of the one or more archived documents to the reference document.

14. The non-transitory computer readable medium of claim 13 , wherein the reference document data comprises text content or non-text content and each archived document comprises text content or non-text content.

15. The non-transitory computer readable medium of claim 14 , further storing instructions for:

obtaining an archived text vector for a portion of text of each of the archived documents, the archived document text vector for the portion of the archived document created by tokenizing the portion of text of the archived document and vectorizing the tokenized portion to obtain the archived document text vector for the portion for that archived document,

tokenizing a portion of the reference document data comprising text content and vectorizing the tokenized portion of the text content of the reference document data to obtain a reference document text vector for the portion of the text content of the reference document data, wherein the similarity analysis between the reference document and the archived document is based on the reference document text vector and the archived document text vector.

16. The non-transitory computer readable medium of claim 15 , wherein the portion of the reference document data comprising text content comprises the entire text content of the reference document data, a word, a linguistic unit, or a paragraph.

17. The non-transitory computer readable medium of claim 16 , further storing instructions for:

generating a reference document path for the portion of the reference document metadata or the portion of the text content of the reference document data;

for each archived document, generating an archived document path for the portion of the archived document metadata or the portion of the text of the archived document, wherein the similarity analysis between the reference document and the archived document is based on the reference document path and the archived document path.

18. The non-transitory computer readable medium of claim 13 , wherein identifying the number of the one or more archived documents comprises ranking the one or more archived documents based on the degree of similarity determined for each of the one or more archived documents and identifying the number of top ranked archived documents or identifying the number of the one or more archived documents having a degree of similarity with the reference document over a threshold.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2023
From: VAN ROTTERDAM, JEROEN MATTIJS; MOHEN, MICHAEL T.; CHEN, CHAO; ZHAO, KUN
To: EMC CORPORATION
Reel/Frame 063731/0114 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2023
From: EMC CORPORATION
To: OPEN TEXT CORPORATION
Reel/Frame 063731/0133 →
Continuity (5)
Continuation 17192498 · Mar 4, 2021
Continuation 16692005 · Nov 22, 2019
Continuation 15811118 · Nov 13, 2017
Continuation 14871501 · Sep 30, 2015
Related Publication 20230282019A1 · Sep 7, 2023
References Cited (24)
US 6182066B1 · Marques · 2001 [cited by examiner]
US 6480835B1 · Light · 2002 [cited by examiner]
US 8761512B1 · Buddemeier · 2014 [cited by examiner]
US 9122747B2 · Inagaki · 2015 [cited by examiner]
US 10061766B2 · Bettersworth · 2018 [cited by examiner]
US 10331785B2 · Arthur · 2019 [cited by examiner]
US 10380151B2 · Miyahara · 2019 [cited by examiner]
US 10810357B1 · Tsypliaev · 2020 [cited by examiner]
US 11561987B1 · Sager · 2023 [cited by examiner]
US 20050055372A1 · Springer · 2005 [cited by examiner]
US 20070237401A1 · Coath · 2007 [cited by examiner]
US 20070266044A1 · Grondin · 2007 [cited by examiner]
US 20100331041A1 · Liao · 2010 [cited by examiner]
US 20110179035A1 · Zhang · 2011 [cited by examiner]
US 20130212090A1 · Sperling · 2013 [cited by examiner]
US 20140280223A1 · Ram · 2014 [cited by examiner]
US 20150206031A1 · Lindsay · 2015 [cited by examiner]
US 20160070732A1 · Bastedo · 2016 [cited by examiner]
US 20160350283A1 · Carus · 2016 [cited by examiner]
US 20160379009A1 · Lee · 2016 [cited by examiner]
US 20210342399A1 · Sisto · 2021 [cited by examiner]
US 20240346093A1 · Allieri · 2024 [cited by examiner]
Guthrie, 0., Allison, B., Liu, W., Guthrie, L., & Wilks, Y. (2006). A closer look at skipgram modelling. In: Proceedings of the fifth international conference on language resources and evaluation (LREC-2006) (pp. 1222-1… [cited by examiner]
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. & Dean, J. (2013). Distributed representations of words and phrases and their compositionality. In Proc. Advances in Neural Information Processing Systems 26 3111-3119. … [cited by examiner]