IP Library › Granted Patent US 10,521,656
Granted Patent B2
US 10,521,656 · App. 15/811,118 · Granted Dec 31, 2019

Method and system for assessing similarity of documents

Inventors: Jeroen Mattijs van Rotterdam (Berkeley, CA); Michael T Mohen (Millington, MD); Chao Chen (Shanghai, CN); Kun Zhao (Shanghai, CN)
Assignee: OPEN TEXT CORPORATION
G06K9/00483G06F16/335G06F16/93G06K9/00469
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,521,656
App. No.
15/811,118
Granted
Dec 31, 2019
Kind
B2
Abstract

Systems and methods for assessing similarity of documents are provided. Embodiments of the systems and methods include extracting a reference document text from a reference document, extracting an archived document text from an archived document, and quantifying the reference document and the archived document. The systems and methods may also include determining a document similarity value of the quantified reference document and the archived document. Determining the document similarity value includes calculating a set of vector similarity values for a set of combinations of a reference document text vector and an archived document text vector, and calculating the document similarity value, including a sum of the plurality of vector similarity values.

Claims (58)

1. A system for assessing similarity of documents, comprising:

a document repository including an archived document and an archived document text vector for a portion of the archived document and archived document metadata vectors for archived document metadata, wherein the archived document text vector for the portion was created by tokenizing the portion of the archived document and vectorizing the tokenized portion to obtain the archived document text vector for the portion and the archived document metadata was created by tokenizing the archived document metadata and vectorizing the tokenized metadata to obtain the archived document metadata vectors;

a non-transitory computer readable medium, comprising instructions for:

extracting a reference document text from a reference document;

quantifying the reference document by:

tokenizing a portion of the reference document and vectorizing the tokenized portion to obtain a reference document text vector for the portion of the reference document, and

tokenizing reference document metadata and vectorizing the tokenized metadata to obtain a reference document metadata vector for the reference document metadata; and

determining a document similarity value of the quantified reference document and the archived document, by:

calculating a plurality of vector similarity values for a plurality of combinations of the reference document text vector and the archived document text vector;

calculating a second vector similarity value for the reference document metadata vector and an archived document metadata vector; and

calculating the document similarity value, by summing the plurality of vector similarity values and the second vector similarity value.

2. The system of claim 1 , wherein the portion comprises the reference document, a word, a linguistic unit, a sentence or an entire paragraph.

3. The system of claim 1 , wherein the instructions are further for reporting the archived document to a user if the document similarity value exceeds a threshold.

4. The system of claim 1 ,

wherein the reference document is one selected from the group consisting of a text document and a document comprising non-text content and text content; and

wherein the archived document is one selected from the group consisting of a text document and a document comprising non-text content and text content.

5. The system of claim 1 , wherein vectorizing the tokenized portion of the reference document and of the archived document is performed using a k-skip-n-gram vectorization,

wherein skip-grams of a length n with a maximum number of skipped tokens k are generated; and

wherein each skip-gram is separately vectorized.

6. The system of claim 1 , wherein calculating the plurality of vector similarity values further comprises multiplying each of the vector similarity values of the plurality of vector similarity values with a corresponding weight.

7. A method for assessing similarity of documents, comprising:

extracting a reference document text from a reference document;

quantifying the reference document by:

tokenizing a portion of the reference document and vectorizing the tokenized portion to obtain a reference document text vector for the portion of the reference document, and

tokenizing reference document metadata and vectorizinq the tokenized metadata to obtain a reference document metadata vector for the reference document metadata;

accessing a document repository including an archived document and an archived document text vector for a portion of the archived document and archived document metadata vectors for archived document metadata, wherein the archived document text vector for the portion was created by tokenizing the portion of the archived document and vectorizing the tokenized portion to obtain the archived document text vector for the portion and the archived document metadata was created by tokenizing the archived document metadata and vectorizing the tokenized metadata to obtain the archived document metadata vectors; and

determining a document similarity value of the quantified reference document and the archived document, by:

calculating a plurality of vector similarity values for a plurality of combinations of the reference document text vector and the archived document text vector;

calculating a second vector similarity value for the reference document metadata vector and an archived document metadata vector; and

calculating the document similarity value, by summing the plurality of vector similarity values and the second vector similarity value.

8. The method of claim 7 , wherein the portion comprises the reference document, a word, a linguistic unit, a sentence or an entire paragraph.

9. The method of claim 7 , further comprising reporting the archived document to a user if the document similarity value exceeds a threshold.

10. The method of claim 7 ,

wherein the reference document is one selected from the group consisting of a text document and a document comprising non-text content and text content; and

wherein the archived document is one selected from the group consisting of a text document and a document comprising non-text content and text content.

11. The method of claim 7 , wherein vectorizing the tokenized portion of the reference document and of the archived document is performed using a k-skip-n-gram vectorization,

wherein skip-grams of a length n with a maximum number of skipped tokens k are generated; and

wherein each skip-gram is separately vectorized.

12. The method of claim 7 , wherein calculating the plurality of vector similarity values further comprises multiplying each of the vector similarity values of the plurality of vector similarity values with a corresponding weight.

13. A non-transitory computer readable medium comprising instructions for:

extracting a reference document text from a reference document;

quantifying the reference document by:

tokenizing a portion of the reference document and vectorizing the tokenized portion to obtain a reference document text vector for the portion of the reference document, and

tokenizing reference document metadata and vectorizing the tokenized metadata to obtain a reference document metadata vector for the reference document metadata;

accessing a document repository including an archived document and an archived document text vector for a portion of the archived document and archived document metadata vectors for archived document metadata, wherein the archived document text vector for the portion was created by tokenizing the portion of the archived document and vectorizing the tokenized portion to obtain the archived document text vector for the portion and the archived document metadata was created by tokenizing the archived document metadata and vectorizing the tokenized metadata to obtain the archived document metadata vectors; and

determining a document similarity value of the quantified reference document and the archived document, by:

calculating a plurality of vector similarity values for a plurality of combinations of the reference document text vector and the archived document text vector;

calculating a second vector similarity value for the reference document metadata vector and an archived document metadata vector; and

calculating the document similarity value, by summing the plurality of vector similarity values and the second vector similarity value.

14. The non-transitory computer readable medium of claim 13 , wherein the portion comprises the reference document, a word, a linguistic unit, a sentence or an entire paragraph.

15. The non-transitory computer readable medium of claim 13 , wherein the instructions are further for reporting the archived document to a user if the document similarity value exceeds a threshold.

16. The non-transitory computer readable medium of claim 13 ,

wherein the reference document is one selected from the group consisting of a text document and a document comprising non-text content and text content; and

wherein the archived document is one selected from the group consisting of a text document and a document comprising non-text content and text content.

17. The non-transitory computer readable medium of claim 13 , wherein vectorizing the tokenized portion of the reference document and of the archived document is performed using a k-skip-n-gram vectorization,

wherein skip-grams of a length n with a maximum number of skipped tokens k are generated; and

wherein each skip-gram is separately vectorized.

18. The non-transitory computer readable medium of claim 13 , wherein calculating the plurality of vector similarity values further comprises multiplying each of the vector similarity values of the plurality of vector similarity values with a corresponding weight.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2017
From: EMC CORPORATION
To: OPEN TEXT CORPORATION
Reel/Frame 044357/0510 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2017
From: VAN ROTTERDAM, JEROEN MATTIJS; MOHEN, MICHAEL T.; CHEN, CHAO; ZHAO, KUN
To: EMC CORPORATION
Reel/Frame 044357/0531 →
Continuity (2)
Continuation 14871501 · Sep 30, 2015
Related Publication 20180068183A1 · Mar 8, 2018