IP Library Granted Patent US 10,083,176
Granted Patent B1
US 10,083,176 · App. 15/056,616 · Granted Sep 25, 2018

Methods and systems to efficiently find similar and near-duplicate emails and files

Inventors: Malay Desai (Freemont, CA); Medha Shewale (San Carlos, CA); Venkat Rangan (San Jose, CA)
Assignee: Veritas Technologies LLC
G06F17/30011G06F17/3069G06F17/30619G06F17/30684
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,083,176
App. No.
15/056,616
Granted
Sep 25, 2018
Kind
B1
Abstract

A set of trigrams can be generated for each document in a plurality of documents processed by an e-discovery system. Each trigram in the set of trigrams for a given document is a sequence of three terms in the given document. A set of trigrams for each similar document is then determined based on the set of trigrams for the original document. To facilitate identification of the similar documents, a full text index is then generated for the plurality of documents and the set of trigrams for each document are indexed into the full text index, as individual terms. Queries can be generated into the full text index based on trigrams of a document to determine other similar or near-duplicate documents. After a set of potentially similar documents are identified, a separate distance criteria can be applied to evaluate the level of similarity between the two documents in an efficient way.

Claims (56)

1. A method for generating and using a semantic space in a computer system, comprising:

receiving, via at least one computer processor, a plurality of documents and a plurality of random document vectors, wherein each random document vector in the plurality of random document vectors is generated based on random indexing and is associated with a corresponding document in the plurality of documents;

selecting, via the at least one computer processor, a term for which to generate a term vector;

for a first document in the plurality of documents, determining, via the at least one computer processor, if the term appears in the document;

if the term appears in the document:

determining, via the at least one computer processor, a frequency of the term in the document; and

adding, via the at least one computer processor, an associated random document vector of the document to the term vector, wherein the associated random document vector is scaled by the term frequency

determining, via the at least one computer processor, if the term appears in any remaining documents in the plurality of documents;

if the term does not appear in any remaining documents in the plurality of documents, generating, via the at least one computer processor, a normalized version of the term vector with the added associated random document vector;

outputting, via the at least one computer processor, the normalized term vector;

receiving a query; and

generating, based at least in part on the semantic space including the normalized version of the term vector, a user interface displaying similar and near-duplicate documents.

2. The method of claim 1 , wherein the term is selected from a plurality of terms from a corpus of the plurality of documents.

3. The method of claim 2 , wherein the plurality of terms do not include terms with low Inverse Document Frequency and do not include terms with low global term frequency associated with the plurality of documents.

4. The method of claim 2 , wherein the plurality of terms do not include language specific characters.

5. The method of claim 2 , wherein the plurality of terms do not include terms with low global term frequency associated with the plurality of documents.

6. The method of claim 1 , wherein the plurality of documents are partitioned and each partition is processed independently.

7. The method of claim 1 , wherein at least one of the plurality of random document vectors include a plurality of floating point values.

8. A system for generating and using a semantic space in a computer system, comprising one or more computer processors configured to:

receive a plurality of documents and a plurality of random document vectors, wherein each random document vector in the plurality of random document vectors is generated based on random indexing and is associated with a corresponding document in the plurality of documents;

select a term for which to generate a term vector;

for a first document in the plurality of documents, determine if the term appears in the document; and

if the term appears in the document:

determine a frequency of the term in the document;

add an associated random document vector of the document to the term vector, wherein the associated random document vector is scaled by the term frequency;

determine if the term appears in any remaining documents in the plurality of documents;

if the term does not appear in any remaining documents in the plurality of documents, generate a normalized version of the term vector with the added associated random document vector;

output the normalized term vector;

receiving a query; and

generating, based at least in part on the semantic space including the normalized version of the term vector, a user interface displaying similar and near-duplicate documents.

9. The system of claim 8 , wherein the term is selected from a plurality of terms from a corpus of the plurality of documents.

10. The system of claim 9 , wherein the plurality of terms do not include terms with low Inverse Document Frequency.

11. The system of claim 9 , wherein the plurality of terms do not include language specific characters.

12. The system of claim 9 , wherein the plurality of terms do not include terms with low global term frequency associated with the plurality of documents.

13. The system of claim 8 , wherein the plurality of documents are partitioned and each partition is processed independently.

14. The system of claim 8 , wherein at least one of the plurality of random document vectors include a plurality of floating point values.

15. An article of manufacture for generating and using a semantic space in a computer system, the article of manufacture comprising:

at least one processor readable storage medium; and

instructions stored on the at least one medium;

wherein the instructions are configured to be readable from the at least one medium by at least one processor and thereby cause the at least one processor to operate so as to:

receive a plurality of documents and a plurality of random document vectors, wherein each random document vector in the plurality of random document vectors is generated based on random indexing and is associated with a corresponding document in the plurality of documents;

select a term for which to generate a term vector;

for a first document in the plurality of documents, determine if the term appears in the document; and

if the term appears in the document:

determine a frequency of the term in the document;

add an associated random document vector of the document to the term vector, wherein the associated random document vector is scaled by the term frequency

determine if the term appears in any remaining documents in the plurality of documents;

if the term does not appear in any remaining documents in the plurality of documents, generate a normalized version of the term vector with the added associated random document vector; and

output the normalized term vector;

receive a query; and

generate, based at least in part on the semantic space including the normalized version of the term vector, a user interface displaying similar and near-duplicate documents.

16. The article of manufacture of claim 15 , wherein the term is selected from a plurality of terms from a corpus of the plurality of documents.

17. The article of manufacture of claim 16 , wherein the plurality of terms do not include terms with low Inverse Document Frequency and do not include terms with low global term frequency associated with the plurality of documents.

18. The article of manufacture of claim 16 , wherein the plurality of terms do not include language specific characters.

19. The article of manufacture of claim 16 , wherein the plurality of terms do not include terms with low global term frequency associated with the plurality of documents.

20. The article of manufacture of claim 15 , wherein the plurality of documents are partitioned and each partition is processed independently.

Assignments (16)
SECURITY INTEREST Recorded Dec 12, 2025
From: ARCTERA US LLC
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073951/0470 →
TERMINATION AND RELEASE OF PATENT SECURITY AGREEMENT AT R/F 070530/0497 Recorded Dec 1, 2025
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: ARCTERA US LLC
Reel/Frame 073833/0730 →
TERMINATION AND RELEASE OF PATENT SECURITY AGREEMENT AT R/F 069585/0150 Recorded Dec 1, 2025
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: ARCTERA US LLC
Reel/Frame 073833/0848 →
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 069632/0613 →
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 069634/0584 →
PATENT SECURITY AGREEMENT Recorded Dec 10, 2024
From: ARCTERA US LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 069585/0150 →
SECURITY INTEREST Recorded Dec 10, 2024
From: ARCTERA US LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 069563/0243 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC
To: ARCTERA US LLC
Reel/Frame 069548/0468 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS AT R/F 052426/0001 Recorded Nov 30, 2020
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 054535/0565 →
SECURITY INTEREST Recorded Aug 20, 2020
From: VERITAS TECHNOLOGIES LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 054370/0134 →
PATENT SECURITY AGREEMENT SUPPLEMENT Recorded Apr 16, 2020
From: VERITAS TECHNOLOGIES, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 052426/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2018
From: SHEWALE, MEDHA; DESAI, MALAY; RANGAN, VENKAT
To: CLEARWELL SYSTEMS, INC.
Reel/Frame 044524/0230 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2018
From: SYMANTEC CORPORATION
To: VERITAS US IP HOLDINGS LLC
Reel/Frame 044991/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2018
From: CLEARWELL SYSTEMS INC.
To: SYMANTEC CORPORATION
Reel/Frame 044524/0153 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2018
From: VERITAS US IP HOLDINGS LLC
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 044524/0084 →
PATENT SECURITY AGREEMENT Recorded Nov 23, 2016
From: VERITAS TECHNOLOGIES LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 040679/0466 →
Continuity (1)
Continuation 13028841 · Feb 16, 2011
Cited By (6)
US 12,265,787 US 12,386,872 US 12,450,295 US 12,481,708 US 12,579,198 US 12,608,414