IP Library Granted Patent US 9,600,568
Granted Patent B2
US 9,600,568 · App. 13/110,806 · Granted Mar 21, 2017

Methods and systems for automatic evaluation of electronic discovery review and productions

Inventor: Venkat Rangan (San Jose, CA)
Assignee: Veritas Technologies LLC
G06F17/3069
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,600,568
App. No.
13/110,806
Granted
Mar 21, 2017
Kind
B2
Abstract

Techniques are provided for automatic sampling evaluation. An automatic sampling evaluation system enables users to evaluate convergence of one or more search processes. For example, given a set of searches that were validated by human review, a system can implement a retrieval process that samples one or more non-retrieved collections. Each individual document's similarity in the one or more non-retrieved collections is automatically evaluated to other documents in any retrieved sets. Given a goal of achieving a high recall, documents with high similarity can then be analyzed for additional noun phrases that may be used for a next iteration of a search. Convergence can be expected if the information gain in the new feedback loop is less than previous iterations, and if the additional documents identified are below a certain threshold document count.

Claims (68)

1. A method for evaluating a search process, the method comprising:

receiving, at a computer system comprising a processor, information identifying, in a collection of documents, a first set of documents that satisfies search criteria associated with a first search;

determining a document feature vector for each document in the first set of documents;

determining a first featured vector comprises at least one of the document feature vectors of the first set of documents;

identifying, in the collection of documents, respective documents that do not satisfy the search criteria associated with the first search;

receiving information identifying, in the respective documents, a second set of documents that satisfy first sampling criteria, wherein the second set of documents does not comprise all of the respective documents;

determining a document feature vector for each document in the second set of documents;

determining a second featured vector representing the second set of documents, wherein the second featured vector comprises at least one of the document feature vectors of the second set of documents;

determining whether a second search within the respective documents that do not satisfy the search criteria associated with the first search causes new document gain relative to the first search based on a measure of similarity between the first featured vector and the second featured vector exceeding a predetermined threshold value, wherein the second search is associated with the criteria of the first search;

generating information indicative of whether the second search of the collection of documents causes new document gain; and

displaying to a user of the generated information.

2. The method of claim 1 , wherein the first featured vector represents all documents in the first set of documents.

3. The method of claim wherein an exceeding measure of similarity compared to the predetermined threshold value indicates a likelihood to increase a number of documents produced in the second search.

4. The method of claim wherein the determining whether the second search within the respective documents causes new document gain comprises determining that the measure of similarity between the first featured vector and a respective document feature vector of at least one respective document in the second set of documents exceeds the predetermined threshold value.

5. The method of claim further comprising:

determining a set of noun phrases associated with the second search based on at least one document in the second set of documents; and

generating search criteria associated with the second search based on the search criteria associated with the first search and the determined set of noun phrases.

6. The method of claim further comprising:

determining whether a third search of the collection of documents causes new document gain based on a document feature vector generated for each document in a third set of documents that satisfies the search criteria associated with the second search and a document feature vector generated for at least one document in a fourth set of documents that does not satisfy the search criteria associated with the second search but satisfies second sampling criteria; and

generating information indicative of whether the third search of the collection of documents causes new document gain.

7. The method of claim wherein the determining the document feature vector for each document in the first set of documents comprises:

determining a plurality of term feature vectors for the document; and

generating the document feature vector for the document based on each term vector in the plurality of term feature vectors.

8. A non-transitory computer-readable medium having instructions that, when executed by a processor, cause the processor to perform operations comprising:

receiving information identifying, in a collection of documents, a first set of documents that satisfies search criteria associated with a first search;

determining a document feature vector for each document in the first set of documents;

determining a first featured vector representing the first set of documents, wherein the first featured vector comprises at least one of the document feature vectors of the first set of documents;

identifying, in the collection of documents, respective documents that do not satisfy the search criteria associated with the first search;

receiving information identifying, in the respective documents, a second set of documents that satisfy first sampling criteria, wherein the second set of documents does not comprise all of the respective documents;

determining a document feature vector for each document in the second set of documents;

determining a second featured vector representing the second set of documents, wherein the second features vector comprises at least one of the document feature vectors of the second set of documents;

determining whether a second search within the respective documents that do not satisfy the search criteria associated with the first search causes new document gain relative to the first search based on a measure of similarity between the first featured vector and the second featured vector exceeding a predetermined threshold value, wherein the second search is associated with the criteria of the first search;

generating information indicative of whether the second search of the collection of documents causes new document gain; and

displaying to a user of the generated information.

9. The non-transitory computer-readable medium of claim 8 , wherein the first featured vector represents all documents in the first set of documents.

10. The non-transitory computer-readable medium of claim 9 , wherein an exceeding measure of similarity compared to the predetermined threshold value indicates a likelihood to increase a number of documents produced in the second search.

11. The non-transitory computer-readable medium of claim wherein the determining whether the second search within the respective documents causes new document gain comprises determining that the measure of similarity between the first featured vector and a respective document feature vector of at least one respective document in the second set of documents exceeds the predetermined threshold value.

12. The non-transitory computer-readable medium of claim 8 , wherein the instructions, when executed by the processing device, cause the processing device to perform further operations comprising:

determining a set of noun phrases associated with the second search based on the at least one document in the second set of documents; and

generating search criteria associated with the second search based on the search criteria associated with the first search and the determined set of noun phrases.

13. The non-transitory computer-readable medium of claim 12 , wherein the instructions, when executed by the processing device, cause the processing device to perform further operations comprising:

determining whether a third search of the collection of documents causes new document gain based on a document feature vector generated for each document in a third set of documents that satisfies the search criteria associated with the second search and a document feature vector generated for at least one document in a fourth set of documents that does not satisfy the search criteria associated with the second search but satisfies second sampling criteria; and

generating information indicative of whether the third search of the collection of documents causes new document gain.

14. The non-transitory computer-readable medium of claim wherein the determining the document feature vector for each document in the first set of documents comprises:

determining a plurality of term feature vectors for the document; and

generating the document feature vector for the document based on each term vector in the plurality of term feature vectors.

15. A system for evaluating a search process of electronic discovery investigations, the system comprising:

a memory; and

a computer processor coupled to the memory, wherein the computer processor is configured to:

receive information identifying, in a collection of documents, a first set of documents that satisfies search criteria associated with a first search;

determine a document feature vector for each document in the first set of documents;

determine a first featured vector representing the first set of documents, wherein the first features vector comprises at least one of the document feature vectors of the first set of documents;

identify, in the collection of documents, respective documents that do not satisfy the search criteria associated with the first search;

receive information identifying, in the respective documents, a second set of documents that satisfy first sampling criteria, wherein the second set of documents does not comprise all of the respective documents;

determine a document feature vector for each document in the second set of documents;

determine a second featured vector representing the second set of documents, wherein the second featured vector comprises at least one of the document feature vectors of the second set of documents;

determine whether a second search within the respective documents that do not satisfy the search criteria associated with the first search causes new document gain relative to the first search based on a measure of similarity between the first a featured vector and the second featured vector exceeding a predetermined threshold value, wherein the second search is associated with the criteria of the first search;

generate information indicative of whether the second search of the collection of documents causes new document gain; and

display to a user of the generated information.

16. The system of claim 15 , wherein the first featured vector represents all documents in the first set of documents.

17. The system of claim 16 , wherein an exceeding measure of similarity compared to the predetermined threshold value indicates a likelihood to increase a number of documents produced in the second search.

18. The system of claim 15 , wherein to determine whether the second search within the respective documents causes new document gain comprises determining that the measure of similarity between the first featured vector and a respective document feature vector of at least one respective document in the second set of documents exceeds the predetermined threshold value.

19. The system of claim 15 , wherein the computer processor is further configured to:

determine a set of noun phrases associated with the second search based on the at least one document in the second set of documents; and

generate search criteria associated with the second search based on the search criteria associated with the first search and the determined set of noun phrases.

20. The system of claim 19 , wherein the computer processor is further configured to:

determine whether a third search of the collection of documents causes new document gain based on a document feature vector generated for each document in a third set of documents that satisfies the search criteria associated with the second search and a document feature vector generated for at least one document in a fourth set of documents that does not satisfy the search criteria associated with the second search but satisfies second sampling criteria; and

generate information indicative of whether the third search of the collection of documents causes new document gain.

Assignments (17)
SECURITY INTEREST Recorded Dec 12, 2025
From: ARCTERA US LLC
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073951/0470 →
TERMINATION AND RELEASE OF PATENT SECURITY AGREEMENT AT R/F 070530/0497 Recorded Dec 1, 2025
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: ARCTERA US LLC
Reel/Frame 073833/0730 →
TERMINATION AND RELEASE OF PATENT SECURITY AGREEMENT AT R/F 069585/0150 Recorded Dec 1, 2025
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: ARCTERA US LLC
Reel/Frame 073833/0848 →
RELEASE OF SECURITY INTEREST Recorded Dec 16, 2024
From: ACQUIOM AGENCY SERVICES LLC, AS COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC (F/K/A VERITAS US IP HOLDINGS LLC)
Reel/Frame 069712/0090 →
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 069634/0584 →
SECURITY INTEREST Recorded Dec 10, 2024
From: ARCTERA US LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 069563/0243 →
PATENT SECURITY AGREEMENT Recorded Dec 10, 2024
From: ARCTERA US LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 069585/0150 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC
To: ARCTERA US LLC
Reel/Frame 069548/0468 →
ASSIGNMENT OF SECURITY INTEREST IN PATENT COLLATERAL Recorded Nov 25, 2024
From: BANK OF AMERICA, N.A., AS ASSIGNOR
To: ACQUIOM AGENCY SERVICES LLC, AS ASSIGNEE
Reel/Frame 069440/0084 →
TERMINATION AND RELEASE OF SECURITY IN PATENTS AT R/F 037891/0726 Recorded Nov 30, 2020
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: VERITAS US IP HOLDINGS, LLC
Reel/Frame 054535/0814 →
SECURITY INTEREST Recorded Aug 20, 2020
From: VERITAS TECHNOLOGIES LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 054370/0134 →
MERGER Recorded Apr 18, 2016
From: VERITAS US IP HOLDINGS LLC
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 038483/0203 →
SECURITY INTEREST Recorded Feb 23, 2016
From: VERITAS US IP HOLDINGS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 037891/0001 →
SECURITY INTEREST Recorded Feb 23, 2016
From: VERITAS US IP HOLDINGS LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 037891/0726 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2016
From: SYMANTEC CORPORATION
To: VERITAS US IP HOLDINGS LLC
Reel/Frame 037693/0158 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2014
From: CLEARWELL SYSTEMS INC.
To: SYMANTEC CORPORATION
Reel/Frame 032381/0682 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2011
From: RANGAN, VENKAT
To: CLEARWELL SYSTEMS, INC.
Reel/Frame 026577/0527 →
Continuity (1)
Related Publication 20120296891A1 · Nov 22, 2012