IP Library Granted Patent US 10,095,747
Granted Patent B1
US 10,095,747 · App. 15/174,135 · Granted Oct 9, 2018

Similar document identification using artificial intelligence

Inventor: Vishalkumar Rajpara (Ashburn, VA)
G06F17/30528G06F17/2705G06F17/2785G06F17/30011G06F17/30554G06N5/047G06N99/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,095,747
App. No.
15/174,135
Granted
Oct 9, 2018
Kind
B1
Abstract

Implementations generally relate to processing similar documents. In some implementations, a method includes receiving a plurality of documents related to e-discovery. The method further includes determining a seed document from the plurality of documents. The method further includes receiving a search request to search at least one selection of text in the seed document. The method further includes identifying other documents from the plurality of documents based on a similarity between text in the other documents and the at least one selection of text in the seed document. The method further includes generating a graphical user interface that includes a similarity panel that provides similarity data between text in the other documents and the at least one selection of text in the seed document.

Claims (69)

1. A method comprising:

receiving a plurality of documents related to e-discovery;

determining a seed document from the plurality of documents;

receiving a search request to search at least one selection of text in the seed document;

identifying other documents from the plurality of documents based on a similarity between text in the other documents and the at least one selection of text in the seed document; and

generating a graphical user interface that includes a similarity panel that provides similarity data between text in the other documents and the at least one selection of text in the seed document;

wherein the similarity panel provides:

a first number of the other documents having text that is identical to the at least one selection of text in the seed document based on a first predetermined similarity threshold;

a second number of the other documents having text that is similar to the at least one selection of text in the seed document based on a second predetermined similarity threshold; and

an option to search for a subset of the other documents based on a similarity percentage between the at least one selection of text in the seed document and text in the other documents.

2. The method of claim 1 , further comprising:

receiving a similarity request to identify the other documents having text that is similar to the at least one selection of text in the seed document; and

identifying the other documents having text that is similar to the at least one selection of text in the seed document.

3. The method of claim 1 , wherein the identifying of the other documents is based on the similarity between the at least one selection of text in the seed document and text in the other documents which includes using pattern recognition.

4. The method of claim 1 , wherein the graphical user interface includes a list of one or more legal issues to associate with any one or more documents of the plurality of documents.

5. The method of claim 1 , further comprising:

associating one or more legal issues with the seed document;

associating one or more other documents with the one or more legal issues associated with the seed document;

receiving a filter request that includes the one or more legal issues; and

filtering one or more other documents from the plurality of documents based on the filter request.

6. The method of claim 1 , wherein the graphical user interface includes, for each document of the plurality of documents, an option to view each document of the plurality of documents in a native format, a graphical format, a text format, a production format, a translated format, or an original format.

7. The method of claim 1 , further comprising:

enabling a user to redact one or more portions of the seed document; and

automatically redacting one or more corresponding portions of one or more other documents.

8. A non-transitory computer-readable storage medium carrying program instructions thereon, the instructions when executed by one or more processors cause the one or more processors to perform operations comprising:

receiving a plurality of documents related to e-discovery;

determining a seed document from the plurality of documents;

receiving a search request to search at least one selection of text in the seed document;

identifying other documents from the plurality of documents based on a similarity between text in the other documents and the at least one selection of text in the seed document; and

generating a graphical user interface that includes a similarity panel that provides similarity data between text in the other documents and the at least one selection of text in the seed document;

wherein the similarity panel provides:

a first number of the other documents having text that is identical to the at least one selection of text in the seed document based on a first predetermined similarity threshold;

a second number of the other documents having text that is similar to the at least one selection of text in the seed document based on a second predetermined similarity threshold; and

an option to search for a subset of the other documents based on a similarity percentage between the at least one selection of text in the seed document and text in the other documents.

9. The computer-readable storage medium of claim 8 , wherein the instructions when executed further cause the one or more processors to perform operations comprising:

receiving a similarity request to identify the other documents having text that is similar to the at least one selection of text in the seed document; and

identifying the other documents having text that is similar to the at least one selection of text in the seed document.

10. The computer-readable storage medium of claim 8 , wherein the identifying of the other documents is based on the similarity between the at least one selection of text in the seed document and text in the other documents which includes using pattern recognition.

11. The computer-readable storage medium of claim 8 , wherein the graphical user interface includes a list of one or more legal issues to associate with any one or more documents of the plurality of documents.

12. The computer-readable storage medium of claim 8 , wherein the instructions when executed further cause the one or more processors to perform operations comprising:

associating one or more legal issues with the seed document;

receiving a filter request that includes the one or more legal issues; and

filtering one or more other documents from the plurality of documents based on the filter request.

13. The computer-readable storage medium of claim 8 , wherein the graphical user interface includes, for each document of the plurality of documents, an option to view each document of the plurality of documents in a native format, a graphical format, a text format, a production format, a translated format, or an original format.

14. The computer-readable storage medium of claim 8 , wherein the instructions when executed further cause the one or more processors to perform operations comprising:

enabling a user to redact one or more portions of the seed document; and

automatically redacting one or more corresponding portions of one or more other documents.

15. A system comprising:

one or more processors; and

logic encoded in one or more non-transitory computer-readable media for execution by the one or more processors and when executed operable to perform operations comprising:

receiving a plurality of documents related to e-discovery;

determining a seed document from the plurality of documents;

receiving a search request to search at least one selection of text in the seed document;

identifying other documents from the plurality of documents based on a similarity between text in the other documents and the at least one selection of text in the seed document; and

generating a graphical user interface that includes a similarity panel that provides similarity data between text in the other documents and the at least one selection of text in the seed document;

wherein the similarity panel provides:

a first number of the other documents having text that is identical to the at least one selection of text in the seed document based on a first predetermined similarity threshold;

a second number of the other documents having text that is similar to the at least one selection of text in the seed document based on a second predetermined similarity threshold; and

an option to search for a subset of the other documents based on a similarity percentage between the at least one selection of text in the seed document and text in the other documents.

16. The system of claim 15 , wherein the logic when executed is further operable to perform operations comprising:

receiving a similarity request to identify the other documents having text that is similar to the at least one selection of text in the seed document; and

identifying the other documents having text that is similar to the at least one selection of text in the seed document.

17. The system of claim 15 , wherein the identifying of the other documents is based on the similarity between the at least one selection of text in the seed document and text in the other documents which includes using pattern recognition.

18. The system of claim 15 , wherein the graphical user interface includes a list of one or more legal issues to associate with any one or more documents of the plurality of documents.

19. The system of claim 15 , wherein the logic when executed is further operable to perform operations comprising:

associating one or more legal issues with the seed document;

receiving a filter request that includes the one or more legal issues; and

filtering one or more other documents from the plurality of documents based on the filter request.

20. The system of claim 15 , wherein the graphical user interface includes, for each document of the plurality of documents, an option to view each document of the plurality of documents in a native format, a graphical format, a text format, a production format, a translated format, or an original format.

Assignments (3)
SECURITY INTEREST Recorded Jan 22, 2025
From: AINS, LLC; CASEPOINT, LLC
To: CCP AGENCY, LLC
Reel/Frame 069972/0130 →
CHANGE OF NAME Recorded Mar 16, 2020
From: @LEGAL DISCOVERY LLC
To: CASEPOINT, LLC
Reel/Frame 052177/0699 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2016
From: RAJPARA, VISHALKUMAR
To: @LEGAL DISCOVERY LLC
Reel/Frame 038831/0056 →
Cited By (2)
US 12,346,432 US 12,585,708