IP Library Granted Patent US 12,019,667
Granted Patent B2
US 12,019,667 · App. 17/851,493 · Granted Jun 25, 2024

Systems and methods for document analysis to produce, consume and analyze content-by-example logs for documents

Inventors: Jeremy Garner Pickens (Bloomville, NY); Andrew Nelson Bye (San Francisco, CA); Thomas Chester Gricks, III (Irwin, PA)
Assignee: OPEN TEXT HOLDINGS, INC.
G06F16/35G06F40/10G06F40/289
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,019,667
App. No.
17/851,493
Granted
Jun 25, 2024
Kind
B2
Abstract

Document analysis systems and methods for the generation of a content-by-example log that expresses withheld documents in terms of a set of disclosed documents are disclosed. Additionally, document analysis systems and methods for the analysis of such a content-by-example log to determine withheld documents of interest without access to those withheld documents are disclosed.

Claims (75)

1. A system for document analysis and characterization, comprising:

a processor;

a non-transitory computer readable medium, comprising instructions for:

obtaining a set of withheld documents, where the withheld documents are accessible by a producing party and the withheld documents are inaccessible to a receiving party;

generating a content-by-example log for the withheld documents, the content-by-example log including an entry for each of the withheld documents by:

obtaining a set of disclosed documents, wherein the disclosed documents are accessible by both the producing party and the receiving party,

comparing each of the withheld documents to the set of disclosed documents using a similarity function, wherein comparing a withheld document to the set of disclosed documents comprises:

generating a feature vector for the withheld document,

generating feature vectors for the set of disclosed documents, and

using the feature vector of the withheld document as a key to rank each of the set of disclosed documents with respect to the withheld document based on the feature vectors of the set of disclosed documents,

determining a set of example documents for each withheld document based on the comparison of each withheld document to the set of disclosed documents, wherein each of the set of example documents is a disclosed document of the set of disclosed documents, and

creating the entry for each of the withheld documents in the content-by-example log based on the determined set of example documents, where the entry for each withheld document associates an identifier for that withheld document with identifiers for the set of example documents determined for that withheld document; and

providing the content-by-example log to the receiving party.

2. The system of claim 1 , wherein the disclosed documents are publicly accessible documents.

3. The system of claim 1 , wherein parameters of the similarity function are determined at least in part by the receiving party.

4. The system of claim 1 , wherein the content-by-example log includes metadata associated with the withheld documents.

5. The system of claim 4 , wherein the metadata includes a reason for withholding the withheld documents.

6. The system of claim 5 , wherein the reasons include a claim of privilege, confidentiality, or security.

7. The system of claim 1 , wherein comparison between a withheld document and a disclosed document is based on at least some of the content of the withheld document and at least some of the content of the disclosed document.

8. A method for document analysis and characterization, comprising:

obtaining a set of withheld documents, where the withheld documents are accessible by a producing party and the withheld documents are inaccessible to a receiving party;

generating a content-by-example log for the withheld documents, the content-by-example log including an entry for each of the withheld documents by:

obtaining a set of disclosed documents, wherein the disclosed documents are accessible by both the producing party and the receiving party, comparing each of the withheld documents to the set of disclosed documents using a similarity function, wherein comparing a withheld document to the set of disclosed documents comprises:

generating a feature vector for the withheld document,

generating feature vectors for the set of disclosed documents, and

using the feature vector of the withheld document as a key to rank each of the set of disclosed documents with respect to the withheld document based on the feature vectors of the set of disclosed documents,

determining a set of example documents for each withheld document based on the comparison of each withheld document to the set of disclosed documents, wherein each of the set of example documents is a disclosed document of the set of disclosed documents, and

creating the entry for each of the withheld documents in the content-by-example log based on the determined set of example documents, where the entry for each withheld document associates an identifier for that withheld document with identifiers for the set of example documents determined for that withheld document; and

providing the content-by-example log to the receiving party.

9. The method of claim 8 , wherein the disclosed documents are publicly accessible documents.

10. The method of claim 8 , wherein parameters of the similarity function are determined at least in part by the receiving party.

11. The method of claim 8 , wherein the content-by-example log includes metadata associated with the withheld documents.

12. The method of claim 11 , wherein the metadata includes a reason for withholding the withheld documents.

13. The method of claim 12 , wherein the reasons include a claim of privilege, confidentiality, or security.

14. The method of claim 8 , wherein comparison between a withheld document and a disclosed document is based on at least some of the content of the withheld document and at least some of the content of the disclosed document.

15. A non-transitory computer readable medium, comprising instructions for:

obtaining a set of withheld documents, where the withheld documents are accessible by a producing party and the withheld documents are inaccessible to a receiving party;

generating a content-by-example log for the withheld documents, the content-by-example log including an entry for each of the withheld documents by:

obtaining a set of disclosed documents, wherein the disclosed documents are accessible by both the producing party and the receiving party,

comparing each of the withheld documents to the set of disclosed documents using a similarity function, wherein comparing a withheld document to the set of disclosed documents comprises:

generating a feature vector for the withheld document,

generating feature vectors for the set of disclosed documents, and

using the feature vector of the withheld document as a key to rank each of the set of disclosed documents with respect to the withheld document based on the feature vectors of the set of disclosed documents,

determining a set of example documents for each withheld document based on the comparison of each withheld document to the set of disclosed documents, wherein each of the set of example documents is a disclosed document of the set of disclosed documents, and

creating the entry for each of the withheld documents in the content-by-example log based on the determined set of example documents, where the entry for each withheld document associates an identifier for that withheld document with identifiers for the set of example documents determined for that withheld document; and

providing the content-by-example log to the receiving party.

16. The non-transitory computer readable medium of claim 15 , wherein the disclosed documents are publicly accessible documents.

17. The non-transitory computer readable medium of claim 15 , wherein parameters of the similarity function are determined at least in part by the receiving party.

18. The non-transitory computer readable medium of claim 15 , wherein the content-by-example log includes metadata associated with the withheld documents.

19. The non-transitory computer readable medium of claim 18 , wherein the metadata includes a reason for withholding the withheld documents.

20. The non-transitory computer readable medium of claim 19 , wherein the reasons include a claim of privilege, confidentiality or security.

21. The non-transitory computer readable medium of claim 15 , wherein comparison between a withheld document and a disclosed document is based on at least some of the content of the withheld document and at least some of the content of the disclosed document.

22. A non-transitory computer readable medium, comprising instructions for:

receiving a corpus of documents for potential production to a receiving party;

segregating, from within the corpus of documents, a first set of documents for production to the receiving party and a second set of documents for non-production to the receiving party;

generating a log of a comparison between each document in the set of non-produced documents and a plurality of documents in the set of produced documents wherein comparing a non-produced document to the plurality of documents in the set of produced documents comprises:

generating a feature vector for the non-produced document,

generating feature vectors for the plurality of documents in the set of produced documents, and

using the feature vector of the non-produced document as a key to rank each of the plurality of documents in the set of produced documents with respect to the non-produced document based on the feature vectors of the plurality of documents in the set of produced documents; and

providing the comparison to the receiving party.

23. The non-transitory computer readable medium of claim 22 , wherein generating the log for the comparison comprises:

comparing each document in the set of non-produced documents to the plurality of documents in the produced set of documents and generating a similarity index based on at least one of:

the similarity index is produced between each non-produced document and each document in the plurality of produced documents;

the similarity index is produced between one or more sections in each of the non-produced documents and one or more sections in each document in the plurality of produced documents;

selecting a particular substantive passage in each of the non-produced documents and a corresponding substantive passage in each of the plurality of produced documents and the similarity index is produced between the selected substantive passages and corresponding substantive produced document passages; and

determining a set of topical regions in each of the non-produced documents and the similarity index is produced between the set of topical regions and regions in each of the plurality of produced documents.

24. The non-transitory computer readable medium of claim 23 , wherein the similarity index comprises a ranking of each of the plurality of produced documents based on similarity, the generated log comprising the ranking.

25. The non-transitory computer readable medium of claim 22 , further comprising:

producing, by a sending party, the corpus of documents; and

receiving the corpus of documents comprising:

receiving the corpus of documents from the sending party.

26. The non-transitory computer readable medium of claim 25 , wherein the document production is performed for electronic document discovery and the sending party and the receiving party are adverse to each other.

27. The non-transitory computer readable medium of claim 22 , wherein the generated log comprises a privilege log for the document production used for electronic discovery.

28. The non-transitory computer readable medium of claim 22 , further comprising:

requesting, by the receiving party, one or more of the non-produced documents based on the generated log of the comparison.

Assignments (2)
MERGER Recorded May 18, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074678/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2022
From: PICKENS, JEREMY GARNER; BYE, ANDREW NELSON; GRICKS, THOMAS CHESTER, III
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 060398/0551 →
Continuity (1)
Related Publication 20230418857A1 · Dec 28, 2023
Cited By (1)
US 12,585,707