IP Library Granted Patent US 12,380,149
Granted Patent B2
US 12,380,149 · App. 18/656,349 · Granted Aug 5, 2025

Systems and methods for document analysis to produce, consume and analyze content-by-example logs for documents

Inventors: Jeremy Garner Pickens (Bloomville, NY); Andrew Nelson Bye (San Francisco, CA); Thomas Chester Gricks, III (Irwin, PA)
Assignee: OPEN TEXT HOLDINGS, INC.
G06F16/35G06F40/10G06F40/289
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,149
App. No.
18/656,349
Granted
Aug 5, 2025
Kind
B2
Abstract

Document analysis systems and methods for the generation of a content-by-example log that expresses withheld documents in terms of a set of disclosed documents are disclosed. Additionally, document analysis systems and methods for the analysis of such a content-by-example log to determine withheld documents of interest without access to those withheld documents are disclosed.

Claims (63)

1. A method, comprising:

receiving a corpus of documents for potential production to a receiving party;

segregating, from within the corpus of documents, a first set of produced documents for production to the receiving party and a second set of non-produced documents for non-production to the receiving party;

generating a log of a comparison between each document in the second set of non-produced documents and a plurality of documents in the first set of produced documents, wherein comparing a non-produced document to the plurality of documents in the first set of produced documents comprises:

generating a feature vector for the non-produced document,

generating feature vectors for the plurality of documents in the first set of produced documents, and

using the feature vector of the non-produced document as a key to rank each of the plurality of documents in the first set of produced documents with respect to the non-produced document, wherein the ranking is based on the feature vectors of the plurality of documents in the first set of produced documents;

storing the ranked documents as an entry in the log of comparison; and

providing the log of comparison to the receiving party.

2. The method of claim 1 , wherein generating the log for the comparison comprises:

comparing each document in the set of non-produced documents to the plurality of documents in the produced set of documents and generating a similarity index based on at least one of:

the similarity index is produced between each non-produced document and each document in the plurality of produced documents;

the similarity index is produced between one or more sections in each of the non-produced documents and one or more sections in each document in the plurality of produced documents;

selecting a particular substantive passage in each of the non-produced documents and a corresponding substantive passage in each of the plurality of produced documents and the similarity index is produced between the selected substantive passages and corresponding substantive produced document passages; and

determining a set of topical regions in each of the non-produced documents and the similarity index is produced between the set of topical regions and regions in each of the plurality of produced documents.

3. The method of claim 2 , wherein the similarity index comprises a ranking of each of the plurality of produced documents based on similarity, the generated log comprising the ranking.

4. The method of claim 1 , further comprising:

producing, by a sending party, the corpus of documents; and

receiving the corpus of documents comprising:

receiving the corpus of documents from the sending party.

5. The method of claim 4 , wherein the document production is performed for electronic document discovery and the sending party and the receiving party are adverse to each other.

6. The method of claim 1 , wherein the generated log comprises a privilege log for the document production used for electronic discovery.

7. The method of claim 1 , further comprising:

requesting, by the receiving party, one or more of the non-produced documents based on the generated log of the comparison.

8. A system, comprising:

a processor;

a non-transitory computer readable medium, comprising instructions for:

receiving a corpus of documents for potential production to a receiving party;

segregating, from within the corpus of documents, a first set of produced documents for production to the receiving party and a second set of non-produced documents for non-production to the receiving party;

generating a log of a comparison between each document in the second set of non-produced documents and a plurality of documents in the first set of produced documents, wherein comparing a non-produced document to the plurality of documents in the first set of produced documents comprises:

generating a feature vector for the non-produced document,

generating feature vectors for the plurality of documents in the first set of produced documents, and

using the feature vector of the non-produced document as a key to rank each of the plurality of documents in the first set of produced documents with respect to the non-produced document, wherein the ranking is based on the feature vectors of the plurality of documents in the first set of produced documents;

storing the ranked documents as an entry in the log of comparison; and

providing the log of comparison to the receiving party.

9. The system of claim 8 , wherein generating the log for the comparison comprises:

comparing each document in the set of non-produced documents to the plurality of documents in the produced set of documents and generating a similarity index based on at least one of:

the similarity index is produced between each non-produced document and each document in the plurality of produced documents;

the similarity index is produced between one or more sections in each of the non-produced documents and one or more sections in each document in the plurality of produced documents;

selecting a particular substantive passage in each of the non-produced documents and a corresponding substantive passage in each of the plurality of produced documents and the similarity index is produced between the selected substantive passages and corresponding substantive produced document passages; and

determining a set of topical regions in each of the non-produced documents and the similarity index is produced between the set of topical regions and regions in each of the plurality of produced documents.

10. The system of claim 9 , wherein the similarity index comprises a ranking of each of the plurality of produced documents based on similarity, the generated log comprising the ranking.

11. The system of claim 8 , wherein the non-transitory computer readable medium further comprises instructions for:

producing, by a sending party, the corpus of documents; and

receiving the corpus of documents comprising:

receiving the corpus of documents from the sending party.

12. The system of claim 11 , wherein the document production is performed for electronic document discovery and the sending party and the receiving party are adverse to each other.

13. The system of claim 8 , wherein the generated log comprises a privilege log for the document production used for electronic discovery.

14. The system of claim 8 , further comprising:

requesting, by the receiving party, one or more of the non-produced documents based on the generated log of the comparison.

15. A non-transitory computer readable medium comprising instructions for:

receiving a corpus of documents for potential production to a receiving party;

segregating, from within the corpus of documents, a first set of produced documents for production to the receiving party and a second set of non-produced documents for non-production to the receiving party;

generating a log of a comparison between each document in the second set of non-produced documents and a plurality of documents in the first set of produced documents, wherein comparing a non-produced document to the plurality of documents in the first set of produced documents comprises:

generating a feature vector for the non-produced document,

generating feature vectors for the plurality of documents in the first set of produced documents, and

using the feature vector of the non-produced document as a key to rank each of the plurality of documents in the first set of produced documents with respect to the non-produced document, wherein the ranking is based on the feature vectors of the plurality of documents in the first set of produced documents;

storing the ranked documents as an entry in the log of comparison; and

providing the log of comparison to the receiving party.

16. The non-transitory computer readable medium of claim 15 , wherein the comparing is based on a similarity function.

17. The non-transitory computer readable medium of claim 16 , wherein the similarity function is determined by the receiving party.

18. The non-transitory computer readable medium of claim 16 , wherein the log of comparison includes metadata associated with the second set of non-produced documents.

19. The non-transitory computer readable medium of claim 18 , wherein comparing second set of non-produced documents to the first set of produced documents is based on the metadata.

Assignments (2)
MERGER Recorded May 18, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074678/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2024
From: PICKENS, JEREMY GARNER; BYE, ANDREW NELSON; GRICKS, THOMAS CHESTER, III
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 067520/0666 →
Continuity (2)
Continuation 17851493 · Jun 28, 2022
Related Publication 20240289372A1 · Aug 29, 2024
References Cited (17)
US 6847984B1 · Midgley · 2005 [cited by examiner]
US 8056128B1 · Dingle · 2011 [cited by examiner]
US 8363961B1 · Avidan · 2013 [cited by examiner]
US 8528084B1 · Dingle · 2013 [cited by examiner]
US 8539591B2 · Eguchi · 2013 [cited by examiner]
US 12147759B2 · Pickens · 2024 [cited by applicant]
US 20140047507A1 · Chang · 2014 [cited by examiner]
US 20150205968A1 · Aiello · 2015 [cited by examiner]
US 20170118271A1 · Reyes · 2017 [cited by examiner]
US 20180373711A1 · Ghatage · 2018 [cited by examiner]
US 20210103629A1 · Kiryu · 2021 [cited by examiner]
US 20250036864A1 · Pickens · 2025 [cited by applicant]
Office Action for U.S. Appl. No. 17/851,506, mailed Sep. 6, 2024, 24 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 17/851,506, mailed May 10, 2024, 22 pgs. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/851,513, mailed Jul. 2, 2024, 16 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 17/851,506, mailed Feb. 3, 2025, 13 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 17/851,506, mailed May 30, 2025, 10 pgs. [cited by applicant]
Cited By (1)
US 12,585,707