IP Library Granted Patent US 12,585,707
Granted Patent B2
US 12,585,707 · App. 17/851,506 · Granted Mar 24, 2026

Systems and methods for document analysis to produce, consume and analyze content-by-example logs for documents

Inventors: Jeremy Garner Pickens (Bloomville, NY); Andrew Nelson Bye (San Francisco, CA); Thomas Chester Gricks, III (Irwin, PA)
Assignee: OPEN TEXT INC.
G06F16/93G06F16/90335
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,707
App. No.
17/851,506
Granted
Mar 24, 2026
Kind
B2
Abstract

Document analysis systems and methods for the generation of a content-by-example log that expresses withheld documents in terms of a set of disclosed documents are disclosed. Additionally, document analysis systems and methods for the analysis of such a content-by-example log to determine withheld documents of interest without access to those withheld documents are disclosed.

Claims (76)

1 . A system for document analysis, comprising:

a processor;

a non-transitory computer readable medium, comprising instructions for:

receiving, by a receiving party, a content-by-example log, the content-by-example log including an entry for each withheld document of a set of withheld documents, wherein the set of withheld documents is inaccessible to the receiving party, wherein the entry for each withheld document associates an identifier for the corresponding withheld document with example identifiers for a set of example documents, wherein the set of example documents exemplify the corresponding withheld document, wherein the set of example documents are disclosed documents accessible to the receiving party;

storing the content-by-example log at a data store;

analyzing the content-by-example log to determine identifiers of withheld documents of interest by:

transforming the content-by-example log into a feature vector index, wherein the feature vector index comprises:

a feature vector associated with each of the identifiers of the withheld documents, wherein the feature vector comprises:

 a set of features determined from the set of example documents for the corresponding withheld document, wherein creating a respective feature vector for each withheld document comprises:

generating a document feature vector for each example document identified in the content-by-example log as being associated with the withheld document, wherein the document feature vector comprises a weighted set of text based features determined from the respective example document,

wherein the feature vector associates with the identifier for the withheld document in the feature vector index by using the document feature vectors generated for each example document of the set of example documents; and

determining the identifiers of withheld documents of interest based on the feature vector index by:

obtaining labels associated with identifiers of withheld documents;

obtaining a supervised machine learning model trained at a first time based on obtained labels for documents;

further training, at a second time after the first time, the supervised machine learning model using newly obtained labels for withheld documents identified in the content-by-example log, wherein the further training at the second time of the supervised machine learning model utilizes features determined from the example documents and provided by the feature vector index, wherein the features are associated with the withheld documents;

ranking identifiers for withheld documents of the content-by-example log based on the feature vector index using the further trained supervised machine learning model; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest; and

generating requests for a producing party having the set of withheld documents, wherein the generated requests correspond to at least some of the set of withheld documents of interest by specifying a subset of identifiers of the at least some of the set of withheld documents of interest.

2 . The system of claim 1 , wherein determining the identifiers of withheld documents of interest comprises:

searching the identifiers for the withheld documents using the feature vector index based on a query to rank the identifiers for the withheld documents; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

3 . The system of claim 2 , wherein the query is determined from content associated with the disclosed documents accessible by the receiving party.

4 . The system of claim 1 , wherein the features of the feature vector are the identifiers of the set of example documents.

5 . The system of claim 1 , wherein determining the identifiers of withheld documents of interest comprises:

generating a set of clusters of identifiers of withheld documents by clustering the identifiers for the withheld documents included in the content-by-example log based on the feature vector index;

selecting an identifier from each of the set of clusters of identifiers of withheld documents as identifiers of the set of withheld documents of interest.

6 . The system of claim 5 , wherein the identifier is selected from a cluster of the set of clusters based on a distance of that identifier from a centroid of that cluster.

7 . A method for document analysis, comprising:

receiving, by a receiving party, a content-by-example log, the content-by-example log including an entry for each withheld document of a set of withheld documents, wherein the set of withheld documents is inaccessible to the receiving party, wherein the entry for each withheld document associates an identifier for the corresponding withheld document with example identifiers for a set of example documents, wherein the set of example documents exemplify the corresponding withheld document, wherein the set of example documents are disclosed documents accessible to the receiving party;

storing the content-by-example log at a data store;

analyzing the content-by-example log to determine identifiers of withheld documents of interest by:

transforming the content-by-example log into a feature vector index, wherein the feature vector index comprises:

a feature vector associated with each of the identifiers of the withheld documents, wherein the feature vector comprises:

a set of features determined from the set of example documents for the corresponding withheld document, wherein creating a respective feature vector for each withheld document comprises:

 generating a document feature vector for each example document identified in the content-by-example log as being associated with the withheld document, wherein the document feature vector comprises a weighted set of text based features determined from the respective example document, wherein the feature vector associates with the identifier for the withheld document in the feature vector index by using the document feature vectors generated for each example document of the set of example documents; and

determining the identifiers of withheld documents of interest based on the feature vector index by:

obtaining labels associated with identifiers of withheld documents;

obtaining a supervised machine learning model trained at a first time based on obtained labels for documents;

further training, at a second time after the first time, the supervised machine learning model using newly obtained labels for withheld documents identified in the content-by-example log, wherein the further training at the second time of the supervised machine learning model utilizes features determined from the example documents and provided by the feature vector index, wherein the features are associated with the withheld documents;

ranking identifiers for withheld documents of the content-by-example log based on the feature vector index using the further trained supervised machine learning model; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest; and

generating requests for a producing party having the set of withheld documents, wherein the generated requests correspond to at least some of the set of withheld documents of interest by specifying a subset of identifiers of the at least some of the set of withheld documents of interest.

8 . The method of claim 7 , wherein determining the identifiers of withheld documents of interest comprises:

searching the identifiers for the withheld documents using the feature vector index based on a query to rank the identifiers for the withheld documents; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

9 . The method of claim 8 , wherein the query is determined from content associated with the disclosed documents accessible by the receiving party.

10 . The method of claim 7 , wherein the features of the feature vector are the identifiers of the set of example documents.

11 . The method of claim 7 , wherein determining the identifiers of withheld documents of interest comprises:

generating a set of clusters of identifiers of withheld documents by clustering the identifiers for the withheld documents included in the content-by-example log based on the feature vector index;

selecting an identifier from each of the set of clusters of identifiers of withheld documents as identifiers of the set of withheld documents of interest.

12 . The method of claim 11 , wherein the identifier is selected from a cluster of the set of clusters based on a distance of that identifier from a centroid of that cluster.

13 . A non-transitory computer readable medium, comprising instructions for:

receiving, by a receiving party, a content-by-example log, the content-by-example log including an entry for each withheld document of a set of withheld documents, wherein the set of withheld documents is inaccessible to the receiving party, wherein the entry for each withheld document associates an identifier for the corresponding withheld document with example identifiers for a set of example documents, wherein the set of example documents exemplify the corresponding withheld document, wherein the set of example documents are disclosed documents accessible to the receiving party;

storing the content-by-example log at a data store;

analyzing the content-by-example log to determine identifiers of withheld documents of interest by:

transforming the content-by-example log into a feature vector index, wherein the feature vector index comprises:

a feature vector associated with each of the identifiers of the withheld documents, wherein the feature vector comprises:

a set of features determined from the set of example documents for the corresponding withheld document, wherein creating a respective feature vector for each withheld document comprises:

generating a document feature vector for each example document identified in the content-by-example log as being associated with the withheld document, wherein the document feature vector comprises a weighted set of text based features determined from the respective example document,

wherein the feature vector associates with the identifier for the withheld document in the feature vector index by using the document feature vectors generated for each example document of the set of example documents; and

determining the identifiers of withheld documents of interest based on the feature vector index by:

obtaining labels associated with identifiers of withheld documents;

obtaining a supervised machine learning model trained at a first time based on obtained labels for documents;

further training, at a second time after the first time, the supervised machine learning model using newly obtained labels for withheld documents identified in the content-by-example log, wherein the further training at the second time of the supervised machine learning model utilizes features determined from the example documents and provided by the feature vector index, wherein the features are associated with the withheld documents;

ranking identifiers for withheld documents of the content-by-example log based on the feature vector index using the further trained supervised machine learning model; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest; and

generating requests for a producing party having the set of withheld documents, wherein the generated requests correspond to at least some of the set of withheld documents of interest by specifying a subset of identifiers of the at least some of the set of withheld documents of interest.

14 . The non-transitory computer readable medium of claim 13 , wherein determining the identifiers of withheld documents of interest comprises:

searching the identifiers for the withheld documents using the feature vector index based on a query to rank the identifiers for the withheld documents; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

15 . The non-transitory computer readable medium of claim 14 , wherein the query is determined from content associated with the disclosed documents accessible by the receiving party.

16 . The non-transitory computer readable medium of claim 13 , wherein the features of the feature vector are the identifiers of the set of example documents.

17 . The non-transitory computer readable medium of claim 13 , wherein determining the identifiers of withheld documents of interest comprises:

generating a set of clusters of identifiers of withheld documents by clustering the identifiers for the withheld documents included in the content-by-example log based on the feature vector index;

selecting an identifier from each of the set of clusters of identifiers of withheld documents as identifiers of the set of withheld documents of interest.

18 . The non-transitory computer readable medium of claim 17 , wherein the identifier is selected from a cluster of the set of clusters based on a distance of that identifier from a centroid of that cluster.

Assignments (2)
MERGER Recorded Jan 23, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 073567/0391 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2022
From: PICKENS, JEREMY GARNER; BYE, ANDREW NELSON; GRICKS, THOMAS CHESTER, III
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 060408/0352 →
Continuity (1)
Related Publication 20230418883A1 · Dec 28, 2023
References Cited (44)
US 5581702A · McArdle · 1996 [cited by applicant]
US 6847984B1 · Midgley · 2005 [cited by applicant]
US 7293033B1 · Tormasov · 2007 [cited by applicant]
US 8056128B1 · Dingle · 2011 [cited by applicant]
US 8290962B1 · Chu · 2012 [cited by applicant]
US 8363961B1 · Avidan · 2013 [cited by applicant]
US 8528084B1 · Dingle · 2013 [cited by applicant]
US 8539591B2 · Eguchi · 2013 [cited by applicant]
US 8640251B1 · Lee · 2014 [cited by applicant]
US 11775522B2 · Sharon · 2023 [cited by examiner]
US 12019667B2 · Pickens · 2024 [cited by applicant]
US 12147759B2 · Pickens · 2024 [cited by applicant]
US 12380149B2 · Pickens · 2025 [cited by applicant]
US 20090157601A1 · Lee · 2009 [cited by applicant]
US 20120310951A1 · Kumar · 2012 [cited by examiner]
US 20130097706A1 · Titonis · 2013 [cited by applicant]
US 20130346401A1 · Karidi · 2013 [cited by applicant]
US 20140047507A1 · Chang · 2014 [cited by applicant]
US 20150205968A1 · Aiello · 2015 [cited by applicant]
US 20170011044A1 · Sandland · 2017 [cited by applicant]
US 20170118271A1 · Reyes · 2017 [cited by applicant]
US 20180373711A1 · Ghatage · 2018 [cited by applicant]
US 20190102400A1 · Kumaran · 2019 [cited by applicant]
US 20200257709A1 · Bull · 2020 [cited by examiner]
US 20200278953A1 · Yang · 2020 [cited by examiner]
US 20210103629A1 · Kiryu · 2021 [cited by applicant]
US 20220019682A1 · Sislow · 2022 [cited by applicant]
US 20220284029A1 · Lal · 2022 [cited by applicant]
US 20230064367A1 · Satish · 2023 [cited by examiner]
US 20230418857A1 · Pickens · 2023 [cited by applicant]
US 20230419026A1 · Pickens · 2023 [cited by applicant]
US 20240289372A1 · Pickens · 2024 [cited by applicant]
US 20250036864A1 · Pickens · 2025 [cited by applicant]
US 20250328573A1 · Pickens · 2025 [cited by applicant]
Abril, Daniel, Guillermo Navarro-Arribas, and Vicenç Torra. “Towards a private vector space model for confidential documents.” Proceedings of the 28th Annual ACM Symposium on Applied Computing. 2013. (Year: 2013). [cited by examiner]
Office Action for U.S. Appl. No. 17/851,493, mailed Nov. 17, 2023, 22 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 17/851,493, mailed Mar. 6, 2024, 7 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 17/851,513, mailed Aug. 17, 2023, 11 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 17/851,513, mailed Dec. 7, 2023, 12 pgs. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/851,513, mailed Jul. 2, 2024, 16 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 18/656,349, mailed Dec. 31, 2024, 19 pgs. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/656,349, mailed Apr. 29, 2025, 5 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 18/904,265, mailed Oct. 1, 2025, 16 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 18/904,265, mailed Jan. 21, 2026, 6 pgs. [cited by applicant]