IP Library Patent Application 18904265
Patent Application
App. No. 18/904,265

Systems and Methods for Document Analysis to Produce, Consume and Analyze Content-By-Example Logs for Documents

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/904,265
Abstract

Document analysis systems and methods for the generation of a content-by-example log that expresses withheld documents in terms of a set of disclosed documents are disclosed. Additionally, document analysis systems and methods for the analysis of such a content-by-example log to determine withheld documents of interest without access to those withheld documents are disclosed.

Claims (59)

1 . A system for document analysis, comprising:

a processor;

a non-transitory computer readable medium, comprising instructions for:

receiving, by a receiving party, a content-by-example log, the content-by-example log including an entry for each of a set of withheld documents inaccessible to the receiving party, wherein the entry for each withheld document associates an identifier for that withheld document with identifiers for a set of example documents for that withheld document, and the set of example documents are disclosed documents accessible to the receiving party;

analyzing the content-by-example log to determine identifiers of withheld documents of interest by:

creating a feature vector index based on the content-by-example log, wherein the feature vector index comprises a feature vector associated with each of the identifiers of the withheld documents, and the feature vector associated with the identifier for a withheld document comprises a set of features determined based on the identified set of example documents associated with that identified withheld document; and

determining the identifiers of withheld documents of interest based on the feature vector index.

2 . The system of claim 1 , wherein determining the identifiers of withheld documents of interest comprises:

searching the identifiers for the withheld documents using the feature vector index based on a query to rank the identifiers for the withheld documents; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

3 . The system of claim 2 , wherein the query is determined from content associated with the disclosed documents accessible by the receiving party.

4 . The system of claim 1 , wherein the features of the feature vector are the identifiers of the set of example documents.

5 . The system of claim 1 , wherein determining the identifiers of withheld documents of interest comprises:

obtaining labels associated with identifiers of withheld documents;

training a supervised machine learning model based on the obtained labels for withheld documents, wherein the supervised machine learning model is trained based on features associated with the labeled withheld documents in the feature vector index;

ranking identifiers for withheld documents based on the feature vector index using the supervised machine learning model; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

6 . The system of claim 1 , wherein determining the identifiers of withheld documents of interest comprises:

generating a set of clusters of identifiers of withheld documents by clustering the identifiers for the withheld documents included in the content-by-example log based on the feature vector index; and

selecting an identifier from each of the set of clusters of identifiers of withheld documents as identifiers of the set of withheld documents of interest.

7 . The system of claim 6 , wherein the identifier is selected from a cluster of the set of clusters based on a distance of that identifier from a centroid of that cluster.

8 . A method for document analysis, comprising:

receiving, by a receiving party, a content-by-example log, the content-by-example log including an entry for each of a set of withheld documents inaccessible to the receiving party, wherein the entry for each withheld document associates an identifier for that withheld document with identifiers for a set of example documents for that withheld document, and the set of example documents are disclosed documents accessible to the receiving party;

analyzing the content-by-example log to determine identifiers of withheld documents of interest by:

creating a feature vector index based on the content-by-example log, wherein the feature vector index comprises a feature vector associated with each of the identifiers of the withheld documents, and the feature vector associated with the identifier for a withheld document comprises a set of features determined based on the identified set of example documents associated with that identified withheld document; and

determining the identifiers of withheld documents of interest based on the feature vector index.

9 . The method of claim 8 , wherein determining the identifiers of withheld documents of interest comprises:

searching the identifiers for the withheld documents using the feature vector index based on a query to rank the identifiers for the withheld documents; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

10 . The method of claim 9 , wherein the query is determined from content associated with the disclosed documents accessible by the receiving party.

11 . The method of claim 8 , wherein the features of the feature vector are the identifiers of the set of example documents.

12 . The method of claim 8 , wherein determining the identifiers of withheld documents of interest comprises:

obtaining labels associated with identifiers of withheld documents;

training a supervised machine learning model based on the obtained labels for withheld documents, wherein the supervised machine learning model is trained based on features associated with the labeled withheld documents in the feature vector index;

ranking identifiers for withheld documents based on the feature vector index using the supervised machine learning model; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

13 . The method of claim 8 , wherein determining the identifiers of withheld documents of interest comprises:

generating a set of clusters of identifiers of withheld documents by clustering the identifiers for the withheld documents included in the content-by-example log based on the feature vector index; and

selecting an identifier from each of the set of clusters of identifiers of withheld documents as identifiers of the set of withheld documents of interest.

14 . The method of claim 13 , wherein the identifier is selected from a cluster of the set of clusters based on a distance of that identifier from a centroid of that cluster.

15 . A non-transitory computer readable medium, comprising instructions for:

receiving, by a receiving party, a content-by-example log, the content-by-example log including an entry for each of a set of withheld documents inaccessible to the receiving party, wherein the entry for each withheld document associates an identifier for that withheld document with identifiers for a set of example documents for that withheld document, and the set of example documents are disclosed documents accessible to the receiving party;

analyzing the content-by-example log to determine identifiers of withheld documents of interest by:

creating a feature vector index based on the content-by-example log, wherein the feature vector index comprises a feature vector associated with each of the identifiers of the withheld documents, and the feature vector associated with the identifier for a withheld document comprises a set of features determined based on the identified set of example documents associated with that identified withheld document; and

determining the identifiers of withheld documents of interest based on the feature vector index.

16 . The non-transitory computer readable medium of claim 15 , wherein determining the identifiers of withheld documents of interest comprises:

searching the identifiers for the withheld documents using the feature vector index based on a query to rank the identifiers for the withheld documents; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

17 . The non-transitory computer readable medium of claim 16 , wherein the query is determined from content associated with the disclosed documents accessible by the receiving party.

18 . The non-transitory computer readable medium of claim 15 , wherein the features of the feature vector are the identifiers of the set of example documents.

19 . The non-transitory computer readable medium of claim 15 , wherein determining the identifiers of withheld documents of interest comprises:

obtaining labels associated with identifiers of withheld documents;

training a supervised machine learning model based on the obtained labels for withheld documents, wherein the supervised machine learning model is trained based on features associated with the labeled withheld documents in the feature vector index;

ranking identifiers for withheld documents based on the feature vector index using the supervised machine learning model; and

selecting a number of top ranked identifiers of withheld documents as identifiers of the set of withheld documents of interest.

20 . The non-transitory computer readable medium of claim 15 , wherein determining the identifiers of withheld documents of interest comprises:

generating a set of clusters of identifiers of withheld documents by clustering the identifiers for the withheld documents included in the content-by-example log based on the feature vector index; and

selecting an identifier from each of the set of clusters of identifiers of withheld documents as identifiers of the set of withheld documents of interest.

21 . The non-transitory computer readable medium of claim 20 , wherein the identifier is selected from a cluster of the set of clusters based on a distance of that identifier from a centroid of that cluster.

Assignments (2)
MERGER Recorded May 18, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074678/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2024
From: PICKENS, JEREMY GARNER; BYE, ANDREW NELSON; GRICKS, THOMAS CHESTER, III
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 069183/0393 →