IP Library Granted Patent US 12,147,759
Granted Patent B2
US 12,147,759 · App. 17/851,513 · Granted Nov 19, 2024

Systems and methods for document analysis to produce, consume and analyze content-by-example logs for documents

Inventors: Jeremy Garner Pickens (Bloomville, NY); Andrew Nelson Bye (San Francisco, CA); Thomas Chester Gricks, III (Irwin, PA)
Assignee: OPEN TEXT HOLDINGS, INC.
G06F40/194G06F16/35G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,147,759
App. No.
17/851,513
Granted
Nov 19, 2024
Kind
B2
Abstract

Document analysis systems and methods for the generation of a content-by-example log that expresses withheld documents in terms of a set of disclosed documents are disclosed. Additionally, document analysis systems and methods for the analysis of such a content-by-example log to determine withheld documents of interest without access to those withheld documents are disclosed.

Claims (61)

1. A method of document production and consumption of a corpus of electronic documents, comprising:

at a producer party computing device:

determining, by a producer party having access to both a set of non-produced documents of the corpus and a set of produced documents of the corpus, a method ology for producing a content-by-example log for the corpus of electronic documents, the methodology comprising a technique for comparing the set of non-produced documents of the corpus to the set of produced documents of the corpus, the content-by-example log comprising an electronic file including a comparison between the set of non-produced documents and the set of produced documents identifying a produced document with one or more of the set of non-produced documents;

producing, by the producer party, the content-by-example log and sending the content-by-example log to a consumer party, wherein producing the content-by-example log comprises applying, at the producer party computing device, a first machine model adapted to implement the methodology for comparing the set of produced documents to the corpus of electronic corpus of documents to identify the produced document with the one or more of the set of non-produced documents; and

providing electronic access to the set of produced document of the electronic corpus to a consumer party such that the consumer party can access the set of produced documents of the corpus at a consumer party computer device but cannot access the set of non-produced documents corpus; and

receiving a request from the consumer party regarding the set of non-produced documents, the request based on a methodology for analyzing the content-by-example log applied to the content-by-example log using a second machine learning model at the consumer party computing device by applying the second machine learning model to evaluate attributes of the set of produced documents identified for each of the non-produced documents identified in the content-by-example log to determine one or more the set of non-produced documents identified in the content-by-example log as documents of interest to the consumer party that cannot access the non-produced documents, the application of the second machine learning model comprising:

determining values for a set of features associated with the second machine learning model for the non-produced document; and

applying the second machine learning model to the values for the set of features to generate output of the second machine learning model associated with the non-produced documents and determining, based on the output of second machine learning model whether non-produced documents are of interest to the consumer party; and

responding, by the producing party, to the request regarding the set of non-produced documents wherein the response comprises production of one or more of the set of non-produced documents or refusal to produce one or more of the set of non-produced documents.

2. The method of claim 1 , wherein the methodology for producing the content-by-example log comprises:

generating a similarity score between each of the non-produced documents and one or more of the produced documents, the content-by-example log comprising a listing of each of the non-produced documents and the similarity score of the one or more set of produced documents.

3. The method of claim 2 , wherein the methodology for producing the content-by-example log comprises selecting a method for generating similarity scores from a set of methods for generating similarity scores.

4. The method of claim 3 , wherein selecting the method for generating similarity scores is based on a standard for producing the content-by-example log.

5. The method of claim 1 , wherein the request from the consumer party regarding the set of non-produced documents comprises:

at least one of: a request by the consumer party to produce at least one of the non-produced documents, or a request for content-by-example criteria of at least one document in the set of non-produced documents.

6. The method of claim 5 , wherein content-by-example criteria for non-produced documents comprises the comparisons within the context-by-example log.

7. The method of claim 1 , wherein the methodology for producing the content-by-example log comprises:

selecting a method for generating similarity scores from a set of methods for generating similarity scores; and

based on the selected method for generating similarity scores, generating a similarity score between each of the non-produced documents and one or more of the produced documents, the content-by-example log comprising a listing of each of the non-produced documents and the similarity score of the one or more set of produced documents; and wherein the request regarding one or more of the documents the set of non-produced documents comprises:

at least one of: a request by the consumer party to produce at least one of the non-produced documents, or a request for content-by-example criteria of at least one document in the set of non-produced documents.

8. A system for document production and consumption, comprising:

a processor;

a memory coupled to the processor, the memory comprising computer executable instructions that, when executed by the processor, perform a method comprising:

determining, by a producer party having access to both a set of non-produced documents of a corpus of electronic documents and a set of produced documents of the corpus of electronic documents, a methodology for producing a content-by-example log for the corpus of electronic documents, the methodology comprising a technique for comparing the set of non-produced documents of the corpus to the set of produced documents of the corpus, the content-by-example log comprising an electronic file including a comparison between the set of non-produced documents and the set of produced documents identifying a produced document with one or more of the set of non-produced documents;

producing, by the producer party, the content-by-example log and sending the content-by-example log to a consumer party, wherein producing the content-by-example log comprises applying, at the producer party computing device, a first machine model adapted to implement the methodology for comparing the set of produced documents to the corpus of electronic corpus of documents to identify the produced document with the one or more of the set of non-produced documents; and

providing electronic access to the set of produced document of the electronic corpus to a consumer party such that the consumer party can access the set of produced documents of the corpus at a consumer party computer device but cannot access the set of non-produced documents corpus; and

receiving a request from the consumer party regarding the set of non-produced documents, the request based on a methodology for analyzing the content-by-example log applied to the content-by-example using a second machine model at the consumer party computing device by applying the second machine learning model to evaluate attributes of the set of produced documents identified for each of the non-produced documents identified in the content-by-example log to determine one or more the set of non-produced documents identified in the content-by-example log as documents of interest to the consumer party that cannot access the non-produced documents, the application of the second machine learning model comprising:

determining values for a set of features associated with the second machine learning model for the non-produced document; and

applying the second machine learning model to the values for the set of features to generate output of the second machine learning model associated with the non-produced documents and determining, based on the output of second machine learning model whether non-produced documents are of interest to the consumer party; and

responding, by the producing party, to the request regarding the set of non-produced documents wherein the response comprises production of one or more of the set of non-produced documents or refusal to produce one or more of the set of non-produced documents.

9. The system of claim 8 , wherein the methodology for producing the content-by-example log comprises:

generating a similarity score between each of the non-produced documents and one or more of the produced documents, the content-by-example log comprising a listing of each of the non-produced documents and the similarity score of the one or more set of produced documents.

10. The system of claim 9 , wherein the methodology for producing the content-by-example log comprises selecting a method for generating similarity scores from a set of methods for generating similarity scores.

11. The system of claim 10 , wherein selecting the method for generating similarity scores is based on a standard for producing the content-by-example log.

12. The system of claim 8 , wherein the request from the consumer party regarding the set of non-produced documents comprises:

a request by the consumer party to produce at least one of the non-produced documents, or a request for content-by-example criteria of at least one document in the set of non-produced documents.

13. The system of claim 12 , wherein content-by-example criteria for non-produced documents comprises the comparisons within the context-by-example log.

14. The system of claim 8 , wherein the methodology for producing the content-by-example log comprises:

selecting a method for generating similarity scores from a set of methods for generating similarity scores; and

based on the selected method for generating similarity scores, generating a similarity score between each of the non-produced documents and one or more of the produced documents, the content-by-example log comprising a listing of each of the non-produced documents and the similarity score of the one or more set of produced documents; and wherein the request regarding one or more of the documents the set of non-produced documents comprises:

at least one of: a request by the consumer party to produce at least one of the non-produced documents, or a request for content-by-example criteria of at least one document in the set of non-produced documents.

15. A computer program product comprising a non-transitory computer readable medium storing instructions executable by a processor to perform a method for document production and consumption, the method comprising:

at a producer party computing device:

determining, by a producer party having access to both a set of non-produced documents of a corpus of electronic documents and a set of produced documents of the corpus of electronic documents, a methodology for producing a content-by-example log for the corpus of electronic documents, the methodology comprising a technique for comparing the set of non-produced documents of the corpus to the set of produced documents of the corpus, the content-by-example log comprising an electronic file including a comparison between the set of non-produced documents and the set of produced documents identifying a produced document with one or more of the set of non-produced documents;

producing, by the producer party, the content-by-example log and sending the content-by-example log to a consumer party, wherein producing the content-by-example log comprises applying, at the producer party computing device, a first machine model adapted to implement the methodology for comparing the set of produced documents to the corpus of electronic corpus of documents to identify the produced document with the one or more of the set of non-produced documents; and

providing electronic access to the set of produced document of the electronic corpus to a consumer party such that the consumer party can access the set of produced documents of the corpus at a consumer party computer device but cannot access the set of non-produced documents corpus; and

receiving a request from the consumer party regarding the set of non-produced documents, the request based on a methodology for analyzing the content-by-example log applied to the content-by-example using a second machine model at the consumer party computing device by applying the second machine learning model to evaluate attributes of the set of produced documents identified for each of the non-produced documents identified in the content-by-example log to determine one or more the set of non-produced documents identified in the content-by-example log as documents of interest to the consumer party that cannot access the non-produced documents, the application of the second machine learning model comprising:

determining values for a set of features associated with the second machine learning model for the non-produced document; and

applying the second machine learning model to the values for the set of features to generate output of the second machine learning model associated with the non-produced documents and determining, based on the output of second machine learning model whether non-produced documents are of interest to the consumer party; and

responding, by the producing party, to the request regarding the set of non-produced documents wherein the response comprises production of one or more of the set of non-produced documents or refusal to produce one or more of the set of non-produced documents.

16. The computer program product of claim 15 , wherein the methodology for producing the content-by-example log comprises:

generating a similarity score between each of the non-produced documents and one or more of the produced documents, the content-by-example log comprising a listing of each of the non-produced documents and the similarity score of the one or more set of produced documents.

17. The computer program product of claim 16 , wherein the request from the consumer party regarding the set of non-produced documents comprises:

at least one of: a request by the consumer party to produce at least one of the non-produced documents, or a request for document criteria of at least one document in the set of non-produced documents.

18. The computer program product of claim 15 , wherein the request from the consumer party regarding the set of non-produced documents comprises:

a request by the consumer party to produce at least one of the non-produced documents, or a request for content-by-example criteria of at least one document in the set of non-produced documents.

19. The computer program product of claim 18 , wherein content-by-example criteria for non-produced documents comprises the comparisons within the context-by-example log.

20. The computer program product of claim 15 , wherein the methodology for producing the content-by-example log comprises:

selecting a method for generating similarity scores from a set of methods for generating similarity scores; and

based on the selected method for generating similarity scores, generating a similarity score between each of the non-produced documents and one or more of the produced documents, the content-by-example log comprising a listing of each of the non-produced documents and the similarity score of the one or more set of produced documents; and wherein the request regarding one or more of the documents the set of non-produced documents comprises:

at least one of: a request by the consumer party to produce at least one of the non-produced documents, or a request for content-by-example criteria of at least one document in the set of non-produced documents.

Assignments (2)
MERGER Recorded May 18, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074678/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2022
From: PICKENS, JEREMY GARNER; BYE, ANDREW NELSON; GRICKS, THOMAS CHESTER, III
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 060401/0125 →
Continuity (1)
Related Publication 20230419026A1 · Dec 28, 2023
Cited By (2)
US 12,380,149 US 12,585,707