IP Library Granted Patent US 7,720,783
Granted Patent B2
US 7,720,783 · App. 11/729,576 · Granted May 18, 2010

Method and system for detecting undesired inferences from documents

Assignee: Palo Alto Research Center Incorporated
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,720,783
App. No.
11/729,576
Granted
May 18, 2010
Kind
B2
Abstract

One embodiment of the present invention provides a system that detects inferences from documents. During operation, the system receives one or more documents and extracts a first set of knowledge relevant to the documents. The system further formulates one or more queries to one or more reference corpora based on the first set of knowledge. The system then extracts a second set of knowledge from results received in response to the queries. Additionally, the system produces a mapping relationship between at least one document and a piece of the second set of knowledge which is not within the first set of knowledge, the mapping relationship indicating an inference from the documents.

Claims (61)

1. A computer-implemented method for detecting inferences from documents, wherein the computer includes a processor, the method comprising:

receiving one or more documents;

extracting a first set of knowledge relevant to the documents;

formulating one or more queries to one or more reference corpora based on the first set of knowledge;

extracting a second set of knowledge from results received in response to the queries; and

detecting an inference by producing a mapping relationship between at least one document and a piece of the second set of knowledge which is not within the first set of knowledge, wherein the mapping relationship indicates an inference from the documents.

2. The method of claim 1 , wherein extracting the first set of knowledge relevant to the document comprises deriving a set of words or phrases relevant to each document.

3. The method of claim 2 , wherein deriving the set of words or phrases further comprises determining a term-frequency inverse-document-frequency (TF.IDF) weight for a word or phrase contained in the document.

4. The method of claim 1 , wherein extracting the first set of knowledge comprises extracting a set of words or phrases from each document; and

wherein formulating the queries comprises constructing single-word queries, multi-word queries, or both from the extracted words or phrases.

5. The method of claim 4 , wherein constructing the queries further comprises applying one or more pre-determined logical combinations of the extracted words or phrases.

6. The method of claim 4 , wherein extracting a set of words or phrases from each document comprises extracting a pre-determined number of words or phrases from the document.

7. The method of claim 1 , further comprising retrieving a pre-determined number of results for each query.

8. The method of claim 1 , further comprising receiving a set of sensitive knowledge relevant to the documents;

wherein the piece of the second set of knowledge mapped to the document is also within the set of sensitive knowledge; and

wherein producing the mapping relationship comprises determining an intersection between the set of sensitive knowledge and the second set of knowledge.

9. The method of claim 1 , wherein producing the mapping relationship comprises presenting one or more words or phrases from the documents, which words or phrases correspond to the first set of knowledge, and one or more sensitive words or phrases extracted from the query results.

10. The method of claim 1 , further comprising:

determining a ratio of:

the number of results returned by a first query based on one or more words or phrases from the documents and one or more sensitive words or phrases, to

the number of results returned by a second query based on one or more words or phrases from the documents; and

determining an inference from the one or more words or phrases used to generate the first query based on the size of the ratio.

11. The method of claim 1 , further comprising:

selecting the one or more reference corpora based on the documents, an intended audience, or both.

12. The method of claim 1 , further comprising:

detecting one or more undesired inferences, wherein the undesired inferences are information intended to remain private, which can be extracted from a union of the document and the corpus and which cannot be extracted from the document or the corpus alone.

13. A computer system for detecting inferences from documents, the computer system comprising:

a processor;

a memory;

a receiving mechanism configured to receive one or more documents;

a first knowledge extraction mechanism configured to extract a first set of knowledge relevant to the documents;

a query formulation mechanism configured to formulate one or more queries to one or more reference corpora based on the first set of knowledge;

a second knowledge extraction mechanism configured to extract a second set of knowledge from results received in response to the queries; and

a detecting mechanism configured to detect an inference by producing a mapping mechanism configured to produce a mapping relationship between at least one document and a piece of the second set of knowledge which is not within the first set of knowledge, wherein the mapping relationship indicates an inference from the documents.

14. The computer system of claim 13 , wherein while extracting the first set of knowledge relevant to the document, the first knowledge extraction mechanism is configured to derive a set of words or phrases relevant to each document.

15. The computer system of claim 14 , wherein while deriving the set of words or phrases, the first knowledge extraction mechanism is further configured to determine a term-frequency inverse-document-frequency (TF.IDF) weight for a word or phrase contained in the document.

16. The computer system of claim 13 , wherein while extracting the first set of knowledge, the first knowledge extraction mechanism is configured to extract a set of words or phrases from each document; and

wherein while formulating the queries, the query formulation mechanism is configured to construct single-word queries, multi-word queries, or both from the extracted words or phrases.

17. The computer system of claim 16 , wherein while constructing the queries, the query formulation mechanism is further configured to apply one or more pre-determined logical combinations of the extracted words or phrases, thereby formulating queries based on a number of words or phrases with logical relationships.

18. The computer system of claim 16 , wherein while extracting a set of words or phrases from each document, the first knowledge extraction mechanism is configured to extract a pre-determined number of words or phrases from the document.

19. The computer system of claim 13 , further comprising a retrieval mechanism configured to retrieve a pre-determined number of results for each query.

20. The computer system of claim 13 ,

wherein the receiving mechanism is further configured to receive a set of sensitive knowledge relevant to the documents;

wherein the piece of the second set of knowledge mapped to the document is also within the set of sensitive knowledge; and

wherein while producing the mapping relationship, the mapping mechanism is configured to determine an intersection between the set of sensitive knowledge and the second set of knowledge.

21. The computer system of claim 13 , wherein while producing the mapping relationship, the mapping mechanism is configured to present one or more words or phrases from the documents, which words or phrases correspond to the first set of knowledge, and one or more sensitive words or phrases extracted from the query results.

22. The computer system of claim 13 , further comprising:

a ranking mechanism configured to determine a ratio of:

the number of results returned by a first query based on one or more words or phrases from the documents and one or more sensitive words or phrases, to

the number of results returned by a second query based on one or more words or phrases from the documents; and

a decision mechanism configured to determine an inference from the one or more words or phrases used to generate the first query based on the size of the ratio.

23. The computer system of claim 13 , further comprising:

a reference selection mechanism configured to select the one or more reference corpora based on the documents, an intended audience, or both.

24. The computer system of claim 13 , further comprising:

a detecting mechanism configured to detect one or more undesired inferences, wherein the undesired inferences are information intended to remain private, which can be extracted from a union of the document and the corpus and which cannot be extracted from the document or the corpus alone.

25. A computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for detecting inferences from documents, the method comprising:

receiving one or more documents;

extracting a first set of knowledge relevant to the documents;

formulating one or more queries to one or more reference corpora based on the first set of knowledge;

extracting a second set of knowledge from results received in response to the queries; and

detecting an inference by producing a mapping relationship between at least one document and a piece of the second set of knowledge which is not within the first set of knowledge, wherein the mapping relationship indicates an inference from the documents.

Assignments (9)
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVAL OF US PATENTS 9356603, 10026651, 10626048 AND INCLUSION OF US PATENT 7167871 PREVIOUSLY RECORDED ON REEL 064038 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064161/0001 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064038/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2007
From: STADDON, JESSICA N.; GOLLE, PHILIPPE J.P.; ZIMNY, BRYCE D.
To: PALO ALTO RESEARCH CENTER INCORPORATED
Reel/Frame 019198/0665 →
Continuity (1)
Related Publication 20080243825A1 · Oct 2, 2008