IP Library › Granted Patent US 9,734,148
Granted Patent B2
US 9,734,148 · App. 14/520,018 · Granted Aug 15, 2017

Information redaction from document data

Inventors: Mike Bendersky (Sunnyvale, CA); Vanja Josifovski (Los Gatos, CA); Amitabh Saikia (Mountain View, CA); Marc-Allen Cartright (Stanford, CA); Jie Yang (Santa Clara, CA); Luis Garcia Pueyo (San Francisco, CA); MyLinh Yang (Saratoga, CA)
Assignee: Google Inc.
G06F17/30011G06F21/6218G06F21/6227G06F21/6254G06F21/64G06Q10/101G06Q30/0254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,734,148
App. No.
14/520,018
Granted
Aug 15, 2017
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for redacting data from a document collection generated for a set of documents that include personal information. The redaction of the data is based in part on a comparison of the document collection to a set of a personal documents of users for which the users have provided explicit approval to use in the processing of the document collection.

Claims (43)

1. A computer-implemented method performed by data processing apparatus, the method comprising:

receiving, by a data processing apparatus, an electronic document data collection generated from a first set of documents, the document data collection including a first set of fixed phrases extracted from the first set of documents, wherein each fixed phrase is a phrase of one or more terms that is determined to not present a personal information exposure risk, and wherein access to the document data collection for examination by a human reviewer is precluded;

receiving, by the data processing apparatus, a second set of documents, the second set of documents including documents that are each a personal document of a user that has personal information of the user and for which the user has provided permission to use the document for processing of the fixed phrases extracted from the first set of documents;

extracting, by the data processing apparatus, candidate phrases from the second set of documents, each candidate phrase being a phrase of one or more terms;

identifying, by the data processing apparatus, fixed phrases extracted from the first set of documents that match candidate phrases extracted from the second set of documents;

generating, from the document data collection, a redacted document data collection in which each fixed phrase that does not match a candidate phrase is redacted, and each fixed phrase that does match a candidate phrase is not redacted; and

providing, by the data processing apparatus, access to the redacted document data collection for examination by the human reviewer.

2. The computer-implemented method of claim 1 , wherein generating a redacted document data collection comprises generating a redacted document data collection in which each fixed phrase that does not match a candidate phrase is removed from the redacted document data collection, and each fixed phrase that does match a candidate phrase is included in the redacted document data collection.

3. The computer-implemented method of claim 1 , wherein generating a redacted document data collection comprises generating an obfuscated document data collection in which each fixed phrase that does not match a candidate phrase is obfuscated, and each fixed phrase that does match a candidate phrase is not obfuscated.

4. The computer-implemented method of claim 3 , wherein generating, from the document data collection, the obfuscated document data collection comprises generating, for each fixed phrase that does not match a candidate phrase, a hash of the fixed phrase to obfuscate the fixed phrase.

5. The computer-implemented method of claim 1 , wherein the document data collection is a template that describes content of the first set of documents in the form of structural data.

6. The computer-implemented method of claim 1 , wherein the document data collection is one of a plurality of document clusters, wherein each document cluster is clustered according to a content characteristic that is different for each document cluster.

7. The computer-implemented method of claim 1 , wherein the data processing apparatus precludes access to each document in the second set of documents for examination by a human reviewer.

8. The computer implemented method of claim 1 , further comprising generating, by the data processing apparatus, the electronic document data collection from the first set of documents, the generating the electronic document data collection comprising:

extracting candidate fixed phrases from the first set of documents, each candidate fixed phrase being a phrase of one or more terms;

for each candidate fixed phrase, determining whether the candidate fixed phrase presents a personal information exposure risk; and

selecting only the candidate fixed phrases that are determined not to present a personal information exposure risk as the fixed phrases.

9. A computer storage medium encoded with a computer program, the program comprising instructions that when executed by data processing apparatus cause the data processing apparatus to perform operations comprising:

receiving an electronic document data collection generated from a first set of documents, the document data collection including a first set of fixed phrases extracted from the first set of documents, wherein each fixed phrase is a phrase of one or more terms that is determined to not present a personal information exposure risk, and wherein access to the document data collection for examination by a human reviewer is precluded;

receiving, by the data processing apparatus, a second set of documents, the second set of documents including documents that are each a personal document of a user that has personal information of the user and for which the user has provided permission to use the document for processing of the fixed phrases extracted from the first set of documents;

extracting candidate phrases from the second set of documents, each candidate phrase being a phrase of one or more terms;

identifying fixed phrases extracted from the first set of documents that match candidate phrases extracted from the second set of documents;

generating, from the document data collection, a redacted document data collection in which each fixed phrase that does not match a candidate phrase is redacted, and each fixed phrase that does match a candidate phrase is not redacted; and

providing access to the redacted document data collection for examination by the human reviewer.

10. A system, comprising:

a data processing apparatus; and

a computer storage medium encoded with a computer program, the program comprising instructions that when executed by the data processing apparatus cause the data processing apparatus to perform operations comprising:

receiving an electronic document data collection generated from a first set of documents, the document data collection including a first set of fixed phrases extracted from the first set of documents, wherein each fixed phrase is a phrase of one or more terms that is determined to not present a personal information exposure risk, and wherein access to the document data collection for examination by a human reviewer is precluded;

receiving, by the data processing apparatus, a second set of documents, the second set of documents including documents that are each a personal document of a user that has personal information of the user and for which the user has provided permission to use the document for processing of the fixed phrases extracted from the first set of documents;

extracting candidate phrases from the second set of documents, each candidate phrase being a phrase of one or more terms;

identifying fixed phrases extracted from the first set of documents that match candidate phrases extracted from the second set of documents;

generating, from the document data collection, a redacted document data collection in which each fixed phrase that does not match a candidate phrase is redacted, and each fixed phrase that does match a candidate phrase is not redacted; and

providing access to the redacted document data collection for examination by the human reviewer.

11. The system of claim 10 , wherein generating a redacted document data collection comprises generating a redacted document data collection in which each fixed phrase that does not match a candidate phrase is removed from the redacted document data collection, and each fixed phrase that does match a candidate phrase is included in the redacted document data collection.

12. The system of claim 10 , wherein generating a redacted document data collection comprises generating an obfuscated document data collection in which each fixed phrase that does not match a candidate phrase is obfuscated, and each fixed phrase that does match a candidate phrase is not obfuscated.

13. The system of claim 12 , wherein generating, from the document data collection, the obfuscated document data collection comprises generating, for each fixed phrase that does not match a candidate phrase, a hash of the fixed phrase to obfuscate the fixed phrase.

14. The system of claim 10 , wherein the document data collection is a template that describes content of the first set of documents in the form of structural data.

15. The system of claim 10 , wherein the document data collection is one of a plurality of document clusters, wherein each document cluster is clustered according to a content characteristic that is different for each document cluster.

16. The system of claim 10 , wherein the data processing apparatus precludes access to each document in the second set of documents for examination by a human reviewer.

17. The system of claim 10 , the operations further comprising generating the electronic document data collection from the first set of documents, the generating the electronic document data collection comprising:

extracting candidate fixed phrases from the first set of documents, each candidate fixed phrase being a phrase of one or more terms;

for each candidate fixed phrase, determining whether the candidate fixed phrase presents a personal information exposure risk; and

selecting only the candidate fixed phrases that are determined not to present a personal information exposure risk as the fixed phrases.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044097/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2014
From: BENDERSKY, MIKE; JOSIFOVSKI, VANJA; SAIKIA, AMITABH; CARTRIGHT, MARC-ALLEN; YANG, JIE; PUEYO, LUIS GARCIA; YANG, MYLINH
To: GOOGLE INC.
Reel/Frame 034495/0927 →
Continuity (1)
Related Publication 20160110352A1 · Apr 21, 2016