IP Library Granted Patent US 12,124,799
Granted Patent B2
US 12,124,799 · App. 18/085,998 · Granted Oct 22, 2024

Method and system for advanced document redaction

Inventor: Shishir Mane (Pune, IN)
Assignee: Genpact USA, Inc.
G06F40/205G06F16/3344G06F21/6254G06F40/253G06F40/289G06F40/40G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,124,799
App. No.
18/085,998
Granted
Oct 22, 2024
Kind
B2
Abstract

A system and method for advanced document redaction are disclosed. According to one embodiment, a system comprises a parser that analyzes documents to identify structured, semi-structured, and unstructured data from a document. A candidates generator generates a list of words for redaction from the structured, semi-structured, and unstructured data. A replacement engine replaces one or more words from the list of words with one or more of a replacement word, random characters, and random numbers.

Claims (39)

1. A system, comprising:

a non-transitory computer readable storage medium having computer program code stored thereon, which when executed by one or more processors implemented on a computing machine, presents a redaction server comprising:

a redaction engine that receives structured, semi-structured, and unstructured content, the redaction engine comprising:

a first candidate generator that generates first redaction candidates for the received structured and semi-structured content, wherein the first candidate generator uses structured and semi-structured metadata based on word length to generate the first redaction candidates;

a second candidate generator configured to:

parse the unstructured content into sentences so that each sentence is subjected to a Parts-of-Speech (POS) tagger to generate POS tags; and

identify second redaction candidates for the generated POS tags from Natural Language Processing (NLP) metadata based on word length, wherein the NLP metadata includes replacements for the generated POS tags;

a replacement engine that generates replacement characters for the first and second redaction candidates based on dictionaries found in replacement metadata; and

a document evaluator that adjusts a font size of the replacement characters generated by the replacement engine so that the replacement characters generated by the replacement engine fit dimensionally in place of corresponding redactable source characters in the received structured, semi-structured, and unstructured content.

2. The system of claim 1 , wherein the redaction server further comprises an information extraction module, and wherein the redaction engine is configured to receive the structured, semi-structured, and unstructured content from the information extraction module.

3. The system of claim 2 , wherein the structured, semi-structured, and unstructured content are extracted from one or more unredacted documents received by a document parser in the information extraction module.

4. The system of claim 3 , wherein the one or more unredacted documents comprise at least one of Portable Document Format (PDF) documents, word processing application documents, spreadsheets, presentations, HyperText Markup Language (HTML) files, or text files.

5. The system of claim 1 , wherein the replacement engine generates one or more redacted documents from the one or more unredacted documents.

6. The system of claim 5 , further comprising a machine learning system that uses the one or more redacted documents for training a model.

7. A method comprising:

receiving, at a redaction engine, structured, semi-structured, and unstructured content extracted from an unredacted document;

generating, with a first candidate generator, first redaction candidates for the received structured and semi-structured content, wherein the first candidate generator uses structured and semi-structured metadata based on word length to generate the first redaction candidates;

parsing, with a second candidate generator, the unstructured content into sentences so that each sentence is subjected to a Parts-of-Speech (POS) tagger to generate POS tags, and identifying second redaction candidates for the generated POS tags from Natural Language Processing (NLP) metadata based on word length, wherein the NLP metadata includes replacements for the generated POS tags;

generating, with a replacement engine, replacement characters for the first and second redaction candidates based on dictionaries found in replacement metadata; and

replacing, with the replacement engine, redactable source characters in the received structured, semi-structured, and unstructured content with the generated replacement characters to produce a redacted document, wherein replacing the redactable source characters comprises adjusting a font size of the replacement characters to ensure that the replacement characters fit dimensionally in place of the redactable source characters.

8. The method of claim 7 , wherein before receiving, at the redaction engine, the structured, semi-structured, and unstructured content:

analyzing the unredacted document at an information extraction module to extract the structured, semi-structured, and unstructured content from the unredacted document prior to sending the structured, semi-structured, and unstructured content to the redaction engine.

9. The method of claim 8 , wherein analyzing the unredacted document to extract the structured, semi-structured, and unstructured content comprises parsing the unredacted document with a document parser.

10. The method of claim 7 , wherein the unredacted document comprises at least one of a Portable Document Format (PDF) document, a word processing application document, a spreadsheet, a presentation, a HyperText Markup Language (HTML) file, or a text file.

11. A system, comprising:

a non-transitory computer readable storage medium having computer program code stored thereon, which when executed by one or more processors implemented on a computing machine, presents a redaction server comprising:

an information extraction engine that receives one or more unredacted documents and extracts structured, semi-structured, and unstructured content from the one or more unredacted documents; and

a redaction engine communicatively coupled to the information extraction engine and comprising:

a first candidate generator that generates first redaction candidates for the structured and semi-structured content, and uses structured and semi-structured metadata based on word length to generate the first redaction candidates;

a second candidate generator configured to:

parse the unstructured content into sentences so that each sentence is subjected to a Parts-of-Speech (POS) tagger to generate POS tags; and

identify second redaction candidates for the generated POS tags from Natural Language Processing (NLP) metadata based on word length, wherein the NLP metadata includes replacements for the generated POS tags;

a replacement engine that generates replacement characters based on dictionaries found in replacement metadata to replace redactable source characters in the structured, semi-structured, and unstructured content thereby producing respective one or more redacted documents from the one or more unredacted documents; and

a document evaluator that adjusts a font size of the replacement characters to ensure that the replacement characters generated by the replacement engine fit dimensionally in place of the redactable source characters.

12. The system of claim 11 , wherein the one or more unredacted documents comprise at least one of a Portable Document Format (PDF) document, a word processing application document, a spreadsheet, a presentation, a HyperText Markup Language (HTML) files, or a text file.

13. The system of claim 11 , wherein the information extraction engine extracts the structured, semi-structured, and unstructured content with a document parser configured to identify structured, semi-structured, and unstructured data.

14. The system of claim 11 , further comprising a machine learning system that uses the one or more redacted documents produced by the redaction server for training a model.

15. The system of claim 14 , further comprising an information extraction system that uses the model trained with the one or more redacted documents to redact unredacted documents.

16. The system of claim 15 , wherein the information extraction system is configured to redact documents by identifying information in the documents and to use the identified information to provide analytics and report generation and to provide services to end users.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYANCE TYPE OF MERGER PREVIOUSLY RECORDED ON REEL 66511 FRAME 683. ASSIGNOR(S) HEREBY CONFIRMS THE CONVEYANCE TYPE OF ASSIGNMENT. Recorded Feb 26, 2024
From: GENPACT LUXEMBOURG S.À R.L. II
To: GENPACT USA, INC.
Reel/Frame 067211/0020 →
MERGER Recorded Feb 7, 2024
From: GENPACT LUXEMBOURG S.À R.L. II
To: GENPACT USA, INC.
Reel/Frame 066511/0683 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2023
From: MANE, SHISHIR
To: GENPACT LUXEMBOURG S.À R.L. II
Reel/Frame 062270/0738 →
Continuity (2)
Continuation 16373216 · Apr 2, 2019
Related Publication 20230205988A1 · Jun 29, 2023