IP Library Granted Patent US 11,562,134
Granted Patent B2
US 11,562,134 · App. 16/373,216 · Granted Jan 24, 2023

Method and system for advanced document redaction

Inventor: Shishir Mane (Pune, IN)
Assignee: Genpact Luxembourg S.à r.l. II
G06F40/205G06F16/3344G06F21/6254G06F40/253G06F40/289G06F40/40G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,134
App. No.
16/373,216
Granted
Jan 24, 2023
Kind
B2
Abstract

A system and method for advanced document redaction are disclosed. According to one embodiment, a system comprises a parser that analyzes documents to identify structured, semi-structured, and unstructured data from a document. A candidates generator generates a list of words for redaction from the structured, semi-structured, and unstructured data. A replacement engine replaces one or more words from the list of words with one or more of a replacement word, random characters, and random numbers.

Claims (27)

1. A system, comprising:

a non-transitory computer readable storage medium having computer program code stored thereon, the computer program code, when executed by one or more processors implemented on a computer machine presents a redaction server comprising:

an information extraction module that receives a document for redaction, the information extraction module comprising a parser that identifies structured, semi-structured, and unstructured content from the received document; and

an advanced redaction engine that receives the structured, semi-structured, and unstructured content from the information extraction module, the advanced redaction engine comprising:

a first candidate generator that generates first redaction candidates for the structured and semi-structured content, wherein the first candidate generator uses structured and semi-structured metadata to generate the first redaction candidates, and wherein the semi-structured metadata is based on word length;

a second candidate generator that generates second redaction candidates for the unstructured content, wherein the first redaction candidates are different from the second redaction candidates, wherein the second candidate generator uses Natural Language Processing (NLP) metadata to generate the second redaction candidates, wherein the NLP metadata is based on word length and Parts-of-Speech (POS) tags, and wherein the second candidate generator compares a parse tree of original text and a parse tree of text with redactions applied to the second redaction candidates and determines whether the parse trees match based on whether a percentage difference between the parse trees meets a threshold;

a replacement engine that generates replacement words or text for the first and second redaction candidates; and

a document evaluator that evaluates whether the generated replacement words or text generated by the replacement engine fit dimensionally in place of corresponding redactable source words or text in the received document.

2. The system of claim 1 , wherein the semi-structured metadata include unique words and less frequently occurring words methodologies, wherein the less frequently occurring words comprise words appearing with a frequency that does not meet a document overlap threshold.

3. The system of claim 1 , wherein the document evaluator changes a font size of the words or text generated by the replacement engine so that the words or text generated by the replacement engine fit dimensionally in place of the corresponding redactable source words or text in the received document.

4. The system of claim 1 , wherein the replacement engine replaces the redactable source words or text in the received document with the generated replacement words or text to produce a redacted document, and wherein the redaction server provides the redacted document to a machine learning (ML) model for training purposes.

5. The system of claim 1 , wherein the replacement engine replaces confidential content using dictionaries, randomizing characters, or combinations thereof.

6. The system of claim 1 , wherein the replacement engine compares a list of keywords against the generated replacement words or text to identify words and/or text to be excluded from the generated replacement words or text.

7. A method for processing a document, the method comprising:

receiving a document for redaction in an information extraction module comprising a parser;

identifying, with the parser, structured, semi-structured, and unstructured content from the received document;

receiving the identified structured, semi-structured, and unstructured content to an advanced redaction engine comprising a first candidate generator and a second candidate generator;

generating, with the first candidate generator, first redaction candidates for the received structured and semi-structured content, wherein the first candidate generator uses structured and semi-structured metadata to generate the first redaction candidates, and wherein the semi-structured metadata is based on word length;

generating, with the second candidate generator, second redaction candidates for the received unstructured content, wherein the first and second redaction candidates are different from one another, wherein the second candidate generator uses Natural Language Processing (NLP) metadata to generate the second redaction candidates, wherein the NLP metadata is based on word length and Parts-of-Speech (POS) tags, and wherein the second candidate generator compares a parse tree of original text and a parse tree of text with redactions applied to the second redaction candidates and determines whether the parse trees match based on whether a percentage difference between the parse trees meets a threshold;

generating, with a replacement engine, replacement words or text for the first and second redaction candidates;

evaluating, with a document evaluator, whether the generated replacement words or text fit dimensionally in place of corresponding redactable source words or text in the received document; and

replacing, with the replacement engine, the redactable source words or text in the received document with the generated replacement words or text to produce a redacted document.

8. The method of claim 7 , further comprising training a ML model with the redacted document.

9. The method of claim 7 , wherein the semi-structured metadata include using both unique words and less frequent words methodologies, wherein the less frequent words comprise words appearing with a frequency that does not meet a document overlap threshold.

10. The method of claim 7 , wherein evaluating whether the generated replacement words or text fit dimensionally in place of the corresponding redactable source words or text in the received document comprises changing a font size of the words or text generated by the replacement engine.

11. The method of claim 7 , wherein replacing the redactable source words or text in the received document comprises replacing confidential content using dictionaries, randomizing characters, or combinations thereof.

12. The method of claim 7 , wherein replacing the redactable source words or text in the received document comprises corn paring a list of keywords against the generated replacement words or text to identify words and/or text to be excluded from the generated replacement words or text.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYANCE TYPE OF MERGER PREVIOUSLY RECORDED ON REEL 66511 FRAME 683. ASSIGNOR(S) HEREBY CONFIRMS THE CONVEYANCE TYPE OF ASSIGNMENT. Recorded Feb 26, 2024
From: GENPACT LUXEMBOURG S.À R.L. II
To: GENPACT USA, INC.
Reel/Frame 067211/0020 →
MERGER Recorded Feb 7, 2024
From: GENPACT LUXEMBOURG S.À R.L. II
To: GENPACT USA, INC.
Reel/Frame 066511/0683 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2021
From: MANE, SHISHIR
To: GENPACT LUXEMBOURG S.À R.L. II
Reel/Frame 058480/0729 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2021
From: GENPACT LUXEMBOURG S.À R.L., A LUXEMBOURG PRIVATE LIMITED LIABILITY COMPANY (SOCIÉTÉ À RESPONSABILITÉ LIMITÉE)
To: GENPACT LUXEMBOURG S.À R.L. II, A LUXEMBOURG PRIVATE LIMITED LIABILITY COMPANY (SOCIÉTÉ À RESPONSABILITÉ LIMITÉE)
Reel/Frame 055104/0632 →
Cited By (2)
US 12,417,312 US 12,536,329