IP Library › Granted Patent US 12,367,250
Granted Patent B2
US 12,367,250 · App. 18/326,947 · Granted Jul 22, 2025

Detection and removal of predefined sensitive information types from electronic documents

Inventors: Amy Hariharan Dang (Orland Park, IL); Sunil Shankar Kadam (Redmond, WA); Hassan Almandil (Bellevue, WA); Yibing Chen (Bellevue, WA); Mark-Gil Parayno (Seattle, WA); Meir Baruch Blachman (Jerusalem, IL); Tomer Cherni (Ganie, IL); Nitzan Frogel (Tel Aviv, IL)
Assignee: Microsoft Technology Licensing, LLC.
G06F16/9532G06F16/957H04L67/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,250
App. No.
18/326,947
Granted
Jul 22, 2025
Kind
B2
Abstract

Automated and semi-automated document redaction technology is disclosed herein. In certain example embodiments, ‘context-aware’ redaction is provided. Automated techniques are used to identify a set of potentially sensitive item(s) within a document. The potentially sensitive item(s) are filtered based on contextual information, such an entity identifier (e.g. person identifier, person group identifier identifying a group of multiple people, organization identifier etc.), resulting in a filtered set of redaction candidate(s). The filtered redaction candidate(s) may, for example, be redacted from the document automatically, or outputted as suggestions in an assisted redaction tool, e.g. via a document redaction graphical user interface. Other example embodiments consider selective redaction when uploading and/or downloading documents via a proxy server, to prevent intended or unintended release of potentially sensitive information, e.g. in a web browsing context.

Claims (65)

1. A computer-implemented method comprising:

obtaining an electronic document and an entity identifier associated with the electronic document;

identifying, in the electronic document, a group of redaction candidates, each redaction candidate of the group of redaction candidates comprising a document portion associated with one of a plurality of sensitive information categories, a first redaction candidate of the group of redaction candidates being associated with a particular sensitive information category of the plurality of sensitive information categories and matches the entity identifier itself and a second redaction candidate of the group of redaction candidates being associated with the particular sensitive information category and does not match the entity identifier itself;

filtering, from the group of redaction candidates, ones of the redaction candidates that match the entity identifier itself, resulting in a group of filtered redaction candidates that excludes at least the first redaction candidate based on the first redaction candidate matching the entity identifier itself;

removing only the group of filtered redaction candidates from the electronic document, resulting in a redacted document comprising the first redaction candidate and not the second redaction candidate; and

outputting the redacted document comprising the first redaction candidate;

wherein:

the method is performed by a proxy server; the obtaining of the electronic document comprises:

receiving, from a client device, a download request associated with the entity identifier;

in response to the receiving of the download request, transmitting a proxied download request from the proxy server to an upstream server; and

receiving from the upstream server at the proxy server the electronic document in response to the proxied download request; and

the method further comprises transmitting the redacted document from the proxy server to the client device.

2. The method of claim 1 , further comprising receiving a document search request comprising the entity identifier, wherein the electronic document is obtained from computer-readable storage via a document search based on the entity identifier.

3. The method of claim 1 , wherein the entity identifier comprises a user identifier associated with the client device.

4. The method of claim 1 , wherein the method is performed by a web proxy server and the download request is received from a web browser executed on the client device.

5. The method of claim 1 , wherein:

the obtaining of the electronic document comprises receiving, from a client device, a message comprising the electronic document, the message being associated with the entity identifier; and

the method further comprises transmitting to an upstream server a proxied message comprising the redacted document.

6. The method of claim 5 , wherein the entity identifier comprises a user identifier associated with the client device.

7. The method of claim 5 , wherein the wherein the method is performed by a web proxy server and the message is received from a web browser executed on the client device.

8. The method of claim 7 , wherein:

the obtaining of the electronic document further comprises:

receiving, at the web proxy server and from the client device, a content request comprising a resource identifier;

in response to the content request, retrieving, at the web proxy server, web content associated with the resource identifier;

generating, based on the web content, modified web content comprising proxy client code, wherein the proxy client code is added to the web content to generate the modified web content; and

transmitting the modified web content to the client device, whereby the proxy client code is executed on the client device; and

the method further comprises:

detecting, in the message comprising the electronic document, marker data inserted by the client proxy code executed on the client device, wherein the identifying of the group of redaction candidates, the filtering of the group of redaction candidates, and the removing of the group of filtered redaction candidates from the electronic document are performed in response to the detecting of the marker data.

9. The method of claim 1 , further comprising:

outputting, via a graphical user interface, an indication of the group of filtered redaction candidates including the second redaction candidate; and

receiving input via the graphical user interface to modify the group of filtered redaction candidates.

10. The method of claim 9 , wherein the outputting of the indication of the group of filtered redaction candidates comprises displaying the electronic document via the graphical user interface, wherein the indication of the group of filtered redaction candidates comprises a visual marker marking the group of filtered candidates within the electronic document.

11. The method of claim 9 , comprising outputting, in association with the indication of the second redaction candidate, an indication of the particular sensitive information category.

12. The method of claim 1 , wherein the entity identifier is a person identifier or a person group identifier, wherein the particular sensitive information category is a predefined personal information category.

13. A proxy server comprising:

at least one memory configured to store computer-readable instructions;

at least one processor coupled to the at least one memory and configured to execute the computer-readable instructions, the computer-readable instructions configured, upon execution on the at least one processor, to cause the at least one processor to:

receive, at the proxy server and from a client device, a content request comprising a resource identifier;

in response to the content request, retrieve web content associated with the resource identifier;

generating, based on the web content, modified web content comprising proxy client code, wherein the proxy client code is added to the web content to generate the modified web content; and

transmitting the modified web content to the client device, whereby the proxy client code is executed on the client device;

receive, at the proxy server and from the client device, an upload request comprising a document;

detect, in the upload request, marker data inserted by the client proxy code executed on the client device;

responsive to the detecting of the marker data in the upload request, cause redaction from the document, resulting in a redacted document;

generate a proxied upload request comprising the redacted document; and

transmit the proxied upload request to an upstream server.

14. The proxy server of claim 13 , wherein the computer-readable instructions are further configured to cause the at least one processor to determine an entity identifier based on the upload request, wherein the redaction of the document is based on the entity identifier.

15. The proxy server of claim 14 , wherein the redaction comprises removing sensitive information from the document that does not match the entity identifier.

16. A non-transitory computer-readable storage medium configured to store computer-readable instructions, the computer-readable instructions configured, upon execution on at least one processor, to cause the at least one processor to implement operations comprising:

receiving a message from a client device;

determining an entity identifier associated with the message;

obtaining an electronic document associated with the message; and

identifying, in the electronic document, a group of redaction candidates, each redaction candidate of the group of redaction candidates comprising a document portion associated with one of a plurality of sensitive information categories, a first redaction candidate of the group of redaction candidates being associated with a particular sensitive information category of the plurality of sensitive information categories and matches the entity identifier itself and a second redaction candidate of the group of redaction candidates being associated with the particular sensitive information category and does not match the entity identifier itself;

filtering, from the group of redaction candidates, ones of the redaction candidates that match the entity identifier itself, resulting in a group of filtered redaction candidates that excludes at least the first redaction candidate based on the first redaction candidate matching the entity identifier itself; and

removing only the group of filtered redaction candidates from the electronic document, resulting in a redacted document comprising the first redaction candidate and not the second redaction candidate; wherein:

the method is performed by a proxy server; the obtaining of the electronic document comprises:

receiving, from a client device, a download request associated with the entity identifier;

in response to the receiving of the download request, transmitting a proxied download request from the proxy server to an upstream server; and

receiving from the upstream server at the proxy server the electronic document in response to the proxied download request; and

the method further comprises transmitting the redacted document from the proxy server to the client device.

17. The computer-readable storage medium of claim 16 , wherein:

the message is a download request, and the obtaining of the electronic document comprises transmitting a proxied download request corresponding to the message to an upstream server, and receiving the electronic document from the upstream server in response to the proxied download request, said operations further comprising transmitting to the client device a response comprising the redacted document; or

the message comprises the document, said operations further comprising transmitting to an upstream server a proxied message comprising the redacted document.

18. The computer-readable storage medium of claim 16 , wherein the message comprises a search request indicating the entity identifier and the electronic document is obtained from document storage via a document search performed based on the search request.

19. The computer-readable storage medium of claim 16 , wherein the entity identifier is a user identifier associated with the message or with the client device, wherein the particular sensitive information category is a predefined personal information category.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2023
From: DANG, AMY HARIHARAN; KADAM, SUNIL SHANKAR; ALMANDIL, HASSAN; CHEN, YIBING; PARAYNO, MARK-GIL; BLACHMAN, MEIR BARUCH; CHERNI, TOMER; FROGEL, NITZAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065318/0039 →
Continuity (1)
Related Publication 20240403374A1 · Dec 5, 2024
References Cited (42)
US 7536635B2 · Racovolis · 2009 [cited by applicant]
US 7913167B2 · Cottrille et al. · 2011 [cited by applicant]
US 8826443B1 · Raman · 2014 [cited by examiner]
US 8973128B2 · Sokolan et al. · 2015 [cited by applicant]
US 10885225B2 · Balzer et al. · 2021 [cited by applicant]
US 10929561B2 · Long et al. · 2021 [cited by applicant]
US 20110239113A1 · Hung · 2011 [cited by examiner]
US 20190129968A1 · Neylan · 2019 [cited by examiner]
US 20200186534A1 · Chanda · 2020 [cited by examiner]
EP 3991389A1 · 2022 [cited by applicant]
“<input type=“hidden”>”, Retrieved From: https://developer.mozilla.org/en-US/docs/Web/HTML/Element/input/hidden, Mar. 13, 2023, 4 Pages. [cited by applicant]
“Apache PDFBox®—A Java PDF Library”, Retrieved From: https://pdfbox.apache.org/, Retrieved On: Nov. 8, 2022, 5 Pages. [cited by applicant]
“Brighter AI's Image & Video Anonymization Solution”, Retrieved From: https://brighter.ai/product/, Retrieved On: Mar. 13, 2023, 14 Pages. [cited by applicant]
“How to Automate your DSAR Process with Discovery & Redaction”, Retrieved From: https://www.onetrust.com/blog/how-to-automate-your-dsar-process-with-discovery-redaction/, Jul. 21, 2021, 7 Pages. [cited by applicant]
“Microsoft Priva Subject Rights Requests”, Retrieved From: https://www.microsoft.com/en-us/security/business/privacy/microsoft-priva-subject-rights-requests, Retrieved On: Mar. 13, 2023, 7 Pages. [cited by applicant]
“OneTrust Acquires DocuVision's Redacted.ai to Expand Automated Data Redaction”, Retrieved From: https://www.onetrust.com/news/onetrust-acquires-docuvisions-redacted-ai-to-expand-automated-data-redaction/, Mar. 1, 2021,… [cited by applicant]
“OneTrust Privacy Rights (DSAR) Automation”, Retrieved From: https://query.prod.cms.rt.microsoft.com/cms/api/am/binary/RE4R7cb, Retrieved On: Mar. 13, 2023, 2 Pages. [cited by applicant]
“Super Redact”, Retrieved From: https://super.ai/redact, Retrieved On: Mar. 13, 2023, 8 Pages. [cited by applicant]
“Take Control of Critical Data to Eliminate Business Risk”, Retrieved From: https://www.datagrail.io/, Retrieved On: Mar. 13, 2023, 5 Pages. [cited by applicant]
“TrustWeek news: OneTrust launches enhanced and automated Data Redaction capabilities”, Retrieved From: https://www.onetrust.com/blog/onetrust-launches-enhanced-and-automated-data-redaction-capabilities/, Oct. 13, 2020,… [cited by applicant]
“Application as Filed in U.S. Appl. No. 17/841,396”, Mailed Date: Jun. 15, 2022, 41 Pages. [cited by applicant]
Dilmegani, Cem, “Guide to Automated Redaction & Its Benefits in 2023”, Retrieved From: https://research.aimultiple.com/automated-redaction/, Dec. 10, 2022, 9 Pages. [cited by applicant]
Foster, Kelsey, “What are the Top PII Redaction APIs and AI Models for 2023?”, Retrieved From: https://www.assemblyai.com/blog/what-are-the-top-pii-redaction-apis-and-ai-models/, Aug. 30, 2022, 17 Pages. [cited by applicant]
Mazzoli, “Sensitive Information Type Entity Definitions”, Retrieved From: https://learn.microsoft.com/en-us/microsoft-365/compliance/sensitive-information-type-entity-definitions?view=o365-worldwide, Feb. 17, 2023, 9 Pa… [cited by applicant]
Moore, Amber, “DataGrail Launches API & Agent to Automate DSR Fulfillment Across All Internal Data Systems, Saving Companies Weeks of Engineering Time”, Retrieved From: https://www.businesswire.com/news/home/20220604005… [cited by applicant]
Rayani, Alym, “Simplify Privacy Protection with Microsoft Priva Subject Rights Requests”, Retrieved From: https://www.microsoft.com/en-us/security/blog/2022/11/10/simplify-privacy-protection-with-microsoft-priva-subject… [cited by applicant]
Sawers, Paul, “OneTrust acquires DocuVision to automatically find and redact sensitive data in documents”, Retrieved From: https://venturebeat.com/business/onetrust-acquires-docuvision-to-automatically-find-and-redact-s… [cited by applicant]
Vandenberg, Steve, “Manage subject rights requests at scale with Microsoft Priva”, Retrieved From: https://www.microsoft.com/en-us/security/blog/2022/03/16/manage-subject-rights-requests-at-scale-with-microsoft-priva/, … [cited by applicant]
Vukos-Walker, et al., “Data Matching for Subject Rights Requests”, Retrieved From: https://learn.microsoft.com/en-us/privacy/priva/subject-rights-requests-data-match, Feb. 22, 2023, 5 Pages. [cited by applicant]
Vukos-Walker, et al., “Integrate and Extend Through Microsoft Graph API and Power Automate”, Retrieved From: https://learn.microsoft.com/en-us/privacy/priva/subject-rights-requests-automate, Feb. 22, 2023, 3 Pages. [cited by applicant]
Vukos-Walker, et al., “Learn About Microsoft Priva”, Retrieved From: https://learn.microsoft.com/en-us/privacy/priva/priva-overview, Apr. 4, 2023, 6 Pages. [cited by applicant]
Vukos-Walker, et al., “Learn about Priva Subject Rights Requests”, Retrieved From: https://learn.microsoft.com/en-us/privacy/priva/subject-rights-requests, Feb. 22, 2023, 2 Pages. [cited by applicant]
“Bleep Private Beta”, Retrieved From: https://game.info.intel.com/bleepbeta-download, Retrieved On: Jun. 1, 2023, 2 Pages. [cited by applicant]
“Bleep Release Notes”, Retrieved From: https://game.info.intel.com/bleepbeta-release-notes, Retrieved On: Jun. 1, 2023, 3 Pages. [cited by applicant]
“Bleep User Guide”, Retrieved From: https://game.info.intel.com/bleepbeta-user-guide, Retrieved On: Jun. 1, 2023, 4 Pages. [cited by applicant]
“Data Loss Protection”, Retrieved From: https://thetalake.com/resource-download-solution-overview-data-loss-protection-for-your-collaboration-platforms/, Retrieved On: Jun. 1, 2023, 4 Pages. [cited by applicant]
“How to Redact a File”, Retrieved From: https://help.accusoft.com/PCC/v8.1/HTML/How%20To%20Redact%20a%20File.html, Retrieved On: Jun. 1, 2023, 1 Page. [cited by applicant]
“Veritone Blog”, Retrieved From: https://www.veritone.com/blog/how-i-built-an-image-proxy-server-to-anonymise-images-in-twenty-minutes/, Jan. 16, 2018, 9 Pages. [cited by applicant]
Porter, John, “Today I Learned About Intel's AI Sliders That Filter Online Gaming Abuse”, Retrieved From: https://www.theverge.com/2021/4/8/22373290/intel-bleep-ai-powered-abuse-toxicity-gaming-filters, Apr. 8, 2021, 9 … [cited by applicant]
Vautier, Nicolas, “Build an SMS Proxy that Redacts PII from Conversation Threads Using Twilio SMS, Pangea Redact Service, and Python”, Retrieved From: https://www.twilio.com/blog/build-sms-proxy-redact-pii-from-sms-conv… [cited by applicant]
Campbell, Oliver-James, “Tech Companies Want to Tackle Harassment in Gaming”, Retrieved From: https://www.wired.com/story/tech-companies-harassment-gaming-riot-intel-microsoft/, Jun. 20, 2021, 18 Pages. [cited by applicant]
“Your top 10 data redaction questions answered”, Retrieved From: https://www.onetrust.com/blog/data-redaction-faq/, Mar. 30, 2021, 10 Pages. [cited by applicant]