IP Library › Granted Patent US 12,743,550
Granted Patent B2
US 12,743,550 · App. 18/766,316 · Granted Sep 22, 2026

Advanced deidentification of information, such as information about a person

Inventors: Lindsay Thomas Mico (Portland, OR); Vivek Tomer (Lake Oswego, OR); Yuqing Guo (Amherst, MA); Amar Nadaa Taiyab (Encinitas, CA)
Assignee: Providence St. Joseph Health
G06F21/6254G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,550
App. No.
18/766,316
Granted
Sep 22, 2026
Kind
B2
Abstract

A facility for the identifying contents of a data object is described. The facility identifies in the data object two or more constituent portions. For each of the constituent portions identified in the data object, the facility: identifies a type of data items occurring within the constituent portion; on the basis of the identified data item type, selects a deidentification operation; and causes the selected deidentification operation to be performed on the data items of the constituent portion, such that these data items are modified to make the data items less identifiable with a person, and/or less-harmfully identifiable with a person. After the causing, the facility assembles the constituent portions containing the modified data items into a modified version of the data object.

Claims (95)

1 . One or more instances of non-transitory computer-readable media collectively having contents configured to cause a computing system to perform a method, the method comprising:

receiving a data object;

identifying in the data object a plurality of constituent portions;

selecting a first constituent portion of the data object containing one or more free text strings;

in a first computing mechanism:

for each free text string:

subjecting the free text string to a trained machine learning model to predict occurrence within the free text string of a personal identifier based on predicting occurring within of certain named entities;

in a repeatable manner, generating a substitute personal identifier from the predicted personal identifier;

creating a copy of the free text string;

in the free text string copy, replacing the predicted personal identifier with the generated substitute personal identifier;

after the replacement, collecting the free text string copies in a modified version of the first constituent portion;

selecting a second constituent portion of the data object distinct from the first constituent portion containing data items of a type other than free text strings;

in a second computing mechanism distinct from the first computing mechanism:

for each data item:

in a repeatable manner, generating a substitute data item from the data item;

collecting the substitute data items in a modified version of the second constituent portion; and

assembling the modified version of the first constituent portion and modified version of the second constituent portion into a modified version of the data object.

2 . The one or more instances of non-transitory computer-readable media of claim 1 , the method further comprising:

in the first computing mechanism:

for each free text string:

before the subjecting:

identifying tokens in the free text string;

determining semantic embeddings representing the identified tokens;

in the free text string, substituting the determined semantic embeddings for the identified tokens.

3 . The one or more instances of non-transitory computer-readable media of claim 1 , the method further comprising:

in the first computing mechanism:

for each free text string:

performing procedural pattern matching in the free text string to predict occurrence within the free text string of a personal identifier based on occurrence within the free text string of substrings matching predetermined text patterns;

in a repeatable manner, generating a substitute personal identifier from the personal identifier predicted based on the pattern matching; and

in the free text string copy, replacing the personal identifier predicted based on the pattern matching with the substitute personal identifier generated from the personal identifier predicted based on the pattern matching.

4 . The one or more instances of non-transitory computer-readable media of claim 1 , the method further comprising:

for each free text string, after the replacement:

repeatably determining a hash value for the free text string;

persistently storing a mapping between the determined hash value and the copy of the free text string; and

for each free text string, before the subjecting, generating, creating, and replacing:

repeatably determining a hash value for the free text string;

determining whether the determined hash value occurs among the persistently-stored mappings;

if the determining determines that the determined hash value occurs among the persistently-stored mappings, collecting the copy of the free text string mapped-to from the determined hash value into the modified version of the first constituent portion without performing the subjecting, generating, creating, and replacing; and

only if the determining determines that the determined hash value does not occur among the persistently-stored mappings, performing the subjecting, generating, creating, and replacing.

5 . The one or more instances of non-transitory computer-readable media of claim 1 wherein the level of computing resources consumed in the first computing mechanism scales linearly with the volume of data objects in which a first constituent portion is selected.

6 . The one or more instances of non-transitory computer-readable media of claim 1 wherein the subjecting is performed by streaming individual free text strings through a swarm of single processing nodes of the first computing mechanism.

7 . A method in a computing system, comprising:

receiving a first data object;

identifying in the first data object a plurality of constituent portions;

for each of the constituent portions identified in the first data object:

identifying a type of data items occurring within the constituent portion;

on the basis of the identified data item type, selecting a deidentification operation;

causing the selected deidentification operation to be performed on the data items of the constituent portion, such that these data items are modified to make the data items less identifiable with a person, and/or less-harmfully identifiable with a person; and

after the causing, assembling the constituent portions containing the modified data items into a modified version of the first data object,

wherein, for each of at least one distinguished constituent portion among the identified constituent portions, a deidentification operation is selected that replaces data objects with artificial data objects of the type identified in the distinguished constituent portions contained by an automatically-generated faker file,

wherein, for at least one of the identified constituent portions, the selected deidentification operation is deterministic, such that each time it is performed on a particular data item, the same modified data item results,

wherein a data item contained by the first data object is a distinguished identifier, for which a distinguished modified data item that is the modified distinguished identifier is included in the modified version of the first data object, and

wherein, when the method is repeated with respect to a second data object containing a data item that is the distinguished identifier, the modified version of the second data object contains the distinguished modified data item,

such that the modified distinguished identifier occurs in both the first and second data objects, and can be correlated between the first and second data objects.

8 . The method of claim 7 wherein, for each of one or more distinguished ones of the identified constituent portions, the deidentification operation that is caused to be performed can be reversed for a data item based upon the corresponding modified data item and additional data,

the method further comprising:

for each distinguished constituent portion:

for each data item of the distinguished constituent portion:

persistently storing the additional data.

9 . The method of claim 8 wherein the deidentification operation that is caused to be performed for one of the distinguished constituent portions is a date shift deidentification operation, and the additional data is a number of days.

10 . The method of claim 9 wherein the number of days is stored for each of a plurality of people.

11 . The method of claim 8 wherein the deidentification operation that is caused to be performed for one of the distinguished constituent portions containing identifiers is a mapping between the identifier and a one-way hash result determined for the identifier.

12 . The method of claim 7 wherein the causing causes the selected the identification operation to be performed by a first computing mechanism for constituent portions whose data items are of data item types among a first set of data item types, and wherein the causing causes the selected the identification operation to be performed by a second computing mechanism distinct from the first computing mechanism for constituent portions whose data items are of data item types among a second set of data item types.

13 . The method of claim 12 wherein the second computing mechanism performs named entity recognition using one or more machine learning models.

14 . The method of claim 7 wherein the selection is performed with reference to an override table.

15 . A method in a computing system, comprising:

receiving a data object;

identifying in the data object a plurality of constituent portions;

selecting a first constituent portion of the data object containing one or more free text strings;

in a first computing mechanism:

for each free text string:

subjecting the free text string to a trained machine learning model to predict occurrence within the free text string of a personal identifier based on predicting occurring within of certain named entities;

in a repeatable manner, generating a substitute personal identifier from the predicted personal identifier;

creating a copy of the free text string;

in the free text string copy, replacing the predicted personal identifier with the generated substitute personal identifier;

after the replacement, collecting the free text string copies in a modified version of the first constituent portion;

selecting a second constituent portion of the data object distinct from the first constituent portion containing data items of a type other than free text strings;

in a second computing mechanism distinct from the first computing mechanism:

for each data item:

in a repeatable manner, generating a substitute data item from the data item;

collecting the substitute data items in a modified version of the second constituent portion; and

assembling the modified version of the first constituent portion and modified version of the second constituent portion into a modified version of the data object.

16 . One or more instances of non-transitory computer-readable media collectively having contents configured to cause in a computing system to perform a method, the method comprising:

receiving a first data object;

identifying in the first data object a plurality of constituent portions;

for each of the constituent portions identified in the first data object:

identifying a type of data items occurring within the constituent portion;

on the basis of the identified data item type, selecting a deidentification operation;

causing the selected deidentification operation to be performed on the data items of the constituent portion, such that these data items are modified to make the data items less identifiable with a person, and/or less-harmfully identifiable with a person; and

after the causing, assembling the constituent portions containing the modified data items into a modified version of the first data object,

wherein, for each of at least one distinguished constituent portion among the identified constituent portions, a deidentification operation is selected that replaces data objects with artificial data objects of the type identified in the distinguished constituent portions contained by an automatically-generated faker file,

wherein, for at least one of the identified constituent portions, the selected deidentification operation is deterministic, such that each time it is performed on a particular data item, the same modified data item results,

wherein a data item contained by the first data object is a distinguished identifier, for which a distinguished modified data item that is the modified distinguished identifier is included in the modified version of the first data object, and

wherein, when the method is repeated with respect to a second data object containing a data item that is the distinguished identifier, the modified version of the second data object contains the distinguished modified data item,

such that the modified distinguished identifier occurs in both the first and second data objects, and can be correlated between the first and second data objects.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2024
From: MICO, LINDSAY THOMAS; TOMER, VIVEK; GUO, YUQING; TAIYAB, AMAR NADAA
To: PROVIDENCE ST. JOSEPH HEALTH
Reel/Frame 067941/0339 →
Continuity (2)
Continuation 18612921 · Mar 21, 2024
Related Publication 20250298923A1 · Sep 25, 2025
References Cited (47)
US 8209549B1 · Bain, III · 2012 [cited by examiner]
US 10691825B2 · Jones · 2020 [cited by examiner]
US 10769305B2 · Lowenberg · 2020 [cited by examiner]
US 10990698B2 · Postnikov · 2021 [cited by examiner]
US 11127403B2 · Medalion · 2021 [cited by examiner]
US 11157652B2 · Basava · 2021 [cited by examiner]
US 11249710B2 · Li · 2022 [cited by examiner]
US 11397831B2 · Lowenberg · 2022 [cited by examiner]
US 11610010B2 · Zargarian · 2023 [cited by examiner]
US 11741103B1 · MacManus · 2023 [cited by examiner]
US 11757626B1 · Rivlin · 2023 [cited by applicant]
US 11816116B2 · Sankaran · 2023 [cited by examiner]
US 11893128B2 · Liu · 2024 [cited by examiner]
US 11983145B2 · Hines · 2024 [cited by examiner]
US 12073001B1 · Mico · 2024 [cited by examiner]
US 12086286B2 · Poe · 2024 [cited by examiner]
US 12184772B1 · Belchee · 2024 [cited by examiner]
US 20200401734A1 · Murdoch · 2020 [cited by examiner]
US 20210004373A1 · Sankaran · 2021 [cited by examiner]
US 20210019374A1 · Donaldson · 2021 [cited by examiner]
US 20210182413A1 · Agarwal · 2021 [cited by examiner]
US 20220207229A1 · Perkins · 2022 [cited by applicant]
US 20220222373A1 · Villax · 2022 [cited by examiner]
US 20230222288A1 · Zhang et al. · 2023 [cited by applicant]
US 20230282322A1 · Rajkumar · 2023 [cited by examiner]
US 20230376858A1 · Tal · 2023 [cited by examiner]
US 20240004623A1 · Groenewegen · 2024 [cited by examiner]
US 20240012885A1 · Sneider et al. · 2024 [cited by applicant]
US 20250036809A1 · Nadav · 2025 [cited by examiner]
US 20250103742A1 · Ma · 2025 [cited by examiner]
US 20250112941A1 · Kling · 2025 [cited by examiner]
US 20250200204A1 · Maikhuri · 2025 [cited by examiner]
US 20250258935A1 · Shintani · 2025 [cited by examiner]
Brown et al., “Interval Estimation for a Binomial Proportion,” [cited by applicant]
Chiu et al., “Named Entity Recognition with Bidirectional LSTM-CNNs,” [cited by applicant]
Cochran, “Sampling Techniques, third edition,” John Wiley & Sons, Inc., New York, Feb. 1977. (442 pages). [cited by applicant]
Dernoncourt et al., “De-identification of patient notes with recurrent neural networks,” [cited by applicant]
Douglass et al., “Computer-Assisted De-Identification of Free Text in the MIMIC II Database,” [cited by applicant]
Emam et al., “Anonymizing Health Data: Case Studies and Methods to Get You Started,” O'Reilly Media, Inc. Sebastopol, CA, Aug. 13, 2014. (228 pages). [cited by applicant]
Ferrández et al., “Evaluating current automatic de-identification methods with Veteran's health administration clinical documents,” [cited by applicant]
Hyndman, “Computing and Graphing Highest Density Regions,” [cited by applicant]
Khin et al., “A Deep Learning Architechture for De-identification of Patient Notes: Implementation and Evaluation,” Oct. 3, 2018, URL=arxiv.org/abs/1810.01570, downloaded on Feb. 21, 2024. (15 pages). [cited by applicant]
Kocaman et al., “Accurate Clinical and Biomedical Named Entity Recognition at Scale,” [cited by applicant]
Liu et al., “De-identification of Clinical Notes via Recurrent Neural Network and Conditional Random Field,” J Biomed Inform 75(S34-S42), Nov. 2017 (HHS Public Access Author Manuscript, available in PMC Nov. 1, 2018). (… [cited by applicant]
Lu et al., “Analysis of regression confidence intervals and Bayesian credible intervals for uncertainty quantification,” [cited by applicant]
Snedecor et al., “Statistical Methods, Eighth Edition,” Iowa State University Press, Ames, Iowa, 1989. (536 pages). [cited by applicant]
Wright, “A Simple Method of Exact Optimal Sample Allocation under Stratification with Any Mixed Constraint Patterns,” Center for Statistical Research & Methodology, Research Report Series (Statistics #2014-07), Aug. 21,… [cited by applicant]