IP Library Granted Patent US 12,216,799
Granted Patent B2
US 12,216,799 · App. 18/381,873 · Granted Feb 4, 2025

Systems and methods for computing with private healthcare data

Inventors: Sankar Ardhanari (Chapel Hill, NC); Karthik Murugadoss (Cambridge, MA); Murali Aravamudan (Andover, MA); Ajit Rajasekharan (West Windsor, NJ)
Assignee: nference, Inc.
G06F21/6254G06N20/00G16H10/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,216,799
App. No.
18/381,873
Granted
Feb 4, 2025
Kind
B2
Abstract

Techniques are provided for computing with private healthcare data. The techniques include a de-identification method including receiving a text sequence; providing the text sequence to a plurality of entity tagging models, each of the plurality of entity tagging models being trained to tag one or more portions of the text sequence having a corresponding entity type; tagging one or more entities in the text sequence using the plurality of entity tagging models; and obfuscating each entity among the one or more tagged entities by replacing the entity with a surrogate, the surrogate being selected based on one or more attributes of the entity and maintaining characteristics similar to the entity being replaced.

Claims (52)

1. A de-identification method comprising:

receiving a plurality of data sets, wherein the plurality of data sets comprises:

a first data set, wherein the first data set comprises a labeled data set for one or more entity types; and

a second data set, wherein the training data set comprises an unlabeled data set for the one or more entity types;

determining one machine-learning model from a plurality of machine-learning models for each of one or more entity types;

fine-tuning the determined machine-learning model for each of the one or more entity types, wherein fine-tuning the determined machine-learning model comprises:

creating a plurality of training data sets, wherein the plurality of training data sets comprises:

a first training data set, wherein the first training data set comprises the first data set; and

a second training data set, wherein the second training data set comprises the second data set;

training the determined machine-learning model using the first training data set;

validating the trained machine-learning model, wherein validating the trained machine learning model further comprises:

generating a recall score for each entity type of the one or more entity types;

comparing the recall score to a threshold for the recall score for each entity type of the one or more entity types; and

updating the trained machine-learning model using the second training data set as a function of the validation; and

obfuscating the second data set using the fine-tuned machine-learning model.

2. The de-identification method of claim 1 , wherein obfuscating the second data set further comprises replacing two or more entities that refer to a common subject with a common surrogate.

3. The de-identification method of claim 2 , further comprising:

selecting the common surrogate based on one or more attributes of the two or more entities.

4. The de-identification method of claim 2 , further comprising:

selecting the common surrogate based on a gender associated with the two or more entities.

5. The de-identification method of claim 2 , further comprising:

selecting the common surrogate based on an ethnicity associated with the two or more entities.

6. The de-identification method of claim 1 , wherein validating the trained machine learning model further comprises:

generating a precision score for each entity type of the one or more entity types; and

comparing the precision score to a threshold for the precision score for each entity type of the one or more entity types.

7. The de-identification method of claim 1 , wherein validating the trained machine learning model further comprises:

generating a F-score for each entity type of the one or more entity types; and

comparing the F-score to a threshold for the F-score for each entity type of the one or more entity types.

8. The de-identification method of claim 1 , further comprising:

updating the trained machine-learning model as a function of a comparison of an average of a F-score, a precision score and the recall score to a threshold success percentage.

9. The de-identification method of claim 1 , wherein:

the one or more entity types comprises two or more personal names; and

obfuscating the second data set further comprises replacing each of the two or more personal names with a different surrogate.

10. The de-identification method of claim 1 , wherein obfuscating the second data set further comprises replacing two or more entities that refer to a common person with surrogates that match a gender associated with the common person.

11. The de-identification method of claim 1 , wherein obfuscating the second data set further comprises replacing two or more entity types that refer to a common person with surrogates that match an ethnicity associated with the common person.

12. The de-identification method of claim 1 , wherein:

obfuscating the second data set further comprises replacing two or more entities that represent dates with surrogate dates, wherein:

the surrogate dates are based on the two or more entity types altered by a random value.

13. The de-identification method of claim 12 , wherein the surrogate dates associated with a common patient are altered by the same random value.

14. The de-identification method of claim 1 , wherein obfuscating the second data set further comprises scrambling two or more entities that represent numeric identifiers with random values to scramble the numeric identifiers.

15. The de-identification method of claim 1 , wherein the one or more entity types comprises at least a portion of an electronic health record.

16. The de-identification method of claim 1 , further comprising:

receiving a text sequence; and

tagging one or more entities in the text sequence.

17. The de-identification method of claim 16 , further comprising:

aggregating the tagged entities from the text sequence; and

passing the aggregated tagged entities through one or more dreg filters, wherein each of the one or more dreg filters is configured to filter a corresponding entity type based on a rule-based template.

18. The de-identification method of claim 17 , further comprising, creating the rule-based template, wherein creating the rule-based template comprises:

mapping each of one or more portions of the text sequence to a corresponding syntax template;

identifying a candidate syntax template based on a second machine learning model that infers one or more candidate syntax templates based on the one or more portions of the text sequence; and

creating the rule-based template from the candidate syntax template by replacing each of the one or more tagged entities in the portion of the text sequence corresponding to the candidate template with a corresponding syntax token.

19. The de-identification method of claim 17 , wherein each of the one or more dreg filters is further configured to filter the corresponding entity type based on a pattern-based filter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2024
From: ARDHANARI, SANKAR; MURUGADOSS, KARTHIK; ARAVAMUDAN, MURALI; RAJASEKHARAN, AJIT
To: NFERENCE, INC.
Reel/Frame 069666/0770 →
Continuity (11)
Continuation 17975489 · Oct 27, 2022
Continuation 17192564 · Mar 4, 2021
Continuation In Part 16908520 · Jun 22, 2020
Provisional Application 63128542 · Dec 21, 2020
Provisional Application 63109769 · Nov 4, 2020
Provisional Application 63012738 · Apr 20, 2020
Provisional Application 62984989 · Mar 4, 2020
Provisional Application 62985003 · Mar 4, 2020
Provisional Application 62962146 · Jan 16, 2020
Provisional Application 62865030 · Jun 21, 2019
Related Publication 20240119176A1 · Apr 11, 2024
References Cited (94)
US 7542969B1 · Rappaport et al. · 2009 [cited by applicant]
US 9514405B2 · Chen et al. · 2016 [cited by applicant]
US 9953095B1 · Scott et al. · 2018 [cited by applicant]
US 10360507B2 · Aravamudan et al. · 2019 [cited by applicant]
US 11062218B2 · Aravamudan et al. · 2021 [cited by applicant]
US 11487902B2 · Ardhanari et al. · 2022 [cited by applicant]
US 11545242B2 · Aravamudan · 2023 [cited by applicant]
US 20040013302A1 · Ma et al. · 2004 [cited by applicant]
US 20070039046A1 · Van Dijk et al. · 2007 [cited by applicant]
US 20080118150A1 · Balakrishnan et al. · 2008 [cited by applicant]
US 20080243825A1 · Staddon et al. · 2008 [cited by applicant]
US 20090116736A1 · Neogi et al. · 2009 [cited by applicant]
US 20110255788A1 · Duggan et al. · 2011 [cited by applicant]
US 20110307460A1 · Vadlamani et al. · 2011 [cited by applicant]
US 20130132331A1 · Kowalczyk et al. · 2013 [cited by applicant]
US 20150161413A1 · Calem · 2015 [cited by examiner]
US 20150254555A1 · Williams, Jr. et al. · 2015 [cited by applicant]
US 20160105402A1 · Soon-Shiong et al. · 2016 [cited by applicant]
US 20160247307A1 · Stoop et al. · 2016 [cited by applicant]
US 20170032243A1 · Corrado et al. · 2017 [cited by applicant]
US 20170061326A1 · Talathe et al. · 2017 [cited by applicant]
US 20170091391A1 · LePendu · 2017 [cited by applicant]
US 20180060282A1 · Kaljurand · 2018 [cited by applicant]
US 20180212971A1 · Costa · 2018 [cited by applicant]
US 20190319982A1 · Durand · 2019 [cited by examiner]
US 20190354883A1 · Aravamudan et al. · 2019 [cited by applicant]
US 20200311300A1 · Callcut · 2020 [cited by examiner]
US 20200402625A1 · Aravamudan et al. · 2020 [cited by applicant]
US 20210019287A1 · Prasad et al. · 2021 [cited by applicant]
US 20210064781A1 · Raphael · 2021 [cited by examiner]
US 20210224264A1 · Barve et al. · 2021 [cited by applicant]
US 20210248268A1 · Ardhanari et al. · 2021 [cited by applicant]
US 20220050921A1 · LaFever et al. · 2022 [cited by applicant]
US 20220138599A1 · Aravamudan et al. · 2022 [cited by applicant]
CN 105938495A · 2016 [cited by applicant]
EP 3987426 · 2022 [cited by applicant]
JP 2019536178A · 2019 [cited by applicant]
WO WO2015084759A1 · 2015 [cited by applicant]
WO WO2015149114 · 2015 [cited by applicant]
WO WO2018057945 · 2018 [cited by applicant]
WO WO2020257783 · 2020 [cited by applicant]
WO WO2021011776 · 2021 [cited by applicant]
WO WO2021146894A1 · 2021 [cited by applicant]
WO WO2021178889 · 2021 [cited by applicant]
AMD Secure Encrypted Virtualization (SEV), https://developer.amd.com/sev/, accessed Sep. 23, 2020 (5 pages). [cited by applicant]
Arora, S. et al., “A Simple but Tough-to-Beat Baseline for Sentence Embeddings”, ICLR, 2017 (16 pages). [cited by applicant]
AWS Key Management Service (KMS), https://aws.amazon.com/kms, accessed Jan. 20, 2021 (3 pages). [cited by applicant]
AWS Key Management Service (KMS): https://aws.amazon.com/kms; accessed Sep. 23, 2020 (6 pages). [cited by applicant]
Bartunov, S. et al., “Breaking Sticks and Ambiguities With Adaptive Skip-Gram”, retrieved online from URL:<https://arxiv.org/pdf/1502.07257.pdf>,[cs CL], Nov. 15, 2015 (15 pages). [cited by applicant]
Bojanowski, P. et al., “Enriching Word Vectors with Subword Information” retrieved online from URL:<https://arxiv.org/pdf/1607.04606.pdf>, [cs CL], Jun. 19, 2017 (12 pages). [cited by applicant]
Confidential Computing Consortium, “What is the Confidential Computing Consortium?”, https://confidentialcomputing.io. accessed Sep. 24, 2020 (2 pages). [cited by applicant]
de Guzman, C.G. al. “Hematopoietic Stem Cell Expansion and Distinct Myeloid Developmental Abnormalities in a Murine Model of the AML1-ETO Translocation”, Molecular and Cellular Biology, 22(15):5506-5517. Aug. 2002 (12 p… [cited by applicant]
Desagulier, G.,“A lesson from associative learning asymmetry and productivity in multiple-slot constructions”, Corpus Linguistic and Linguistic Theory, 12(2):173-219, 2016, submitted Aug. 13, 2015, <http://www.degruyter… [cited by applicant]
Devlin. J. et al., “Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv:1810.04805v2 [cs.CL], May 24, 2019 (16 pages). [cited by applicant]
Divatia, A., “The Fact and Fiction of Homomorphic Encryption”, Dark Reading, www.darkreading.com/attacks-breaches/the-fact-and-fiction-of-homomorphic-encryption/a/d-ld/1333691 Jan. 22, 2019 (3 pages). [cited by applicant]
Dwork, C., “Differential Privacy: A Survey of Results”, Lecture Notes in Computer Science, vol. 4978, pp. 1-19, 2008 (19 pages). [cited by applicant]
Ferraiuolo, A. et al. “Kornodo: Using verification to disentangle secure-enclave hardware from software”, SOSP 17, Shanghai, China, pp. 287-305, Oct. 28. 2017 (19 pages). [cited by applicant]
Garten, Y. et al., “Pharmspresso; a text mining tool for extraction of pharmacogenomic concepts and relationships from full text”, BMC Bioinformatics, 10(Suppl. 2):S6, February 5. 2009 (9 pages). [cited by applicant]
Genkin, D. et al., “Privacy in Decentralized Cryptocurrencies”, Communications of the ACM, 61(6):78-88, Jun. 2018 (11 pages). [cited by applicant]
Hageman, G.S. et al., “A common haplotype in the complement regulatory gene factor H (HF1 /CFH) predisposes individuals to age-related macular degeneration”, PNAS, 102(20):7227-7232. May 17, 2005 (6 pages). [cited by applicant]
Ikeda, T. et al., “Anticorresponding mutations of the KRAS and PTEN genes in human endometrial cancer”, Oncology Reports, 7:567-570, published online May 1, 2000 (4 pages). [cited by applicant]
Intel, “What is intel® SGX?”, http://www.intel.com/content/www/us/en/architecture-and-technology/software-guard-extensions.html, accessed Sep. 23, 2020 (8 pages). [cited by applicant]
International Preliminary Report on Patentability issued in International Application No. PCT/US2021/020906, Sep. 15, 2022 (9 pages). [cited by applicant]
International Search Report and Written Opinion issued by the European Patent Office as International Searching Authority in International Application PCT/US2017/053039, dated Dec. 20, 2017 (15 pages). [cited by applicant]
International Search Report and Written Opinion issued by the U.S. Patent and Trademark Office as International Searching Authority Issued in: International Application No. PCT/US21/20906, dated May 19, 2021 (10 pages). [cited by applicant]
International Search Report and Written Opinion issued by U.S. Patent and Trademark Office as International Searching Authority in International Application No. PCT/US20/42336, dated Sep. 30, 2020 (10 pages). [cited by applicant]
International Search Report and Written Opinion Issued by U.S. Patent and Trademark Office as International Searching Authority, for International Application No. PCT/US20/38987, dated Nov. 9, 2020 (26 pages). [cited by applicant]
Joulin, A. et al., “Bag of Tricks for Efficient Text Classification”, retrieved online from URL:<https://arXiv.org/pdf/1607.01759v3.pdf>, [cs CL], Aug. 9, 2016 (5 pages). [cited by applicant]
Kiros, R. et al., “Skip-Thought Vectors” retrieved online from URL:<https://arXiv.org/abs/1506.06726v.1>, [cs CL], Jun. 22, 2015 (11 pages). [cited by applicant]
Kolte, P. “Why is Homomorphic Encryption Not Ready for Primetime?”, Baffle, https://baffle.io/blog/why-is-hornomorphic-encryption-not-ready-for-primetime/, Mar. 17, 2017 (4 pages). [cited by applicant]
Korger, C., “Clustering of Distributed Word Representations and its Applicability for Enterprise Search”, Doctoral Thesis, Dresden University of Technology, Faculty of Computer Science, Institute of Software and Multime… [cited by applicant]
Kutuzov, A. et al., “Cross-lingual Trends Detection for Named Entities in News Texts with Dynamic Neural Embedding Models”, Proceedings of the NewsIR 16 Workshop at ECIR, Padua, Italy, Mar. 20, 2016 (6 pages). [cited by applicant]
Le, Q. et al., “Distributed Representations of Sentences and Documents”, Proceedings of the 31st International Conference of Machine Learning, Beijing, China, vol. 32, 2014 (9 pages). [cited by applicant]
Li, H. et al., “Cheaper and Better: Selecting Good Workers for Crowdsourcing,” retrieved online from URL: https://arXiv.org/abs/1502.00725v.1, pp. 1-16, Feb. 3, 2015 (16 pages). [cited by applicant]
Ling, W. et al., “Two/Too Simple Adaptations of Word2Vec for Syntax Problems”, retrieved online from URL:<https://cs.cmu.edu/˜lingwang/papers/naacl2015.pdf>, 2015 (6 pages). [cited by applicant]
Maxwell, K.N. et al., “Adenoviral-mediated expression of Pcsk9 in mice results in a low-density lipoprotein receptor knockout phenotype”, PNAS, 101(18):7100-7105, May 4, 2004 (6 pages). [cited by applicant]
Mikolov, T. et al., “Distributed Representations for Words and Phrases and their Compositionality”, retrieved online from URL:https://arXiv.org/abs/1310.4546.v1 [cs CL], Oct. 16, 2013 (9 pages). [cited by applicant]
Mikolov, T. et al., “Efficient Estimation of Word Representations in Vector Space”, retrieved online from URL: https://arXiv.org/abs/1301.3781v3 [cs CL] Sep. 7, 2013 (12 pages). [cited by applicant]
Murray, K., “A Semantic Scan Statistic for Novel Disease Outbreak Detection”, Master's Thesis, Carnegie Mellon University, Aug. 16, 2013 (68 pages). [cited by applicant]
Neelakantan, A., et al., “Efficient Non-parametric Estimation of Multiple Embeddings per Word in Vector Space,” Department of Computer Science, University of Massachusetts, (2015) (11 pages). [cited by applicant]
Pennington, J. et al., “GloVe: Global Vectors for Word Representation”; retrieved online from URL:<https://nlp.stanford.edu/projects/glove.pdf>,2014 (12 pages). [cited by applicant]
Rajagopalan, H., et al., “Tumorigenesis: RAF/RAS oncogenes and mismatch-repair status”, Nature, 418:934, Aug. 29, 2002 (1 page). [cited by applicant]
Rajasekharan, “Unsupervised NER using BERT”, [accessed Jun. 10, 2021], Toward Data Science, <URL: https://towardsdatascience.com/unsupervised-ner-using-bert-2d7af5f90b8a>, Feb. 28, 2020 (27 pages). [cited by applicant]
Shamir, A. “How to Share a Secret”, Communications of the ACM, 22(11):612-613, Nov. 1979 (2 pages). [cited by applicant]
Shweta, Fnu et al., “Augmented Curation of Unstructured Clinical Notes from a Massive EHR System Reveals Specific Phenotypic Signature of impending COVID-19 Diagnosis”, https://www.medrxiv.org/content/10.1101/2020.04.19… [cited by applicant]
Van Mulligen, E.M. et al., “The EU-ADR corpus: Annotated drugs, diseases, targets, and their relationships”, Journal of Biomedical Informatics, 45:879-884, published online Apr. 25, 2012 (6 pages). [cited by applicant]
Wieting, J. et al., “Revisiting Recurrent Networks for Paraphrastic Sentence Embeddings”, retrieved online from URL:<https://arXiv.org/pdf/1705.00364v1.pdf>, [cs CL], Apr. 30, 2017 (12 pages). [cited by applicant]
Yao, Z. et al., “Dynamic Word Embeddings for Evolving Semantic Discovery”, WSDM 2018, Marina Del Rey, CA, USA, Feb. 5-9, 2018 (9 pages). [cited by applicant]
Zuccon, G., et al., “Integrating and Evaluating Neural Word Embeddings in Information Retrieval”, ADCS, Parramatta, NSW, Australia, Dec. 8-9, 2015 (8 pages). [cited by applicant]
Daniel Genkin et al: “Privacy in decentralized cryptocurrencies”, Communications of the ACM, Association for Computing Machinery, Inc, United States, vol. 61, No. 6, May 23, 2018 (May 23, 2018), pp. 78-88, XP058407634, … [cited by applicant]
Bradley Malin: “Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule”, Nov. 26, 2012 (Nov. 26, … [cited by applicant]
Devlin Jacob et al: “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Oct. 11, 2018 (Oct. 11, 2018), XP055968792, Retrieved from the Internet: URL:https://www.arxiv.org/pdf/1810.04805vl… [cited by applicant]
Rajasekharan Ajit: “Unsupervised NER using BERT”, Towards Data Science, Feb. 28, 2020 (Feb. 28, 2020), XP093054321, Retrieved from the Internet: URL:https://towardsdatascience.com/unsuper vised-ner-using-bert-2d7af5f90b… [cited by applicant]
European Search Report; EP 24 17 9988; By: Maenpaa, Jari; Sep. 2, 2024. [cited by applicant]