IP Library Granted Patent US 11,568,080
Granted Patent B2
US 11,568,080 · App. 15/036,537 · Granted Jan 31, 2023

Systems and method for obfuscating data using dictionary

Inventors: Brian J. Stankiewicz (Mahtomedi, MN); Eric C. Lobner (Woodbury, MN); Richard H. Wolniewicz (Longmont, CO); William L. Schofield (Silver Spring, MD)
Assignee: 3M Innovative Properties Company
G06F21/6245G06F16/24568G16H50/70G06F2221/2101
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,080
App. No.
15/036,537
Granted
Jan 31, 2023
Kind
B2
Abstract

At least some aspects of the present disclosure feature systems and methods for obfuscating data. The method includes the steps of receiving an input data stream including a sequence of n-grams, mapping at least some of the sequence of n-grams to corresponding dictionary terms using a dictionary, and disposing the corresponding tokens to an output data stream.

Claims (45)

1. A method for obfuscating data using a computer system having one or more processors and memories, the method comprising:

obtaining, by the one or more processors, a dictionary that includes a set of predefined dictionary terms and a set of n-grams and that maps the set of n-grams to the predefined dictionary terms;

receiving, by the one or more processors, a first digital data stream comprising a sequence of n-grams, wherein each respective n-gram represents either a respective single word or a respective contiguous sequence of words, wherein ‘n’ represents a number of words in the respective contiguous sequence, and wherein the first data stream is from multiple documents;

parsing, by the one or more processors, the first data stream to distinguish each respective n-gram of the sequence of n-grams received in the first data stream;

comparing, by the one or more processors, each respective n-gram of the sequence of n-grams received in the first data stream with the predefined dictionary terms included in the dictionary;

when a respective n-gram is identified in the dictionary based on the comparing, disposing, by the one or more processors, a corresponding dictionary term in a second digital data stream, wherein the second data stream comprises a predetermined data structure including one or more descriptors, wherein the corresponding dictionary term is an obfuscated token representing the respective n-gram, wherein the obfuscated token comprises a part-of-speech identifier that is appended to the obfuscated token, and after each n-gram in the first data stream has been compared, the second data stream comprises only a plurality of obfuscated tokens;

applying, by the one or more processors, a statistical process on the second data stream, wherein the statistical process includes at least determining one or more usage patterns of one or more of the obfuscated tokens in the second data stream;

generating, by the one or more processors, one or more feature vectors based on the second data stream;

providing, by the one or more processors, the one or more feature vectors to a machine learning model having been trained to perform a clustering analysis on any number of received feature vectors;

performing, by the one or more processors and using the machine learning model, the clustering analysis of the one or more feature vectors;

identifying, by the one or more processors, documents based on the results of the clustering analysis; and

predicting, by the one or more processors and using the identified documents, one or more potentially preventable conditions.

2. The method of claim 1 , wherein at least one of the one or more descriptors describes a category of dictionary terms.

3. The method of claim 1 , further comprising:

predicting, by the one or more processors, a relationship between the one or more potentially preventable conditions and a combination of n-grams from the subset of n-grams disposed in the second data stream.

4. The method of claim 1 , further comprising:

encrypting and tokenizing, by the one or more processors, the first data stream using a random seed; and

providing, by the one or more processors, the random seed to the user to reverse the obfuscation of one or more obfuscated tokens in the second data stream.

5. The method of claim 1 , wherein n is greater than 1.

6. The method of claim 1 , further comprising:

applying, by the one or more processors, one or more machine learning algorithms on the one or more determined usage patterns associated with the obfuscated tokens of the second data stream.

7. A computer system comprising:

a memory storing a dictionary that maps a set of n-grams to predefined dictionary terms; and

one or more processors in communication with the memory, the one or more processors being configured to:

obtain the dictionary from the memory;

receive a first digital data stream comprising a sequence of n-grams, wherein each respective n-gram represents either a respective single word or a respective contiguous sequence of words, wherein ‘n’ represents a number of words in the respective contiguous sequence, and wherein the first data stream is from multiple documents;

parse the first data stream to distinguish each respective n-gram of the sequence of n-grams received in the first data stream;

compare each respective n-gram of the sequence of n-grams received in the first data stream with the predefined dictionary terms included in the dictionary;

when a respective n-gram is identified in the dictionary based on the comparing, dispose a corresponding dictionary term in a second digital data stream, wherein the second data stream comprises a predetermined data structure including one or more descriptors, wherein the corresponding dictionary term is an obfuscated token representing the respective n-gram, wherein the obfuscated token comprises a part-of-speech identifier that is appended to the obfuscated token, and after each n-gram in the first data stream has been compared, the second data stream comprises only a plurality of obfuscated tokens;

apply a statistical process on the second data stream, wherein the statistical process includes at least determining one or more usage patterns of one or more of the obfuscated tokens in the second data stream;

generate one or more feature vectors based on the second data stream;

provide the one or more feature vectors to a machine learning model having been trained to perform a clustering analysis on provided feature vectors;

perform, using the machine learning model, the clustering analysis of the one or more feature vectors;

identify documents based on the results of the clustering analysis; and

predict, using the identified documents, one or more potentially preventable conditions.

8. The system of claim 7 , wherein the one or more processors are further configured to:

receive a request for data from a user,

wherein to receive the first data stream, the one or more processors are configured to retrieve the first data stream from the memory in response to receipt of the request for data received from the user.

9. The system of claim 7 , wherein the one or more processors are further configured to generate one or more n-gram-based statistics using the statistical process.

10. The system of claim 7 , wherein at least one of the first data stream or the second data stream comprises at least one of a document, a data set, a database record, or a medical record.

11. The system of claim 7 , wherein the one or more processors are further configured to predict a relationship between the one or more potentially preventable conditions and a combination of n-grams from the subset of n-grams disposed in the second data stream.

12. The system of claim 7 , wherein the one or more processors are further configured to:

encrypt and tokenize the first data stream using a random seed; and

provide the random seed to the user to reverse the obfuscation of one or more obfuscated tokens in the second data stream.

13. The system of claim 7 , wherein the one or more processors are further configured to apply one or more machine learning algorithms on the one or more determined usage patterns associated with the obfuscated tokens in the second data stream.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: 3M INNOVATIVE PROPERTIES COMPANY
To: SOLVENTUM INTELLECTUAL PROPERTIES COMPANY
Reel/Frame 066433/0105 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2016
From: STANKIEWICZ, BRIAN J; LOBNER, ERIC C; WOLNIEWICZ, RICHARD H.; SCHOFIELD, WILLIAM L.
To: 3M INNOVATIVE PROPERTIES COMPANY
Reel/Frame 038584/0131 →
Cited By (1)
US 12,244,577