IP Library Granted Patent US 10,503,709
Granted Patent B2
US 10,503,709 · App. 14/204,187 · Granted Dec 10, 2019

Data content identification

Inventors: Ben Lorenz (La Crosse, WI); Sophie Beutler (La Crosse, WI)
Assignee: SAP SE
G06F16/215
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,503,709
App. No.
14/204,187
Granted
Dec 10, 2019
Kind
B2
Abstract

The subject matter disclosed herein provides methods for identifying the type of content found in a database or source file having data records. A source file having one or more data records may be accessed. The data records may be associated with one or more data values arranged into columns. One or more data types may be proposed for at least one column by examining the data values in the column. A confidence score may be calculated for each proposed data type. The proposed data types may be arranged into a prioritized list based on each data type's confidence score. One or more rules may be applied to the column to finalize priorities of the proposed data types. The rules may be applied without referring to the data values in the column. Results may be provided based on the finalized priorities. Related apparatus, systems, techniques, and articles are also described.

Claims (40)

1. A method comprising:

accessing a source file having one or more data records, the one or more data records associated with one or more data values arranged into at least one column having an unknown data type, the unknown data type including a mislabeled data type; and

proposing one or more data types for the at least one column having the unknown data type by at least:

examining the one or more data values in the at least one column, the examining including examining, based on pattern matching, the one or more data values in the at least one column and examining, based on data directory matching, the one or more data values with one or more entries in at least one data directory, wherein the examining the one or more data values in the at least one column is based on at least a match between a format of the one or more data values with one or more patterns and/or a match between the one or more data values with one or more entries in the at least one data directory,

calculating, for the at least one column, one or more confidence scores as one or more percentages indicating, based on pattern matching and the data directory matching, how many of the one or more data values in the at least one column match and/or the at least one data directory match, wherein the one or more confidence scores are based on a first percentage of data values in the at least one column having a format that matches the one or more patterns and/or a second percentage of data values in the at least one column that match the one or more entries in the at least one data directory,

in response to the one or more confidence scores being over a threshold score, assigning the one or more proposed data types for the at least one column, wherein the one or more confidence scores prioritize the one or more proposed data types for the at least one column,

in response to the one or more confidence scores being below the threshold score, applying one or more context rules to the at least one column to determine, based on a neighboring data type contained in a neighboring column proximate to the at least one column, the one or more proposed data types, and

in response to the one or more confidence scores being below the threshold score, assigning, based on the applying of the context rules, the determined one or more proposed data types for the at least one column, wherein the one or more proposed data types have been adjusted, due to the applied one or more context rules, one or more confidence scores and corresponding priorities,

wherein the accessing and the proposing are performed by at least one processor.

2. The method of claim 1 , wherein the assigned data type is selected from the one or more proposed data types, and

wherein the assigning is performed by at least one processor.

3. The method of claim 1 , wherein the one or more context rules comprise a comparison of the one or more proposed data types for the at least one column with one or more data types of all other columns in the source file.

4. The method of claim 3 , wherein the one or more context rules comprise a first proximity rule that examines one or more name components associated with the neighboring column, and

wherein the one or more name components comprise a given name, a middle initial, or a family name.

5. The method of claim 1 , wherein the one or more context rules comprise a second proximity rule that examines one or more address components associated with the neighboring column, and

wherein the one or more address components comprise a street, a city, a state, a country, or a zip code.

6. The method of claim 1 , wherein the neighboring column is adjacent to the at least one column.

7. A non-transitory computer-readable medium containing instructions to configure at least one processor to perform operations comprising:

accessing a source file having one or more data records, the one or more data records associated with one or more data values arranged into at least one column having an unknown data type, the unknown data type including a mislabeled data type; and

proposing one or more data types for the at least one column having the unknown data type by at least:

examining the one or more data values in the at least one column, the examining including examining, based on pattern matching, the one or more data values in the at least one column and examining, based on data directory matching, the one or more data values with one or more entries in at least one data directory, wherein the examining the one or more data values in the at least one column is based on at least a match between a format of the one or more data values with one or more patterns and/or a match between the one or more data values with one or more entries in the at least one data directory,

calculating, for the at least one column, one or more confidence scores as one or more percentages indicating, based on pattern matching and the data directory matching, how many of the one or more data values in the at least one column match and/or the at least one data directory match, wherein the one or more confidence scores are based on a first percentage of data values in the at least one column having a format that matches the one or more patterns and/or a second percentage of data values in the at least one column that match the one or more entries in the at least one data directory,

in response to the one or more confidence scores being over a threshold score, assigning the one or more proposed data types for the at least one column, wherein the one or more confidence scores prioritize the one or more proposed data types for the at least one column,

in response to the one or more confidence scores being below the threshold score, applying one or more context rules to the at least one column to determine, based on a neighboring data type contained in a neighboring column proximate to the at least one column, the one or more proposed data types, and

in response to the one or more confidence scores being below the threshold score, assigning, based on the applying of the context rules, the determined one or more proposed data types for the at least one column, wherein the one or more proposed data types have been adjusted, due to the applied one or more context rules, one or more confidence scores and corresponding priorities,

wherein the accessing and the proposing are performed by at least one processor.

8. The non-transitory computer-readable medium of claim 7 , wherein the one or more context rules comprise a comparison of the one or more proposed data types for the at least one column with one or more data types of all other columns in the source file.

9. A system comprising:

at least one processor; and

at least one memory, wherein the at least one processor and the at least one memory are configured to perform operations comprising:

accessing a source file having one or more data records, the one or more data records associated with one or more data values arranged into at least one column having an unknown data type, the unknown data type including a mislabeled data type; and

proposing one or more data types for the at least one column having the unknown data type, the proposed one or more data types determined via a two-stage analysis comprising:

a first-stage analysis comprising:

examining the one or more data values in the at least one column, the examining including examining, based on pattern matching, the one or more data values in the at least one column and examining, based on data directory matching, the one or more data values with one or more entries in at least one data directory, wherein the examining the one or more data values in the at least one column is based on at least a match between a format of the one or more data values with one or more patterns and/or a match between the one or more data values with one or more entries in the at least one data directory,

calculating, for the at least one column, one or more confidence scores as one or more percentages indicating, based on pattern matching and the data directory matching, how many of the one or more data values in the at least one column match and/or the at least one data directory match, wherein the one or more confidence scores are based on a first percentage of data values in the at least one column having a format that matches the one or more patterns and/or a second percentage of data values in the at least one column that match the one or more entries in the at least one data directory, and

in response to the one or more confidence scores being over a threshold score, assigning the one or more proposed data types for the at least one column, wherein the one or more confidence scores prioritize the one or more proposed data types for the at least one column, and

a second-stage comprising:

in response to the one or more confidence scores being below the threshold score, applying one or more context rules to the at least one column to determine, based on a neighboring data type contained in a neighboring column proximate to the at least one column, the one or more proposed data types, and

in response to the one or more confidence scores being below the threshold score, assigning, based on the applying of the context rules, the determined one or more proposed data types for the at least one column, wherein the one or more proposed data types have been adjusted, due to the applied one or more context rules, one or more confidence scores and corresponding priorities.

10. The system of claim 9 , wherein the one or more context rules comprise a comparison of the one or more proposed data types for the at least one column with one or more data types of all other columns in the source file.

Assignments (2)
CHANGE OF NAME Recorded Aug 26, 2014
From: SAP AG
To: SAP SE
Reel/Frame 033625/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2014
From: LORENZ, BEN; BEUTLER, SOPHIE
To: SAP AG
Reel/Frame 032407/0463 →