IP Library Granted Patent US 10,467,244
Granted Patent B2
US 10,467,244 · App. 15/583,966 · Granted Nov 5, 2019

Automatic generation of structured data from semi-structured data

Inventors: Ravikiran Krishnan (San Mateo, CA); Ayush Parashar (Foster City, CA); Sudeep Sarkar (Tampa, FL)
Assignee: UNIFI SOFTWARE, INC.
G06F16/258G06F16/211G06F16/22G06F16/86
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,467,244
App. No.
15/583,966
Granted
Nov 5, 2019
Kind
B2
Abstract

A method and system for generating structured data from semi-structured data are provided. The method includes reading a plurality of records from a data file including semi-structured data. Further, the method includes obtaining aligned delimiters in a list for every record that has been read. The method also includes selecting a most occurring delimiter from the list. The method then includes constructing a regular expression using the selected delimiter to split the records into different fields. The method also includes reconstructing the records for the regular expression to fit and split into fields. In addition, the method includes displaying the records split into the fields.

Claims (95)

1. A method for generating structured data from semi-structured data, the method comprising:

reading a plurality of records from a data file comprising semi-structured data;

obtaining aligned delimiters in a list for every record that has been read;

selecting a most occurring delimiter from the list;

constructing a regular expression using the selected delimiter to split the records into different fields;

reconstructing the records for the regular expression to fit and split into the different fields; and

displaying the records split into the different fields.

2. The method as claimed in claim 1 , wherein the obtaining comprises:

aligning read records using dynamic time warping.

3. The method as claimed in claim 2 , wherein the obtaining further comprises:

identifying matching non-alpha characters as initial delimiters list D1.

4. The method as claimed in claim 3 , wherein the obtaining further comprises iteratively performing:

picking a test record (test);

aligning the test record with aligned record (A);

obtaining common delimiters list D;

aligning the common delimiters list D with initial delimiters list D1 using dynamic time warping;

obtaining aligned records delimiters list D2;

adding D2 to a delimiter set S; and

repeating above steps till N number of test records are processed.

5. The method as claimed in claim 1 , wherein the semi-structured data comprises at least one of:

end of line delimiter as a new line character;

column delimiter as alpha numeric;

column delimiter as multi-byte or multi-character;

different column delimiter for every column; and

any number of column delimiters in every row.

6. The method as claimed in claim 1 and further comprising:

identifying missing delimiters and missing values.

7. The method as claimed in claim 6 , wherein the reconstructing comprises:

reconstructing the records for the regular expression to fit and subsequently filling in NULL for the missing values.

8. A method for generating structured data from semi-structured data, the method comprising:

reading a plurality of records from a data file comprising semi-structured data;

obtaining aligned delimiters in a list for every record that has been read;

selecting a most occurring delimiter from the list;

constructing a regular expression using the selected delimiter to split the records into different fields;

identifying missing delimiters and missing values;

reconstructing the records for the regular expression to fit and split into the different fields, and subsequently filling in NULL for the missing values; and

displaying the records in a split tabulated form.

9. The method as claimed in claim 8 , wherein the obtaining comprises:

aligning read records using dynamic time warping.

10. The method as claimed in claim 9 , wherein the obtaining further comprises:

identifying matching non-alpha characters as initial delimiters list D1.

11. The method as claimed in claim 10 , wherein the obtaining further comprises iteratively performing:

picking a test record (test);

aligning the test record with aligned record (A);

obtaining common delimiters list D;

aligning the common delimiters list D with initial delimiters list D1 using dynamic time warping;

obtaining aligned records delimiters list D2;

adding D2 to a delimiter set S; and

repeating above steps till N number of test records are processed.

12. The method as claimed in claim 8 , wherein the semi-structured data comprises at least one of:

end of line delimiter as a new line character;

column delimiter as alphanumeric;

column delimiter as multi-byte or multi-character;

different column delimiter for every column; and

any number of column delimiters in every row.

13. A system comprising:

a processor; and

a memory coupled to the processor, the memory storing instructions which when executed by the processor cause the system to perform a method for providing information to at user, the method comprising

reading a plurality of records from a data file comprising semi-structured data;

obtaining aligned delimiters in a list for every record that has been read;

selecting a most occurring delimiter from the list;

constructing a regular expression using the selected delimiter to split the records into different fields;

reconstructing the records for the regular expression to fit and split into the different fields; and

displaying the records split into the different fields.

14. The system as claimed in claim 13 , wherein the obtaining comprises:

aligning read records using dynamic time warping.

15. The system as claimed in claim 14 , wherein the obtaining further comprises:

identifying matching non-alpha characters as initial delimiters list D1.

16. The system as claimed in claim 15 , wherein the obtaining further comprises iteratively performing:

picking a test record (test);

aligning the test record with aligned record (A);

obtaining common delimiters list D;

aligning the common delimiters list D with initial delimiters list D1 using dynamic time warping;

obtaining aligned records delimiters list D2;

adding D2 to a delimiter set S; and

repeating above steps till N number of test records are processed.

17. The system as claimed in claim 13 , wherein the semi-structured data comprises at least one of:

end of line delimiter as a new line character;

column delimiter as alpha numeric;

column delimiter as multi-byte or multi-character;

different column delimiter for every column; and

any number of column delimiters in every row.

18. The system as claimed in claim 13 and further comprising:

identifying missing delimiters and missing values.

19. The system as claimed in claim 18 , wherein the reconstructing comprises:

reconstructing the records for the regular expression to fit and subsequently filling in NULL for the missing values.

20. The system as claimed in claim 13 , wherein the reconstructing comprises: applying symmetric delimiters to split the records into the different fields.

21. A method for generating structured data from semi-structured data, the method comprising:

reading a plurality of records from a data file comprising semi-structured data;

aligning read records using dynamic time warping;

obtaining aligned delimiters in a list for every record that has been read;

selecting a most occurring delimiter from the list;

constructing a regular expression using the selected delimiter to split the records into different fields;

reconstructing the records for the regular expression to fit and split into the different fields; and

displaying the records split into the different fields.

Assignments (7)
SECURITY INTEREST Recorded Nov 12, 2024
From: SIXTH STREET SPECIALTY LENDING, INC.
To: BLUE OWL CAPITAL CORPORATION
Reel/Frame 069342/0406 →
SECURITY INTEREST Recorded Oct 1, 2021
From: BOOMI, LLC
To: SIXTH STREET SPECIALTY LENDING, INC., AS COLLATERAL AGENT
Reel/Frame 057679/0908 →
MERGER Recorded Aug 10, 2020
From: UNIFI SOFTWARE, INC.
To: BOOMI, INC.
Reel/Frame 053444/0292 →
RELEASE OF SECURITY INTEREST Recorded Dec 20, 2019
From: PACIFIC WESTERN BANK
To: UNIFI SOFTWARE, INC.
Reel/Frame 051347/0977 →
SECURITY INTEREST Recorded Oct 24, 2019
From: UNIFI SOFTWARE, INC.
To: PACIFIC WESTERN BANK
Reel/Frame 050812/0420 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2017
From: SARKAR, SUDEEP
To: UNIVERSITY OF SOUTH FLORIDA
Reel/Frame 042417/0800 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2017
From: KRISHNAN, RAVIKIRAN; PARASHAR, AYUSH
To: UNIFI SOFTWARE, INC.
Reel/Frame 042417/0813 →
Continuity (2)
Provisional Application 62329982 · Apr 29, 2016
Related Publication 20170316070A1 · Nov 2, 2017