IP Library Granted Patent US 11,238,035
Granted Patent B2
US 11,238,035 · App. 16/814,855 · Granted Feb 1, 2022

Personal information indexing for columnar data storage format

Inventors: Hamed Ahmadi (Coquitlam, CA); Jian Wen (Hollis, NH); Shrikumar Hariharasubrahmanian (Palo Alto, CA); Sanjay Jinturkar (Santa Clara, CA); Nipun Agarwal (Saratoga, CA)
Assignee: ORACLE INTERNATIONAL CORPORATION
G06F16/245G06F16/221G06F16/2228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,238,035
App. No.
16/814,855
Granted
Feb 1, 2022
Kind
B2
Abstract

Techniques are described herein for indexing personal information in columnar data storage format based files. In an embodiment, row groups of rows that comprise a plurality of columns are stored in a set of files. Each column of a row group is stored in a chunk of column pages in the set of files. A regular expression index that indexes a particular column in the set of files is stored for each row group. The regular expression index identifies column pages in the chunk of the particular column that include a particular column value that satisfies a regular expression specified in a query. The regular expression specified in the query in evaluated against the particular column using the regular expression index.

Claims (44)

1. A computer-implemented method comprising:

storing in a set of files a plurality row groups of rows that comprise a plurality of columns, wherein for each row group, each column of said plurality of columns is stored in a respective chunk of column pages in the set of files;

for each row group of said plurality of row groups: generating a respective regular expression index that indexes a particular column, determining whether the size of the respective regular expression index is below a threshold, and storing in the set of files said respective regular expression index based on the determination, wherein said regular expression index identifies one or more column pages in the respective chunk of said particular column and said each row group that include a particular column value that satisfies a regular expression.

2. The method of claim 1 , further comprising:

receiving a request to evaluate a particular regular expression against the particular column.

3. The method of claim 2 , further comprising:

for each row group of said plurality of row groups, evaluating the particular regular expression against said particular column using the respective regular expression index that indexes the particular column.

4. The method of claim 3 , wherein evaluating the regular expression against the particular column includes:

for each row group of said plurality of row groups:

determining a respective set of column pages that the regular expression index identifies as satisfying said regular expression; and

evaluating the regular expression against the respective set of column pages.

5. The method of claim 2 , further comprising:

determining that a regular expression index exists for the particular column and in response, evaluating the particular regular expression against the particular column.

6. The method of claim 2 , further comprising:

determining that a regular expression index does not exist for the particular column and in response, skipping evaluating the particular regular expression against the particular column.

7. The method of claim 2 , wherein each data node of a plurality of data nodes stores a row group of the plurality of row groups.

8. The method of claim 7 , further comprising:

each data node of the plurality of data nodes determining, for the respective stored row group of the plurality of row groups, a respective set of column pages that the regular expression index identifies as satisfying the particular regular expression;

each data node of the plurality of data nodes transmitting the respective set of column pages to a compute node for evaluation.

9. The method of claim 1 , further comprising:

for each row group of the plurality of row groups, scanning metadata associated with the respective row group to determine whether a regular expression index exists for the particular column.

10. The method of claim 1 , wherein each file of the set of files is an Apache Parquet file.

11. One or more non-transitory computer-readable media storing instructions which, when executed by one or more processors, cause:

storing in a set of files a plurality row groups of rows that comprise a plurality of columns, wherein for each row group, each column of said plurality of columns is stored in a respective chunk of column pages in the set of files;

for each row group of said plurality of row groups: generating a respective regular expression index that indexes a particular column, determining whether the size of the respective regular expression index is below a threshold, and storing in the set of files said respective regular expression index based on the determination, wherein said regular expression index identifies one or more column pages in the respective chunk of said particular column and said each row group that include a particular column value that satisfies a regular expression.

12. The one or more non-transitory computer-readable media of claim 11 , further comprising instructions which, when executed by the one or more processors, cause:

receiving a request to evaluate a particular regular expression against the particular column.

13. The one or more non-transitory computer-readable media of claim 12 , further comprising instructions which, when executed by the one or more processors, cause:

for each row group of said plurality of row groups, evaluating the particular regular expression against said particular column using the respective regular expression index that indexes the particular column.

14. The one or more non-transitory computer-readable media of claim 13 , wherein evaluating the regular expression against the particular column includes:

for each row group of said plurality of row groups:

determining a respective set of column pages that the regular expression index identifies as satisfying said regular expression; and

evaluating the regular expression against the respective set of column pages.

15. The one or more non-transitory computer-readable media of claim 12 , further comprising instructions which, when executed by the one or more processors, cause:

determining that a regular expression index exists for the particular column and in response, evaluating the particular regular expression against the particular column.

16. The one or more non-transitory computer-readable media of claim 12 , further comprising instructions which, when executed by the one or more processors, cause:

determining that a regular expression index does not exist for the particular column and in response, skipping evaluating the particular regular expression against the particular column.

17. The one or more non-transitory computer-readable media of claim 12 , wherein each data node of a plurality of data nodes stores a row group of the plurality of row groups.

18. The one or more non-transitory computer-readable media of claim 17 , further comprising instructions which, when executed by the one or more processors, cause:

each data node of the plurality of data nodes determining, for the respective stored row group of the plurality of row groups, a respective set of column pages that the regular expression index identifies as satisfying the particular regular expression;

each data node of the plurality of data nodes transmitting the respective set of column pages to a compute node for evaluation.

19. The one or more non-transitory computer-readable media of claim 11 , further comprising instructions which, when executed by the one or more processors, cause:

for each row group of the plurality of row groups, scanning metadata associated with the respective row group to determine whether a regular expression index exists for the particular column.

20. The one or more non-transitory computer-readable media of claim 11 , wherein each file of the set of files is an Apache Parquet file.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2020
From: AHMADI, HAMED; WEN, JIAN; HARIHARASUBRAHMANIAN, SHRIKUMAR; JINTURKAR, SANJAY; AGARWAL, NIPUN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 052074/0164 →
Continuity (1)
Related Publication 20210286806A1 · Sep 16, 2021