IP Library Granted Patent US 11,704,331
Granted Patent B2
US 11,704,331 · App. 16/926,537 · Granted Jul 18, 2023

Dynamic generation of data catalogs for accessing data

Inventors: Andrew Edward Caldwell (Santa Clara, CA); Anurag Windlass Gupta (Atherton, CA); Mehul A. Shah (Seratoga, CA); Prajakta Datta Damle (San Jose, CA); George Steven McPherson (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G06F16/254G06F16/2358G06F16/283G06F16/951
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,704,331
App. No.
16/926,537
Granted
Jul 18, 2023
Kind
B2
Abstract

Dynamic generation of data catalogs may be implemented for accessing data sets in different storage locations. Data sets may be accessed in order to extract portions of data. Structure recognition techniques may be applied to the extracted data in order to determine structural information for the data sets. The structural information may then be stored as part of a data catalog for the data sets. Requests to access the data catalog from different clients may be received and the requested structural data supplied so that the clients may access different data sets utilizing the supplied structural data. Data catalogs may be updated as changes to data sets are made.

Claims (40)

1. A system, comprising:

a plurality of computing devices, respectively comprising at least one processor and a memory, configured to implement a data catalog service as part of a provider network, wherein the data catalog service is configured to:

identify a plurality of data sets maintained in different storage locations;

receive respective selections of one or more recognizers of a plurality of different recognizers corresponding to a plurality of different file formats supported by the data catalog service to apply to data scanned from the plurality of data sets;

add respective structural data for the plurality of data sets to a data catalog that provides a centralized location for searching for desired data sets, wherein the respective structural data makes a client capable of interpreting between different items within the plurality of data sets when the client connects to data sources for individual ones of the plurality of data sets, and wherein the adding comprises:

access the different storage locations to apply the selected one or more recognizers to the data scanned from the plurality of data sets to determine the respective structural data; and

store the respective structural data as part of the data catalog; and

provide access to the respective structural data in the data catalog in response to one or more requests to access the data catalog.

2. The system of claim 1 , wherein access to the respective structural data in the data catalog is provided according to one or more access policies.

3. The system of claim 1 , wherein the selected one or more recognizers recognize respective delimiters between items in the plurality of data sets.

4. The system of claim 1 , wherein the addition further comprises the inclusion of respective data lineage for respective changes to the plurality of data sets.

5. The system of claim 1 , wherein the addition further comprises obtain respective access credentials to access the plurality of data sets.

6. The system of claim 1 , wherein the respective structural data comprises columns of the plurality of data sets and respective data types for the columns.

7. The system of claim 1 , wherein the access to the different storage locations to apply the selected one or more recognizers to the data scanned from the plurality of data sets to determine the respective structural data is performed according to a schedule.

8. A method, comprising:

performing by one or more computing devices comprising one or more respective processors and memory and implementing a data catalog service of a provider network:

identifying, by the data catalog service of a provider network, a plurality of data sets maintained in different storage locations;

receiving, via an interface of the data catalog service, respective selections of one or more recognizers of a plurality of different recognizers corresponding to a plurality of different file formats supported by the data catalog service to apply to data scanned from the plurality of data sets;

adding, by the data catalog service, respective structural data for the plurality of data sets to a data catalog that provides a centralized location for searching for desired data sets, wherein the respective structural data makes a client capable of interpreting between different items within the plurality of data sets when the client connects to data sources for individual ones of the plurality of data sets, and wherein the adding comprises:

accessing the different storage locations to apply the selected one or more recognizers to the data scanned from the plurality of data sets to determine the respective structural data; and

storing the respective structural data as part of the data catalog; and

providing, by the data catalog service, access to the respective structural data in the data catalog in response to one or more requests to access the data catalog.

9. The method of claim 8 , wherein access to the respective structural data in the data catalog is provided according to one or more access policies.

10. The method of claim 8 , wherein the selected one or more recognizers recognize respective delimiters between items in the plurality of data sets.

11. The method of claim 8 , wherein the adding further comprises including respective data lineage for respective changes to the plurality of data sets.

12. The method of claim 8 , wherein the adding further comprises obtain respective access credentials to access the plurality of data sets.

13. The method of claim 8 , wherein the respective structural data comprises columns of the plurality of data sets and respective data types for the columns.

14. The method of claim 8 , wherein the access to the different storage locations to apply the selected one or more recognizers to the data scanned from the plurality of data sets to determine the respective structural data is performed according to a schedule.

15. One or more non-transitory, computer-readable storage media, storing program instructions that when executed on or across one or more computing devices, cause the one or more computing devices to implement:

identifying, by a data catalog service of a provider network, a plurality of data sets maintained in different storage locations;

receiving, via an interface of the data catalog service, respective selections of one or more recognizers of a plurality of different recognizers corresponding to a plurality of different file formats supported by the data catalog service to apply to data scanned from the plurality of data sets;

adding, by the data catalog service, respective structural data for the plurality of data sets to a data catalog that provides a centralized location for searching for desired data sets, wherein the respective structural data makes a client capable of interpreting between different items within the plurality of data sets when the client connects to data sources for individual ones of the plurality of data sets, and wherein the adding comprises:

accessing the different storage locations to apply the selected one or more recognizers to the data scanned from the plurality of data sets to determine the respective structural data; and

storing the respective structural data as part of the data catalog; and

providing, by the data catalog service, access to the respective structural data in the data catalog in response to one or more requests to access the data catalog.

16. The or more non-transitory, computer-readable storage media of claim 15 , wherein access to the respective structural data in the data catalog is provided according to one or more access policies.

17. The or more non-transitory, computer-readable storage media of claim 15 , wherein the selected one or more recognizers recognize respective delimiters between items in the plurality of data sets.

18. The or more non-transitory, computer-readable storage media of claim 15 , wherein the adding further comprises including respective data lineage for respective changes to the plurality of data sets.

19. The or more non-transitory, computer-readable storage media of claim 15 , wherein the adding further comprises obtain respective access credentials to access the plurality of data sets.

20. The or more non-transitory, computer-readable storage media of claim 15 , wherein the respective structural data comprises columns of the plurality of data sets and respective data types for the columns.

Continuity (2)
Continuation 15199505 · Jun 30, 2016
Related Publication 20200409967A1 · Dec 31, 2020