IP Library Granted Patent US 11,694,112
Granted Patent B2
US 11,694,112 · App. 16/808,598 · Granted Jul 4, 2023

Machine learning data management system and data management method

Inventor: Yuichi Taguchi (Tokyo, JP)
Assignee: HITACHI, LTD.
G06N20/00G06F16/219G06F16/25G06F16/254G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,694,112
App. No.
16/808,598
Granted
Jul 4, 2023
Kind
B2
Abstract

The data analysis server manages first data configuration information for managing a correspondence of the learned data model, a raw data bucket that stores the raw data for generating the learned data model, and a curated data bucket that stores the curated data for generating the learned data model and second data configuration information for managing correspondences of the learned data model and the first storage area, the curated data bucket and the second storage area, and the raw data bucket and the third storage area, and gives an instruction to acquire snapshots of the second storage area that stores the curated data for generating the learned data model and the third storage area that stores the raw data for generating the learned data model to the data storage via the storage management server when the learned data model is generated.

Claims (57)

1. A data management system, comprising:

a data analysis server that generates curated data from raw data of various kinds of data, learns the curated data, and generates a learned data model;

a data storage including a first storage area that stores the learned data model, a second storage area that stores the curated data, and a third storage area that stores the raw data; and

a storage management server that manages the data storage,

wherein the data analysis server:

manages first data configuration information for managing a correspondence of the learned data model, a raw data bucket that stores the raw data for generating the learned data model, and a curated data bucket that stores the curated data for generating the learned data model;

manages second data configuration information for managing correspondences of the learned data model and the first storage area, the curated data bucket and the second storage area, and the raw data bucket and the third storage area;

sends an instruction to acquire snapshots of the second storage area that stores the curated data for generating the learned data model and the third storage area that stores the raw data for generating the learned data model to the data storage via the storage management server when the learned data model is generated;

receives storage configuration management information and snapshot management information from the storage management server, and generates third data configuration information for managing snapshot acquisition times of the first storage area, the second storage area, and the third storage area; and

a data management relation is generated on the basis of the first data configuration information, the second data configuration information, and the third data configuration information.

2. The data management system according to claim 1 , wherein

the raw bucket is a data lake,

the curated data bucket is a learning data set, and

the learned data model is a learning model.

3. The data management system according to claim 1 , wherein the data configuration relation generated by the data analysis server indicates:

a relation of the raw data bucket, the curated data bucket, and the learned data model,

relations of the raw data bucket and the third storage area, the curated data bucket and the second storage area, and the learned data model and the third storage area, and

a relation with at least one snapshot of the first storage area, the second storage area, and the third storage area.

4. The data management system according to claim 2 , wherein

the data analysis server receives storage configuration management information and snapshot management information from the storage management server, and generates third data configuration information for managing snapshot acquisition times of the first storage area, the second storage area, and the third storage area, and

at least two of snapshots of the learned data model, the curated data bucket, and the raw data bucket acquired at the same time are managed as a consistent snapshot group on the basis of the second data configuration information and the third data configuration information.

5. The data management system according to claim 2 , wherein

the data analysis server includes a processing unit and a display device,

the processing unit displays a plurality of job objects in a node list part of the display device, and

workflow definition information of a data preparation process is generated by selecting the job object from the node list part and dropping the job object onto a workflow display area.

6. The data management system according to claim 5 , wherein the plurality of job objects displayed in the node list part include:

a load data object for reading data from the raw data bucket,

a data curation object for curating, for data of an associated job object, the data,

a storage job object of storing, for data of an associated job object, the data, and

a snapshot acquisition job object for acquiring, for data of an associated job object, a snapshot of the data.

7. A data management system, comprising:

a data analysis server that generates curated data from raw data of various kinds of data, learns the curated data, and generates a learned data model;

a data storage including a first storage area that stores the learned data model, a second storage area that stores the curated data, and a third storage area that stores the raw data; and

a storage management server that manages the data storage,

wherein the data analysis server:

manages first data configuration information for managing a correspondence of the learned data model, a raw data bucket that stores the raw data for generating the learned data model, and a curated data bucket that stores the curated data for generating the learned data model;

manages second data configuration information for managing correspondences of the learned data model and the first storage area, the curated data bucket and the second storage area, and the raw data bucket and the third storage area;

sends an instruction to acquire snapshots of the second storage area that stores the curated data for generating the learned data model and the third storage area that stores the raw data for generating the learned data model to the data storage via the storage management server when the learned data model is generated;

receives storage configuration management information and snapshot management information from the storage management server, and generates third data configuration information for managing snapshot acquisition times of the first storage area, the second storage area, and the third storage area, and

generates a data chronological relation on the basis of the first data configuration information, the second data configuration information, and the third data configuration information.

8. The data management system according to claim 7 , wherein the data chronological relation generated by the data analysis server indicates:

a relation of the raw data bucket, the curated data bucket, and the learned data model,

relations of the raw data bucket and the third storage area, the curated data bucket and the second storage area, and the learned data model and the third storage area, and

a relation with snapshot acquisition times of the first storage area, the second storage area, and the third storage area.

9. A data management method in a data management system including a data analysis server that generates curated data from raw data of various kinds of data, learns the curated data, and generates a learned data model, a data storage including a first storage area that stores the learned data model, a second storage area that stores the curated data, and a third storage area that stores the raw data, and a storage management server that manages the data storage, the data management method comprising:

managing, by the data analysis server, a correspondence of the learned data model, a raw data bucket that stores the raw data for generating the learned data model, and a curated data bucket that stores the curated data for generating the learned data model as first data configuration information;

managing, by the data analysis server, second data configuration information for managing correspondences of the learned data model and the first storage area, the curated data bucket and the second storage area, and the raw data bucket and the third storage area;

sending, by the data analysis server, an instruction to acquire snapshots of the second storage area that stores the curated data for generating the learned data model and the third storage area that stores the raw data for generating the learned data model to the data storage via the storage management server when the learned data model is generated;

receiving, by the data analysis server, storage configuration management information and snapshot management information from the storage management server, and generating third data configuration information for managing snapshot acquisition times of the first storage area, the second storage area, and the third storage area; and

managing at least two of snapshots of the learned data model, the curated data bucket, and the raw data bucket acquired at the same time as a consistent snapshot group on the basis of the second data configuration information and the third data configuration information.

10. The data management method according to claim 9 , wherein

the data analysis server notifies, if an instruction to generate the learned data model is received, the storage management server of a snapshot acquisition instruction of the third storage area corresponding to the referred raw data bucket with reference to the raw data bucket,

the storage management server instructs the data storage to acquire the snapshot of the third storage area,

the data storage acquires the snapshot of the third storage area and gives a notification to the storage management server,

the data analysis server executes curation of the referred raw data bucket if the acquisition notification of the snapshot of the referred third storage area is received from the storage management server,

data obtained by executing the curation is stored in the curated data bucket, and an instruction to acquire a snapshot of the second storage area corresponding to the curated data bucket is given to the storage management server, and

the data storage acquires the snapshot of the second storage area in accordance with the instruction from the storage management server and gives a notification to the storage management server.

Assignments (2)
COMPANY SPLIT Recorded Aug 20, 2024
From: HITACHI, LTD.
To: HITACHI VANTARA, LTD.
Reel/Frame 069518/0761 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2020
From: TAGUCHI, YUICHI
To: HITACHI, LTD.
Reel/Frame 052008/0131 →
Priority Claims (1)
JP 2019-192920 · Oct 23, 2019 · national
Continuity (1)
Related Publication 20210125099A1 · Apr 29, 2021