IP Library › Granted Patent US 11,562,294
Granted Patent B2
US 11,562,294 · App. 16/670,185 · Granted Jan 24, 2023

Apparatus and method for analyzing time-series data based on machine learning

Inventors: Ji-Hyeon Seo (Seoul, KR); Jeong-Hyung Park (Seoul, KR); Wang-Geun Park (Seoul, KR)
Assignee: SAMSUNG SDS CO., LTD.
G06N20/00G06K9/6256G06K9/6267G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,294
App. No.
16/670,185
Granted
Jan 24, 2023
Kind
B2
Abstract

An apparatus and method for analyzing time series data on the basis of machine learning are provided. According to the disclosed embodiments, it is possible to effectively augment time series data, which is a target to be learned, according to characteristics of the time series data, thereby solving a problem of overfitting a machine learning model due to limited training data and a problem of deterioration of prediction accuracy due to imbalance of distribution of time series data and improving reliability of time series data analysis. In addition, according to the disclosed embodiments, it is possible to effectively set an optimal parameter for augmenting time series data.

Claims (22)

1. An apparatus for analyzing time series data, comprising:

a training data generation module configured to generate one or more pieces of a first training data by extracting a predetermined length of time period from raw time series data which comprises one or more columns and a plurality of observations measured for the one or more columns at consistent time intervals;

a first data augmentation module configured to, when a target value for predicting from the observations of the first training data satisfies a specific condition, generate a first augmented training data by further generating one or more pieces of a second training data which have the same target value as the target value but have a different length of time period to be extracted from the predetermined length of time, wherein the first augmented training data includes the first training data and the second training data;

a feature extractor configured to receive the first augmented training data and extract one or more feature values from the first augmented training data; and

a classifier configured to classify the first augmented training data on the basis of the one or more extracted feature values.

2. The apparatus of claim 1 , further comprising a second data augmentation module configured to augment one or more of the first augmented training data and the feature values according to a characteristic of the one or more columns.

3. The apparatus of claim 2 , wherein when the one or more columns comprise one or more numeric columns, the second data augmentation module is further configured to generate a second augmented training data by augmenting one or more pieces of training data by applying a predetermined data augmentation scheme to observations that correspond to the numeric columns of the first augmented training data, and the feature extractor is further configured to receive the second augmented training data and extract one or more feature values from the second augmented training data.

4. The apparatus of claim 2 , wherein when the one or more columns comprise one or more categorical columns, the second data augmentation module is further configured to generate a second augmented training data by augmenting one or more feature values by applying a predetermined data augmentation scheme to the one or more feature values extracted by the feature extractor and the classifier is further configured to classify the second augmented training data on the basis of the augmented feature values.

5. The apparatus of claim 2 , further comprising a policy recommendation module configured to determine an optimal data augmentation policy for augmenting one or more of the first augmented training data and the feature values.

6. The apparatus of claim 5 , wherein the policy recommendation module is further configured to perform learning by applying a plurality of different data augmentation policies to the first augmented training data and the feature values and determine the optimal data augmentation policy by comparing training results for each data augmentation policy.

7. The apparatus of claim 6 , wherein the policy recommendation module is further configured to terminate learning in accordance with a specific data augmentation policy and delete the specific data augmentation policy when a value of loss function increases during a learning process to which the specific data augmentation policy is applied or when a difference between a mean of values of the loss function during a specific past period and a current value of loss function is smaller than a set threshold.

8. A method of analyzing time series data, which is performed by a computing device comprising one or more processors and a memory in which one or more programs to be executed by the one or more processors are stored, the method comprising:

generating one or more pieces of a first training data by extracting a predetermined length of time period from raw time series data which comprises one or more columns and a plurality of observations measured for the one or more columns at consistent time intervals;

when a target value for predicting from the observations of the first training data satisfies a specific condition, generating a first augmented training data by further generating one or more pieces of a second training data which have the same target value as the target value but have a different length of time period to be extracted from the predetermined length of time, wherein the first augmented training data includes the first training data and the second training data;

receiving the first augmented training data and extracting one or more feature values from the first augmented training data; and

classifying the first augmented training data on the basis of the one or more extracted feature values.

9. The method of claim 8 , further comprising augmenting one or more of the first augmented training data and the feature values according to a characteristic of the columns.

10. The method of claim 9 , wherein when the columns comprise one or more numeric columns, the augmenting comprises generating a second augmented training data by augmenting the one or more pieces of the first training data by applying a predetermined data augmentation scheme to observations that correspond to the numeric columns of the first augmented training data, and the extracting of the one or more feature values comprises receiving the second augmented training data and extracting the one or more feature values from the second augmented training data.

11. The method of claim 9 , wherein when the columns comprise one or more categorical columns, the augmenting comprises generating a second augmented training data by augmenting the one or more feature values by applying a predetermined data augmentation scheme to the one or more extracted feature values and the classifying comprises classifying the second augmented training data on the basis of the augmented feature values.

12. The method of claim 9 , further comprising determining an optimal data augmentation policy for augmenting one or more of the first augmented training data and the feature values.

13. The method of claim 12 , wherein the determining of the optimal data augmentation policy comprises performing learning by applying a plurality of different data augmentation policies to the first augmented training data and the feature values and determining the optimal data augmentation policy by comparing training results for each data augmentation policy.

14. The method of claim 12 , wherein the determining of the optimal data augmentation policy comprises terminating learning in accordance with a specific data augmentation policy and deleting the specific data augmentation policy when a value of loss function increases during a learning process to which the specific data augmentation policy is applied or when a difference between a mean of values of the loss function during a specific past period and a current value of loss function is smaller than a set threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2019
From: SEO, JI-HYEON; PARK, JEONG-HYUNG; PARK, WANG-GEUN
To: SAMSUNG SDS CO., LTD.
Reel/Frame 050899/0357 →
Priority Claims (1)
KR 10-2019-0062881 · May 29, 2019 · national
Continuity (2)
Provisional Application 62853757 · May 29, 2019
Related Publication 20200380409A1 · Dec 3, 2020
Cited By (3)
US 12,326,918 US 12,327,193 US 12,640,271