IP Library › Granted Patent US 12,406,182
Granted Patent B2
US 12,406,182 · App. 17/338,707 · Granted Sep 2, 2025

Machine learning method and machine learning system involving data augmentation

Inventors: Chih-Yang Chen (Taoyuan, TW); Che-Han Chang (Taoyuan, TW); Edward Chang (Taoyuan, TW)
Assignee: HTC Corporation
G06N3/08G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,182
App. No.
17/338,707
Granted
Sep 2, 2025
Kind
B2
Abstract

A machine learning method includes steps of: (a) obtaining initial values of hyperparameters and hypernetwork parameters; (b) generating first classification model parameters according to the hyperparameters and the hypernetwork parameters, and updating the hypernetwork parameters according to a classification result based on the first classification model parameters relative to a training sample; (c) generating second classification model parameters according to the hyperparameters and the updated hypernetwork parameters, and updating the hyperparameters according to another classification result based on the second classification model parameters relative to a verification sample; and (d) repeating the steps (b) and (c) for updating the hypernetwork parameters and the hyperparameters.

Claims (58)

1. A machine learning method, comprising:

(a) obtaining initial values of a hyperparameter and a hypernetwork parameter;

(b) generating a first classification model parameter according to the hyperparameter and the hypernetwork parameter, and updating the hypernetwork parameter according to a classification result based on the first classification model parameter relative to a training sample, wherein the step (b) comprises:

(b1) performing data augmentation, by a data augmentation model based on the hyperparameter, on the training sample for generating an augmented training sample;

(b2) converting the hyperparameter, by a hypernetwork based on the hypernetwork parameter and a plurality of exploration values, into a plurality of exploration classification model parameters; and

(b3) forming a plurality of exploration classification models by the classification model based on the exploration classification model parameters respectively, and performing classification on the augmented training sample by the exploration classification models respectively for generating a plurality of first prediction labels corresponding to the augmented training sample;

(c) generating a second classification model parameter according to the hyperparameter and the hypernetwork parameter updated in the step (b), and updating the hyperparameter according to another classification result based on the second classification model parameter relative to a verification sample; and

(d) repeating the steps (b) and (c) for updating the hypernetwork parameter and the hyperparameter,

wherein each of the exploration classification models comprises a plurality of neural network structural layers, the neural network structural layers are divided into a first structural layer portion and a second structural layer portion after the first structural layer portion, each of the exploration classification model parameters for forming the exploration classification models comprises a first weight parameter content and a second weight parameter content, the first weight parameter content is configured to determine operations of the first structural layer portion, and the second weight parameter content is configured to determine operations of the second structural layer portion, each one of the first structure layer portions in front of the second structural layer portions within the exploration classification models has the first weight parameter content independent from the first weight parameter content of another exploration classification model within the exploration classification models, the second weight parameter contents applied to the second structural layer portion of a first exploration classification model of the exploration classification models are the same as the second weight parameter contents applied to the second structural layer portion of a second exploration classification model of the exploration classification models, the second structural layer portion of the first exploration classification model is operating with the same logic as the second structural layer portion of the second exploration classification model.

2. The machine learning method of claim 1 , wherein the step (b) further comprises:

(b4) updating the hypernetwork parameter according to a plurality of first losses generated by comparing the first prediction labels with a training label of the training sample.

3. The machine learning method of claim 2 , wherein

the step (b4) comprises:

calculating the first losses by comparing the first prediction labels with the training label of the training sample; and

updating the hypernetwork parameter according to the exploration classification models and the first losses corresponding to the exploration classification models.

4. The machine learning method of claim 3 , wherein in the step (b4):

calculating the first losses by a cross-entropy calculation between the first prediction labels of the exploration classification models and the training label respectively.

5. The machine learning method of claim 1 , wherein the first structural layer portion in each of the exploration classification models comprises at least one first convolutional layer, the first convolutional layers among the exploration classification models have different weight parameters from each other.

6. The machine learning method of claim 1 , wherein the second structural layer portion in each of the exploration classification models comprises at least one second convolutional layer and at least one fully connected layer, the at least one second convolutional layer and the at least one fully connected layer among the exploration classification models have same weight parameters across the exploration classification models.

7. The machine learning method of claim 1 , wherein the step (c) comprises:

(c1) converting the hyperparameter, by a hypernetwork based on the hypernetwork parameter updated in the step (b), into the second classification model parameter;

(c2) performing classification, by a classification model based on the second classification model parameter, on the verification sample for generating a second prediction label corresponding to the verification sample; and

(c3) updating the hyperparameter according to a second loss generated by comparing the second prediction label with a verification label of the verification sample.

8. The machine learning method of claim 7 , wherein the step (c3) comprises:

calculating the second loss by a cross-entropy calculation between the second prediction label and the verification label.

9. A machine learning system, comprising:

a memory unit, configured for storing initial values of a hyperparameter and a hypernetwork parameter;

a processing unit, coupled with the memory unit, wherein the processing unit is configured to run a hypernetwork, a data augmentation model and a classification model, the processing unit is configured to execute operations of:

(a) generating a first classification model parameter by the hypernetwork according to the hyperparameter and the hypernetwork parameter, generating a classification result by the classification model based on the first classification model parameter relative to a training sample, and updating the hypernetwork parameter according to the classification result, wherein the operation (a) comprises:

(a1) performing data augmentation, by the data augmentation model based on the hyperparameter, on the training sample for generating an augmented training sample;

(a2) converting the hyperparameter, by the hypernetwork based on the hypernetwork parameter and a plurality of exploration values, into a plurality of exploration classification model parameters; and

(a3) forming a plurality of exploration classification models by the classification model based on the exploration classification model parameters respectively, and performing classification on the augmented training sample by the exploration classification models respectively for generating a plurality of first prediction labels corresponding to the augmented training sample;

(b) generating a second classification model parameter by the hypernetwork according to the hyperparameter and the hypernetwork parameter updated in the step (a), generating another classification result by the classification model based on the second classification model parameter relative to a verification sample, and updating the hyperparameter according to the another classification result; and

(c) repeating the operations (a) and (b) for updating the hypernetwork parameter and the hyperparameter,

wherein each of the exploration classification models comprises a plurality of neural network structural layers, the neural network structural layers are divided into a first structural layer portion and a second structural layer portion after the first structural layer portion, each of the exploration classification model parameters for forming the exploration classification models comprises a first weight parameter content and a second weight parameter content, the first weight parameter content is configured to determine operations of the first structural layer portion, and the second weight parameter content is configured to determine operations of the second structural layer portion, each one of the first structure layer portions in front of the second structural layer portions within the exploration classification models has the first weight parameter content independent from the first weight parameter content of another exploration classification model within the exploration classification models, the second weight parameter contents applied to the second structural layer portion of a first exploration classification model of the exploration classification models are the same as the second weight parameter contents applied to the second structural layer portion of a second exploration classification model of the exploration classification models, the second structural layer portion of the first exploration classification model is operating with the same logic as the second structural layer portion of the second exploration classification model.

10. The machine learning system of claim 9 , wherein the operation (a) executed by the processing unit further comprises:

(a4) updating the hypernetwork parameter according to a plurality of first losses generated by comparing the first prediction labels with a training label of the training sample.

11. The machine learning system of claim 10 , wherein

the operation (a4) executed by the processing unit comprises:

calculating the first losses by comparing the first prediction labels with the training label of the training sample; and

updating the hypernetwork parameter according to the exploration classification models and the first losses corresponding to the exploration classification models.

12. The machine learning system of claim 11 , wherein the operation (a2) executed by the processing unit comprises:

calculating the first losses by a cross-entropy calculation between the first prediction labels of the exploration classification models and the training label respectively.

13. The machine learning system of claim 9 , wherein the first structural layer portion in each of the exploration classification models comprises at least one first convolutional layer, the first convolutional layers among the exploration classification models have different weight parameters from each other.

14. The machine learning system of claim 9 , wherein the second structural layer portion in each of the exploration classification models comprises at least one second convolutional layer and at least one fully connected layer, the at least one second convolutional layer and the at least one fully connected layer among the exploration classification models have same weight parameters across the exploration classification models.

15. The machine learning system of claim 9 , wherein the operation (b) executed by the processing unit comprises:

(b1) converting the hyperparameter, by the hypernetwork based on the hypernetwork parameter updated in the step (a), into the second classification model parameter;

(b2) performing classification, by the classification model based on the second classification model parameter, on the verification sample for generating a second prediction label corresponding to the verification sample; and

(b3) updating the hyperparameter according to a second loss generated by comparing the second prediction label with a verification label of the verification sample.

16. A non-transitory computer-readable storage medium, storing at least one instruction program executed by a processor to perform a machine learning method, the machine learning method comprising:

(a) obtaining initial values of a hyperparameter and a hypernetwork parameter;

(b) generating a first classification model parameter according to the hyperparameter and the hypernetwork parameter, and updating the hypernetwork parameter according to a classification result based on the first classification model parameter relative to a training sample, wherein the step (b) comprises:

(b1) performing data augmentation, by a data augmentation model based on the hyperparameter, on the training sample for generating an augmented training sample;

(b2) converting the hyperparameter, by a hypernetwork based on the hypernetwork parameter and a plurality of exploration values, into a plurality of exploration classification model parameters; and

(b3) forming a plurality of exploration classification models by the classification model based on the exploration classification model parameters respectively, and performing classification on the augmented training sample by the exploration classification models respectively for generating a plurality of first prediction labels corresponding to the augmented training sample;

(c) generating a second classification model parameter according to the hyperparameter and the hypernetwork parameter updated in the step (b), and updating the hyperparameter according to another classification result based on the second classification model parameter relative to a verification sample; and

(d) repeating the steps (b) and (c) for updating the hypernetwork parameter and the hyperparameter,

wherein each of the exploration classification models comprises a plurality of neural network structural layers, the neural network structural layers are divided into a first structural layer portion and a second structural layer portion after the first structural layer portion, each of the exploration classification model parameters for forming the exploration classification models comprises a first weight parameter content and a second weight parameter content, the first weight parameter content is configured to determine operations of the first structural layer portion, and the second weight parameter content is configured to determine operations of the second structural layer portion, each one of the first structure layer portions in front of the second structural layer portions within the exploration classification models has the first weight parameter content independent from the first weight parameter content of another exploration classification model within the exploration classification models, the second weight parameter contents applied to the second structural layer portion of a first exploration classification model of the exploration classification models are the same as the second weight parameter contents applied to the second structural layer portion of a second exploration classification model of the exploration classification models, the second structural layer portion of the first exploration classification model is operating with the same logic as the second structural layer portion of the second exploration classification model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2021
From: CHEN, CHIH-YANG; CHANG, CHE-HAN; CHANG, EDWARD
To: HTC CORPORATION
Reel/Frame 057469/0072 →
Continuity (2)
Provisional Application 63034993 · Jun 5, 2020
Related Publication 20210383224A1 · Dec 9, 2021
References Cited (53)
US 20190066313A1 · Kim · 2019 [cited by examiner]
US 20190251439A1 · Zoph et al. · 2019 [cited by applicant]
US 20200210773A1 · Li · 2020 [cited by examiner]
CN 110110861A · 2019 [cited by applicant]
JP 2020087103A · 2020 [cited by applicant]
KR 102336295B1 · 2016 [cited by examiner]
TW I675335B · 2019 [cited by applicant]
WO 2020070876A1 · 2020 [cited by applicant]
Wu, Jia, et al. “Hyperparameter optimization for machine learning models based on Bayesian optimization.” Journal of Electronic Science and Technology 17.1 (2019): 26-40. (Year: 2019). [cited by examiner]
Kim, Juyong, et al. “SplitNet: Learning to semantically split deep networks for parameter reduction and model parallelization.” International conference on machine learning. PMLR, 2017. (Year: 2017). [cited by examiner]
Ekin D. Cubuk et al., “AutoAugment: Learning Augmentation Strategies from Data”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. [cited by applicant]
Sungbin Lim et al., “Fast AutoAugment”, NIPS'19: Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS), 2019. [cited by applicant]
Qizhe Xie et al., “Unsupervised Data Augmentation for Consistency Training”, arXiv:1904.12848, 2019. [cited by applicant]
David Berthelot et al., “MixMatch: A Holistic Approach to Semi-Supervised Learning”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019. [cited by applicant]
Ting Chen et al., “A Simple Framework for Contrastive Learning of Visual Representations”, Proceedings of the 37th International Conference on Machine Learning, PMLR 119, 2020. [cited by applicant]
Ilya Kostrikov et al., “Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels”, arXiv:2004.13649, 2020. [cited by applicant]
Sergey Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, ICML'15: Proceedings of the 32nd International Conference on International Conference on Machine Learn… [cited by applicant]
Hieu Pham et al., “Efficient Neural Architecture Search via Parameter Sharing”, arXiv:1802.03268, 2018. [cited by applicant]
Andrew Brock et al., “SMASH: One-Shot Model Architecture Search through HyperNetworks”, ICLR 2018, 2018. [cited by applicant]
Gabriel Bender et al., “Understanding and Simplifying One-Shot Architecture Search”, Proceedings of the 35th International Conference on Machine Learning (ICML), PMLR 80, 2018. [cited by applicant]
Alex Krizhevsky et al., “Learning Multiple Layers of Features from Tiny Images”, Technical Report, 2009. [cited by applicant]
Sergey Zagoruyko et al., “Wide Residual Networks”, BMVC, 2016. [cited by applicant]
Yuval Netzer et al., “Reading Digits in Natural Images with Unsupervised Feature Learning”, NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011. [cited by applicant]
Jia Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database”, 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009. [cited by applicant]
Kaiming He et al., “Deep Residual Learning for Image Recognition”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. [cited by applicant]
Connor Shorten et al., “A survey on Image Data Augmentation for Deep Learning”, Journal of Big Data, 2019. [cited by applicant]
Alex Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks”, Advances in Neural Information Processing Systems, 2012. [cited by applicant]
Terrance Devries et al., “Improved Regularization of Convolutional Neural Networks with Cutout”, arXiv:1708.04552, 2017. [cited by applicant]
Hongyi Zhang et al., “mixup: Beyond Empirical Risk Minimization”, ICLR 2018. [cited by applicant]
Sangdoo Yun et al., “CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features”, 2019 EEE/CVF International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
Barret Zoph et al., “Neural Architecture Search With Reinforcement Learning”, ICLR 2017. [cited by applicant]
Hanxiao Liu et al., “Darts: Differentiable Architecture Search”, ICLR 2019. [cited by applicant]
Chen Lin et al., “Online Hyper-parameter Learning for Auto-Augmentation Strategy”, 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
Ekin D. Cubuk et al., “RandAugment: Practical automated data augmentation with a reduced search space”, arXiv:1909.13719, 2019. [cited by applicant]
Yonggang Li et al., “DADA: Differentiable Automatic Data Augmentation”, arXiv:2003.03780, 2020. [cited by applicant]
Tong Yu et al., “Hyper-Parameter Optimization: A Review of Algorithms and Applications”, arXiv:2003.05689, 2020. [cited by applicant]
Nitish Srivastava et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”, Journal of Machine Learning Research vol. 15, pp. 1929-1958, 2014. [cited by applicant]
Kinyu Zhang et al., “Adversarial Autoaugment”, ICLR 2020, 2020. [cited by applicant]
Diederik P. Kingma et al., “Adam: a Method for Stochastic Optimization”, ICLR 2015, 2015. [cited by applicant]
Adam Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019. [cited by applicant]
Xavier Gastaldi et al., “Shake-Shake regularization”, ICLR 2017. [cited by applicant]
Yoshihiro Yamada et al., “Shakedrop Regularization for Deep Residual Learning”, IEEE Access, 2019. [cited by applicant]
Felipe Petroski Such et al., “Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training Data”, arXiv:1912.07768, 2019. [cited by applicant]
David Ha et al.,“ HYPERNETWORKS”, arXiv:1609.09106 (ICLR 2017), 2016. [cited by applicant]
European Search Report of the European application No. 21177789.1 issued on Nov. 5, 2021. [cited by applicant]
Corresponding Taiwan office action issued on Aug. 23, 2022. [cited by applicant]
Jonathan Lorraine et al., “Stochastic Hyperparameter Optinization through Hypernetworks”, arXiv:1802.09419v2, Mar. 8, 2018. [cited by applicant]
The office action of the corresponding Korean application No. KR10-2021-0072866 issued on Jul. 15, 2024. [cited by applicant]
Daniel Ho et al., “Population Based Augmentation: Efficient Learning of Augmentation Policy Schedules”, Proceedings of the 36th International Conference on Machine Learning (ICML), 2019. [cited by applicant]
Max Jaderberg et al., “Population Based Training of Neural Networks”, arXiv:1711.09846v2, 2017. [cited by applicant]
Matthew Mackay et al., “Self-Tuning Networks:Bilevel Optimization of Hyperparameters Using Structured Best-Response Functions”, ICLR, 2019. [cited by applicant]
The office action of the corresponding Taiwanese application No. TW110120430 issued on May 11, 2023. [cited by applicant]
Corresponding Japan office action issued on Jun. 14, 2022. [cited by applicant]