IP Library Granted Patent US 12,646,503
Granted Patent B2
US 12,646,503 · App. 18/766,302 · Granted Jun 2, 2026

Method, device and non-transitory computer readable medium for training machine learning models

Inventors: Chi-Chun Lee (Hsinchu, TW); Ya-Tse Wu (Hsinchu, TW)
Assignee: National Tsing Hua University
G10L15/063G10L21/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,503
App. No.
18/766,302
Granted
Jun 2, 2026
Kind
B2
Abstract

A method for training a machine learning model is provided. In this method, samples are drawn from a first training set according to multiple weights for training the machine learning model in a first epoch. The first training set includes multiple levels corresponding to the multiple weights. After the first epoch, the machine learning model is evaluated using a validation set for multiple first performances at the multiple levels. Additionally, the weight that corresponds to each level is updated based on the first performances, and samples are redrawn from the first training set according to the updated weights for training the machine learning model in a second epoch that follows the first epoch. Moreover, an electronic device and a non-transitory computer-readable medium for utilizing the above method are also provided.

Claims (51)

1 . A computer-implemented method for training a machine learning model comprising:

sampling from a first training set, based on a plurality of weights, to train the machine learning model in a first epoch, the first training set comprising a plurality of levels corresponding to the plurality of weights;

evaluating the machine learning model using a validation set after the first epoch to obtain a plurality of first performances at the plurality of levels;

updating the plurality of weights corresponding to the plurality of levels based on the plurality of first performances to obtain a plurality of updated weights; and

resampling from the first training set, based on the plurality of updated weights, to train the machine learning model in a second epoch, wherein the second epoch follows the first epoch.

2 . The computer-implemented method of claim 1 , further comprising:

evaluating the machine learning model using the validation set after the second epoch to obtain a plurality of second performances at the plurality of levels; and

updating the plurality of updated weights corresponding to the plurality of levels based on the plurality of second performances.

3 . The computer-implemented method of claim 1 , wherein the plurality of updated weights corresponding to the plurality of levels is negatively correlated with the plurality of first performances of the machine learning model at the plurality of levels.

4 . The computer-implemented method of claim 3 , wherein each of the plurality of updated weights is equal to or more than a predetermined minimum value.

5 . The computer-implemented method of claim 1 , further comprising:

mixing noise data into a noise-free training set to generate the first training set.

6 . The computer-implemented method of claim 5 , wherein the sampling from the first training set, based on the plurality of weights, to train the machine learning model in the first epoch comprises:

sampling from the plurality of levels of the first training set based on the plurality of weights to obtain a first sample training set;

merging the first sample training set with the noise-free training set to generate a second training set; and

using the second training set to train the machine learning model in the first epoch.

7 . The computer-implemented method of claim 1 , wherein the first training set comprises an ordered data set.

8 . The computer-implemented method of claim 1 , further comprising:

dividing the first training set into the plurality of levels based on a distortion index.

9 . The computer-implemented method of claim 8 , wherein the machine learning model comprises a speech recognition model.

10 . The computer-implemented method of claim 9 , wherein the dividing of the first training set into the plurality of levels, based on the distortion index, comprises:

dividing the first training set into the plurality of levels based on at least one of a Perceptual Evaluation of Speech Quality (PESQ), a Short-Time Objective Intelligibility (STOI), and a Frequency-Weighted Signal-to-Noise Ratio Segmental (fwSNRseq).

11 . An electronic device comprising:

one or more memories storing at least one instruction; and

one or more processors coupled to the one or more memories, wherein the at least one instruction, when executed by the one or more processors, causes the electronic device to:

sample from a first training set, based on a plurality of weights, to train a machine learning model in a first epoch, the first training set comprising a plurality of levels corresponding to the plurality of weights;

evaluate the machine learning model using a validation set after the first epoch to obtain a plurality of first performances at the plurality of levels;

update the plurality of weights corresponding to the plurality of levels based on the plurality of first performances to obtain a plurality of updated weights; and

resample from the first training set, based on the plurality of updated weights, to train the machine learning model in a second epoch, wherein the second epoch follows the first epoch.

12 . The electronic device of claim 11 , wherein the at least one instruction, when executed by the one or more processors, further causes the electronic device to:

evaluate the machine learning model using the validation set after the second epoch to obtain a plurality of second performances at the plurality of levels; and

update the plurality of updated weights corresponding to the plurality of levels based on the plurality of second performances.

13 . The electronic device of claim 11 , wherein the plurality of updated weights corresponding to the plurality of levels is negatively correlated with the plurality of first performances of the machine learning model at the plurality of levels.

14 . The electronic device of claim 13 , wherein each of the plurality of weights and the plurality of updated weights is equal to or more than a predetermined minimum value.

15 . The electronic device of claim 11 , wherein the at least one instruction, when executed by the one or more processors, further cause the electronic device to:

mix noise data into a noise-free training set to generate the first training set.

16 . The electronic device of claim 15 , wherein the sampling from the first training set, based on the plurality of weights, to train the machine learning model in the first epoch comprises:

sampling from the plurality of levels of the first training set based on the plurality of weights to obtain a first sample training set;

merging the first sample training set with the noise-free training set to generate a second training set; and

using the second training set to train the machine learning model in the first epoch.

17 . The electronic device of claim 11 , wherein the first training set comprises an ordered data set.

18 . The electronic device of claim 11 , wherein the at least one instruction, when executed by the one or more processors, further cause the electronic device to:

divide the first training set into the plurality of levels based on a distortion index.

19 . The electronic device of claim 18 , wherein the machine learning model comprises a speech recognition model.

20 . The electronic device of claim 19 , wherein the dividing of the first training set into the plurality of levels, based on the distortion index, comprises:

dividing the first training set into the plurality of levels based on at least one of a Perceptual Evaluation of Speech Quality (PESQ), a Short-Time Objective Intelligibility (STOI), and a Frequency-Weighted Signal-to-Noise Ratio Segmental (fwSNRseq).

21 . A non-transitory computer-readable medium, comprising at least one instruction, when executed by a processor of an electronic device, causes the electronic device to:

sample from a first training set, based on a plurality of weights, to train a machine learning model in a first epoch, the first training set comprising a plurality of levels corresponding to the plurality of weights;

evaluate the machine learning model using a validation set after the first epoch to obtain a plurality of first performances at the plurality of levels;

update the plurality of weights corresponding to the plurality of levels based on the plurality of first performances to obtain a plurality of updated weights; and

resample from the first training set, based on the plurality of updated weights, to train the machine learning model in a second epoch, wherein the second epoch follows the first epoch.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2024
From: LEE, CHI-CHUN; WU, YA-TSE
To: NATIONAL TSING HUA UNIVERSITY
Reel/Frame 067979/0167 →
Priority Claims (1)
TW 112131571 · Aug 22, 2023 · national
Continuity (1)
Related Publication 20250069590A1 · Feb 27, 2025
References Cited (14)
US 11650968B2 · Nair · 2023 [cited by examiner]
US 12026974B2 · Zhao · 2024 [cited by examiner]
US 20200372342A1 · Nair · 2020 [cited by examiner]
US 20220126864A1 · Moustafa · 2022 [cited by examiner]
US 20230095685A1 · Greving · 2023 [cited by examiner]
US 20230222326A1 · Jafari · 2023 [cited by examiner]
US 20250069590A1 · Lee · 2025 [cited by examiner]
US 20250209336A1 · Lee · 2025 [cited by examiner]
US 20260051316A1 · Jukic · 2026 [cited by examiner]
CN 115176254A · 2022 [cited by examiner]
TW 202135529A · 2021 [cited by examiner]
TW 202240735A · 2022 [cited by applicant]
WO WO2022051855A1 · 2022 [cited by examiner]
WO WO2026024858A1 · 2026 [cited by examiner]