IP Library Granted Patent US 12700398
Granted Patent B2
US 12700398 · App. 18/272,246 · Granted Aug 4, 2026

Learning method, learning system and learning program

Inventors: Hosana Kamiyama (Tokyo, JP); Yoshikazu Yamaguchi (Tokyo, JP)
Assignee: NTT, Inc.
G10L15/063G10L15/16G10L19/00G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12700398
App. No.
18/272,246
Granted
Aug 4, 2026
Kind
B2
Abstract

A learning method includes the following processes. A shuffling process acquires learning data arranged in a time series and rearranges the learning data in an order different from the order of the time series. A learning process trains an acoustic model using the learning data rearranged through the shuffling process.

Claims (59)

1 . A learning method comprising:

acquiring learning data in sequence based on a first sequence in a time series;

sequencing the learning data in a second sequence that is distinct from the first sequence in the time series;

generating noise data, wherein the noise data includes an amount by which the learning data is not able to be restored, in which the learning data is sequenced in the time series, by using information on each time in the time series and a feature amount including a minute noise;

adding the noise data to the learning data in the second sequence; and

training an acoustic model using the acquired learning data in the second sequence.

2 . The learning method according to claim 1 , wherein the acquiring further comprises:

acquiring the learning data including a feature amount and a phoneme label;

generating new learning data by combining the feature amount in a predetermined period around a predetermined time and the phoneme label at the predetermined time; and

sequencing the new learning data in the second sequence that is distinct from the first sequence in the time series,

wherein the training further comprises training the acoustic model using the new learning data.

3 . The learning method according to claim 1 , further comprising:

generating noise data, wherein the noise data include an amount by which the learning data is not able to be restored to a state in which the learning data is sequenced in the time series even if the learning data is rearranged; and

adding the noise data to the learning data in the second sequence.

4 . The learning method according to claim 3 , wherein the generating the noise data further comprises generating the noise data based on a predetermined base acoustic model.

5 . The learning method according to claim 3 , wherein the generating noise data further comprises generating the noise data that prevents deviating from fluctuation in speech.

6 . A learning system comprising a processor configured to execute operations comprising:

acquiring learning data in sequence based on a first sequence in a time series;

sequencing the learning data in a second sequence that is distinct from the first sequence in the time series;

generating noise data, wherein the noise data includes an amount by which the learning data is not able to be restored, in which the learning data is sequenced in the time series, by using information on each time in the time series and a feature amount including a minute noise;

adding the noise data to the learning data in the second sequence; and

training an acoustic model using the acquired learning data.

7 . A computer-readable non-transitory recording medium storing computer-executable program instructions that when executed by a processor cause a computer system to execute operations comprising:

acquiring learning data in sequence based on a first sequence in a time series;

sequencing the learning data in a second sequence that is distinct from the first sequence in the time series;

generating noise data, wherein the noise data includes an amount by which the learning data is not able to be restored, in which the learning data is sequenced in the time series, by using information on each time in the time series and a feature amount including a minute noise;

adding the noise data to the learning data in the second sequence; and

training an acoustic model using the acquired learning data.

8 . The learning method according to claim 1 , wherein the acquired learning data includes protected data according to the first sequence in the time series, and

the training the acoustic model using the acquired learning data sequenced in the second sequence prevents exposing the protected data.

9 . The learning method according to claim 2 , further comprising:

generating noise data, wherein the noise data include amount by which the learning data is not able to be restored to a state in which the learning data is sequenced in the time series even if the learning data is rearranged; and

adding the noise data to the learning data in the second sequence.

10 . The learning method according to claim 4 , wherein the generating noise data further comprises generating the noise data that prevents deviating from fluctuation in speech.

11 . The learning system according to claim 6 , wherein the acquiring further comprises:

acquiring the learning data including a feature amount and a phoneme label;

generating new learning data by combining the feature amount in a predetermined period around a predetermined time and the phoneme label at the predetermined time; and

sequencing the new learning data in the second sequence that is distinct from the first sequence in the time series, wherein the training further comprises training the acoustic model using the new learning data.

12 . The learning system according to claim 6 , the processor further configured to execute operations comprising:

generating noise data, wherein the noise data include an amount by which the learning data is not able to be restored to a state in which the learning data is sequenced in the time series even if the learning data is rearranged; and

adding the noise data to the learning data in the second sequence.

13 . The learning system according to claim 6 , wherein the acquired learning data includes protected data according to the first sequence in the time series, and

the training the acoustic model using the acquired learning data sequenced in the second sequence prevents exposing the protected data.

14 . The learning system according to claim 11 , the processor further configured to execute operations comprising:

generating noise data, wherein the noise data include amount by which the learning data is not able to be restored to a state in which the learning data is sequenced in the time series even if the learning data is rearranged; and

adding the noise data to the learning data in the second sequence.

15 . The learning system according to claim 12 , wherein the generating the noise data further comprises generating the noise data based on a predetermined base acoustic model.

16 . The learning system according to claim 12 , wherein the generating noise data further comprises generating the noise data that prevents deviating from fluctuation in speech.

17 . The computer-readable non-transitory recording medium according to claim 7 , wherein the acquiring further comprises:

acquiring the learning data including a feature amount and a phoneme label;

generating new learning data by combining the feature amount in a predetermined period around a predetermined time and the phoneme label at the predetermined time;

sequencing the new learning data in the second sequence that is distinct from the first sequence in the time series; and

the training further comprises training the acoustic model using the new learning data.

18 . The computer-readable non-transitory recording medium according to claim 7 , the computer-executable program instructions when executed further causing the computer system to execute operations comprising:

generating noise data, wherein the noise data include an amount by which the learning data is not able to be restored to a state in which the learning data is sequenced in the time series even if the learning data is rearranged; and

adding the noise data to the learning data in the second sequence.

19 . The computer-readable non-transitory recording medium according to claim 7 , wherein the acquired learning data includes protected data according to the first sequence in the time series, and

the training the acoustic model using the acquired learning data sequenced in the second sequence prevents exposing the protected data.

20 . The computer-readable non-transitory recording medium according to claim 18 , wherein the generating the noise data further comprises generating the noise data based on a predetermined base acoustic model.