IP Library Granted Patent US 12,542,128
Granted Patent B2
US 12,542,128 · App. 18/690,377 · Granted Feb 3, 2026

Learning apparatus, conversion apparatus, learning method and program

Inventors: Daisuke Niizumi (Musashino, JP); Kunio Kashino (Musashino, JP); Yasunori Oishi (Musashino, JP); Daiki Takeuchi (Musashino, JP); Noboru Harada (Musashino, JP)
Assignee: NTT, Inc.
G10L15/063G10L15/02G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,542,128
App. No.
18/690,377
Granted
Feb 3, 2026
Kind
B2
Abstract

An aspect of the present invention provides a learning device including: a neural network that converts an acoustic time series that is a time series representing a sound into feature data expressed in a predetermined format required by a downstream task; and an update unit that updates the neural network on the basis of an execution result of the downstream task using the feature data, in which the neural network includes: a feature extraction unit that converts an input acoustic time series into an intermediate feature tensor that is a third-order tensor indicating features of the acoustic time series and is a tensor having time, frequency, and channel; and an intermediate network that executes processing of converting a representation of the intermediate feature tensor into a representation of a second-order tensor having time and direct product amounts that are amounts indicating a direct product of frequency and channel, and processing of acquiring, for each direct product amount of the second-order tensor, a one-dimensional vector indicating a statistic in a time axis direction of each direct product amount as the feature data.

Claims (44)

1 . A learning device comprising:

a processor includes a neural network that converts an acoustic time series that is a time series representing a sound into feature data that is information regarding the sound obtained on the basis of the acoustic time series and is information expressed in a predetermined format required by a downstream task; and

a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of:

updating the neural network on the basis of an execution result of the downstream task using the feature data,

in which the neural network:

converts an input acoustic time series into an intermediate feature tensor that is a third-order tensor that indicates features of the acoustic time series and has time, frequency, and channel; and

the neural network includes

an intermediate network that executes non-equivalent second-order tensor processing of converting a representation of the intermediate feature tensor into a representation of a planar tensor that is a second-order tensor having time and direct product amounts that are amounts indicating a direct product of frequency and channel, and vectorization processing of acquiring, for each direct product amount of the planar tensor, a one-dimensional vector indicating a statistic in a time axis direction of each direct product amount as the feature data.

2 . The learning device according to claim 1 , wherein

the statistic is either an average value or a maximum value.

3 . The learning device according to claim 1 , wherein

the feature data is a one-dimensional vector of a vector sum of a one-dimensional vector indicating, for each direct product amount of the planar tensor, an average value in the time axis direction of each direct product amount and a one-dimensional vector indicating, for each direct product amount of the planar tensor, a maximum value in the time axis direction of each direct product amount.

4 . The learning device according to claim 1 , wherein

a non-equivalent second-order tensor processing execution unit updates the planar tensor with a mathematical model of machine learning.

5 . A conversion device comprising:

a processor and

a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of:

acquiring an acoustic time series that is a time series representing a sound; and

converting, into feature data, the acoustic time series acquired by the acquiring using a trained neural network obtained by learning by a learning device, a learning device comprising: a processor includes a neural network that converts an acoustic time series that is a time series representing a sound into feature data that is information regarding the sound obtained on the basis of the acoustic time series and is information expressed in a predetermined format required by a downstream task; and a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of: updating the neural network on the basis of an execution result of the downstream task using the feature data, in which the neural network converts an input acoustic time series into an intermediate feature tensor that is a third-order tensor that indicates features of the acoustic time series and has time, frequency, and channel; and the neural network includes an intermediate network that executes non-equivalent second-order tensor processing of converting a representation of the intermediate feature tensor into a representation of a planar tensor that is a second-order tensor having time and direct product amounts that are amounts indicating a direct product of frequency and channel, and vectorization processing of acquiring, for each direct product amount of the planar tensor, a one-dimensional vector indicating a statistic in a time axis direction of each direct product amount as the feature data.

6 . A learning method executed by a learning device comprising: a processor includes a neural network that converts an acoustic time series that is a time series representing a sound into feature data that is information regarding the sound obtained on the basis of the acoustic time series and is information expressed in a predetermined format required by a downstream task; and a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of: updating the neural network on the basis of an execution result of the downstream task using the feature data, in which the neural network converts an input acoustic time series into an intermediate feature tensor that is a third-order tensor that indicates features of the acoustic time series and has time, frequency, and channel; and the neural network includes an intermediate network that executes non-equivalent second-order tensor processing of converting a representation of the intermediate feature tensor into a representation of a planar tensor that is a second-order tensor having time and direct product amounts that are amounts indicating a direct product of frequency and channel, and vectorization processing of acquiring, for each direct product amount of the planar tensor, a one-dimensional vector indicating a statistic in a time axis direction of each direct product amount as the feature data,

the learning method comprising:

converting an input acoustic time series into an intermediate feature tensor that is a third-order tensor that indicates features of the acoustic time series and has time, frequency, and channel;

converting a representation of the intermediate feature tensor into a representation of a planar tensor that is a second-order tensor having time and direct product amounts that are amounts indicating a direct product of frequency and channel; and

acquiring, for each direct product amount of the planar tensor, a one-dimensional vector indicating a statistic in a time axis direction of each direct product amount as the feature data.

7 . A non-transitory computer readable medium which stores a program for causing a computer to function as the learning device according to claim 1 .

8 . A non-transitory computer readable medium which stores a program for causing a computer to function as the conversion device according to claim 5 .

9 . A learning device comprising:

a processor includes a neural network that converts an acoustic time series that is a time series representing a sound into feature data that is information regarding the sound obtained on the basis of the acoustic time series and is information expressed in a predetermined format required by a downstream task; and

a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of:

updating the neural network on the basis of an execution result of the downstream task using the feature data,

in which the neural network:

acquires, on the basis of an input acoustic time series, an intermediate feature tensor that is a third-order tensor that indicates features of the acoustic time series and has time, frequency, and channel; and

the neural network includes an intermediate network that executes non-equivalent second-order tensor processing of converting a representation of the intermediate feature tensor into a representation of a planar tensor that is a second-order tensor having time and direct product amounts that are amounts indicating a direct product of frequency and channel.

10 . The learning device according to claim 9 , wherein

a plurality of intermediate feature tensors are acquired by the acquiring, and

the intermediate network acquires the planar tensor for each one of a plurality of the intermediate feature tensors.

11 . The learning device according to claim 10 , wherein

the intermediate network executes:

combining processing of combining a plurality of the planar tensors; and

vectorization processing of acquiring, for each direct product amount of a combined planar tensor obtained by the combining processing, a one-dimensional vector indicating a statistic in a time axis direction of each direct product amount as the feature data.

12 . The learning device according to claim 10 , wherein

the intermediate network executes processing of making a plurality of the intermediate feature tensors to have the same number of elements in a time direction.

13 . The learning device according to claim 10 , wherein

the intermediate network executes processing of making a plurality of the planar tensors to be subjected to the combining processing to have the same number of elements in a time direction.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072998/0094 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2024
From: NIIZUMI, DAISUKE; KASHINO, KUNIO; OISHI, YASUNORI; TAKEUCHI, DAIKI; HARADA, NOBORU
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 066694/0904 →
Priority Claims (1)
WO PCT/JP2021/034318 · Sep 17, 2021 · international
Continuity (1)
Related Publication 20250131914A1 · Apr 24, 2025
References Cited (2)
S. Hershey et al., “CNN architectures for large-scale audio classification”, 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 131-135, doi: 10.1109/ICASSP.2017.7952132. [cited by applicant]
Kong et al., “PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition”, in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2880-2894, 2020, doi: 10.1109/TASLP.2020… [cited by applicant]