IP Library Patent Application 18858453
Patent Application
App. No. 18/858,453

LEARNING APPARATUS, CONVERTING APPARATUS, METHODS AND PROGRAMS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/858,453
Abstract

A learning device includes: spectrogram generation circuitry 1 that generates a spectrogram from a first sound signal and generates a target spectrogram from a second sound signal; patch generation circuitry 2 that divides the spectrogram to generate a plurality of patches and divides the target spectrogram to generate a plurality of target patches; mask processing circuitry 3 that selects some patches as masked patches; reconstruction circuitry 4 that obtains a plurality of reconstructed patches by reconstructing the plurality of patches by processing of an encoder and a decoder by using visible patches other than some patches among the plurality of patches and mask tokens; and parameter update circuitry 5 that updates parameters such that target patches corresponding to the masked patches among the plurality of target patches approach reconstructed patches corresponding to the masked patches among the plurality of reconstructed patches.

Claims (35)

1 . A learning device comprising:

spectrogram generation circuitry that generates a spectrogram from an input first sound signal and generates a target spectrogram from an input second sound signal;

patch generation circuitry that divides the generated spectrogram to generate a plurality of patches and divides the generated target spectrogram to generate a plurality of target patches;

mask processing circuitry that selects some patches from among the plurality of patches as masked patches;

reconstruction circuitry that obtains a plurality of reconstructed patches by reconstructing the plurality of patches by processing of an encoder and a decoder in a transformer serving as a deep learning model by using visible patches other than the some patches among the plurality of patches and mask tokens corresponding to the masked patches; and

parameter update circuitry that updates a parameter of the encoder and a parameter of the decoder such that target patches corresponding to the masked patches among the plurality of target patches approach reconstructed patches corresponding to the masked patches among the plurality of reconstructed patches.

2 . A transform device comprising:

a storage that stores the encoder and the decoder learned by the learning device according to claim 1 ;

spectrogram generation circuitry that generates a spectrogram from an input sound signal;

patch generation circuitry that divides the generated spectrogram to generate a plurality of patches;

mask processing circuitry that selects some patches from among the plurality of patches as masked patches; and

reconstruction circuitry that obtains a plurality of reconstructed patches by reconstructing the plurality of patches by processing of the encoder and the decoder read from the storage by using visible patches other than the some patches among the plurality of patches and mask tokens corresponding to the masked patches.

3 . The transform device according to claim 2 , wherein:

processing of the mask processing circuitry and the reconstruction circuitry is repeatedly performed;

in the repetitive processing, the mask processing circuitry selects some or all of unselected patches from among the plurality of patches as the masked patches; and

the transform device further includes integration circuitry that performs processing of integrating reconstructed patches corresponding to the masked patches among the plurality of reconstructed patches to generate a reconstructed spectrogram after the repetitive processing is completed, and time domain transform circuitry that transforms the reconstructed spectrogram into a time domain sound signal.

4 . The transform device according to claim 3 , wherein:

the plurality of patches is arranged on a two-dimensional plane; and

the mask processing circuitry selects masked patches in a checkered pattern.

5 . The transform device according to claim 3 , wherein:

the plurality of patches is arranged on a two-dimensional plane; and

the mask processing circuitry selects every other masked patch in one coordinate axis direction of a coordinate system of the two-dimensional plane.

6 . A learning method comprising:

a spectrogram generation step of causing spectrogram generation circuitry to generate a spectrogram from an input first sound signal and generate a target spectrogram from an input second sound signal;

a patch generation step of causing patch generation circuitry to divide the generated spectrogram to generate a plurality of patches and divide the generated target spectrogram to generate a plurality of target patches;

a mask processing step of causing mask processing circuitry to select some patches from among the plurality of patches as masked patches;

a reconstruction step of causing reconstruction circuitry to obtain a plurality of reconstructed patches by reconstructing the plurality of patches by processing of an encoder and a decoder in a transformer serving as a deep learning model by using visible patches other than the some patches among the plurality of patches and mask tokens corresponding to the masked patches; and

a parameter update step of causing parameter update circuitry to update a parameter of the encoder and a parameter of the decoder such that target patches corresponding to the masked patches among the plurality of target patches approach reconstructed patches corresponding to the masked patches among the plurality of reconstructed patches.

7 . A transform method comprising:

a spectrogram generation step of causing spectrogram generation circuitry to generate a spectrogram from an input sound signal;

a patch generation step of causing patch generation circuitry to divide the generated spectrogram to generate a plurality of patches;

a mask processing step of causing mask processing circuitry to select some patches from among the plurality of patches as masked patches; and

a reconstruction step of causing reconstruction circuitry to obtain a plurality of reconstructed patches by reconstructing the plurality of patches by processing of an encoder and a decoder read from a storage unit storing the encoder and the decoder learned by the learning method according to claim 6 by using visible patches other than the some patches among the plurality of patches and mask tokens corresponding to the masked patches.

8 . A non-transitory computer readable medium that stores a program for causing a computer to perform each step of the learning method according to claim 6 .

9 . A non-transitory computer readable medium that stores a program for causing a computer to perform each step of the transform method according to claim 7 .

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2024
From: NIIZUMI, DAISUKE; KASHINO, KUNIO; OHISHI, YASUNORI; TAKEUCHI, DAIKI; HARADA, NOBORU
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 069562/0180 →