Sound source separation method, sound source separation apparatus, and program
A mixed acoustic signal including sound emitted from a plurality of sound sources and sound source video signals representing at least one video of the plurality of sound sources are received as inputs, and at least a separated signal including a signal representing a target sound emitted from one sound source represented by the video is acquired. However, at least the separated signal is acquired using properties of the sound source that affects sound emitted by the sound source acquired from the video and/or features of a structure used for the sound source to emit the sound.
1 . A method for separating a sound source, comprising:
estimating a separated signal including a signal representing a target sound emitted from a sound source among a plurality of sound sources by applying a mixed acoustic signal and sound source video signals to a model, wherein the mixed acoustic signal represents a mixed sound of sound emitted from the plurality of sound sources, the sound source video signals represent videos of at least some of the plurality of sound sources to the model,
the model is obtained by learning using a neural network, wherein the model is learned based on differences between features of the separated signal and features of teaching data of sound source video signals, the features of the separated signal are obtained by applying teaching mixed acoustic signals as teaching data of the mixed acoustic signal and teaching sound source video signals as teaching data of the sound source video signals to the model, wherein
the learning is further based on:
a similarity between a first element representing the features of the teaching sound source video signals corresponding to a first sound source among the plurality of sound sources and a second element representing the features of the separated signal corresponding to a second sound source different from the first sound source decreases, and
a similarity between the first element representing the features of the teaching sound source video signals corresponding to the first sound source and a third element representing the features of the separated signal corresponding to the first sound source increases.
2 . The according to claim 1 , wherein
the model is further obtained by learning based on differences between the separated signal obtained by applying the teaching mixed acoustic signals and the teaching sound source video signals to the model and teaching sound source signals which are teaching data of the separated signal corresponding to the teaching mixed acoustic signals and the teaching sound source video signals.
3 . The method according to claim 1 , wherein
the sound source video signals represent a video of each of the plurality of sound sources.
4 . The method according to claim 1 , wherein
the plurality of sound sources includes a plurality of different speakers, the mixed acoustic signal includes a speech signal, and the sound source video signals represent videos of the speakers.
5 . The method according to claim 1 , wherein
the separated signal includes a signal representing a target sound emitted from a certain sound source among the plurality of sound sources and a signal representing a sound emitted from another sound source.
6 . The method according to claim 1 , wherein
the model is further obtained by learning based on differences between the separated signal obtained by applying the teaching mixed acoustic signals and the teaching sound source video signals to the model and teaching sound source signals which are teaching data of the separated signal corresponding to the teaching mixed acoustic signals and the teaching sound source video signals.
7 . The method according to claim 1 , wherein
the sound source video signals represent a video of each of the plurality of sound sources.
8 . The method according to claim 1 , wherein
the plurality of sound sources includes a plurality of different speakers, the mixed acoustic signal includes a speech signal, and the sound source video signals represent videos of the speakers.