Information processing apparatus and method that use learning models to alter watermark information based on mixed audio signals
An information processing apparatus includes a decoder that extracts, using a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together. A change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible.
1 . An information processing apparatus comprising:
a decoder that extracts, using a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together,
wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;
wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.
2 . The information processing apparatus according to claim 1 ,
wherein the additional information is information added to the mixed sound signal where the plurality of sound source signals is mixed together.
3 . The information processing apparatus according to claim 1 , wherein the additional information is digital watermark information.
4 . The information processing apparatus according to claim 1 ,
wherein the decoder extracts the additional information from separation signals obtained by performing sound source separation on the mixed sound signal.
5 . The information processing apparatus according to claim 4 , further comprising
a sound source separation unit that performs the sound source separation.
6 . The information processing apparatus according to claim 1 ,
wherein at least one of the sound source signals constituting the mixed sound signal includes the additional information.
7 . The information processing apparatus according to claim 6 ,
wherein each of the sound source signals constituting the mixed sound signal includes the additional information.
8 . An information processing apparatus comprising a concealer that includes, using a certain learning model, additional information in a plurality of sound source signals or a mixed sound signal where the plurality of sound source signals is mixed together, wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;
wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.
9 . The information processing apparatus according to claim 8 ,
wherein the concealer adds the additional information to the mixed sound signal where the plurality of sound source signals is mixed together.
10 . The information processing apparatus according to claim 8 ,
wherein the learning model to be used when the additional information is included is selectable.
11 . The information processing apparatus according to claim 8 , wherein the additional information is digital watermark information.
12 . The information processing apparatus according to claim 8 ,
wherein the concealer adds the additional information to at least one of the plurality of sound source signals.
13 . The information processing apparatus according to claim 12 ,
wherein the concealer adds the additional information to all the plurality of sound source signals.
14 . An information processing method comprising:
extracting, using a decoder and a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together,
wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;
wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.
15 . A non-transitory computer-readable medium storing a program causing a computer to perform an information processing method comprising:
extracting, using a decoder and a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together,
wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;
wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.
16 . An information processing method comprising:
including, using a concealer and a certain learning model, additional information in a plurality of sound source signals or a mixed sound signal where the plurality of sound source signals is mixed together,
wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;
wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.
17 . A non-transitory computer-readable medium storing a program causing a computer to perform an information processing method comprising:
including, using a concealer and a certain learning model, additional information in a plurality of sound source signals or a mixed sound signal where the plurality of sound source signals is mixed together,
wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;
wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.