Masking apparatus, masking method, and program
A masking technology is provided for curbing discomfort at the time of a change in a masking sound by presenting a video corresponding to the masking sound at the time of the change in the masking sound. A masking device includes: a spoken voice volume evaluation unit configured to generate an evaluation value for a volume of a spoken voice (hereinafter referred to as a spoken voice volume evaluation value) from a spoken voice signal by using, as the spoken voice signal, a sound collection signal output by a microphone installed for collecting the spoken voice which is a voice of a speaking person; a masking sound signal generation unit configured to generate a signal for emitting a masking sound from a speaker (hereinafter referred to as a masking sound signal) corresponding to the spoken voice volume evaluation value, the masking sound preventing the spoken voice from being heard by surrounding persons other than the speaking person; and a masking video signal generation unit configured to generate a signal for presenting a video corresponding to the masking sound (hereinafter referred to as a masking video signal) from a video presentation device.
1 . A masking device comprising:
a spoken voice volume evaluation circuitry configured to generate an evaluation value for a volume of a spoken voice (hereinafter referred to as a spoken voice volume evaluation value) from a spoken voice signal by using, as the spoken voice signal, a sound collection signal output by a microphone installed for collecting the spoken voice which is a voice of a speaking person;
a masking sound signal generation circuitry configured to generate a signal for emitting a masking sound from a speaker (hereinafter referred to as a masking sound signal) corresponding to the spoken voice volume evaluation value, the masking sound preventing the spoken voice from being heard by surrounding persons other than the speaking person; and
a masking video signal generation circuitry configured to generate a signal for presenting a video corresponding to the masking sound (hereinafter referred to as a masking video signal) from a video presentation device,
wherein the masking sound signal generation circuitry selects a masking sound signal with a small volume when the spoken voice volume elevation value is a value indicating that the spoken voice volume is small, and a masking sound signal with a large volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is large, from among multiple masking sound signals recorded in the recording circuitry, and
the masking video signal generation circuitry selects the masking video signal using meta-information of the masking sound signal.
2 . The masking device according to claim 1 , further comprising:
a masking sound erasing circuitry configured to generate a signal in which a component caused by the masking sound included in the sound collection signal is erased by using the sound collection signal and the masking sound signal, and to use the signal as the spoken voice signal.
3 . A masking device comprising:
a microphone array processing circuitry configured to generate an integrated sound collection signal from N (where N is an integer of 2 or more) sound collection signals output by a microphone array including N microphones installed for collecting a spoken voice that is a voice of a speaking person and to set the integrated sound collection signal as a spoken voice signal;
a spoken voice volume evaluation circuitry configured to generate an evaluation value for a volume of the spoken voice (hereinafter referred to as a spoken voice volume evaluation value) from the spoken voice signal;
a masking sound signal generation circuitry configured to generate a signal for emitting a masking sound (hereinafter referred to as a masking sound signal) corresponding to the spoken voice volume evaluation value from a speaker array including M (where M is an integer of 2 or more) speakers, the masking sound preventing the spoken voice from being heard by surrounding persons other than the speaking person;
a masking video signal generation circuitry configured to generate a signal for presenting a video corresponding to the masking sound (hereinafter referred to as a masking video signal) from a video presentation device; and
a speaker array processing circuitry configured to generate M individual masking sound signals for emitting sound from the speakers included in the speaker array from the masking sound signal,
wherein the masking sound signal generation circuitry selects a masking sound signal with a small volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is small, and a masking sound signal with a large volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is large, from among multiple masking sound signals recorded in the recording circuitry, and
the masking video signal generation circuitry selects the masking video signal using meta-information of the masking sound signal.
4 . The masking device according to claim 3 , further comprising:
a masking sound erasing circuitry configured to generate a signal in which a component caused by the masking sound included in the integrated sound collection signal is erased by using the integrated sound collection signal and the masking sound signal, and to use the signal as the spoken voice signal.
5 . A masking device comprising:
a microphone array processing circuitry configured to generate an integrated sound collection signal from N (where N is an integer of 2 or more) sound collection signals output by a microphone array including N microphones installed for collecting a spoken voice that is a voice of a speaking person and to set the integrated sound collection signal as a spoken voice signal;
a spoken voice volume evaluation circuitry configured to generate an evaluation value for a volume of the spoken voice (hereinafter referred to as a spoken voice volume evaluation value) from the spoken voice signal;
a masking sound signal generation circuitry configured to generate a signal for emitting a masking sound (hereinafter referred to as a masking sound signal) corresponding to the spoken voice volume evaluation value from a speaker array including M (where M is an integer of 2 or more) speakers, the masking sound preventing the spoken voice from being heard by surrounding persons other than the speaking person; and
a speaker array processing circuitry configured to generates M individual masking sound signals for emitting sound from the speakers included in the speaker array from the masking sound signal,
wherein, of the M individual masking sound signals, an individual masking sound signal directed to a direction of the speaking person is a signal such that the higher the spoken voice volume evaluation value indicates, the greater the sound emitted by the signal is and,
the masking sound signal generation circuitry selects a masking sound signal with a small volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is small, and a masking sound signal with a large volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is large, from among multiple masking sound signals recorded in the recording circuitry.
6 . A masking method comprising:
a spoken voice volume evaluation step of generating, by a masking device, an evaluation value for a volume of a spoken voice (hereinafter referred to as a spoken voice volume evaluation value) from a spoken voice signal by using, as the spoken voice signal, a sound collection signal output by a microphone installed for collecting the spoken voice which is a voice of a speaking person;
a masking sound signal generation step of generating, by the masking device, a signal for emitting a masking sound from a speaker (hereinafter referred to as a masking sound signal) corresponding to the spoken voice volume evaluation value, the masking sound preventing the spoken voice from being heard by surrounding persons other than the speaking person; and
a masking video signal generation step of generating, by the masking device, a signal for presenting a video corresponding to the masking sound (hereinafter referred to as a masking video signal) from a video presentation device,
wherein, in the masking sound signal generation step, the masking device selects a masking sound signal with a small volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is small, and a masking sound signal with a large volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is large, from among multiple masking sound signals recorded in the recording circuitry, and
in the masking video signal generation step, the masking device selects the masking video signal using meta-information of the masking sound signal.
7 . A masking method comprising:
a microphone array processing step of generating, by a masking device, an integrated sound collection signal from N (where N is an integer of 2 or more) sound collection signals output by a microphone array including N microphones installed for collecting a spoken voice that is a voice of a speaking person and to set the integrated sound collection signal as a spoken voice signal;
a spoken voice volume evaluation step of generating, by the masking device, an evaluation value for a volume of the spoken voice (hereinafter referred to as a spoken voice volume evaluation value) from the spoken voice signal;
a masking sound signal generation step of generating, by the masking device, a signal for emitting a masking sound (hereinafter referred to as a masking sound signal) corresponding to the spoken voice volume evaluation value from a speaker array including M (where M is an integer of 2 or more) speakers, the masking sound preventing the spoken voice from being heard by surrounding persons other than the speaking person;
a masking video signal generation step of generating, by the masking device, a signal for presenting a video corresponding to the masking sound (hereinafter referred to as a masking video signal) from a video presentation device; and
a speaker array processing step of generating, by the masking device, M individual masking sound signals for emitting sound from the speakers included in the speaker array from the masking sound signal,
wherein, in the masking sound signal generation step, the masking device selects a masking sound signal with a small volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is small, and a masking sound signal with a large volume when the spoken voice volume evaluation value is a value indicating that the spoken voice volume is large, from among multiple masking sound signals recorded in the recording circuitry, and
in the masking video signal generation step, the masking device selects the masking video signal using meta-information of the masking sound signal.
8 . A non-transitory computer-readable recording medium recording a program causing a computer to function as the masking device according to claim 1 .
9 . A non-transitory computer-readable recording medium recording a program causing a computer to function as the masking device according to claim 3 .