Sound collection control method and sound collection apparatus
A sound collection control method recognizes a speaker from an image, detects a position of the recognized speaker, sets a first collection beam based on the position of the recognized speaker, recognizes a specific object other than the recognized speaker from an image, detects a position of the recognized specific object, and sets a second collection beam based on the detected position of the recognized specific object.
1 . A sound collection control method comprising:
recognizing a speaker from an image;
detecting a position of the recognized speaker;
setting a first collection beam based on the detected position of the recognized speaker;
recognizing a specific object that is not a person and that does not generate sound-from the image;
detecting a position of the recognized specific object; and
setting a second collection beam based on the detected position of the recognized specific object,
wherein the first collection beam is a sound collection beam, and
wherein, in a case where no person speaks, only the second collection beam is set without setting the first sound collection beam.
2 . The sound collection control method according to claim 1 , wherein the second collection beam includes a non-sound collection beam having a lower sensitivity than other directions.
3 . The sound collection control method according to claim 1 , wherein:
the recognizing of the speaker recognizes a plurality of speakers, including the speaker, from the image,
the detecting of the position of the recognized speaker detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker,
the setting of the first collection beam sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers,
the recognizing of the specific object recognizes a plurality of specific objects, including the specific object, from the image,
the detecting of the position of the recognized specific object detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object,
the setting of the second collection beam sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects, and
a total number of the first collection beams and the second collection beams has a maximum.
4 . The sound collection control method according to claim 3 , wherein, in a state where the total number exceeds the maximum, the setting of the first collection beam does not set a newest first collection beam or sets the newest first collection beam in place of a previously set first collection beam.
5 . The sound collection control method according to claim 3 , wherein, in a state where the total number exceeds the maximum, the setting of the first collection beam sets a newest first collection beam in place of a previously set first or second collection beam, which is selectable by a user.
6 . The sound collection control method according to claim 1 , further comprising:
receiving a mute operation on the first collection beam or the second collection beam; and
muting the first collection beam or the second collection beam that receives the mute operation.
7 . A sound collection control method comprising:
recognizing a speaker from an image;
detecting a position of the recognized speaker;
setting a first collection beam based on the detected position of the recognized speaker;
recognizing a specific object other than the recognized speaker from the image;
detecting a position of the recognized specific object; and
setting a second collection beam based on the detected position of the recognized specific object,
wherein:
the first collection beam is a sound collection beam;
the recognizing of the speaker recognizes a plurality of speakers, including the speaker, from the image;
the detecting of the position of the recognized speaker detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker;
the setting of the first collection beam sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers;
the recognizing of the specific object recognizes a plurality of specific objects, including the specific object, from the image;
the detecting of the position of the recognized specific object detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object;
the setting of the second collection beam sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects;
a total number of the first collection beams and the second collection beams has a maximum; and
in a state where the total number exceeds the maximum, the setting of the first collection beam sets a newest first collection beam in place of a previously set first or second collection beam based on predefined priority.
8 . A sound collection apparatus comprising:
an array microphone; and
a hardware controller configured to execute a plurality of tasks, including:
a speaker recognizing task that recognizes a speaker from an image;
a speaker position detecting task that detects a position of the recognized speaker;
a first collection beam setting task that sets a first collection beam to the array microphone based on the detected position of the speaker;
an object recognizing task that recognizes a specific object that is not a person and that does not generate sound from the image;
an object position detecting task that detects a position of the recognized specific object; and
a second collection beam setting task that sets a second collection beam to the array microphone based on the detected position of the recognized specific object,
wherein the first collection beam is a sound collection beam, and
wherein, in a case where no person speaks, only the second collection beam is set without setting the first sound collection beam.
9 . The sound collection apparatus according to claim 8 , wherein the second collection beam includes a non-sound collection beam having a lower sensitivity than other directions.
10 . The sound collection apparatus according to claim 8 , wherein:
the speaker recognizing task recognizes a plurality of speakers, including the speaker, from the image,
the speaker position detecting task detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker,
the first collection beam setting task sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers,
the object recognizing task recognizes a plurality of specific objects, including the specific object, from the image,
the object position detecting task detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object,
the second collection beam setting task sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects, and
a total number of the first collection beams and the second collection beams has a maximum.
11 . The sound collection apparatus according to claim 10 , wherein, in a state where the total number exceeds the maximum, the first collection beam setting task does not set a newest first collection beam or sets the newest first collection beam in place of a previously set first collection beam.
12 . The sound collection apparatus according to claim 10 , wherein, in a state where the total number exceeds the maximum, the first collection beam setting task sets a newest first collection beam in place of a previously set first or second collection beam, which is selectable by a user.
13 . The sound collection apparatus according to claim 8 , further comprising a mute controller that:
receives a mute operation on the first collection beam or the second collection beam; and
mutes the first collection beam or the second collection beam that receives the mute operation.
14 . A sound collection apparatus comprising:
an array microphone; and
a hardware controller configured to execute a plurality of tasks, including:
a speaker recognizing task that recognizes a speaker from an image;
a speaker position detecting task that detects a position of the recognized speaker;
a first collection beam setting task that sets a first collection beam to the array microphone based on the detected position of the speaker;
an object recognizing task that recognizes a specific object other than the speaker from the image;
an object position detecting task that detects a position of the recognized specific object; and
a second collection beam setting task that sets a second collection beam to the array microphone based on the detected position of the recognized specific object,
wherein:
the first collection beam is a sound collection beam;
the speaker recognizing task recognizes a plurality of speakers, including the speaker, from the image;
the speaker position detecting task detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker;
the first collection beam setting task sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers;
the object recognizing task recognizes a plurality of specific objects, including the specific object, from the image;
the object position detecting task detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object;
the second collection beam setting task sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects;
a total number of the first collection beams and the second collection beams has a maximum; and
in a state where the total number exceeds the maximum, the first beam setting task sets a newest first collection beam in place of a previously set first or second collection beam based on predefined priority.