IP Library Granted Patent US 12,581,232
Granted Patent B2
US 12,581,232 · App. 18/366,078 · Granted Mar 17, 2026

Sound collection control method and sound collection apparatus

Inventor: Yoshifumi Oizumi (Hamamatsu, JP)
Assignee: YAMAHA CORPORATION
H04R1/403G06T7/70H04R1/406G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,581,232
App. No.
18/366,078
Granted
Mar 17, 2026
Kind
B2
Abstract

A sound collection control method recognizes a speaker from an image, detects a position of the recognized speaker, sets a first collection beam based on the position of the recognized speaker, recognizes a specific object other than the recognized speaker from an image, detects a position of the recognized specific object, and sets a second collection beam based on the detected position of the recognized specific object.

Claims (84)

1 . A sound collection control method comprising:

recognizing a speaker from an image;

detecting a position of the recognized speaker;

setting a first collection beam based on the detected position of the recognized speaker;

recognizing a specific object that is not a person and that does not generate sound-from the image;

detecting a position of the recognized specific object; and

setting a second collection beam based on the detected position of the recognized specific object,

wherein the first collection beam is a sound collection beam, and

wherein, in a case where no person speaks, only the second collection beam is set without setting the first sound collection beam.

2 . The sound collection control method according to claim 1 , wherein the second collection beam includes a non-sound collection beam having a lower sensitivity than other directions.

3 . The sound collection control method according to claim 1 , wherein:

the recognizing of the speaker recognizes a plurality of speakers, including the speaker, from the image,

the detecting of the position of the recognized speaker detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker,

the setting of the first collection beam sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers,

the recognizing of the specific object recognizes a plurality of specific objects, including the specific object, from the image,

the detecting of the position of the recognized specific object detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object,

the setting of the second collection beam sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects, and

a total number of the first collection beams and the second collection beams has a maximum.

4 . The sound collection control method according to claim 3 , wherein, in a state where the total number exceeds the maximum, the setting of the first collection beam does not set a newest first collection beam or sets the newest first collection beam in place of a previously set first collection beam.

5 . The sound collection control method according to claim 3 , wherein, in a state where the total number exceeds the maximum, the setting of the first collection beam sets a newest first collection beam in place of a previously set first or second collection beam, which is selectable by a user.

6 . The sound collection control method according to claim 1 , further comprising:

receiving a mute operation on the first collection beam or the second collection beam; and

muting the first collection beam or the second collection beam that receives the mute operation.

7 . A sound collection control method comprising:

recognizing a speaker from an image;

detecting a position of the recognized speaker;

setting a first collection beam based on the detected position of the recognized speaker;

recognizing a specific object other than the recognized speaker from the image;

detecting a position of the recognized specific object; and

setting a second collection beam based on the detected position of the recognized specific object,

wherein:

the first collection beam is a sound collection beam;

the recognizing of the speaker recognizes a plurality of speakers, including the speaker, from the image;

the detecting of the position of the recognized speaker detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker;

the setting of the first collection beam sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers;

the recognizing of the specific object recognizes a plurality of specific objects, including the specific object, from the image;

the detecting of the position of the recognized specific object detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object;

the setting of the second collection beam sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects;

a total number of the first collection beams and the second collection beams has a maximum; and

in a state where the total number exceeds the maximum, the setting of the first collection beam sets a newest first collection beam in place of a previously set first or second collection beam based on predefined priority.

8 . A sound collection apparatus comprising:

an array microphone; and

a hardware controller configured to execute a plurality of tasks, including:

a speaker recognizing task that recognizes a speaker from an image;

a speaker position detecting task that detects a position of the recognized speaker;

a first collection beam setting task that sets a first collection beam to the array microphone based on the detected position of the speaker;

an object recognizing task that recognizes a specific object that is not a person and that does not generate sound from the image;

an object position detecting task that detects a position of the recognized specific object; and

a second collection beam setting task that sets a second collection beam to the array microphone based on the detected position of the recognized specific object,

wherein the first collection beam is a sound collection beam, and

wherein, in a case where no person speaks, only the second collection beam is set without setting the first sound collection beam.

9 . The sound collection apparatus according to claim 8 , wherein the second collection beam includes a non-sound collection beam having a lower sensitivity than other directions.

10 . The sound collection apparatus according to claim 8 , wherein:

the speaker recognizing task recognizes a plurality of speakers, including the speaker, from the image,

the speaker position detecting task detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker,

the first collection beam setting task sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers,

the object recognizing task recognizes a plurality of specific objects, including the specific object, from the image,

the object position detecting task detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object,

the second collection beam setting task sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects, and

a total number of the first collection beams and the second collection beams has a maximum.

11 . The sound collection apparatus according to claim 10 , wherein, in a state where the total number exceeds the maximum, the first collection beam setting task does not set a newest first collection beam or sets the newest first collection beam in place of a previously set first collection beam.

12 . The sound collection apparatus according to claim 10 , wherein, in a state where the total number exceeds the maximum, the first collection beam setting task sets a newest first collection beam in place of a previously set first or second collection beam, which is selectable by a user.

13 . The sound collection apparatus according to claim 8 , further comprising a mute controller that:

receives a mute operation on the first collection beam or the second collection beam; and

mutes the first collection beam or the second collection beam that receives the mute operation.

14 . A sound collection apparatus comprising:

an array microphone; and

a hardware controller configured to execute a plurality of tasks, including:

a speaker recognizing task that recognizes a speaker from an image;

a speaker position detecting task that detects a position of the recognized speaker;

a first collection beam setting task that sets a first collection beam to the array microphone based on the detected position of the speaker;

an object recognizing task that recognizes a specific object other than the speaker from the image;

an object position detecting task that detects a position of the recognized specific object; and

a second collection beam setting task that sets a second collection beam to the array microphone based on the detected position of the recognized specific object,

wherein:

the first collection beam is a sound collection beam;

the speaker recognizing task recognizes a plurality of speakers, including the speaker, from the image;

the speaker position detecting task detects a plurality of positions of the recognized plurality of speakers, including the position of the recognized speaker;

the first collection beam setting task sets a plurality of first collection beams, including the first collection beam, based on the detected plurality of positions of the recognized plurality of speakers;

the object recognizing task recognizes a plurality of specific objects, including the specific object, from the image;

the object position detecting task detects a plurality of positions of the recognized plurality of specific objects, including the position of the recognized specific object;

the second collection beam setting task sets a plurality of second collection beams, including the second collection beam, based on the detected plurality of positions of the recognized plurality of specific objects;

a total number of the first collection beams and the second collection beams has a maximum; and

in a state where the total number exceeds the maximum, the first beam setting task sets a newest first collection beam in place of a previously set first or second collection beam based on predefined priority.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2023
From: OIZUMI, YOSHIFUMI
To: YAMAHA CORPORATION
Reel/Frame 064507/0801 →
Priority Claims (1)
JP 2022-134670 · Aug 26, 2022 · national
Continuity (1)
Related Publication 20240073596A1 · Feb 29, 2024
References Cited (16)
US 5940118A · Van Schyndel · 1999 [cited by applicant]
US 10341784B2 · Recker · 2019 [cited by examiner]
US 10499164B2 · Li · 2019 [cited by examiner]
US 11172122B2 · Welbourne · 2021 [cited by examiner]
US 11321866B2 · Kim · 2022 [cited by examiner]
US 11375309B2 · Hirose · 2022 [cited by examiner]
US 11778368B2 · Veselinovic · 2023 [cited by examiner]
US 12010490B1 · Delikaris Manias · 2024 [cited by examiner]
US 12288566B1 · Ganguly · 2025 [cited by examiner]
US 20150022636A1 · Savransky · 2015 [cited by examiner]
US 20200265860A1 · Mouncer · 2020 [cited by examiner]
US 20210120333A1 · Hirose · 2021 [cited by examiner]
US 20240296821A1 · Joshi · 2024 [cited by examiner]
JP 2021197658A · 2021 [cited by applicant]
Extended European Search Report issued in European Appln. No. 23193180.9, mailed Jan. 17, 2024. [cited by applicant]
Communication Pursuant to Article 94(3) EPC issued in European Appln. No. 23 193 180.9 mailed Nov. 14, 2025. [cited by applicant]