IP Library Granted Patent US 11,462,235
Granted Patent B2
US 11,462,235 · App. 17/176,322 · Granted Oct 4, 2022

Surveillance camera system for extracting sound of specific area from visualized object and operating method thereof

Inventor: Yongmin Kang (Seongnam-Si, KR)
Assignee: HANWHA TECHWIN CO., LTD.
G10L25/57G10L21/10G10L25/78G10L25/90H04N5/76H04N7/18H04R1/406H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,462,235
App. No.
17/176,322
Granted
Oct 4, 2022
Kind
B2
Abstract

A camera system for extracting a sound of a specific area includes: a camera device configured to receive video signals and audio signals from an area; at least one memory configured to store information about the area including data corresponding to the video signals and the audio signals from the area; and a processor configured to change an audio zooming point of the camera device from a first point, in an image of the area captured by the camera device, to a second point based on the information about the area, and perform a beam-forming on an audio signal corresponding to the second point.

Claims (37)

1. A camera system for extracting sounds of a specific area, comprising:

a camera device configured to receive video signals and audio signals from an area;

at least one memory configured to store information about the area including data corresponding to the video signals and the audio signals from the area; and

a processor configured to change an audio zooming point of the camera device from a first point, in an image of the area captured by the camera device, to a second point based on the information about the area, and perform a beam-forming on an audio signal corresponding to the second point,

wherein the processor is configured to implement:

a data collector configured to generate a sound-based heatmap for the area based on the information about the area;

a user area selector configured to select a user-designated area to perform the beam-forming thereon, and select the first point in the user-designated area as the audio zooming point;

a calculator configured to select the second point as the audio zooming point in the user-designated area, based on the first point and the sound-based heatmap; and

a corrector configured to perform the beam-forming in a direction corresponding to the second point.

2. The system of claim 1 , wherein the audio signal on which the beam-forming is performed is limited to a human voice signal.

3. The system of claim 2 , wherein the processor is further configured to determine the audio signal as the human voice signal based on determining that the audio signal includes a vowel among language components.

4. The system of claim 1 , wherein the sound-based heatmap displays positions of a pitch range corresponding to human voice signals by a sound source localization based on the video signals and the audio signals received by the camera device.

5. The system of claim 1 , wherein the area is split into a plurality of areas, and the user area selector is configured to select at least one of the plurality of areas as the user-designated area.

6. The system of claim 5 , wherein the user area selector is configured to specify an object corresponding to a target of the beam-forming based on a motion detection algorithm and/or a face recognition algorithm.

7. The system of claim 1 , wherein both the first and second points are displayed on the image of the area.

8. The system of claim 7 , wherein the second point is selected from a plurality of second points displayed by the camera device.

9. The system of claim 1 , wherein the camera device comprises:

a camera configured to collect the video signals having a specific viewing angle; and

a microphone array comprising a plurality of microphones spaced apart from one another, and configured to collect the audio signals.

10. The system of claim 9 , wherein the camera and the microphone array are disposed on different surfaces of the camera device.

11. The system of claim 1 , wherein the memory is further configured to manage data output from the calculator, and store a measurement time and a date corresponding the video signals and the audio signals received by the camera device together with the data output from the calculator.

12. A method for operating a camera system for extracting sounds of a specific area, the method comprising:

receiving video signals and audio signals from an area;

storing information about the area including the video signals and the audio signals;

generating a sound-based heatmap for the area based on the information about the area;

selecting a user-designated area to perform beam-forming thereon and selecting a first point in the user-designated area as an audio zooming point;

changing the audio zooming point to a second point in the user-designated area based on the first point and the sound-based heatmap; and

performing the beam-forming in a direction corresponding to the second point.

13. The method of claim 12 , wherein an audio signal on which the beam-forming is performed is limited to a human voice signal.

14. The method of claim 13 , further comprising:

determining the audio signal as the human voice signal based on determining that the audio signal includes a vowel among language components.

15. The method of claim 12 , wherein the sound-based heatmap displays positions of a pitch range corresponding to human voice signals by a sound source localization based on the video signals and the audio signals.

16. The method of claim 12 , wherein the area is split into a plurality of areas, and the selecting the user-designated area comprises selecting at least one of the plurality of split areas as the user-designated area.

17. The method of claim 16 , wherein the selecting the user-designated area further comprises:

specifying an object corresponding to a target of the beam-forming based on a motion detection algorithm and/or a face recognition algorithm.

18. The method of claim 12 , wherein both the first and second audio zooming points are displayed on an image of the area.

19. The method of claim 18 , wherein the second point is selected from a plurality of second points displayed by a camera device.

Assignments (2)
CHANGE OF NAME Recorded Aug 10, 2023
From: HANWHA TECHWIN CO., LTD.
To: HANWHA VISION CO., LTD.
Reel/Frame 064549/0075 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2021
From: KANG, YONGMIN
To: HANWHA TECHWIN CO., LTD.
Reel/Frame 055269/0802 →
Priority Claims (2)
KR 10-2018-0095425 · Aug 16, 2018 · national
KR 10-2019-0090302 · Jul 25, 2019 · national
Continuity (2)
Continuation PCTKR2019010094 · Aug 9, 2019
Related Publication 20210201933A1 · Jul 1, 2021