IP Library Granted Patent US 12,256,191
Granted Patent B2
US 12,256,191 · App. 17/964,818 · Granted Mar 18, 2025

Electronic apparatus generating a sound of a plurality of channels and controlling method thereof

Inventors: Nabil Ibtehaz (Dacca, BD); Golam Rahman Chowdhury (Dacca, BD); Md Abdullah Al Hadi (Dacca, BD)
Assignee: Samsung Electronics Co., Ltd.
H04R1/326G06F3/16G10L19/00H04S5/00H04S7/30H04S2400/11H04S2400/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,256,191
App. No.
17/964,818
Granted
Mar 18, 2025
Kind
B2
Abstract

An electronic apparatus and/or a controlling method are provided. The electronic apparatus may include a camera for photographing an image, a microphone for receiving an input of a sound of a first channel, and a processor for generating sounds of a plurality of channels based on the input sound, wherein the processor is configured to identify an object and the location of the object from the photographed image, classify the input sound based on an audio source, and allot the sound to the corresponding identified object, copy the classified sound and generate sounds of two channels, adjust characteristics of the generated sounds of two channels based on the audio source allotted to the identified object and the location of the identified object, and mix the sounds of two channels wherein the characteristics were adjusted according to the audio source and generate a stereo sound of two channels.

Claims (58)

1. An electronic apparatus comprising:

a camera configured to capture an image;

a display:

an input interface comprising circuitry;

a microphone configured to receive an input of a sound of a first channel;

a Time of Flight (ToF) sensor; and

a processor, comprising processing circuitry, configured to cause generation of sounds of a plurality of channels based on the input sound and to:

identify an object and the location of the object from the image by the ToF sensor,

classify the input sound based on an audio source, and allot the classified sound to the corresponding identified object,

copy the classified sound and control to generate sounds of at least two channels,

control to adjust characteristics of the generated sounds of at least two channels based on the audio source allotted to the identified object and the location of the identified object using at least one of sound panning to a predetermined location, time delay, phase delay, strength adjustment, amplitude adjustment and spectral modification, the location of the identified object being sensed by the ToF sensor,

control to mix the sounds of at least two channels wherein the characteristics were adjusted according to the audio source and generate a stereo sound of at least two channels,

based on failing to identify an object corresponding to the classified sound, control the display to display a predetermined indicator and the identified object on the classified sound, and

match the classified sound on which the predetermined indicator is displayed to the identified object based on a user's command input through the input interface.

2. The electronic apparatus of claim 1 ,

wherein the processor is configured to:

separate the input sound into respective sounds, identify audio sources corresponding to the respective separated sounds based on frequency characteristics, and classify the respective sounds based on the identified audio sources.

3. The electronic apparatus of claim 1 ,

wherein the processor is configured to:

identify the object and the location of the object based on an image processing artificial intelligence model, and adjust the characteristics of the generated sounds of two channels based on a sound processing artificial intelligence model.

4. An electronic apparatus comprising:

a camera configured to capture an image;

a display:

an input interface comprising circuitry:

a microphone configured to receive an input of a sound;

a Time of Flight (ToF) sensor and

a processor, comprising processing circuitry, configured to cause generation of sounds of a plurality of channels based on the input sound and to:

identify an object and the location of the object from the image,

classify the input sound based on an audio source, and allot the sound to the corresponding identified object,

extract a base sound based on the input sound, assume a rear sound, and cluster the allotted sound based on the location of the identified object,

control to adjust characteristics of the extracted base sound, the assumed rear sound, and the clustered sound using at least one of sound panning to a predetermined location, time delay, phase delay, strength adjustment, amplitude adjustment and spectral modification, the location of the identified object being sensed by the ToF sensor,

allot the sounds wherein the characteristics were adjusted to respective channels according to the audio sources and generate a surround sound,

based on failing to identify an object corresponding to the classified sound, control the display to display a predetermined indicator and the identified object on the classified sound, and

match the classified sound on which the predetermined indicator is displayed to the identified object based on a user's command input through the input interface.

5. The electronic apparatus of claim 4 ,

wherein the processor is configured to:

assume a sound other than the sound allotted to the identified object in the image among the input sound as the rear sound.

6. The electronic apparatus of claim 4 ,

wherein the processor is configured to:

divide the image into predetermined areas, and cluster sounds allotted to objects included in the same area among the respective divided areas in the same group.

7. The electronic apparatus of claim 4 ,

wherein the processor is configured to:

perform low-pass filtering of the input sound and extract the base sound.

8. A controlling method of an electronic apparatus, the method comprising:

capturing an image, and receiving an input of a sound; and

generating sounds of a plurality of channels based on the input sound,

wherein the generating the sounds of a plurality of channels comprises:

identifying an object and the location of the object from the image by ToF sensor of the electronic apparatus;

classifying the input sound based on an audio source, and allotting the classified sound to the corresponding identified object;

copying the classified sound and generating sounds of at least two channels;

adjusting characteristics of the generated sounds of at least two channels based on the audio source allotted to the identified object and the location of the identified object by method(s) using at least one of sound panning to a predetermined location, time delay, phase delay, strength adjustment, amplitude adjustment and spectral modification, the location of the identified object being sensed by the ToF sensor; and

mixing the sounds of at least two channels wherein the characteristics were adjusted according to the audio source and generating a stereo sound of at least two channels,

wherein the generating the sounds of a plurality of channels comprises:

based on failing to identify an object corresponding to the classified sound, displaying a predetermined indicator and the identified object on the classified sound, and

matching the sound on which the predetermined indicator is displayed to the identified object according to a user's command input.

9. The controlling method of an electronic apparatus of claim 8 ,

wherein the generating the sounds of a plurality of channels comprises:

separating the input sound into respective sounds, identifying audio sources corresponding to the respective separated sounds based on frequency characteristics, and classifying the respective sounds based on the identified audio sources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: IBTEHAZ, NABIL; CHOWDHURY, GOLAM RAHMAN; HADI, MD ABDULLAH AL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061402/0569 →
Priority Claims (1)
KR 10-2021-0162574 · Nov 23, 2021 · national
Continuity (2)
Continuation PCTKR2022012729 · Aug 25, 2022
Related Publication 20230164482A1 · May 25, 2023
References Cited (26)
US 9621122B2 · Kim et al. · 2017 [cited by applicant]
US 10178490B1 · Sheaffer et al. · 2019 [cited by applicant]
US 20130106997A1 · Kim et al. · 2013 [cited by applicant]
US 20130272527A1 · Oomen et al. · 2013 [cited by applicant]
US 20140211969A1 · Kim et al. · 2014 [cited by applicant]
US 20170188176A1 · Jang · 2017 [cited by applicant]
US 20180262747A1 · Meirlaen · 2018 [cited by applicant]
US 20210218878A1 · Kim · 2021 [cited by examiner]
US 20210329405A1 · Eubank et al. · 2021 [cited by applicant]
US 20220386061A1 · Shi · 2022 [cited by applicant]
JP 2014505420A · 2014 [cited by applicant]
KR 1020120130496 · 2012 [cited by applicant]
KR 1020130045553A · 2013 [cited by applicant]
KR 1020140096774A · 2014 [cited by applicant]
KR 1020170058839A · 2017 [cited by applicant]
KR 1020210131422A · 2021 [cited by applicant]
WO WO2021197020A1 · 2021 [cited by applicant]
“How do you guys convert your mono signals into stereo?” Reddit, 2017, www.reddit.com/r/AdvancedProduction/comments/5wkirq/how_do_you_guys_convert_your_mono_signals_into/. (Year: 2017). [cited by examiner]
English machine translation of KR20120130496 (Year: 2012). [cited by examiner]
Ross Girshick, “Fast R-CNN”, Proceedings of the IEEE international conference on computer vision, 2015, pp. 1440-1448. [cited by applicant]
Joseph Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection”, Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 779-788, at URL: http://pjreddie.com/yolo/. [cited by applicant]
Carsten Klein, “Separating mixed signals with Independent Component Analysis”, Towards Data Science, Apr. 20, 2019, 10 pp., at URL: https://towardsdatascience.com/separating-mixed-signals-with-independent-component-anal… [cited by applicant]
Jing Zhu et al., “Learning Object-Specific Distance From a Monocular Image”, Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 3839-3848. [cited by applicant]
Byung-Hak Kim et al., “LumièreNet: Lecture Video Synthesis from Audio”, arXiv preprint arXiv:1907.02253, 2019, pp. 1-9. [cited by applicant]
PCT Search Report dated Dec. 19, 2022 for PCT Application No. PCT/KR2022/012729. [cited by applicant]
PCT Written Opinion dated Dec. 19, 2022 for PCT Application No. PCT/KR2022/012729. [cited by applicant]