IP Library › Granted Patent US 12,322,412
Granted Patent B2
US 12,322,412 · App. 17/953,883 · Granted Jun 3, 2025

Method for providing video and electronic device supporting the same

Inventors: Donghwan Seo (Suwon-si, KR); Sungoh Kim (Suwon-si, KR); Dasom Lee (Suwon-si, KR); Sanghun Lee (Suwon-si, KR); Sungsoo Choi (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L25/57G06V20/46G10L21/028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,412
App. No.
17/953,883
Filed
Sep 27, 2022
Granted
Jun 3, 2025
Kind
B2
Art Unit
2656
USPC
704/270
Abstract

An electronic device is provided. The electronic device includes a memory, and at least one processor electrically connected to the memory, wherein the at least one processor is configured to obtain a video including an image and an audio, obtain information on at least one object included in the image from the image, obtain a visual feature of the at least one object, based on the image and the information on the at least one object, obtain a spectrogram of the audio, obtain an audio feature of the at least one object from the spectrogram of the audio, combine the visual feature and the audio feature, obtain, based on the combined visual feature and audio feature, information on a position of the at least one object the information indicating the position of the at least one object in the image, obtain an audio part corresponding to the at least one object in the audio, based on the combined visual feature and audio feature, and store, in the memory, the information on the position of the at least one object and the audio part corresponding to the at least one object.

Claims (67)

1. An electronic device comprising:

at least one processor including processing circuitry; and

memory storing instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to:

obtain a video including an image and an audio,

obtain information on at least one object included in the image from the image, the information on the at least one object including a map in which the at least one object included in the image is masked,

obtain a visual feature of the at least one object, based on the image and the information on the at least one object,

obtain a spectrogram of the audio,

obtain an audio feature of the at least one object from the spectrogram of the audio,

combine the visual feature and the audio feature,

obtain, based on the combined visual feature and audio feature, information on a position of the at least one object, the information indicating the position of the at least one object in the image, by obtaining a mask having a value of possibility that each pixel represents the at least one object, based on the combined visual feature and audio feature,

obtain an audio part corresponding to the at least one object in the audio, based on the combined visual feature and audio feature,

store, in the memory, the information on the position of the at least one object and the audio part corresponding to the at least one object,

display the video, and

while displaying the video, displaying an audio volume indicator and a time interval indicator for the audio part, the time interval indicator having a first portion displayed having a first visual attribute and a remaining portion of the time interval indicator displayed having a second visual attribute, the first visual attribute being a same as a visual attribute of the at least one object while the audio part corresponding to the at least one object is output, wherein the first visual attribute is different from the second visual attribute.

2. The electronic device of claim 1 , wherein the instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to combine the visual feature and the audio feature by performing an add operation, a multiplication operation, or a concatenation operation for the visual feature and the audio feature.

3. The electronic device of claim 1 , wherein the instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to:

obtain an image having a value of possibility that each pixel represents an audio corresponding to the at least one object, based on the combined visual feature and audio feature, and

obtain an audio part corresponding to the at least one object in the audio, based on the obtained image.

4. The electronic device of claim 3 , wherein the instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to obtain an audio part corresponding to the at least one object in the audio, based on performing an AND operation for the spectrogram of the audio and the obtained image.

5. The electronic device of claim 1 ,

wherein the instructions that, when executed by the at least one processor individually or collectively, further cause the electronic device to perform training to generate an artificial intelligence model, and

wherein, to perform the training, the instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to:

obtain multiple videos,

obtain information on at least one object from each of images of the multiple videos,

obtain a visual feature of each of the at least one object,

obtain a spectrogram of an audio corresponding to the at least one object,

obtain an audio feature of each of the at least one object,

combine the audio feature and the visual feature for each of the at least one object,

obtain, based on the combined visual feature and audio feature, information on a position of the at least one object, the information indicating the position of the at least one object in each of the images, and

obtain an audio part corresponding to the at least one object in the audio, based on the combined visual feature and audio feature.

6. The electronic device of claim 5 , wherein the instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to generate an artificial intelligence model related to a segmentation artificial intelligence network, based on images of the multiple videos, and an image part of the at least one object in each of the images as a ground truth.

7. The electronic device of claim 5 , wherein, to obtain the information on the position of the at least one object, the instructions that, when executed by the at least one processor individually or collectively, further cause the electronic device to correct an error such that a distance between the audio feature and the visual feature is minimized by using a loss function based on metric learning.

8. The electronic device of claim 5 , wherein, to obtain the audio part corresponding to the at least one object in the audio, the instructions that, when executed by the at least one processor individually or collectively, further cause the electronic device to correct an error such that by using a loss function and for each of the at least one object, a pixel-specific distance between a feature obtained by combining the visual feature and the audio feature, and the spectrogram of the audio corresponding to the at least one object is minimized.

9. A method for providing a video by an electronic device, the method comprising:

obtaining a video including an image and an audio;

obtaining information on at least one object included in the image from the image, the information on the at least one object including a map in which the at least one object included in the image is masked;

obtaining a visual feature of the at least one object, based on the image and the information on the at least one object;

obtaining a spectrogram of the audio;

obtaining an audio feature of the at least one object from the spectrogram of the audio;

combining the visual feature and the audio feature;

obtaining, based on the combined visual feature and audio feature, information on a position of the at least one object, the information indicating the position of the at least one object in the image, by obtaining a mask having a value of possibility that each pixel represents the at least one object, based on the combined visual feature and audio feature;

obtaining an audio part corresponding to the at least one object in the audio, based on the combined visual feature and audio feature;

storing, in a memory of the electronic device, the information on the position of the at least one object and the audio part corresponding to the at least one object,

display the video; and

while displaying the video, displaying an audio volume indicator and a time interval indicator for the audio part, the time interval indicator having a first portion displayed having a first visual attribute and a remaining portion of the time interval indicator is displayed having a second visual attribute, the first visual attribute being a same as a visual attribute of the at least one object while the audio part corresponding to the at least one object is output, wherein the first visual attribute is different from the second visual attribute.

10. The method of claim 9 , wherein the combining of the visual feature and the audio feature comprises combining the visual feature and the audio feature by performing an add operation, a multiplication operation, or a concatenation operation for the visual feature and the audio feature.

11. The method of claim 10 , wherein the obtaining of the audio part corresponding to the at least one object comprises:

obtaining an image having a value of possibility that each pixel represents an audio corresponding to the at least one object, based on the combined visual feature and audio feature; and

obtaining an audio part corresponding to the at least one object in the audio, based on the obtained image.

12. The method of claim 11 , wherein the obtaining of the audio part corresponding to the at least one object in the audio comprises obtaining an audio part corresponding to the at least one object in the audio, based on performing an AND operation for the spectrogram of the audio and the obtained image.

13. The method of claim 9 , further comprising:

performing training to generate an artificial intelligence model,

wherein the performing of the training comprises:

obtaining multiple videos;

obtaining information on at least one object from each of images of the multiple videos;

obtaining a visual feature of each of the at least one object;

obtaining a spectrogram of an audio corresponding to the at least one object;

obtaining an audio feature of each of the at least one object;

combining the audio feature and the visual feature for each of the at least one object;

obtaining, based on the combined visual feature and audio feature, information on a position of the at least one object, the information indicating the position of the at least one object in each of the images; and

obtaining an audio part corresponding to the at least one object in the audio, based on the combined visual feature and audio feature.

14. The method of claim 13 , wherein the obtaining of the information on the at least one object comprises generating an artificial intelligence model related to a segmentation artificial intelligence network, based on images of the multiple videos, and an image part of the at least one object in each of the images as a ground truth.

15. The method of claim 13 , wherein the obtaining of the information on the position of the at least one object further comprises correcting an error such that a distance between the audio feature and the visual feature is minimized by using a loss function based on metric learning.

16. The method of claim 9 , wherein the obtaining of the audio part corresponding to the at least one object in the audio further comprises correcting an error such that by using a loss function and for each of the at least one object, a pixel-specific distance between a feature obtained by combining the visual feature and the audio feature, and the spectrogram of the audio corresponding to the at least one object is minimized.

17. The method of claim 9 , further comprising:

while displaying the video, displaying a second object, the time interval indicator having a third portion displayed having a third visual attribute, the second object being displayed having the third visual attribute.

18. The method of claim 9 , wherein the first visual attribute comprises at least one of an effect, a color, or a highlight.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2022
From: SEO, DONGHWAN; KIM, SUNGOH; LEE, DASOM; LEE, SANGHUN; CHOI, SUNGSOO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061228/0261 →
Priority Claims (1)
KR 10-2021-0131180 · Oct 1, 2021 · national
Continuity (2)
Continuation PCTKR2022013980 · Sep 19, 2022
Related Publication 20230124111A1 · Apr 20, 2023
References Cited (39)
US 11996900B2 · Cella · 2024 [cited by examiner]
US 20020078446A1 · Dakss · 2002 [cited by examiner]
US 20090097670A1 · Jeong et al. · 2009 [cited by applicant]
US 20110022361A1 · Sekiya et al. · 2011 [cited by applicant]
US 20120099732A1 · Msser · 2012 [cited by applicant]
US 20120128165A1 · Visser et al. · 2012 [cited by applicant]
US 20120316869A1 · Xiang et al. · 2012 [cited by applicant]
US 20130141439A1 · Kryzhanovsky et al. · 2013 [cited by applicant]
US 20140314391A1 · Kim et al. · 2014 [cited by applicant]
US 20160054895A1 · Lee et al. · 2016 [cited by applicant]
US 20160071526A1 · Wingate et al. · 2016 [cited by applicant]
US 20160180865A1 · Citerin et al. · 2016 [cited by applicant]
US 20170265016A1 · Oh et al. · 2017 [cited by applicant]
US 20180047407A1 · Mitsufuji · 2018 [cited by applicant]
US 20190222798A1 · Honma et al. · 2019 [cited by applicant]
US 20190253828A1 · Honma et al. · 2019 [cited by applicant]
US 20200143838A1 · Peleg · 2020 [cited by examiner]
US 20200288256A1 · Jung et al. · 2020 [cited by applicant]
US 20210096810A1 · Kim et al. · 2021 [cited by applicant]
US 20210174817A1 · Grauman · 2021 [cited by examiner]
US 20210201933A1 · Kang · 2021 [cited by applicant]
US 20210319321A1 · Krishnamurthy · 2021 [cited by examiner]
US 20210350135A1 · Salamon · 2021 [cited by examiner]
JP 2001242898A · 2001 [cited by applicant]
JP 2011027825A · 2011 [cited by applicant]
JP 2013545137A · 2013 [cited by applicant]
KR 1020090037692A · 2009 [cited by applicant]
KR 1020130084298A · 2013 [cited by applicant]
KR 101373020B1 · 2014 [cited by applicant]
KR 1020190118994A · 2019 [cited by applicant]
KR 1020200020590A · 2020 [cited by applicant]
KR 1020200054344A · 2020 [cited by applicant]
KR 1020210022600A · 2021 [cited by applicant]
KR 1020210043958A · 2021 [cited by applicant]
WO 2017208821A1 · 2017 [cited by applicant]
Weakly-Supervised Audio-Visual Sound Source Detection, Mar. 25, 2021. [cited by applicant]
International Search Report dated Dec. 23, 2022, issued in International Patent Application No. PCT/KR2022/013980. [cited by applicant]
Relja et al., Objects that Sound, Jul. 25, 2018. [cited by applicant]
Ariel et al., Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation, Aug. 2018. [cited by applicant]