IP Library › Granted Patent US 10,565,435
Granted Patent B2
US 10,565,435 · App. 15/992,398 · Granted Feb 18, 2020

Apparatus and method for determining video-related emotion and method of generating data for learning video-related emotion

Inventors: Jee Hyun Park (Daejeon, KR); Jung Hyun Kim (Daejeon, KR); Yong Seok Seo (Daejeon, KR); Won Young Yoo (Daejeon, KR); Dong Hyuck Im (Daejeon, KR)
Assignee: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
G06K9/00302G06F16/71G06F16/7834G06K9/6267G06K9/66
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,565,435
App. No.
15/992,398
Granted
Feb 18, 2020
Kind
B2
Abstract

A method for determining a video-related emotion and a method of generating data for learning video-related emotions include separating an input video into a video stream and an audio stream; analyzing the audio stream to detect a music section; extracting at least one video clip matching the music section; extracting emotion information from the music section; tagging the video clip with the extracted emotion information and outputting the video clip; learning video-related emotions by using the at least one video clip tagged with the emotion information to generate a video-related emotion classification model; and determining an emotion related to an input query video by using the video-related emotion classification model to provide the emotion.

Claims (51)

1. A method of determining a video-related emotion, the method comprising:

separating an input video into a video stream and an audio stream;

analyzing the audio stream to detect a music section;

extracting at least one video clip matching the music section;

extracting emotion information from the music section;

tagging the video clip with the extracted emotion information and outputting the video clip;

learning video-related emotions by using the at least one video clip tagged with the emotion information to generate a video-related emotion classification model; and

determining an emotion related to an input query video by using the video-related emotion classification model to provide the emotion.

2. The method of claim 1 , wherein the extracting of the emotion information from the music section comprises:

determining whether the music section includes voice;

when voice is included in the music section, removing the voice from the music section and acquiring a music signal; and

extracting music-related emotion information from the acquired music signal.

3. The method of claim 1 , wherein the extracting of the at least one video clip matching the music section comprises:

selecting a video section corresponding to time information of the detected music section; and

separating the selected video section to generate at least one video clip.

4. The method of claim 1 , wherein the tagging of the video clip with the extracted emotion information and the outputting of the video clip comprise inserting the emotion information into a metadata area of the video clip and outputting the video clip.

5. The method of claim 1 , wherein the tagging of the video clip with the extracted emotion information and the outputting of the video clip comprise storing the emotion information in a separate file and outputting the video clip and the separate file.

6. The method of claim 1 , wherein the tagging of the video clip with the extracted emotion information and the outputting of the video clip comprise inputting the emotion information to an emotion information database.

7. A method of generating video-related emotion training data, the method comprising:

separating an input video into a video stream and an audio stream;

detecting music sections in the separated audio stream;

removing voice from the music sections to acquire music signals;

extracting music-related emotion information from the acquired music signals; and

tagging a corresponding video clip with music-related emotion information corresponding to each music section.

8. The method of claim 7 , wherein a plurality of video clips tagged with the music-related emotion information are provided as video-related emotion training data.

9. The method of claim 7 , further comprising detecting a video section matching each music section, and separating the detected video section to generate a video clip related to at least one scene.

10. The method of claim 7 , wherein the music emotion information tagged to the video clip is stored in a metadata area of the video clip, a separate file, or an emotion information database.

11. An apparatus for determining a video-related emotion, the apparatus comprising:

a processor, and

a memory configured to store at least one command executed by the processor,

wherein the at least one command includes:

a command to separate an input video into a video stream and an audio stream;

a command to detect music sections by analyzing the audio stream;

a command to extract at least one video clip matching the music sections;

a command to extract emotion information from the music sections;

a command to tag the video clip with the extracted emotion information and output the video clip;

a command to learn video-related emotions by using a plurality of video clips tagged with emotion information and generate a video-related emotion classification model; and

a command to determine an emotion related to an input query video by using the video-related emotion classification model and provide the emotion.

12. The apparatus of claim 11 , further comprising a database configured to store the plurality of video clips and the plurality of pieces of emotion information related to the plurality of video clips.

13. The apparatus of claim 11 , wherein the command to extract the emotion information from the music sections includes:

a command to determine whether the music sections include voice;

a command to, when voice is included in the music sections, remove the voice from the music sections and acquire music signals; and

a command to extract music-related emotion information from the acquired music signals.

14. The apparatus of claim 11 , wherein the command to extract the at least one video clip matching the music section includes:

a command to select video sections corresponding to time information of the detected music sections; and

a command to generate at least one video clip by separating the selected video sections.

15. The apparatus of claim 11 , wherein the command to extract the at least one video clip matching the music sections includes a command to detect a video section matching each music section and generate a video clip related to at least one scene by separating the detected video section.

16. The apparatus of claim 11 , wherein the emotion information tagged to the video clip is stored in a metadata area of the video clip or in a separate file.

17. The apparatus of claim 11 , wherein the command to tag the video clip with the extracted emotion information and output the video clip includes:

a command to determine a position in a metadata area to which the emotion information will be inserted by parsing and analyzing the metadata area of the input video; and

a command to insert the emotion information to the determined position in the metadata area.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2018
From: PARK, JEE HYUN; KIM, JUNG HYUN; SEO, YONG SEOK; YOO, WON YOUNG; IM, DONG HYUCK
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 045932/0512 →
Priority Claims (1)
KR 10-2018-0027637 · Mar 8, 2018 · national
Continuity (1)
Related Publication 20190278978A1 · Sep 12, 2019