Computer system for detecting target video from sports video based on voice recognition and method of the same
A computer system for detecting a target video from a sports video based on voice recognition is configured to convert a relay voice corresponding to a sports video to text; and to detect a target video related to a preset event from the sports video based on the text.
1 . A method of detecting a target video within a sports video, the method comprising:
extracting, by a processor, chunk videos from the sports video, wherein each of the chunk videos is based on a predetermined time unit defined by a unit of time;
generating a partial video from at least one of the chunk videos within a sports video of a sports game based on text relay data of the sports game by recognizing a score board from each of the chunk videos and mapping the text relay data to the chuck videos, wherein the partial video includes an audio signal and a noise signal;
extracting a relay voice from the audio signal of the partial video by removing the noise signal using a noise filter;
converting the relay voice of the partial video to text;
excluding the partial video as the target video without recognizing an action of at least one object from the partial video, when the text in the partial video does not include a preset number of keywords;
determining that the partial video is the target video when the text includes the preset number of keywords; and
in response to the determining that the text includes the preset number of keywords, confirming that the partial video is the target video by:
recognizing the action of the at least one object from the partial video; and
determining that the action of the at least one object corresponds to a preset motion related to an event.
2 . The method of claim 1 , wherein the detecting of the target video comprises determining the partial video is the target video when the text includes at least one keyword related to a preset event.
3 . The method of claim 2 , wherein the detecting of the partial video as the target video comprises:
recognizing an action of at least one object from the partial video when the text includes the at least one keyword; and
determining the partial video is the target video when the action is related to the preset event.
4 . The method of claim 1 , wherein the converting of the relay voice corresponding to the partial video to the text comprises:
recognizing an action of at least one object from the partial video; and
converting the relay voice corresponding to the partial video to the text when the action is related to the preset event.
5 . The method of claim 1 , wherein the generating of the partial video comprises:
detecting the partial video from a location in the sports video that is determined based on the text relay data,
wherein the text relay data includes identification information and timestamp information for each event that occurs in the sports game.
6 . The method of claim 1 wherein the detecting of the partial video comprises detecting the partial video related to a play involving an out from a baseball game.
7 . A non-transitory computer-readable recording medium storing instructions that, when executed by the processor, cause the processor to perform the method of claim 1 .
8 . A computer system comprising:
a memory; and
a processor configured to connect to the memory and to execute at least one instruction stored in the memory,
wherein the processor is configured to:
extract chunk videos from a sports video of a sports game, wherein each of the chunk videos is based on a predetermined time unit defined by a unit of time;
generate a partial video from at least one of the chunk videos based on text relay data of the sports game by recognizing a score board from each of the chuck videos and mapping the text relay data to the chunk videos, wherein the partial video includes an audio signal and a noise signal;
extract a relay voice from the audio signal of the partial video by removing the noise signal using a noise filter;
convert the relay voice of the partial video to text, and
exclude the partial video as the target video without recognizing an action of at least one object from the partial video, when the text in the partial video does not include a preset number of keywords;
determine that the partial video is the target video when the text includes the preset number of keywords;
in response to the determining that the text includes the preset number of keywords, confirming that the partial video is the target video by:
recognizing the action of the at least one object from the partial video; and
determining that the action of the at least one object corresponds to a preset motion related to an event.
9 . The computer system of claim 8 , wherein the processor is configured to determine that the partial video is the target video when the text includes at least one keyword related to a preset event.
10 . The computer system of claim 9 , wherein the processor is configured to recognize an action of at least one object from the partial video when the text includes the at least one keyword, and to determine that the partial video is the target video when the action is related to the preset event.
11 . The computer system of claim 8 , wherein the processor is configured to recognize an action of at least one object from the partial video, and to convert the relay voice corresponding to the partial video to the text when the action is related to the preset event.
12 . The computer system of claim 8 , wherein the processor is configured to detect the partial video from a location in the sports video that is determined based on the text relay data, wherein the text relay data includes identification information and timestamp information for each event that occurs in the sports video.
13 . The computer system of claim 8 , wherein the processor is configured to detect the partial video related to a play involving an out from a baseball game.
14 . The method of claim 1 , wherein the partial video is detected as the target video by:
determining whether the text includes at least one keyword related to the event;
recognizing the action of at least one object from the partial video in response to the determination that the text includes the keyword related to the event;
determining that the action corresponds to the preset motion related to the event; and
confirming the partial video as the target video.
15 . The method of claim 1 , wherein the text relay data includes identification information and timestamp information for each event that occurs in the sports game.
16 . The method of claim 1 , further comprising verifying that the partial video is the target video after determining that the partial video is the target video.
17 . The computer system of claim 8 , wherein the process verifies that the partial video is the target video after determining that the partial video is the target video.