DEVICE AND METHOD FOR QUESTION ANSWERING
A device receives a question, which can be related to an event, from a user through a user interface, and at least one hardware processor generates a first summary of at least one dialogue of the event; and output a result obtained by processing the summary, the question and possible answers to the question.
1 . A device comprising:
a user interface configured to receive, from a user, a question about an event in at least one video scene; and
at least one hardware processor configured to:
generate a first summary of at least one dialogue of the at least one video scene; and
output a first result obtained by processing the first summary, the question and possible answers to the question.
2 . The device of claim 1 , wherein the processing is performed using a transformer.
3 . The device of claim 1 , wherein the first result is a score for each possible answer or the possible answer with the highest score.
4 . The device of claim 1 , wherein the at least one hardware processor is further configured to:
generate a second summary of video information related to the at least one video scene;
process the second summary, the question and the possible answers to the question to obtain a second result; and
output a final result obtained by processing the first result and the second result.
5 . The device of claim 4 , wherein the at least one hardware processor is configured to process the first result and the second result using a fusion mechanism.
6 . The device of claim 5 , wherein the fusion mechanism is a modality attention mechanism that weights the first result and the second result.
7 . The device of claim 1 , wherein the dialogue obtained from at least one video or through capture using a microphone and an audio-to-text function.
8 . The device of claim 1 , wherein the device is one of a home assistant and a TV.
9 . (canceled)
10 . A method comprising:
receiving, from a user through a user interface, a question about an event in at least one video scene;
generating, by at least one hardware processor, a first summary of at least one dialogue of the at least one video scene; and
outputting, by the at least one hardware processor, a first result, the first result obtained by processing the first summary, the question and possible answers to the question.
11 . The method of claim 10 , wherein the processing is performed using a transformer.
12 . The method of claim 10 , wherein the first result is a score for each possible answer or the possible answer with the highest score.
13 . The method of claim 10 , further comprising:
generating, by the at least one hardware processor, a second summary of information related to the at least one video scene;
processing, by the at least one hardware processor, the second summary, the question and the possible answers to the question to obtain a second result; and
outputting, by the at least one hardware processor, a final result obtained by processing the first result and the second result.
14 . The method of claim 13 , wherein the at least one hardware processor is configured to process the first result and the second result using a fusion mechanism.
15 . The method of claim 14 , wherein the fusion mechanism is a modality attention mechanism that weights the first result and the second result.
16 . A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one hardware processor perform the method of claim 10 .