IP Library › Granted Patent US 12,050,632
Granted Patent B2
US 12,050,632 · App. 17/270,514 · Granted Jul 30, 2024

Question answering apparatus and method

Inventors: Byoung-Tak Zhang (Seoul, KR); Seongho Choi (Seoul, KR); Kyoung-Woon On (Seoul, KR); Yu-Jung Heo (Anyang-si, KR); You Won Jang (Seoul, KR); Ahjeong Seo (Sejong-si, KR); Seungchan Lee (Seoul, KR); Minsu Lee (Seongnam-si, KR)
Assignee: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
G06F16/3329G06F16/7328G06F16/783G06F40/126G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,050,632
App. No.
17/270,514
Granted
Jul 30, 2024
Kind
B2
Abstract

A question answering method that is performed by a question answering apparatus includes: receiving a data set including video content and question-answer pairs; generating input time-series sequences from the video content of the input data set and also generating a question-answer time-series sequence from the question-answer pair of the input data set; calculating weights by associating the input time-series sequence with the question-answer time-series sequence and also calculating first result values by performing operations on the calculated weights and the input time-series sequences; calculating second result values by paying attention to portions of the input time-series sequences that are directly related to characters appearing in questions and answers; and calculating third result values by concatenating the time-series sequences, the first result values, the second result values, and Boolean flags and selecting a final answer based on the third result values.

Claims (30)

1. A question answering method that is performed by a question answering apparatus, the question answering method comprising:

a step of receiving a data set including video content and question-answer pairs based on a user input;

a preprocessing step of generating input time-series sequences from the video content of the input data set and also generating a question-answer time-series sequence from the question-answer pair of the input data set;

a step of calculating weights by associating the input time-series sequence with the question-answer time-series sequence and also calculating first result values by performing operations on the calculated weights and the input time-series sequences;

a step of calculating second result values by paying attention to portions of the input time-series sequences that are directly related to characters appearing in questions and answers;

a step of calculating third result values by concatenating the time-series sequences, the first result values, the second result values, and Boolean flags and selecting a final answer based on the third result values; and

a step of outputting the final answer,

wherein the preprocessing step further comprises:

a step of generating time-series data by concatenating the pieces of data included in the data set in sequence;

a step of generating a feature vector including a related character by extracting a word vector and an image feature vector from the time-series data and concatenating the extracted vectors with related character information of the time-series data as an one-hot vector; and

a step of generating a time-series sequence having a contextual flow by inputting the feature vector including a related character to a bidirectional Long/Short Term Memory (bi-LSTM) model.

2. The question answering method of claim 1 , wherein the data set comprises question-answer pairs, a script in which utterers are indicated, visual metadata (behaviors and emotions), and visual bounding boxes.

3. The question answering method of claim 1 , wherein the step of calculating second result values comprises a step of calculating third result values using dot-product attention and multi-head attention.

4. A question answering apparatus comprising:

a storage unit configured to store a program that performs question answering; and

a control unit including at least one processor;

wherein when a data set including video content and question-answer pairs is received by executing the program based on a user input, the control unit:

generates input time-series sequences from the video content of the input data set, and also generates a question-answer time-series sequence from the question-answer pair of the input data set;

calculates weights by associating the input time-series sequence with the question-answer time-series sequence, and also calculates first result values by performing operations on the calculated weights and the input time-series sequences;

calculates second result values by paying attention to portions of the input time-series sequences that are directly related to characters appearing in questions and answers;

calculates third result values by concatenating the time-series sequences, the first result values, the second result values, and Boolean flags, and selects a final answer based on the third result values; and

outputs the final answer,

wherein when generating the input time-series sequences and the question-answer time-series sequence from the input data set, the control unit:

generates time-series data by concatenating the pieces of data included in the data set in sequence;

generates a feature vector including a related character by extracting a word vector and an image feature vector from the time-series data and concatenating the extracted vectors with related character information of the time-series data as an one-hot vector; and

generates a time-series sequence having a contextual flow by inputting the feature vector including a related character to a bidirectional Long/Short Term Memory (bi-LSTM) model.

5. The question answering apparatus of claim 4 , wherein the data set comprises question-answer pairs, a script in which utterers are indicated, visual metadata (behaviors and emotions), and visual bounding boxes.

6. The question answering apparatus of claim 4 , wherein when calculating second result values, the control unit calculates third result values using dot-product attention and multi-head attention.

7. A non-transitory computer-readable storage medium having stored thereon a program that performs the method set forth in claim 1 .

8. A computer program that is executed by a question answering apparatus and stored in a non-transitory computer-readable medium to perform the method set forth in claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: ZHANG, BYOUNG-TAK; CHOI, SEONGHO; ON, KYOUNG-WOON; HEO, YU-JUNG; JANG, YOU WON; SEO, AHJEONG; LEE, SEUNGCHAN; LEE, MINSU
To: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 055367/0681 →
Priority Claims (1)
KR 10-2020-0131339 · Oct 12, 2020 · national
Continuity (1)
Related Publication 20220350826A1 · Nov 3, 2022
Cited By (1)
US 12,749,311