IP Library › Granted Patent US 12,067,759
Granted Patent B2
US 12,067,759 · App. 17/615,662 · Granted Aug 20, 2024

Method of constructing transformer model for answering questions about video story and computing apparatus for performing the same

Inventors: Byoung-Tak Zhang (Seoul, KR); Seongho Choi (Seoul, KR)
Assignee: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
G06V10/44G06V10/25G06V10/764G06V40/20H04N21/4884
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,067,759
App. No.
17/615,662
Filed
Dec 1, 2021
Granted
Aug 20, 2024
Kind
B2
Art Unit
2671
USPC
382/190
Abstract

A method of constructing a transformer model for answering questions about a video story according to an embodiment includes: extracting feature vectors related to each character of a video from video data including vision data and subtitle data and question data for video questions and answers, and generating an input embedding using the feature vectors related to the character; and training a transformer model using the input embedding.

Claims (29)

1. A method of constructing a transformer model for answering questions about a video story, the method comprising:

extracting feature vectors related to each character of a video from video data including vision data and subtitle data and question data for video questions and answers, and generating an input embedding using the feature vectors related to the character; and

training a transformer model using the input embedding,

wherein generating the input embedding comprises:

classifying the vision data, the subtitle data, and the question data into a plurality of categories;

extracting feature vectors for the plurality of respective categories;

generating a feature embedding, a segment embedding, and a position embedding using the extracted feature vectors; and

generating the input embedding by summing the feature embedding, the segment embedding, and the position embedding, and

wherein the plurality of categories includes one or more categories related to features of the character.

2. The method of claim 1 , wherein the categories related to the features of the character comprise a bounding box including the character in an image frame included in the video, behavior of the character, and emotion of the character.

3. The method of claim 1 , wherein generating the feature embedding, the segment embedding, and the position embedding using the extracted feature vectors comprises:

generating the feature embedding by concatenating all the feature vectors extracted for the plurality of respective categories;

generating the segment embedding by performing embedding lookups using a learnable embedding matrix for the plurality of respective categories; and

generating the position embedding by generating vectors including position information related to the feature vectors extracted for the plurality of respective categories.

4. The method of claim 1 , wherein training the transformer model is performed via multi-task learning including masked language modeling, masked frame modeling, and response language modeling.

5. A non-transitory computer-readable storage medium having stored therein a program for performing the method set forth in claim 1 .

6. A computer program that is executed by a computing apparatus and stored in a non-transitory storage medium to perform the method set forth in claim 1 .

7. A computing apparatus for constructing a transformer model for answering questions about a video story, the computing apparatus comprising:

an input/output unit configured to receive video data including vision data and subtitle data and question data for video questions and answers, and to output video story question and answer results;

a storage unit configured to store a program and data for answering questions about a video story; and

a control unit comprising at least one processor, and configured to construct a transformer model for answering the questions about the video story by executing the stored program;

wherein the control unit extracts feature vectors related to each character of a video from the video data and the question data, generates an input embedding using the feature vectors related to the character, and trains the transformer model using the input embedding,

wherein when generating the input embedding, the control unit generates the input embedding by classifying the vision data, the subtitle data, and the question data into a plurality of categories, extracting feature vectors for the plurality of respective categories, generating a feature embedding, a segment embedding, and a position embedding using the extracted feature vectors, and summing the feature embedding, the segment embedding, and the position embedding;

wherein the plurality of categories comprises one or more categories related to features of the character.

8. The method of claim 1 ,

wherein the feature embedding, the segment embedding, and the position embedding include position information of the character in an image frame included in the video.

9. The computing apparatus of claim 7 , wherein the categories related to the features of the character comprise a bounding box including the character in an image frame included in the video, behavior of the character, and emotion of the character.

10. The computing apparatus of claim 7 , wherein when generating the feature embedding, the segment embedding, and the position embedding using the extracted feature vectors, the control unit generates the feature embedding by concatenating all the feature vectors extracted for the plurality of respective categories, generates the segment embedding by performing embedding lookups using a learnable embedding matrix for the plurality of respective categories, and generates the position embedding by generating vectors including position information related to the feature vectors extracted for the plurality of respective categories.

11. The computing apparatus of claim 7 , wherein when training the transformer model, the control unit performs multi-task learning including masked language modeling, masked frame modeling, and response language modeling.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2021
From: ZHANG, BYOUNG-TAK; CHOI, SEONGHO
To: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 059331/0292 →
Priority Claims (1)
KR 10-2021-0093486 · Jul 16, 2021 · national
Continuity (1)
Related Publication 20240037896A1 · Feb 1, 2024
Cited By (1)
US 12,327,084