IP Library Patent Application 18843684
Patent Application
App. No. 18/843,684

SUMMARY GENERATION APPARATUS, SUMMARY MODEL LEARNING APPARATUS, SUMMARY GENERATION METHOD, SUMMARY MODEL LEARNING METHOD, AND PROGRAM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/843,684
Abstract

A summary generation device includes: an image processing unit that receives an input of an image related to a moving image, and extracts at least a text from the image; a sound processing unit that receives an input of sound in the moving image, and extracts at least a text from the sound; and a summary generation unit that generates a summary text of the moving image from information extracted from the image and information extracted from the sound, using a trained summary model.

Claims (51)

1 . A device comprising a processor configured to execute operations comprising:

receiving an input of an image related to a moving image, and extracts at least a text from the image;

receiving an input of sound in the moving image, and extracts at least a text from the sound; and

generating a summary text of the moving image from information extracted from the image and information extracted from the sound, using a trained summary model.

2 . A device comprising a processor configured to execute operations comprising:

receiving an input of an image of a moving image;

extracting at least a first training text from the image;

receiving an input of sound in the moving image;

extracting at least a second training text from the sound; and

training a summary model, using the first training text, the second training text, and a correct summary text of the moving image.

3 . The device according to claim 2 , the processor further configured to execute operations comprising:

acquiring the moving image and the correct summary text from a server in a network.

4 . The device according to claim 2 , the processor further configured to execute operations comprising:

performing pre-training on the summary model, using a pre-training text in a field associated with the moving image and a correct summary text of the pre-training text.

5 . The device according to claim 2 , the processor further configured to execute operations comprising:

generating at least one further training data set from a training data set,

wherein the training data set comprises first information extracted from the image, second information extracted from the sound, and the correct summary text of the training moving image,

the first information extracted from the image comprises the first training text, and

the second information extracted from the sound comprises the second training text.

6 . A method implemented by a computer, comprising:

a first image processing step of extracting at least a text from an image related to a moving image;

a first sound processing step of extracting at least a text from sound in the moving image; and

a summary generating step of generating a summary text of the moving image from information extracted from the image related to the moving image and information extracted from the sound, using a trained summary model.

7 . The method according to claim 6 , comprising:

a second image processing step of extracting at least a first training text from a training image related to a training moving image;

a second sound processing step of extracting at least a second training text from sound in the training moving image; and

a summary model training step of training a summary model, using the first training text, the second training text, and a correct summary text of the training moving image.

8 . (canceled)

9 . The device according to claim 1 , wherein the image related to the moving image is based on a frame of the moving image.

10 . The device according to claim 1 , the processor further configured to execute operations comprising:

performing pre-training on a summary model, using a pre-training text in a field associated with a pre-training moving image and a correct summary text of the pre-training text.

11 . The device according to claim 10 , the processor further configured to execute operations comprising:

receiving an input of a training image of a training moving image;

extracting at least a first training text from the training image;

receiving an input of training sound in the training moving image;

extracting at least a second training text from the training sound; and

training the summary model, using the first training text, the second training text, and a correct summary text of the training moving image.

12 . The device according to claim 11 , wherein a first amount of training data based on a first combination comprising the pre-training text and the correct summary text of the pre-training text is larger than a second amount of training data based on a second combination comprising the first training text, the second training text, and the correct summary text of the training moving image, and the training of the summary model represents a fine-tuning of the summary model.

13 . The device according to claim 11 , the processor further configured to execute operations comprising:

acquiring the training moving image and the correct summary text from a server in a network.

14 . The method according to claim 6 , further comprising:

performing pre-training on a summary model, using a pre-training text in a field associated with a pre-training moving image and a correct summary text of the pre-training text.

15 . The method according to claim 6 , further comprising:

wherein the image related to the moving image is based on a frame of the moving image.

16 . The method according to claim 7 , further comprising:

acquiring the training moving image and the correct summary text from a server in a network.

17 . The method according to claim 7 , further comprising:

performing pre-training on the summary model, using a pre-training text in a field associated with a pre-training moving image and a correct summary text of the pre-training text.

18 . The method according to claim 17 ,

wherein a first amount of training data comprising a first combination of the pre-training text and the correct summary text of the pre-training text is larger than a second amount of training data comprising a second combination of the first training text, the second training text, and the correct summary text of the training moving image, and

the training of the summary model represents a fine-tuning of the summary model.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0725 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2024
From: SAITO, ITSUMI; NISHIDA, KYOSUKE; YOSHIDA, SEN
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 069706/0952 →