IP Library Granted Patent US 12,646,301
Granted Patent B2
US 12,646,301 · App. 18/387,515 · Granted Jun 2, 2026

Learning device, parameter adjustment method and recording medium

Inventors: Tomokazu Kaneko (Tokyo, JP); Soma Shiraishi (Tokyo, JP); Ryosuke Sakai (Tokyo, JP)
Assignee: NEC CORPORATION
G06V10/776G06V10/764G06V10/7715G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,301
App. No.
18/387,515
Granted
Jun 2, 2026
Kind
B2
Abstract

A learning device includes a learning model for a still image. The learning model includes a mask generation means configured to generate a first object mask identifying an area in which an object exists in a still image, for each individual object. A first parameter including at least one parameter used for processing of generating the first object mask is adjusted based on a first loss. The first loss indicates a difference of the first object mask with respect to a second object mask identifying an area in which an object exists in a moving image including the still image, for each individual object.

Claims (28)

1 . A learning device comprising:

a memory configured to store instructions; and

a processor functioning as a learning model for a still image, the processor configured to execute the instructions to:

generate a first object mask identifying an area in which an object exists in the still image, for each individual object,

wherein a first parameter including at least one parameter used for processing of generating the first object mask is adjusted based on a first loss, the first loss indicating a difference of the first object mask with respect to a second object mask identifying an area in which an object exists in a moving image including the still image, for each individual object,

wherein the processor is further configured to execute the instructions to extract a first feature quantity representing a feature of the entire still image,

wherein a second parameter including at least one parameter used for processing of extracting the first feature quantity is adjusted based on a second loss, the second loss indicating a difference of the first feature quantity with respect to a second feature quantity representing a feature of the entire moving image.

2 . The learning device according to claim 1 , wherein the first parameter is adjusted in a manner that the first loss becomes 0.

3 . The learning device according to claim 1 , wherein the first parameter and the second parameter are adjusted in a manner that both the first loss and the second loss become 0.

4 . The learning device according to claim 1 ,

wherein the processor is further configured to execute the instructions to acquire a first object representation for each individual object in the still image,

wherein a third parameter including at least one parameter used for processing of acquiring the first object representation is adjusted based on a third loss, the third loss indicating a difference of the first object representation with respect to a second object representation, for each individual object in the moving image.

5 . The learning device according to claim 4 , wherein the first parameter and the third parameter are adjusted in a manner that both the first loss and the third loss become 0.

6 . The learning device according to claim 1 ,

wherein the processor is configured to execute the instructions to:

extract the first feature quantity;

acquire a first object representation for each individual object in the still image based on the first feature quantity; and

generate the first object mask based on the first object representation,

wherein a third parameter including at least one parameter used for processing of acquiring the first object representation is adjusted based on a third loss, the third loss indicating a difference of the first object representation with respect to a second object representation, for each individual object in the moving image.

7 . The learning device according to claim 6 , wherein the first parameter, the second parameter and the third parameter are adjusted in a manner that all the first loss, the second loss and the third loss become 0.

8 . A parameter adjustment method applied to a learning model for a still image that generates a first object mask identifying an area where an object exists in the still image for each individual object, the method comprising:

adjusting a first parameter including at least one parameter used for processing of generating the first object mask based on a first loss, the first loss indicating a difference of the first object mask with respect to a second object mask identifying an area where an object exists in a moving image including the still image, for each individual object,

wherein further executing the instructions to extract a first feature quantity representing a feature of the entire still image,

wherein a second parameter including at least one parameter used for processing of extracting the first feature quantity is adjusted based on a second loss, the second loss indicating a difference of the first feature quantity with respect to a second feature quantity representing a feature of the entire moving image.

9 . A non-transitory computer-readable recording medium storing a program, the program causing a computer to execute processing for a learning model for a still image that generates a first object mask identifying an area where an object exists in the still image for each individual object, the processing comprising:

adjusting a first parameter including at least one parameter used for processing of generating the first object mask based on a first loss, the first loss indicating a difference of the first object mask with respect to a second object mask identifying an area where an object exists in a moving image including the still image, for each individual object,

wherein further executing the instructions to extract a first feature quantity representing a feature of the entire still image,

wherein a second parameter including at least one parameter used for processing of extracting the first feature quantity is adjusted based on a second loss, the second loss indicating a difference of the first feature quantity with respect to a second feature quantity representing a feature of the entire moving image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: KANEKO, TOMOKAZU; SHIRAISHI, SOMA; SAKAI, RYOSUKE
To: NEC CORPORATION
Reel/Frame 065479/0095 →
Priority Claims (1)
JP 2022-179283 · Nov 9, 2022 · national
Continuity (1)
Related Publication 20240153255A1 · May 9, 2024
References Cited (7)
US 20200250436A1 · Lee · 2020 [cited by examiner]
US 20210042963A1 · Takagi · 2021 [cited by examiner]
US 20220198783A1 · Yoshida · 2022 [cited by examiner]
US 20230191605A1 · Handa · 2023 [cited by examiner]
US 20240127440A1 · Hatsutani · 2024 [cited by examiner]
WO WO2020240727A1 · 2020 [cited by examiner]
Francesco Locatello,et.al, “Object-Centric Learning with Slot Attention”, [online], Oct. 14, 2020, arXiv, [Search on Oct. 28, 2022], Internet<URL:https://arxiv.org/pdf/2006.15055.pdf>. [cited by applicant]