IP Library Patent Application 18482310
Patent Application
App. No. 18/482,310

LEARNING DEVICE, INFERENCE DEVICE, LEARNING METHOD, AND INFERENCE METHOD

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/482,310
Abstract

A learning device includes a convolutional neural network configured to output action, re-identification, size, and position feature maps in response to respective image frames constituting a video sequence being input; a processor; and a memory storing program instructions that cause the processor to: receive the action feature map, the re-identification feature map, and the size feature map to output an action feature, a re-identification feature, and a size feature; receive the action feature and the re-identification feature to output a feature obtained by considering an interaction between features; output a group activity classification result based on the output feature; output an action classification result based on the output feature; and update model parameters of the convolutional neural network so as to minimize an error between the position feature map, the size feature, the re-identification feature, the group activity classification result, and the action classification result; and correct data.

Claims (38)

1 . A learning device that performs learning for activity recognition, comprising:

a convolutional neural network configured to output an action feature map, a re-identification feature map, a size feature map, and a position feature map in response to respective image frames constituting a video sequence being input;

a processor; and

a memory storing program instructions that cause the processor to:

receive the action feature map, the re-identification feature map, and the size feature map to output an action feature, a re-identification feature, and a size feature;

receive the action feature and the re-identification feature to output a feature obtained by considering an interaction between features;

output a group activity classification result based on the output feature;

output an action classification result based on the output feature; and

update model parameters of the convolutional neural network so as to minimize an error between the position feature map, the size feature, the re-identification feature, the group activity classification result, and the action classification result; and correct data.

2 . The learning device as claimed in claim 1 , wherein the processor converts the action feature by using the re-identification feature as auxiliary information or converts the re-identification feature by using the action feature as the auxiliary information.

3 . The learning device as claimed in claim 1 , wherein the processor converts the action feature by using the re-identification feature as auxiliary information, and converts the re-identification feature by using the action feature as the auxiliary information.

4 . An inference device that performs inference for activity recognition, comprising:

a convolutional neural network configured to output an action feature map, a re-identification feature map, a size feature map, and a position feature map in response to respective image frames constituting an image sequence being input;

a processor; and

a memory storing program instructions that cause the processor to:

receive point position data obtained based on the position feature map, the action feature map, the re-identification feature map, and the size feature map to output an action feature, a re-identification feature, and a size feature;

receive a detection result and the re-identification feature to output a track result, the detection result being obtained based on the point position data and the size feature;

receive the action feature and the re-identification feature to output a feature obtained by considering an interaction between features;

output a group activity classification result based on the output feature; and

output an action classification result based on the output feature and the track result.

5 . The inference device as claimed in claim 4 , wherein the processor converts the action feature by using the re-identification feature as auxiliary information, or converts the re-identification feature by using the action feature as the auxiliary information.

6 . The inference device as claimed in claim 4 , wherein the processor converts the action feature by using the re-identification feature as auxiliary information, and converts the re-identification feature by using the action feature as the auxiliary information.

7 . A learning method executed by a learning device that performs learning for activity recognition, the learning method comprising:

inputting respective image frames constituting a video sequence into a convolutional neural network to output an action feature map, a re-identification feature map, a size feature map, and a position feature map;

receiving the action feature map, the re-identification feature map, and the size feature map to output an action feature, a re-identification feature, and a size feature;

receiving the action feature and the re-identification feature to output a feature obtained by considering an interaction between features;

outputting a group activity classification result based on the output feature;

outputting an action classification result based on the output feature; and

updating model parameters of the convolutional neural network so as to minimize an error between the position feature map, the size feature, the re-identification feature, the group activity classification result, and the action classification result; and correct data.

8 . An inference method executed by an inference device that performs inference for activity recognition, the inference method comprising:

inputting respective image frames constituting a video sequence into a convolutional neural network to output an action feature map, a re-identification feature map, a size feature map, and a position feature map;

receiving point position data obtained based on the position feature map, the action feature map, the re-identification feature map, and the size feature map to output an action feature, a re-identification feature, and a size feature;

receiving a detection result and the re-identification feature to output a track result, the detection result being obtained based on the point position data and the size feature;

receiving the action feature and the re-identification feature to output a feature obtained by considering an interaction between features;

outputting a group activity classification result based on the output feature; and

outputting an action classification result based on the output feature and the track result.

9 . A non-transitory computer-readable recording medium storing a program for causing a computer to perform the learning method as claimed in claim 7 .

10 . A non-transitory computer-readable recording medium storing a program for causing a computer to function as the inference method as claimed in claim 8 .

Assignments (2)
CHANGE OF NAME Recorded Oct 1, 2025
From: NTT COMMUNICATIONS CORPORATION
To: NTT DOCOMO BUSINESS, INC.
Reel/Frame 072987/0475 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: TARASHIMA, SHUHEI
To: NTT COMMUNICATIONS CORPORATION
Reel/Frame 065147/0533 →