IP Library Granted Patent US 12,417,546
Granted Patent B2
US 12,417,546 · App. 18/021,296 · Granted Sep 16, 2025

Learning device, learning method, tracking device, and storage medium

Inventor: Yasunori Babazaki (Tokyo, JP)
Assignee: NEC CORPORATION
G06T7/248G06T7/73G06T2207/20076G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,546
App. No.
18/021,296
Granted
Sep 16, 2025
Kind
B2
Abstract

A learning device 1 X includes an acquisition means 15 X, an estimation result matching means 16 X, and a learning means 18 X. The acquisition means 15 X acquires tracking training data in which first and second training images in time series, tracking target position information regarding position or posture of a tracking target shown in each first and second training images, and identification information of the tracking target are associated. The estimation result matching means 16 X compares the tracking target position information with posture information indicating posture of the tracking target estimated from the first and second training images and associates the posture information with the identification information. The learning means 18 X learns an inference engine, which infers correspondence information indicating correspondence relation of the tracking target between the training images when information based on the posture information is inputted, the correspondence information based on the posture information and the identification information.

Claims (50)

1. A learning device comprising:

at least one memory configured to store instructions; and

at least one processor configured to execute the instructions to:

acquire tracking training data in which a first training image and a second training image which are training images captured in time series, tracking target object position information regarding a position or a posture of a tracking target object shown in each of the first training image and the second training image, and identification information of the tracking target object are associated;

compare the tracking target object position information with posture information indicative of an estimated posture of the tracking target object estimated from each of the first training image and the second training image and

associate the posture information with the identification information; and

learn an inference engine based on the posture information and the identification information,

the inference engine being configured to infer correspondence information when information based on the posture information is inputted to the inference engine,

the correspondence information indicating correspondence relation of the tracking target object between the first training image and the second training image.

2. The learning device according to claim 1 ,

wherein the at least one processor is configured to further execute the instructions to convert the posture information into feature information which is information indicating a position of each feature point of the tracking target objects detected from the training images captured in time series, and

wherein the at least one processor is configured to execute the instructions to input the feature information to the inference engine as the information based on the posture information.

3. The learning device according to claim 2 ,

wherein the at least one processor is configured to execute the instructions to generate the feature information according to a format that is based on:

the number of detections of the tracking target objects in the first training image;

the number of detections of the tracking target objects in the second training image; and

the number of the feature points.

4. The learning device according to claim 1 ,

wherein the inference engine is a neural network with a convolution layer.

5. The learning device according to claim 1 ,

wherein the correspondence information indicates a matrix including each element indicating a probability that each of the tracking target objects in the first training image corresponds to each of the tracking target objects in the second training image.

6. The learning device according to claim 5 ,

wherein the matrix further comprises a row or a column indicating a probability of appearance or disappearance of the tracking target object in the first training image and the second training image.

7. The learning device according to claim 6 ,

wherein the at least one processor is configured to execute the instructions to learn the inference engine so that a sum of elements for each row or each column of the matrix is set to be a predetermined value.

8. The learning device according to claim 5 ,

wherein the matrix includes:

a first channel including each element indicating the probability that each of the tracking target objects in the first training image corresponds to each of the tracking target objects in the second training image; and

a second channel including each element indicating a probability that each of the tracking target objects in the first training image does not correspond to each of the tracking target objects in the second training image, and

wherein the at least one processor is configured to execute the instructions to learn the inference engine so that a sum of the each element in channel direction is set to be a predetermined value.

9. The learning device according to claim 1 ,

wherein the tracking target object position information is

information indicating an existence area of each tracking target object in the first training image and the second training image, or

information indicating a position of each feature point of each tracking target object in the first training image and the second training image.

10. The learning device according to claim 1 ,

wherein the at least one processor is configured to further execute the instructions to generate the posture information by estimating the posture of each tracking target object shown in the first training image and the second training image based on the first training image and the second training image.

11. A tracking device comprising:

at least one memory configured to store instructions; and

at least one processor configured to execute the instructions to:

acquire a first captured image and a second captured image captured in time series;

generate, based on the first captured image and the second captured image, posture information indicative of an estimation result of a posture of a tracking target object in each of the first captured image and the second captured image; and

generate correspondence information based on the posture information and an inference engine, the inference engine being configured to infer the correspondence information when information based on the posture information is inputted to the inference engine, the correspondence information indicating a correspondence relation of the tracking target object between the first captured image and the second captured image.

12. The tracking device according to claim 11 ,

wherein the at least one processor is configured to further execute the instructions to manage identification information assigned to the tracking target object based on the correspondence information.

13. A learning method executed by a computer, the learning method comprising:

acquiring tracking training data in which a first training image and a second training image which are training images captured in time series, tracking target object position information regarding a position or a posture of a tracking target object shown in each of the first training image and the second training image, and identification information of the tracking target object are associated;

comparing the tracking target object position information with posture information indicative of an estimated posture of the tracking target object estimated from each of the first training image and the second training image and associating the posture information with the identification information; and

learning an inference engine based on the posture information and the identification information,

the inference engine being configured to infer correspondence information when information based on the posture information is inputted to the inference engine,

the correspondence information indicating correspondence relation of the tracking target object between the first training image and the second training image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2023
From: BABAZAKI, YASUNORI
To: NEC CORPORATION
Reel/Frame 062694/0280 →
Continuity (1)
Related Publication 20230326041A1 · Oct 12, 2023
References Cited (12)
US 20130050502A1 · Saito et al. · 2013 [cited by applicant]
US 20190066313A1 · Kim et al. · 2019 [cited by applicant]
US 20190147292A1 · Watanabe et al. · 2019 [cited by applicant]
US 20200294373A1 · Srinivasan · 2020 [cited by examiner]
US 20210158566A1 · Ogawa · 2021 [cited by examiner]
CN 107563313A · 2018 [cited by applicant]
JP 2018026108A · 2018 [cited by applicant]
JP 2019091138A · 2019 [cited by applicant]
WO 2011102416A1 · 2011 [cited by applicant]
International Search Report for PCT Application No. PCT/JP2020/032459, mailed on Nov. 24, 2020. [cited by applicant]
Henschel, Roberto et al., “Multiple People Tracking using Body and Joint Detections”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) [online], IEEE, Apr. 9, 2020, pp. 770-779. [cited by applicant]
Leal-Taixe, Laura et al., “Learning by tracking: Siamese CNN for robust target association”, arXiv.org [online], arXiv:1604.07866v3, Cornel University, 2016, <URL : https://arxiv.org/pdf/1604.07866v3>, pp. 1-9. [cited by applicant]