IP Library › Granted Patent US 12,646,179
Granted Patent B2
US 12,646,179 · App. 18/374,722 · Granted Jun 2, 2026

Image processing apparatus, image processing method, and computer readable recording medium

Inventor: Youki Sada (Tokyo, JP)
Assignee: NEC CORPORATION
G06T7/194G06V10/761G06V10/771G06T2207/20084G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,179
App. No.
18/374,722
Granted
Jun 2, 2026
Kind
B2
Abstract

An image processing apparatus includes: a foreground information generating unit that generates, from a frame constituting video frame, foreground information indicating a region of the frame in which a target is present; and a model applying unit that applies the frame and the foreground information generated from the frame to a neural network model that has performed machine learning of an image feature map of the target.

Claims (34)

1 . An image processing apparatus comprising:

a memory storing program instructions;

at least one processor configured to execute the instructions to implement:

a foreground information generating unit configured to generate, from a frame constituting video frame, foreground information indicating a region of the frame in which a target is present; and

a model applying unit configured to apply the frame and the foreground information generated from the frame to a neural network model that has performed machine learning of an image feature map of the target,

wherein, if the frame is a non-key frame, which is a frame other than a key frame that is designated in advance, the foreground information generating unit generates the foreground information from the frame, and further generates difference information indicating a difference between the frame and the key frame that is closest in time to the frame, and

the model applying unit applies, to the neural network model, the frame, the foreground information generated from the frame, and the difference information generated for the frame,

the least one processor configured to execute the instructions to further implement:

an intermediate feature map acquiring unit configured to, in a case in which the neural network model is a neural network, acquire an intermediate feature map from an intermediate layer of the neural network,

wherein the foreground information generating unit generates the foreground information and the difference information using the intermediate feature map when a frame earlier than the frame was applied to the neural network model.

2 . The image processing apparatus according to claim 1 ,

wherein the foreground information generating unit generates the foreground information by applying the difference information to a second neural network model that has performed machine learning of an inter-frame difference and the feature map of the target.

3 . The image processing apparatus according to claim 1 ,

wherein the model applying unit applies, to the neural network model, one or both of the foreground information generated from the frame and the difference information generated for the frame in accordance with a state of a camera capturing the video frame.

4 . The image processing apparatus according to claim 1 ,

wherein the model applying unit calculates a logical AND of the foreground information generated from the frame and the difference information generated for the frame, and applies the result of the calculation of the logical AND to the neural network model.

5 . The image processing apparatus according to claim 2 ,

wherein the second neural network model has been constructed by machine learning in which outputs from the neural network model are used as training data.

6 . The image processing apparatus according to claim 1 ,

wherein the model applying unit calculates a logical OR from pieces of difference information generated for a plurality of the frames, and applies the result of the calculation of the logical OR to the neural network model.

7 . An image processing method comprising:

generating, from a frame constituting video frame, foreground information indicating a region of the frame in which a target is present; and

applying the frame and the foreground information generated from the frame to a neural network model that has performed machine learning of an image feature map of the target,

wherein, if the frame is a non-key frame, which is a frame other than a key frame that is designated in advance, generating the foreground information from the frame, and further generating difference information indicating a difference between the frame and the key frame that is closest in time to the frame, and

applying to the neural network model, the frame, the foreground information generated from the frame, and the difference information generated for the frame,

in a case in which the neural network model is a neural network, acquiring an intermediate feature map from an intermediate layer of the neural network,

wherein the foreground information and the difference information are generated using the intermediate feature map when a frame earlier than the frame was applied to the neural network model.

8 . A non-transitory computer readable recording medium that includes a program recorded thereon, the program including instructions that causes a computer to carry out the steps of:

generating, from a frame constituting video frame, foreground information indicating a region of the frame in which a target is present; and

applying the frame and the foreground information generated from the frame to a neural network model that has performed machine learning of an image feature map of the target,

wherein, if the frame is a non-key frame, which is a frame other than a key frame that is designated in advance, generating the foreground information from the frame, and further generating difference information indicating a difference between the frame and the key frame that is closest in time to the frame, and

applying to the neural network model, the frame, the foreground information generated from the frame, and the difference information generated for the frame,

in a case in which the neural network model is a neural network, acquiring an intermediate feature map from an intermediate layer of the neural network,

wherein the foreground information and the difference information are generated using the intermediate feature map when a frame earlier than the frame was applied to the neural network model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2023
From: SADA, YOUKI
To: NEC CORPORATION
Reel/Frame 065070/0068 →
Priority Claims (1)
JP 2022-161355 · Oct 6, 2022 · national
Continuity (1)
Related Publication 20240119601A1 · Apr 11, 2024
References Cited (11)
US 11816895B2 · Chen · 2023 [cited by examiner]
US 20190311202A1 · Lee et al. · 2019 [cited by applicant]
US 20210256266A1 · Chen · 2021 [cited by examiner]
US 20240119601A1 · Sada · 2024 [cited by examiner]
JP 2020013206A · 2020 [cited by applicant]
JP 2021128705A · 2021 [cited by applicant]
JP 2022018173A · 2022 [cited by applicant]
JP 2022020353A · 2022 [cited by applicant]
JP 2022043651A · 2022 [cited by applicant]
Mathias Parger, et al., “DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in Videos”, Graz University of Technology, Meta Reality Labs, Mar. 8, 2022. [cited by applicant]
JP Office Action for JP Application No. 2022-161355, mailed on Apr. 21, 2026 with English Translation. [cited by applicant]