IP Library Granted Patent US 12,579,840
Granted Patent B2
US 12,579,840 · App. 18/205,005 · Granted Mar 17, 2026

Behavior estimation device, behavior estimation method, and recording medium

Inventors: Ryuhei Ando (Tokyo, JP); Yasunori Babazaki (Tokyo, JP)
Assignee: NEC CORPORATION
G06V40/20G06V10/806G06V20/46G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,840
App. No.
18/205,005
Granted
Mar 17, 2026
Kind
B2
Abstract

In the behavior estimation device, a person feature extraction means extracts a feature of a person detected from a plurality of images in a time series. An object feature extraction means extracts a feature of an object detected from the plurality of images. A peripheral feature extraction means extracts a feature of a periphery of the person in the plurality of images. A feature aggregation means executes aggregation processing for aggregating the feature of the person, the feature of the object, and the feature of the periphery of the person. A behavior estimation processing means executes estimation processing for estimating the person's behavior included in the plurality of images based on information including a processing result of the aggregation processing.

Claims (26)

1 . A behavior estimation device comprising:

a memory configured to store instructions; and

one or more processors configured to execute the instructions to:

extract a first feature quantity corresponding to at least one person area detected from a plurality of images in a time series using a neural network;

extract a second feature quantity corresponding of an object area detected from the plurality of images using a neural network, wherein the object area is different from the person area;

extract a third feature quantity corresponding to a background area in the plurality of images using a neural network, wherein the background area is different from the person area and the object area;

acquire a fourth feature quantity by weighted addition of the first feature quantity and the third feature quantity;

acquire a fifth feature quantity by weighted addition of the second feature quantity and the fourth feature quantity;

acquire a sixth feature quantity by weighted addition of the third feature quantity and the fifth feature quantity;

acquire a first estimation result of the behavior of a person included in the person area by applying the fourth feature quantity to a first estimation model;

acquire a second estimation result of the behavior of the person by applying the fifth feature quantity to a second estimation model;

acquire a third estimation result of the behavior of the person by applying the sixth feature quantity to a third estimation model; and

acquire a final estimation result of the behavior of the person based on the first estimation result, the second estimation result, and the third estimation result.

2 . The behavior estimation device according to claim 1 , wherein the one or more processors acquire the values according to weights used in the weighted addition in acquiring the fifth feature as relevance degree information that is information indicating a relevance degree of an object included in the object area to the behavior of a person included in the person area, and set weights to be used in the weighted addition when acquiring the sixth feature using the acquired relevance degree information.

3 . A behavior estimation method comprising:

extracting a first feature quantity corresponding to at least one person area detected from a plurality of images in a time series using a neural network;

extracting a second feature quantity corresponding to an object area detected from the plurality of images using a neural network, wherein the object area is different from the person area;

extracting a third feature quantity corresponding to a background area in the plurality of images using a neural network, wherein the background area is different from the person area and the object area;

acquiring a fourth feature quantity by weighted addition of the first feature quantity and the third feature quantity;

acquiring a fifth feature quantity by weighted addition of the second feature quantity and the fourth feature quantity;

acquiring a sixth feature quantity by weighted addition of the third feature quantity and the fifth feature quantity;

acquiring a first estimation result of the behavior of a person included in the person area by applying the fourth feature quantity to a first estimation model;

acquiring a second estimation result of the behavior of the person by applying the fifth feature quantity to a second estimation model;

acquiring a third estimation result of the behavior of the person by applying the sixth feature quantity to a third estimation model; and

acquiring a final estimation result of the behavior of the person based on the first estimation result, the second estimation result, and the third estimation result.

4 . A non-transitory recording medium recording a program, the program causing a computer to execute the behavior estimation method according to claim 3 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2023
From: ANDO, RYUHEI; BABAZAKI, YASUNORI
To: NEC CORPORATION
Reel/Frame 063836/0912 →
Priority Claims (1)
JP 2022-093414 · Jun 9, 2022 · national
Continuity (1)
Related Publication 20230401894A1 · Dec 14, 2023
References Cited (41)
US 9355368B2 · Djugash · 2016 [cited by examiner]
US 10475185B1 · Raghavan · 2019 [cited by examiner]
US 11172189B1 · Elmieh · 2021 [cited by examiner]
US 11567572B1 · Pratt · 2023 [cited by examiner]
US 11735017B2 · Albero · 2023 [cited by examiner]
US 11954990B2 · Albero · 2024 [cited by examiner]
US 20140279733A1 · Djugash · 2014 [cited by examiner]
US 20150294193A1 · Tate et al. · 2015 [cited by applicant]
US 20150302310A1 · Wernevi · 2015 [cited by examiner]
US 20170015318A1 · Scofield · 2017 [cited by examiner]
US 20180178808A1 · Zhao · 2018 [cited by examiner]
US 20190012603A1 · Miller · 2019 [cited by examiner]
US 20190163982A1 · Block · 2019 [cited by applicant]
US 20200222949A1 · Murad · 2020 [cited by examiner]
US 20210064127A1 · Park · 2021 [cited by examiner]
US 20210248885A1 · Huang et al. · 2021 [cited by applicant]
US 20220405501A1 · Chowdhury · 2022 [cited by examiner]
US 20220415138A1 · Albero · 2022 [cited by examiner]
US 20220415146A1 · Albero · 2022 [cited by examiner]
US 20220415149A1 · Albero · 2022 [cited by examiner]
US 20230252814A1 · Kim · 2023 [cited by examiner]
US 20230386305A1 · Albero · 2023 [cited by examiner]
CN 110334607A · 2019 [cited by examiner]
CN 111414797A · 2020 [cited by examiner]
CN 112149616A · 2020 [cited by examiner]
CN 112784765A · 2021 [cited by examiner]
CN 114445741A · 2022 [cited by applicant]
JP 2015204030A · 2015 [cited by applicant]
JP 2022073882A · 2022 [cited by applicant]
JP 2022083232A · 2022 [cited by applicant]
WO 2018163555A1 · 2018 [cited by applicant]
CN-110334607-A (machine translation) (Year: 2019). [cited by examiner]
CN-111414797-A (machine translation) (Year: 2020). [cited by examiner]
CN-112149616-A (machine translation) (Year: 2020). [cited by examiner]
CN-112784765-A (machine translation) (Year: 2021). [cited by examiner]
Chen et al., “An Efficient Recommendation Filter Model on Smart Home Big Data Analytics for Enhanced Living Environments.” Sensors (Basel). Oct. 15, 2016;16(10):1706. doi: 10.3390/s16101706. PMID: 27754456; PMCID: PMC50… [cited by examiner]
Liu et al., “Human Object Interaction Detection using Two-Direction Spatial Enhancement and Exclusive Object Prior.” arXiv preprint arXiv:2105.03089 (2021). (Year: 2021). [cited by examiner]
JP Office Action for JP Application No. 2022-093414, mailed on Jan. 27, 2026 with English Translation. [cited by applicant]
Nazli Ikizler-Cinbis et al., “Object, Scene and Actions: Combining Multiple Features for Human Action Recognition”, Lecture Notes in Computer Science, Springer, 2010, pp. 494-07, DOI: 10.1007/978-3-642-15549-9_36, Retri… [cited by applicant]
Kobayashi et al., “Recognition of Detailed Actions in Working Images of Production Lines”, SSII2019 [USB], Symposium on Sending via Image Information, Jun. 12, 2019, published by the Image Sensing Technology Workshop, J… [cited by applicant]
Do, Hang Nga et al., “Mining Specific actions from Youtube video with spatio-temporal features”, IEICE Technical Report, PRMU, vol. 110, No. 414, pp. 159-164, Report No. PRMU2010-233, Feb. 10, 2011, The Institute of Ele… [cited by applicant]