IP Library Granted Patent US 12,249,138
Granted Patent B2
US 12,249,138 · App. 17/769,269 · Granted Mar 11, 2025

Context-driven learning of human-object interactions

Inventors: Mert Kilickaya (Amsterdam, NL); Noureldien Mahmoud Elsayed Hussein (Amsterdam, NL); Efstratios Gavves (Amsterdam, NL); Arnold Wilhelmus Maria Smeulders (Amsterdam, NL)
Assignee: QUALCOMM Technologies, Inc.
G06V20/00G06V10/454G06V10/751G06V10/764G06V10/778G06V10/806G06V10/82G06V20/52G06V40/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,249,138
App. No.
17/769,269
Granted
Mar 11, 2025
Kind
B2
Abstract

A method for classifying a human-object interaction includes identifying a human-object interaction in the input. Context features of the input are identified. Each identified context feature is compared with the identified human-object interaction. An importance of the identified context feature is determined for the identified human-object interaction. The context feature is fused with the identified human-object interaction when the importance is greater than a threshold.

Claims (54)

1. A processor-implemented method for classifying an input, the processor implemented method being performed by one or more processors and comprising:

identifying a human-object interaction in the input;

identifying context features of the input;

comparing each identified context feature with the identified human-object interaction;

determining an importance score of the identified context feature for the identified human-object interaction based on the comparing;

fusing the context feature with the identified human-object interaction to obtain a fused representation responsive to the importance score being greater than a threshold; and

classifying the human-object interaction based on the fused representation.

2. The processor-implemented method of claim 1 , in which the context features comprise one or more of clothing, human body part features, human-object spatial relations, scene locality, scene geometry, or co-occurring objects.

3. The processor-implemented method of claim 1 , further comprising determining a probability of a verb and noun pair based on the fused representation of the human-object interaction and the context features.

4. The processor-implemented method of claim 1 , further comprising adding noise during a training forward pass to robustify the importance score determination.

5. The processor-implemented method of claim 4 , further comprising removing the noise during a training backward pass.

6. The processor-implemented method of claim 1 , in which a classifier for classifying the input is robust to horizontal and vertical flipping.

7. The processor-implemented method of claim 1 , further comprising jittering a location of one or more of a human bounding box or object bounding box during training.

8. An apparatus, comprising:

at least one memory; and

at least one processor coupled to the at least one memory, the at least one processor configured to:

identify a human-object interaction in an input;

identify context features of the input;

perform a comparison of each identified context feature with the identified human-object interaction;

determine an importance score of the identified context feature for the identified human-object interaction based on the comparison;

fuse the context feature with the identified human-object interaction to obtain a fused representation responsive to the importance score being greater than a threshold; and

classify the human-object interaction based on the fused representation.

9. The apparatus of claim 8 , in which the context features comprise one or more of clothing, human-object spatial relations, scene locality, scene geometry, or co-occurring objects.

10. The apparatus of claim 8 , in which the at least one processor is further configured to determine a probability of a verb and noun pair based on the fused representation of the human-object interaction and the context features.

11. The apparatus of claim 8 , in which the at least one processor is further configured to add noise during a training forward pass to robustify the importance score determination.

12. The apparatus of claim 11 , in which the at least one processor is further configured to remove the noise during a training backward pass.

13. The apparatus of claim 8 , in which a classifier for classifying the input is robust to horizontal and vertical flipping.

14. The apparatus of claim 8 , in which the at least one processor is further configured to jitter a location of one or more of a human bounding box or object bounding box during training.

15. An apparatus, comprising:

means for identifying a human-object interaction in an input;

means for identifying context features of the input;

means for performing a comparison of each identified context feature with the identified human-object interaction;

means for determining an importance score of the identified context feature for the identified human-object interaction based on the comparison;

means for fusing the context feature with the identified human-object interaction to obtain a fused representation responsive to the importance score being greater than a threshold; and

means for classifying the human-object interaction based on the fused representation.

16. The apparatus of claim 15 , in which the context features comprise one or more of clothing, human body part features, human-object spatial relations, scene locality, scene geometry, or co-occurring objects.

17. The apparatus of claim 15 , further comprising means for determining a probability of a verb and noun pair based on the fused representation of the human-object interaction and the context features.

18. The apparatus of claim 15 , further comprising means for adding noise during a training forward pass to robustify the importance score determination.

19. The apparatus of claim 18 , further comprising means for removing the noise during a training backward pass.

20. The apparatus of claim 15 , in which a classifier for classifying the input is robust to horizontal and vertical flipping.

21. The apparatus of claim 15 , further comprising means for jittering a location of one or more of a human bounding box or object bounding box during training.

22. A non-transitory computer-readable medium having program code recorded thereon, the program code executed by a processor and comprising:

program code to identify a human-object interaction in an input;

program code to identify context features of the input;

program code to perform a comparison of each identified context feature with the identified human-object interaction;

program code to determine an importance score of the identified context feature for the identified human-object interaction based on the comparison;

program code to fuse the context feature with the identified human-object interaction to obtain a fused representation responsive to the importance score being greater than a threshold; and

program code to classify the human-object interaction based on the fused representation.

23. The non-transitory computer-readable medium of claim 22 , in which the context features comprises one or more of clothing, human-object spatial relations, scene locality, scene geometry, or co-occurring objects.

24. The non-transitory computer-readable medium of claim 22 , further comprising program code to determine a probability of a verb and noun pair based on the fused representation of the human-object interaction and the context features.

25. The non-transitory computer-readable medium of claim 22 , further comprising program code to add noise during a training forward pass to robustify the importance score determination.

26. The non-transitory computer-readable medium of claim 25 , further comprising program code to remove the noise during a training backward pass.

27. The non-transitory computer-readable medium of claim 22 , in which a classifier for classifying the input is robust to horizontal and vertical flipping.

28. The non-transitory computer-readable medium of claim 22 , further comprising program code to jitter a location of one or more of a human bounding box or object bounding box during training.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: KILICKAYA, MERT; HUSSEIN, NOURELDIEN MAHMOUD ELSAYED; GAVVES, EFSTRATIOS; SMEULDERS, ARNOLD WILHELMUS MARIA
To: UNIVERSITEIT VAN AMSTERDAM
Reel/Frame 060510/0566 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: UNIVERSITEIT VAN AMSTERDAM
To: QUALCOMM TECHNOLOGIES, INC.
Reel/Frame 060510/0589 →
Priority Claims (1)
GR 20190100515 · Nov 15, 2019 · national
Continuity (2)
Related Publication 20240135712A1 · Apr 25, 2024
Related Publication 20240233365A9 · Jul 11, 2024
References Cited (18)
US 10628679B1 · Queen · 2020 [cited by examiner]
US 20170357877A1 · Lin · 2017 [cited by examiner]
US 20180101955A1 · Varadarajan et al. · 2018 [cited by applicant]
US 20190180090A1 · Jiang et al. · 2019 [cited by applicant]
US 20190286892A1 · Li et al. · 2019 [cited by applicant]
US 20200012924A1 · Ma · 2020 [cited by examiner]
US 20200054306A1 · Mehanian · 2020 [cited by examiner]
CN 108805080A · 2018 [cited by applicant]
CN 109716354A · 2019 [cited by applicant]
CN 110263872A · 2019 [cited by examiner]
WO 2017007626A1 · 2017 [cited by applicant]
WO 2019109972A1 · 2019 [cited by applicant]
Gkioxari (“Detecting and Recognizing Human-Object Interactions, 2018 IEEE/CVF Conference on Computer Vision and Plattern Recognition, IEEE, Jun. 18, 2018 (Jun. 18, 2018), pp. 8359-8367, XP033473759, DOI: 10.1109/CVPR.20… [cited by examiner]
Gkioxari G., et al., “Detecting and Recognizing Human-Object Interactions”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Jun. 18, 2018 (Jun. 18, 2018), pp. 8359-8367, XP033473759, DOI: 10.1… [cited by applicant]
International Search Report and Written Opinion—PCT/US2020/060626—ISA/EPO—Feb. 25, 2021. [cited by applicant]
Kilickaya M., et al., “Self-Selective Context for Interaction Recognition”, ARXIV.ORG, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 17, 2020 (Oct. 17, 2020), XP081789790, 8 Page… [cited by applicant]
Nguyen T H., et al., “Automatic Video Captioning Using Deep Neural Network”, Master's Thesis of Rochester Institute of Technology RIT Scholar Works, Apr. 1, 2017 (Apr. 1, 2017), XP055549091, 91 Pages, ISBN: 978-0-355-16… [cited by applicant]
Zhuang B., et al., “Towards Context-Aware Interaction Recognition for Visual Relationship Detection”, 2017 IEEE International Conference on Computer Vision (ICCV), IEEE, Oct. 22, 2017 (Oct. 22, 2017), pp. 589-598, KP033… [cited by applicant]