IP Library › Granted Patent US 12,536,836
Granted Patent B2
US 12,536,836 · App. 18/152,627 · Granted Jan 27, 2026

Human object interaction detection using compositional model

Inventors: David Nguyen (Newark, CA); Hailing Zhou (Shenzhen, CN); Nan Ke (Xiamen, CN)
Assignee: ACCENTURE GLOBAL SOLUTIONS LIMITED
G06V40/20G06F40/30G06V10/761G06V10/774G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,836
App. No.
18/152,627
Granted
Jan 27, 2026
Kind
B2
Abstract

Implementations include actions of receiving an image; extracting a visual HOI and a set of visual embeddings, the visual HOI indicating a subject and an object; obtaining, using a vector library, a set of semantic HOIs and sets of semantic embeddings based on the subject, the object and a set of verbs included in the vector library, each set of semantic embeddings corresponding to a semantic HOI; processing, by a compositional model, the set of visual embeddings to provide a set of transition visual embeddings; processing the sets of semantic embeddings to provide respective sets of transition semantic embeddings; determining a set of scores based on the set of transition visual embeddings and the sets of transition semantic embeddings, each score representing a degree of similarity between the visual HOI and a semantic HOI; and determining at least one predicted HOI represented within the image based on the scores.

Claims (59)

1 . A computer-implemented method for determining human-object interactions (HOIs) in images, the method comprising:

receiving an image;

extracting, by a feature embedding model, a visual HOI and a set of visual embeddings, the visual HOI indicating a subject and an object;

obtaining, using a vector library, a set of semantic HOIs and sets of semantic embeddings based on the subject, the object and a set of verbs included in the vector library, each set of semantic embeddings corresponding to a semantic HOI;

processing, by a compositional model, the set of visual embeddings to provide a set of transition visual embeddings;

processing, by the compositional model, the sets of semantic embeddings to provide respective sets of transition semantic embeddings;

determining a set of scores based on the set of transition visual embeddings and the sets of transition semantic embeddings, each score representing a degree of similarity between the visual HOI and a semantic HOI of the set of semantic HOIs, wherein determining the set of scores based on the set of transition visual embeddings and the sets of transition semantic embeddings comprises, for each set of transition semantic embeddings:

determining a first distance between a transition visual embedding representative of the subject and a transition semantic embedding representative of the subject,

determining a second distance between a transition visual embedding representative of the object and a transition semantic embedding representative of the object,

determining a third distance between a transition visual embedding representative of bounding boxes and a transition semantic embedding representative of a verb, the bounding boxes bounding the subject and the object within the image,

determining an aggregate distance using the first distance, the second distance, and the third distance; and

determining at least one predicted HOI represented within the image based on the scores.

2 . The method of claim 1 , wherein the compositional model comprises a set of visual models and a set of language models, the set of visual models comprising a subject visual model, an object visual model, and a union visual model, and the set of language models comprising a subject language model, an object language model, and a verb language model.

3 . The method of claim 1 , wherein the vector library comprises a set of subject word embeddings, a set of object word embeddings, and a set of verb word embeddings.

4 . The method of claim 1 , wherein the vector library comprises word embeddings generated by processing labels of training data using a word embedding model, the training data being used to train the compositional model.

5 . The method of claim 1 , wherein determining at least one predicted HOI comprises:

identifying a semantic HOI as having a highest score; and

providing the semantic HOI with the highest score as the at least one predicted HOI.

6 . The method of claim 1 , wherein each of the first distance, the second distance, and the third distance is at least partially determined as a cosine distance.

7 . A system, comprising:

one or more processors; and

a computer-readable storage device coupled to the one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for determining human-object interactions (HOIs) in images, the operations comprising:

receiving an image;

extracting, by a feature embedding model, a visual HOI and a set of visual embeddings, the visual HOI indicating a subject and an object;

obtaining, using a vector library, a set of semantic HOIs and sets of semantic embeddings based on the subject, the object and a set of verbs included in the vector library, each set of semantic embeddings corresponding to a semantic HOI;

processing, by a compositional model, the set of visual embeddings to provide a set of transition visual embeddings;

processing, by the compositional model, the sets of semantic embeddings to provide respective sets of transition semantic embeddings;

determining a set of scores based on the set of transition visual embeddings and the sets of transition semantic embeddings, each score representing a degree of similarity between the visual HOI and a semantic HOI of the set of semantic HOIs, wherein determining the set of scores based on the set of transition visual embeddings and the sets of transition semantic embeddings comprises, for each set of transition semantic embeddings:

determining a first distance between a transition visual embedding representative of the subject and a transition semantic embedding representative of the subject,

determining a second distance between transition visual embedding representative of the object and a transition semantic embedding representative of the object,

determining a third distance between a transition visual embedding representative of bounding boxes and a transition semantic embedding representative of a verb, the bounding boxes bounding the subject and the object within the image,

determining an aggregate distance using the first distance, the second distance, and the third distance; and

determining at least one predicted HOI represented within the image based on the scores.

8 . The system of claim 7 , wherein the compositional model comprises a set of visual models and a set of language models, the set of visual models comprising a subject visual model, an object visual model, and a union visual model, and the set of language models comprising a subject language model, an object language model, and a verb language model.

9 . The system of claim 7 , wherein the vector library comprises a set of subject word embeddings, a set of object word embeddings, and a set of verb word embeddings.

10 . The system of claim 7 , wherein the vector library comprises word embeddings generated by processing labels of training data using a word embedding model, the training data being used to train the compositional model.

11 . The system of claim 7 , wherein determining at least one predicted HOI comprises:

identifying a semantic HOI as having a highest score; and

providing the semantic HOI with the highest score as the at least one predicted HOI.

12 . The system of claim 7 , wherein each of the first distance, the second distance, and the third distance is at least partially determined as a cosine distance.

13 . A non-transitory computer-readable storage media coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for determining human-object interactions (HOIs) in images, the operations comprising:

receiving an image;

extracting, by a feature embedding model, a visual HOI and a set of visual embeddings, the visual HOI indicating a subject and an object;

obtaining, using a vector library, a set of semantic HOIs and sets of semantic embeddings based on the subject, the object and a set of verbs included in the vector library, each set of semantic embeddings corresponding to a semantic HOI;

processing, by a compositional model, the set of visual embeddings to provide a set of transition visual embeddings;

processing, by the compositional model, the sets of semantic embeddings to provide respective sets of transition semantic embeddings;

determining a set of scores based on the set of transition visual embeddings and the sets of transition semantic embeddings, each score representing a degree of similarity between the visual HOI and a semantic HOI of the set of semantic HOIs, wherein determining the set of scores based on the set of transition visual embeddings and the sets of transition semantic embeddings comprises, for each set of transition semantic embeddings:

determining a first distance between a transition visual embedding representative of the subject and a transition semantic embedding representative of the subject,

determining a second distance between transition visual embedding representative of the object and a transition semantic embedding representative of the object,

determining a third distance between a transition visual embedding representative of bounding boxes and a transition semantic embedding representative of a verb, the bounding boxes bounding the subject and the object within the image,

determining an aggregate distance using the first distance, the second distance, and the third distance; and

determining at least one predicted HOI represented within the image based on the scores.

14 . The non-transitory computer-readable storage media of claim 13 , wherein the compositional model comprises a set of visual models and a set of language models, the set of visual models comprising a subject visual model, an object visual model, and a union visual model, and the set of language models comprising a subject language model, an object language model, and a verb language model.

15 . The non-transitory computer-readable storage media of claim 13 , wherein the vector library comprises a set of subject word embeddings, a set of object word embeddings, and a set of verb word embeddings.

16 . The non-transitory computer-readable storage media of claim 13 , wherein the vector library comprises word embeddings generated by processing labels of training data using a word embedding model, the training data being used to train the compositional model.

17 . The non-transitory computer-readable storage media of claim 13 , wherein determining at least one predicted HOI comprises:

identifying a semantic HOI as having a highest score; and

providing the semantic HOI with the highest score as the at least one predicted HOI.

18 . The non-transitory computer-readable storage media of claim 13 , wherein each of the first distance, the second distance, and the third distance is at least partially determined as a cosine distance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 11, 2023
From: NGUYEN, DAVID; ZHOU, HAILING; KE, NAN
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 062350/0014 →
Continuity (1)
Related Publication 20240233439A1 · Jul 11, 2024
References Cited (28)
US 11410449B2 · Wang · 2022 [cited by examiner]
US 11734924B1 · Nguyen · 2023 [cited by examiner]
US 11748571B1 · Glavaš · 2023 [cited by examiner]
US 11784964B2 · Khasanova · 2023 [cited by examiner]
US 20210049236A1 · Nguyen · 2021 [cited by examiner]
US 20210295091A1 · Li · 2021 [cited by examiner]
US 20210374553A1 · Li · 2021 [cited by examiner]
US 20220108195A1 · Kehler · 2022 [cited by examiner]
US 20220139384A1 · Wu · 2022 [cited by examiner]
US 20220180101A1 · Rai · 2022 [cited by examiner]
US 20220279220A1 · Khavronin · 2022 [cited by examiner]
US 20220358005A1 · Saha · 2022 [cited by examiner]
US 20220374595A1 · Gotmare · 2022 [cited by examiner]
US 20220391640A1 · Xing · 2022 [cited by examiner]
US 20220391755A1 · Li · 2022 [cited by examiner]
US 20220405501A1 · Chowdhury · 2022 [cited by examiner]
US 20230102422A1 · Zhou · 2023 [cited by examiner]
US 20230177384A1 · Nagrani · 2023 [cited by examiner]
US 20240233439A1 · Nguyen · 2024 [cited by examiner]
Hou, Zhi et al., Visual Compositional Learning for Human-Object Interaction Detection, Oct. 4, 2020 (Year: 2020). [cited by examiner]
Li, Zhimin et al., Improving Human-Object Interaction Detection via Phrase Learning and Label Composition, Jan. 15, 2022 (Year: 2022). [cited by examiner]
Hou et al., “Affordance Transfer Learning for Human-Object Interaction Detection,” Presented at Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, Jun. 20-… [cited by applicant]
Hou et al., “Detecting Human-Object Interaction via Fabricated Compositional Learning,” Presented at Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, Jun… [cited by applicant]
Hou et al., “Visual Compositional Learning for Human-Object Interaction Detection,” Presented at Proceedings of the European Conference on Computer Vision 2020, Glasgow, UK, Aug. 23-28, 2020, 17 pages. [cited by applicant]
Kato et al., “Compositional Learning for Human Object Interaction,” Presented at Proceedings of the European Conference on Computer Vision 2018, Munich, Germany, Sep. 8-14, 2018; Computer Vision—ECCV 2018, Oct. 9, 2018,… [cited by applicant]
Li et al., “HOI Analysis: Integrating and Decomposing Human-Object Interaction,” Presented at Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, Dec. 6-12, 2020… [cited by applicant]
Liu et al., “ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction Detection,” Presented at Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, Oct. 12-16, 2020, 42… [cited by applicant]
Maraghi et al., “Scaling Human-Object Interaction Recognition in the Video through Zero-Shot Learning,” Comput. Intell. Neuroscience, Jun. 9, 2021, 2021:9922697, 15 pages. [cited by applicant]