IP Library › Granted Patent US 12,288,391
Granted Patent B2
US 12,288,391 · App. 17/743,661 · Granted Apr 29, 2025

Image grounding with modularized graph attentive networks

Inventors: Zhenfang Chen (Cambridge, MA); Chuang Gan (Cambridge, MA); Bo Wu (Cambridge, MA); Pin-Yu Chen (White Plains, NY)
Assignee: International Business Machines Corporation
G06V10/82G06V10/422G06V30/262
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,391
App. No.
17/743,661
Granted
Apr 29, 2025
Kind
B2
Abstract

A system may include a memory and a processor in communication with the memory. The processor may be configured to perform operations. The operations may include receiving an input, extracting features from the input, and mining object relations using the features. The operations may include determining feature vectors using the object relations and generating, using the feature vectors, an output indicating a target region, wherein the target region corresponds to the input.

Claims (51)

1. A system, said system comprising:

a memory; and

a processor in communication with said memory, said processor being configured to perform operations, said operations comprising:

receiving an input;

extracting features from said input, wherein extracting features from said input comprises:

using an attention network to extract textual features for a plurality of specific modules, the plurality of specific modules including at least a subject module, a location module, and a relation module, wherein the attention network parses different components of the input for each specific module, including parsing a subject component for the subject module, a location component for the location module, and a relation component for the relation module;

mining object relations using said features;

determining feature vectors using said object relations; and

generating, using said feature vectors, an output indicating a target region, wherein said target region corresponds to said input.

2. The system of claim 1 , wherein:

said input includes a textual statement with a plurality of textual modules.

3. The system of claim 1 , said operations further comprising:

obtaining visual features for a plurality of regions in said input by performing message passing.

4. The system of claim 1 , said operations further comprising:

applying a text-guided graph attentive neural network to a module of said input.

5. The system of claim 1 , said operations further comprising:

determining visual regional node aggregate context information based on neighborhood nodes.

6. The system of claim 1 , said operations further comprising:

matching, bilinearly, modularized similarity between modularized textual features of said input and modularized visual features of said input.

7. A computer-implemented method, said method comprising:

receiving an input;

extracting features from said input, wherein extracting features from said input comprises:

using an attention network to extract textual features for a plurality of specific modules, the plurality of specific modules including at least a subject module, a location module, and a relation module, wherein the attention network parses different components of the input for each specific module, including parsing a subject component for the subject module, a location component for the location module, and a relation component for the relation module;

mining object relations using said features;

determining feature vectors using said object relations; and

generating, using said feature vectors, an output indicating a target region, wherein said target region corresponds to said input.

8. The computer-implemented method of claim 7 , wherein:

said input includes a textual statement with a plurality of textual modules.

9. The computer-implemented method of claim 8 , further comprising:

obtaining visual features for a plurality of regions in said input by performing message passing.

10. The computer-implemented method of claim 7 , further comprising:

applying a text-guided graph attentive neural network to a module of said input.

11. The computer-implemented method of claim 7 , further comprising:

determining visual regional node aggregate context information based on neighborhood nodes.

12. The computer-implemented method of claim 7 , further comprising:

matching, bilinearly, modularized similarity between modularized textual features of said input and modularized visual features of said input.

13. The computer-implemented method of claim 7 , further comprising:

aggregating, linearly, similarities over a plurality of modules to obtain a final similarity, wherein said final similarity is used in generating said output.

14. A computer program product, said computer program product comprising a computer readable storage medium having program instructions embodied therewith, said program instructions executable by a processor to cause said processor to perform a function, said function comprising:

receiving an input;

extracting features from said input, wherein extracting features from said input comprises:

using an attention network to extract textual features for a plurality of specific modules, the plurality of specific modules including at least a subject module, a location module, and a relation module, wherein the attention network parses different components of the input for each specific module, including parsing a subject component for the subject module, a location component for the location module, and a relation component for the relation module;

mining object relations using said features;

determining feature vectors using said object relations; and

generating, using said feature vectors, an output indicating a target region, wherein said target region corresponds to said input.

15. The computer program product of claim 14 , said function further comprising:

applying a text-guided graph attentive neural network to a module of said input.

16. The computer program product of claim 14 , said function further comprising:

determining visual regional node aggregate context information based on neighborhood nodes.

17. The computer program product of claim 14 , said function further comprising:

matching, bilinearly, modularized similarity between modularized textual features of said input and modularized visual features of said input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2022
From: CHEN, ZHENFANG; GAN, CHUANG; WU, BO; CHEN, PIN-YU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059899/0535 →
Continuity (1)
Related Publication 20230368510A1 · Nov 16, 2023
References Cited (14)
US 20210224601A1 · Chen · 2021 [cited by examiner]
US 20210248376A1 · Zhao · 2021 [cited by examiner]
CN 110110043A · 2019 [cited by applicant]
Antol et al., “VQA: Visual Question Answering”, arXiv:1505.00468v7 [cs.CL] Oct. 27, 2016, 25 pages, doi.org/10.48550/arXiv.1505.00468. [cited by applicant]
He et al., “Mask R-CNN”, arxiv.org, Computer Vision and Pattern Recognition (cs.CV), submitted on Mar. 20, 2017 (v1), last revised Jan. 24, 2018 (v3), retrieved from the Internet < Mar. 10, 2022 at 1:23 PM EST>, arXiv:1… [cited by applicant]
Lee et al., “Stacked Cross Attention for Image-Text Matching”, Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 201-216, Abstract only, 1 page. [cited by applicant]
Liu et al., “Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing”, arXiv:1903.00839v2 [cs.CV] Apr. 2, 2019, 10 pages, doi.org/10.48550/arXiv.1903.00839. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, Recommendations of the National Institute of Standards and Technology, NIST Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Pal et al., “Image Retrieval: A Literature Review”, International Journal of Advanced Research in Computer Engineering and Technology (IJARCET), vol. 2, Issue 6, Jun. 2013, doi.org/10.13140/2.1.1051.4565, 5 pages. [cited by applicant]
Qi et al., “REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments”, arXiv:1904.10151v2 [cs.CV] Jan. 6, 2020, 16 pages, doi.org/10.48550/arXiv.1904.10151. [cited by applicant]
Rohrbach et al., “Grounding of Textual Phrases in Images by Reconstruction”, 18 pages, arXiv:1511.03745v4 [cs.CV] Feb. 17, 2017. [cited by applicant]
Shrestha et al., “MAGNet: Multi-Region Attention-Assisted Grounding of Natural Language Queries at Phrase Level”, arXiv:2006.03776v1 [cs.CV] Jun. 6, 2020, 13 pages, doi.org/10.48550/arXiv.2006.03776. [cited by applicant]
Yu et al., “MAttNet: Modular Attention Network for Referring Expression Comprehension”, arXiv:1801.08186v3 [cs.CV] Mar. 27, 2018, doi.org/10.48550/arXiv.1801.08186, 15 pages. [cited by applicant]
Yu et al., “Modeling Context in Referring Expressions”, arXiv:1608.00272v3 [cs.CV] Aug. 10, 2016, doi.org/10.48550/arXiv.1608.00272, 19 pages. [cited by applicant]
Cited By (1)
US 12,423,524