IP Library Granted Patent US 12,450,441
Granted Patent B2
US 12,450,441 · App. 18/182,068 · Granted Oct 21, 2025

Methods and systems for generating a semantic computation graph for understanding and grounding referring expressions

Inventors: Zhe Lin (Fremont, CA); Walter W. Chang (San Jose, CA); Scott Cohen (Sunnyvale, CA); Khoi Viet Pham (Hyattsville, MD); Jonathan Brandt (Santa Cruz, CA); Franck Dernoncourt (Sunnyvale, CA)
Assignee: Adobe Inc.
G06F40/30G06F16/532G06F16/55G06F40/205G06F40/295G06N5/02G06N5/04G06N20/00G06V10/40G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,441
App. No.
18/182,068
Granted
Oct 21, 2025
Kind
B2
Abstract

Embodiments of the present invention provide systems, methods, and non-transitory computer storage media for parsing a given input referring expression into a parse structure and generating a semantic computation graph to identify semantic relationships among and between objects. At a high level, when embodiments of the preset invention receive a referring expression, a parse tree is created and mapped into a hierarchical subject, predicate, object graph structure that labeled noun objects in the referring expression, the attributes of the labeled noun objects, and predicate relationships (e.g., verb actions or spatial propositions) between the labeled objects. Embodiments of the present invention then transform the subject, predicate, object graph structure into a semantic computation graph that may be recursively traversed and interpreted to determine how noun objects, their attributes and modifiers, and interrelationships are provided to downstream image editing, searching, or caption indexing tasks.

Claims (30)

1. A computer-implemented method comprising:

receiving input text containing at least one referring expression referencing a first image object, a second image object, and a spatial relationship between the first image object and the second image object;

generating a semantic computation graph comprising a predicate node that represents the spatial relationship between the first image object and the second image object; and

performing an image task based on traversing the semantic computation graph to match the at least one referring expression with a computer vision label associated with an image that contains a detected instance of the first and second image objects.

2. The computer-implemented method of claim 1 , wherein the semantic computation graph comprises a modifier node that represents an object modifier in the referring expression and classifies the object modifier into a modifier type and a value of the modifier type.

3. The computer-implemented method of claim 1 , wherein the predicate node represents the spatial relationship between the first image object and the second image object using a spatial preposition.

4. The computer-implemented method of claim 1 , wherein the image task comprises associating the semantic computation graph with object information from a computer vision system for image editing requests.

5. The computer-implemented method of claim 1 , wherein the image task comprises generating a query intention model to represent the at least one referring expression in an image query.

6. The computer-implemented method of claim 1 , wherein the image task comprises extracting information from an image caption based on the semantic computation graph to create a semantic index for an image search system.

7. The computer-implemented method of claim 1 , wherein the semantic computation graph comprises an object node that represents the first image object and stores or identifies a plurality of hypernyms or synonyms of the first image object, wherein traversing the semantic computation graph to match the at least one referring expression with the computer vision label comprises matching the computer vision label with one of the plurality of hypernyms or synonyms of the first image object.

8. The computer-implemented method of claim 1 , wherein the semantic computation graph comprises an object node that represents the first image object, wherein generating the semantic computation graph comprises using an extensible grounding ontology that expands over time to expand the object node to represent hypernyms or synonyms of the first image object.

9. A system comprising:

one or more hardware processors; and

one or more non-transitory computer storage media storing computer-useable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to execute operations comprising:

receiving input text containing at least one referring expression referencing a first image object, a second image object, and a spatial relationship between the first image object and the second image object;

generating a semantic computation graph comprising a predicate node that represents the spatial relationship between the first image object and the second image object; and

performing an image task based on traversing the semantic computation graph.

10. The system of claim 9 , wherein the semantic computation graph comprises a modifier node that represents an object modifier in the referring expression and classifies the object modifier into a modifier type and a value of the modifier type.

11. The system of claim 9 , wherein the predicate node represents the spatial relationship between the first image object and the second image object using a spatial preposition.

12. The system of claim 9 , wherein the image task comprises associating the semantic computation graph with object information from a computer vision system for image editing requests.

13. The system of claim 9 , wherein the image task comprises generating a query intention model to represent the at least one referring expression in an image query.

14. The system of claim 9 , wherein the image task comprises extracting information from an image caption based on the semantic computation graph to create a semantic index for an image search system.

15. The system of claim 9 , wherein the semantic computation graph comprises an object node that represents the first image object and stores or identifies a plurality of hypernyms or synonyms of the first image object, wherein traversing the semantic computation graph is to match the at least one referring expression with a computer vision label associated with an image that contains a detected instance of the first and second image objects based at least on matching the computer vision label with one of the plurality of hypernyms or synonyms of the first image object.

16. The system of claim 9 , wherein the semantic computation graph comprises an object node that represents the first image object, wherein generating the semantic computation graph comprises using an extensible grounding ontology that expands over time to expand the object node to represent hypernyms or synonyms of the first image object.

17. One or more non-transitory computer storage media having computer-executable instructions embodied thereon that, when executed, by one or more processors, cause the one or more processors to perform operations comprising:

receiving input text containing at least one referring expression referencing a first image object, a second image object, and a spatial relationship between the first image object and the second image object; and

generating a semantic computation graph comprising a predicate node that represents the spatial relationship between the first image object and the second image object.

18. The one or more non-transitory computer storage media of claim 17 , wherein the semantic computation graph comprises a modifier node that represents an object modifier in the referring expression and classifies the object modifier into a modifier type and a value of the modifier type.

19. The one or more non-transitory computer storage media of claim 17 , wherein the predicate node represents the spatial relationship between the first image object and the second image object using a spatial preposition.

20. The one or more non-transitory computer storage media of claim 17 , wherein the semantic computation graph comprises an object node that represents the first image object and stores or identifies a plurality of hypernyms or synonyms of the first image object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2023
From: LIN, ZHE; CHANG, WALTER W.; COHEN, SCOTT; PHAM, KHOI VIET; BRANDT, JONATHAN; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 062949/0820 →
Continuity (2)
Continuation 16775697 · Jan 29, 2020
Related Publication 20230214600A1 · Jul 6, 2023
References Cited (28)
US 7574652B2 · Lennon · 2009 [cited by examiner]
US 10282389B2 · Liang et al. · 2019 [cited by applicant]
US 11263277B1 · Podgorny · 2022 [cited by examiner]
US 11636270B2 · Lin · 2023 [cited by examiner]
US 20070016863A1 · Qu · 2007 [cited by examiner]
US 20100070448A1 · Omoigui · 2010 [cited by examiner]
US 20120166373A1 · Sweeney et al. · 2012 [cited by applicant]
US 20130330008A1 · Zadeh · 2013 [cited by applicant]
US 20140142922A1 · Liang · 2014 [cited by examiner]
US 20140188844A1 · Kogan · 2014 [cited by examiner]
US 20140201126A1 · Zadeh · 2014 [cited by examiner]
US 20150169758A1 · Assom et al. · 2015 [cited by applicant]
US 20160078059A1 · Kang · 2016 [cited by examiner]
US 20160378861A1 · Eledath et al. · 2016 [cited by applicant]
US 20170262412A1 · Liang et al. · 2017 [cited by applicant]
US 20170329760A1 · Rachevsky · 2017 [cited by examiner]
US 20180011611A1 · Pagaime da Silva · 2018 [cited by examiner]
US 20180204111A1 · Zadeh · 2018 [cited by examiner]
US 20180225032A1 · Jones · 2018 [cited by examiner]
US 20180329879A1 · Galitsky · 2018 [cited by examiner]
US 20190318405A1 · Hu · 2019 [cited by examiner]
US 20200175303A1 · Bhat · 2020 [cited by examiner]
US 20210232770A1 · Lin · 2021 [cited by examiner]
US 20230214600A1 · Lin · 2023 [cited by examiner]
WO 2019050968A1 · 2019 [cited by applicant]
Cirik, V., et al., “Using Syntax To Ground Referring Expressions in Natural Images”, In Thirty-Second AAAI Conference on Artificial Intelligence, pp. 1-10 (2018). [cited by applicant]
Elhoseiny, M., et al., “Automatic Annotation of Structured Facts in Images”, arXiv preprint arXiv:1604.00466, pp. 1-19 (Apr. 8, 2016). [cited by applicant]
Yu, L., et al., “MAttNet: Modular Attention Network for Referring Expression Comprehension”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1307-1315 (2018). [cited by applicant]