IP Library › Granted Patent US 12,399,890
Granted Patent B2
US 12,399,890 · App. 17/087,943 · Granted Aug 26, 2025

Scene graph modification based on natural language commands

Inventors: Quan Tran (San Jose, CA); Zhe Lin (Bellevue, WA); Xuanli He (Clayton, AU); Walter Chang (San Jose, CA); Trung Bui (San Jose, CA); Franck Dernoncourt (Sunnyvale, CA)
Assignee: ADOBE INC.
G06F16/243G06F40/56G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,399,890
App. No.
17/087,943
Granted
Aug 26, 2025
Kind
B2
Abstract

Systems and methods for natural language processing are described. Embodiments are configured to receive a structured representation of a search query, wherein the structured representation comprises a plurality of nodes and at least one edge connecting two of the nodes, receive a modification expression for the search query, wherein the modification expression comprises a natural language expression, generate a modified structured representation based on the structured representation and the modification expression using a neural network configured to combine structured representation features and natural language expression features, and perform a search based on the modified structured representation.

Claims (55)

1. A method for natural language processing, comprising:

receiving a query comprising an initial natural language expression including a plurality of elements;

generating a structured representation of the query, wherein the structured representation comprises a plurality of nodes corresponding to the plurality of elements and at least one edge connecting two of the plurality of nodes;

generating, using a graph encoder, a plurality of node embeddings corresponding to the plurality of nodes based at least in part on the at least one edge;

receiving a modification command for the query, wherein the modification command comprises a natural language expression describing a change to an element of the plurality of elements in the query, and wherein the modification command specifies a symbolic change to the structured representation of the query;

generating, using a text encoder, a plurality of text embeddings representing the modification command;

generating, using a feature fusion network, a combined embedding representing the query and the modification command based on the plurality of node embeddings and the plurality of text embeddings; and

generating a modified structured representation by decoding the combined embedding, wherein the modified structured representation includes an updated plurality of nodes and an updated plurality of edges representing the query with the change indicated by the modification command.

2. The method of claim 1 , wherein:

the plurality of node embeddings and the plurality of text embeddings are located in a common embedding space.

3. The method of claim 1 , wherein:

the structured representation comprises a scene graph and the modified structured representation comprises a modified scene graph.

4. The method of claim 1 , further comprising:

performing a search based on the modified structured representation.

5. The method of claim 1 , wherein generating the modified structured representation further comprises:

applying a transformer network to a plurality of combined embeddings including the plurality of node embeddings and the plurality of text embeddings.

6. The method of claim 1 , wherein generating the modified structured representation further comprises:

modifying the plurality of node embeddings based on the modification command to obtain a plurality of updated node embeddings; and

modifying the plurality of text embeddings based on the structured representation to obtain a plurality of updated text embeddings, wherein a plurality of combined embeddings include the plurality of updated node embeddings and the plurality of updated text embeddings.

7. The method of claim 1 , further comprising:

performing a search based on the modified structured representation; and

retrieving a plurality of images corresponding to the modified structured representation based on the search.

8. An apparatus for natural language processing, comprising:

a graph generator configured to generate a structured representation of a query comprising an initial natural language expression including a plurality of elements, wherein the structured representation comprises a plurality of nodes corresponding to the plurality of elements and at least one edge connecting two of the plurality of nodes;

a graph encoder configured to generate a plurality of node embeddings corresponding to the plurality of nodes based at least in part on the at least one edge;

a text encoder configured to generate a plurality of text embeddings representing a modification command for the query, wherein the modification command comprises a natural language expression describing a change to an element of the plurality of elements in the query, and wherein the modification command specifies a symbolic change to the structured representation of the query;

a feature fusion network configured to generate a combined embedding representing the query and the modification command based on the plurality of node embeddings and the plurality of text embeddings, and generate a modified structured representation by decoding the combined embedding;

a node decoder configured to generate an updated plurality of nodes based on a plurality of combined embeddings that combines the plurality of node embeddings and the plurality of text embeddings; and

an edge decoder configured to generate an updated plurality of edges based on the plurality of combined embeddings and the updated plurality of nodes.

9. The apparatus of claim 8 , further comprising:

a search component configured to perform a search based on the modified structured representation comprising the updated plurality of nodes and the updated plurality of edges.

10. The apparatus of claim 8 , wherein:

the graph encoder comprises a sparsely connected transformer network.

11. The apparatus of claim 8 , wherein:

the text encoder comprises a transformer network.

12. The apparatus of claim 8 , wherein:

the feature fusion network comprises a stage gating mechanism.

13. The apparatus of claim 8 , wherein:

the feature fusion network comprises a cross-attention network.

14. The apparatus of claim 8 , wherein:

the node decoder comprises a recurrent neural network (RNN).

15. The apparatus of claim 8 , wherein:

the edge decoder comprises an adjacency matrix decoder.

16. The apparatus of claim 8 , wherein:

the edge decoder comprises a flat edge-level decoder.

17. A method for training a neural network, the method comprising:

identifying training data including a plurality of annotated training examples, wherein each of the annotated training examples comprises a source structured representation of a query, a target structured representation, and at least one modification command for the query, wherein the query comprises an initial natural language expression including a plurality of elements, wherein the source structured representation comprises a plurality of nodes corresponding to the plurality of elements and at least one edge connecting two of the plurality of nodes, wherein the at least one modification command comprises a natural language expression describing a change to an element of the plurality of elements in the query, and wherein the at least one modification command specifies a symbolic change to the source structured representation of the query;

generating a plurality of node embeddings corresponding to the plurality of nodes based at least in part on the at least one edge for the source structured representation using a graph encoder;

generating a plurality of text embeddings representing the at least one modification command using a text encoder;

generating, using a feature fusion network, a combined embedding representing the query and the at least one modification command based on the plurality of node embeddings and the plurality of text embeddings;

generating a modified structured representation by decoding the combined embedding, wherein the modified structured representation includes an updated plurality of nodes and an updated plurality of edges representing the query with the change indicated by the at least one modification command;

comparing the updated plurality of nodes and the updated plurality of edges to the target structured representation; and

updating the neural network based on the comparison.

18. The method of claim 17 , wherein:

the neural network is updated using an end-to-end training technique wherein parameters of the graph encoder, the text encoder, the feature fusion network, a node decoder, and an edge decoder are updated during each training iteration.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2020
From: TRAN, QUAN; LIN, ZHE; HE, XUANLI; CHANG, WALTER; BUI, TRUNG; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 054256/0246 →
Continuity (1)
Related Publication 20220138185A1 · May 5, 2022
References Cited (68)
US 11720346B2 · Wu · 2023 [cited by examiner]
US 20200133952A1 · Sheinin · 2020 [cited by examiner]
US 20200272915A1 · Tata · 2020 [cited by examiner]
US 20210294781A1 · Fernández Musoles · 2021 [cited by examiner]
US 20220108188A1 · Wu · 2022 [cited by examiner]
CN 109145763A · 2019 [cited by applicant]
CN 110909673A · 2020 [cited by applicant]
CN 111062865A · 2020 [cited by applicant]
“A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation,” Yin et al; Yin, arXiv, Jul. 17, 2020 (Year: 2020). [cited by examiner]
“Variational Lossy Autoencoder,” Chen et al; Chen, arXiv, Mar. 4, 2017 (Year: 2017). [cited by examiner]
“Auto-Encoding Scene Graphs for Image Captioning,” Yang, arXiv, Dec. 11, 2018 (Year: 2018). [cited by examiner]
“Iterative Visual Reasoning Beyond Convolutions,” Chen, arXiv, Mar. 29, 2018 (Year: 2018). [cited by examiner]
10.8 Beam Search—Dive Into Deep Learning (Year: 2020). [cited by examiner]
“Text GNN—Improving Text Encoder via Graph Neural Network in Sponsored Search,” WWW '21, Apr. 19-23, 2021 (Year: 2021). [cited by examiner]
“Iterative GNN-based Decoder for Question Generation,” Fei et al; Fei (Year: 2021). [cited by examiner]
“Autoencoder,” Wikipedia (Year: 2019). [cited by examiner]
“Attention (machine learning),” Wikipedia (Year: 2024). [cited by examiner]
“Adjacency Matrix of Directed Graphs,” GeeksforGeeks (Year: 2024). [cited by examiner]
“Generating Long Sequences with Sparse Transformers,” arXiv, Child et al; Child (Year: 2019). [cited by examiner]
“Generating Semantically Precise Scene Graphs from Textual Descriptions for Improved Image Retrieval,” Stanford University, Schuster et al; Schuster (Year: 2018). [cited by examiner]
What Is a Transformer Model? Nvidia (Year: 2024). [cited by examiner]
“Text Generation from Knowledge Graphs with Graph Transformers,” Proceedings of NAACL-HLT, pp. 2284-2293, Koncel-Kedziorski et al; Koncel-Kedziorski. (Year: 2019). [cited by examiner]
“Scene Graph Generation by Iterative Message Passing,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5410-5419, Xu et al; Xu, (Year: 2017). [cited by examiner]
“Stacked squeeze-and-excitation recurrent residual network for visual-semantic matching,” Pattern Recognition, vol. 105, Wang et al; Wang. (Year: 2020). [cited by examiner]
Peter Anderson, et al, “SPICE: Semantic Propositional Image Caption Evaluation”, In European Conference on Computer Vision, 2016, pp. 382-398. Springer. [cited by applicant]
Dzmitry Bahdanau, et al, “Neural Machine Translation by Jointly Learning to Align and Translate”, arXiv preprint arXiv:1409.0473, 2014, 15 pages. [cited by applicant]
Joost Bastings, et al, “Graph Convolutional Encoders for Syntax-aware Neural Machine Translation”, In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 1957-1967, 2017, Copenhag… [cited by applicant]
Daniel Beck, et al., “Graph-to-Sequence Learning using Gated Graph Neural Networks”, arXiv preprint arXiv:1806.09835, 2018, 13 pages. [cited by applicant]
Jan Buys, et al., “Robust Incremental Neural Semantic Graph Parsing”, In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 2017, pp. 1215-1226. [cited by applicant]
Deng Cai, et al., “Graph Transformer for Graph-to-Sequence Learning”, arXiv preprint arXiv:1911.07470, 2019, 9 pages. [cited by applicant]
Danqi Chen, et al, “A Fast and Accurate Dependency Parser using Neural Networks”, In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 740-750. [cited by applicant]
Rewon Child, et al., “Generating Long Sequences with Sparse Transformers”. arXiv preprint arXiv:1904.10509, 2019, 10 pages. [cited by applicant]
Kyunghyun Cho, et al., “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation”, arXiv preprint arXiv:1406.1078, 2014, 15 pages. [cited by applicant]
Kevin Clark, et al., “Semi-Supervised Sequence Modeling with Cross-View Training”, arXiv preprint arXiv:1809.08370, 2018, 17 pages. [cited by applicant]
Li Dong, et al., “Coarse-to-Fine Decoding for Neural Semantic Parsing”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018, pp. 731-742. [cited by applicant]
Timothy Dozat, et al., “Deep Biaffine Attention for Neural Dependency Parsing”, arXiv preprint arXiv:1611.01734, 2016, 8 pages. [cited by applicant]
Sergey Edunov, et al., “Understanding Back-Translation at Scale”, arXiv preprint arXiv:1808.09381, 2018, 12 pages. [cited by applicant]
Matt Gardner, et al., “AllenNLP: A Deep Semantic Natural Language Processing Platform”, 2017, 6 pages. [cited by applicant]
Zhijiang Guo, et., “Densely Connected Graph Convolutional Networks for Graph-to-Sequence Learning”, Transactions of the Association for Computational Linguis- tics, vol. 7, 2019, pp. 297-312. [cited by applicant]
Xuanli He, et al, “A Pointer Network Architecture for Context-Dependent Semantic Parsing”, In Proceedings of the 17th Annual Workshop of the Australasian Language Technology Association, 2019, pp. 94-99. [cited by applicant]
Srinivasan Iyer, et al., “Learning a Neural Semantic Parser from User Feedback”, In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 2017, pp. 963-973. [cited by applicant]
Mohit Iyyer, et al., “Search-based Neural Structured Learning for Sequential Question Answering”, In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), 2017, p… [cited by applicant]
Sébastien Jean, et al, “Montreal Neural Machine Translation Systems for WMT'15”, In Proceedings of the Tenth Workshop on Statistical Machine Translation, 2015, pp. 134-140. [cited by applicant]
Justin Johnson, et al., “Image Retrieval using Scene Graphs”, In Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3668-3678. [cited by applicant]
Thomas N. Kipf, et al., “Semi-Supervised Classification with Graph Convolutional Networks”, CoRR, abs/1609.02907, 2016, 14 pages. [cited by applicant]
Ranjay Krishna, et al., “Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations”, International Journal of Computer Vision, 123(1):32-73, 2017. [cited by applicant]
Yujia Li, et al., “Learning Deep Generative Models of Graphs”, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 2018, 21 pages. [cited by applicant]
Tsung-Yi Lin, et al. “Microsoft COCO: Common Objects in Context”, In European Conference on Computer Vision, 2014, pp. 740-755. Springer. [cited by applicant]
Minh-Thang Luong, et al., “Effective Approaches to Attention-Based Neural Machine Translation”, arXiv preprint arXiv:1508.04025, 2015, 11 pages. [cited by applicant]
Ramesh Manuvinakurike, et al., “Edit me: A Corpus and a Framework for Understanding Natural Language Image Editing”, In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 201… [cited by applicant]
Julien Perez, et al. “Dialog state tracking, a machine reading approach using Memory Network”, arXiv preprint arXiv:1606.04052, 2016, 10 pages. [cited by applicant]
Liliang Ren, et al., “Towards Universal Dialogue State Tracking”, arXiv preprint arXiv:1810.09587, 2018, 7 pages. [cited by applicant]
Sebastian Schuster, et al., “Generating Semantically Precise Scene Graphs from Textual Descriptions for Improved Image Retrieval”, In Proceedings of the 2015 Workshop on Vision and Language, 2016, pp. 70-80. [cited by applicant]
Piyush Sharma, et al., “Conceptual Captions: A Cleaned, Hypernymed, Image Alt-Text Dataset for Automatic Image Captioning”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol… [cited by applicant]
Martin Simonovsky, et al., “GraphVAE: Towards Generation of Small Graphs Using Variational Autoencoders”, In International Conference on Artificial Neural Networks, pp. 412-422, 2018, Springer. [cited by applicant]
Linfeng Song, et al., “A Graph-to-Sequence Model for AMR-to-Text Generation”, arXiv preprint arXiv:1805.02473, 2018, 11 pages. [cited by applicant]
Shashank Srivastava, et al., “Parsing Natural Language Conversations using Contextual Cues”, IJCAI, 2017, pp. 4089-4095. [cited by applicant]
Alane Suhr, et al., “Learning to Map Context-Dependent Sentences to Executable Formal Queries”, arXiv preprint arXiv:1804.06868, 2018, 21 pages. [cited by applicant]
Ashish Vaswani, et al., “Attention Is All You Need”, In Advances in neural information processing systems, 2017, pp. 5998-6008. [cited by applicant]
Ivan Vendrov, et al., “Order-Embeddings of Images and Language”, arXiv preprint arXiv:1511.06361, 2015, 12 pages. [cited by applicant]
Yu-Siang Wang, et al., “Scene Graph Parsing as Dependency Parsing”, arXiv preprint arXiv:1803.09189, 2018, 11 pages. [cited by applicant]
Zonghan Wu, et al., “A Comprehensive Survey on Graph Neural Networks”, IEEE Transactions on Neural Networks and Learning Systems, 2020, 22 pages. [cited by applicant]
Xu Yang, et al., “Auto-Encoding Scene Graphs for Image Captioning”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 10685-10694. [cited by applicant]
Jiaxuan You, et al., GraphRNN: Generating Realistic Graphs with Deep Auto-regressive Models, arXiv preprint arXiv:1802.08773, 2018, 12 pages. [cited by applicant]
Junru Zhou , et al., “Head-Driven Phrase Structure Grammar Parsing on Penn Treebank”, arXiv preprint arXiv:1907.02684, 2019, 13 pages. [cited by applicant]
OpenAI, et al., “Solving Rubik's Cube With A Robot Hand”, 2019. [cited by applicant]
Khalil Mrini, et al., “Rethinking Self-attention: An Interpretable Self-attentive Encoder-Decoder Parser”, arXiv preprint arXiv:1911.03875, 2019, 8 pages. [cited by applicant]
Office Action issued by the CNIPA on Apr. 9, 2025 in related Chinese Patent Application No. 202110955848.X, 14 pages, in Chinese. [cited by applicant]