IP Library › Granted Patent US 12,505,345
Granted Patent B2
US 12,505,345 · App. 17/824,571 · Granted Dec 23, 2025

Method and system for scene graph generation

Inventors: Efthymia Tsamoura (Chertsey, GB); Davide Buffelli (Chertsey, GB)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/08G06N3/04G06V10/751G06V10/771G06V10/774G06V10/82G06V10/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,345
App. No.
17/824,571
Granted
Dec 23, 2025
Kind
B2
Abstract

Broadly speaking, the disclosure generally relates to relates to a computer-implemented methods and systems for scene graph generation, and in particular for training a machine learning, ML, model to generate a scene graph. The method includes inputting training a training image into a machine learning model, outputting a predicted label for at least two objects in the training image and a predicted label for a relationship between the at least two objects. The training method includes calculating a loss, which takes into account both a supervised loss calculated by comparing the predicted labels to the actual labels for the training image, and a logic-based loss calculated by comparing the predicted labels to stored integrity constraints comprising common-sense knowledge. Advantageously, this means that the performance of the model is improved without increasing processing at inference-time.

Claims (61)

1 . A computer-implemented method for training a machine learning model comprising a plurality of neural network layers for scene graph generation, SGG, comprising:

inputting a training image into the machine learning model, the training image depicting a scene comprising at least two objects;

obtaining, from the machine learning model, a predicted label for the each of the at least two objects and a predicted label for a relationship between the at least two objects;

determining a loss; and

updating, based on the determined loss, weights for the plurality of neural network layers of the machine learning model,

wherein the determining the loss comprises:

determining a supervised loss by comparing the predicted label for the each of the at least two objects and the predicted label for the relationship between the at least two objects to actual labels corresponding to the training image; and

determining a logic-based loss by comparing the predicted label for the each of the at least two objects and the predicted label for the relationship between the at least two objects to a plurality of stored integrity constraints, wherein the integrity constraints comprise positive integrity constraints expressing permissible relationships between objects, and negative integrity constraints expressing impermissible relationships between objects, and

wherein the determining the logic-based loss comprises:

selecting a maximally non-satisfied subset of the plurality of stored integrity constraints, and

determining the logic-based loss based on the maximally non-satisfied subset, and

wherein the selecting the maximally non-satisfied subset comprises:

obtaining output predictions of the machine learning model with highest confidences; ordering the obtained output predictions by decreasing confidence;

removing any prediction of the obtained output predictions for which there is no corresponding negative integrity constraint in the plurality of stored integrity constraints;

selecting random subsets of the obtained output predictions;

determining a loss for each of the random subsets; and

selecting a maximum loss corresponding to the random subsets as the logic-based loss.

2 . The method of claim 1 , wherein the logic-based loss is inversely proportional to a satisfaction function, which measures how closely the output prediction for the each of the at least two objects and the output prediction for the relationship between the at least two objects satisfies the plurality of stored integrity constraints.

3 . The method of claim 1 , wherein the logic-based loss is determined using probabilistic logic.

4 . The method of claim 1 , wherein the logic-based loss is determined using fuzzy logic.

5 . A system comprising:

at least one processor, coupled to memory, configured to train a machine learning model comprising a plurality of neural network layers for scene graph generation, SGG, by:

inputting a training image into the machine learning model, the training image depicting a scene comprising at least two objects;

obtaining, from the machine learning model, a predicted label for the each of the at least two objects and a predicted label for a relationship between the at least two objects;

determining a loss; and

updating, based on the determined loss, weights for the plurality of neural network layers of the machine learning model,

wherein the determining the loss comprises:

determining a supervised loss by comparing the predicted label for the each of the at least two objects and the predicted label for the relationship between the at least two objects to actual labels corresponding to the training image; and

determining a logic-based loss by comparing the predicted label for the each of the at least two objects and the predicted label for the relationship between the at least two objects to a plurality of stored integrity constraints, wherein the integrity constraints comprise positive integrity constraints expressing permissible relationships between objects, and negative integrity constraints expressing impermissible relationships between objects,

wherein the at least one processor is configured to train the machine learning model by:

selecting a maximally non-satisfied subset of the plurality of stored integrity constraints, and

determining the logic-based loss based on the maximally non-satisfied subset, and

wherein the at least one processor is configured to train the machine learning model by:

obtaining output predictions of the machine learning model with highest confidences: ordering the obtained output predictions by decreasing confidence;

removing any prediction of the obtained output predictions for which there is no corresponding negative integrity constraint in the plurality of stored integrity constraints;

selecting random subsets of the obtained output predictions;

determining a loss for each of the random subsets; and

selecting a maximum loss corresponding to the random subsets as the logic-based loss.

6 . The system of claim 5 , wherein the logic-based loss is inversely proportional to a satisfaction function, which measures how closely the output prediction for the each of the at least two objects and the output prediction for the relationship between the at least two objects satisfies the plurality of stored integrity constraints.

7 . The system of claim 5 , wherein the logic-based loss is determined using probabilistic logic.

8 . The system of claim 5 , wherein the logic-based loss is determined using fuzzy logic.

9 . A non-transitory machine-readable medium comprising instructions that, when executed by at least one processor of an electronic device, cause the at least one processor to train a machine learning model comprising a plurality of neural network layers for scene graph generation, SGG, by:

inputting a training image into the machine learning model, the training image depicting a scene comprising at least two objects;

obtaining, from the machine learning model, a predicted label for the each of the at least two objects and a predicted label for a relationship between the at least two objects;

determining a loss; and

updating, based on the determined loss, weights for the plurality of neural network layers of the machine learning model,

wherein calculating the determining the loss comprises:

determining a supervised loss by comparing the predicted label for the each of the at least two objects and the predicted label for the relationship between the at least two objects to actual labels corresponding to the training image; and

determining a logic-based loss by comparing the predicted label for the each of the at least two objects and the predicted label for the relationship between the at least two objects to a plurality of stored integrity constraints, wherein the integrity constraints comprise positive integrity constraints expressing permissible relationships between objects, and negative integrity constraints expressing impermissible relationships between objects,

wherein the instructions when executed cause the at least one processor to train the machine learning model by:

selecting a maximally non-satisfied subset of the plurality of stored integrity constraints, and

determining the logic-based loss based on the maximally non-satisfied subset, and

wherein the instructions when executed cause the at least one processor to train the machine learning model by:

obtaining output predictions of the machine learning model with highest confidences; ordering the obtained output predictions by decreasing confidence;

removing any prediction of the obtained output predictions for which there is no corresponding negative integrity constraint in the plurality of stored integrity constraints;

selecting random subsets of the obtained output predictions;

determining a loss for each of the random subsets; and

selecting a maximum loss corresponding to the random subsets as the logic-based loss.

10 . The non-transitory machine-readable medium of claim 9 , wherein the logic-based loss is inversely proportional to a satisfaction function, which measures how closely the output prediction for the each of the at least two objects and the output prediction for the relationship between the at least two objects satisfies the plurality of stored integrity constraints.

11 . The non-transitory machine-readable medium of claim 9 , wherein the logic-based loss is determined using probabilistic logic.

12 . The non-transitory machine-readable medium of claim 9 , wherein the logic-based loss is determined using fuzzy logic.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2022
From: TSAMOURA, EFTHYMIA; BUFFELLI, DAVIDE
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 060448/0448 →
Priority Claims (2)
GR 20210100345 · May 25, 2021 · national
GB 2117872 · Dec 10, 2021 · national
Continuity (1)
Related Publication 20220391704A1 · Dec 8, 2022
References Cited (66)
US 11361550B2 · Herdade · 2022 [cited by examiner]
US 12026226B2 · Savvides · 2024 [cited by examiner]
US 20170132526A1 · Cohen · 2017 [cited by examiner]
US 20180089540A1 · Merler · 2018 [cited by examiner]
US 20200364553A1 · Amos · 2020 [cited by examiner]
US 20200401835A1 · Zhao · 2020 [cited by examiner]
US 20210374489A1 · Prakash · 2021 [cited by examiner]
Hai Wan, et al., Iterative Visual Relationship Detection via Commonsense Knowledge Graph, Big Data Research, vol. 23, Feb. 15, 2021, 100175, ISSN 2214-5796, https://doi.org/10.1016/j.bdr.2020.100175. (https://www.scienc… [cited by examiner]
Communication dated May 24, 2022, issued by the Great Britain Intellectual Property Office in Great Britain Patent Application No. GB2117872.8. [cited by applicant]
Hou et al., “Relational Reasoning using Prior Knowledge for Visual Captioning,” arXiv:1906.01290v1 [cs.CV], Jun. 4, 2019, Total 10 pages. [cited by applicant]
Zareian et al., “Learning Visual Commonsense for Robust Scene Graph Generation,” arXiv:2006.09623v2 [cs.CV], Jul. 18, 2020, Total 20 pages. [cited by applicant]
Somak Aditya et al., “Explicit Reasoning over End-to-End Neural Architectures for Visual Question Answering”, The Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18), 2018, 9 pages total. [cited by applicant]
Somak Aditya et al., “Integrating Knowledge and Reasoning in Image Understanding”, Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI-19), 2019, 8 pages total. [cited by applicant]
Somak Aditya et al., “Combining Knowledge and Reasoning through Probabilistic Soft Logic for Image Puzzle Solving”, 2018, 11 pages total. [cited by applicant]
Somak Aditya et al., “Visual Commonsense for Scene Understanding Using Perception, Semantic Parsing and Reasoning”, Logical Formalizations of Commonsense Reasoning: Papers from the 2015 AAAI Spring Symposium, 2015, 8 pa… [cited by applicant]
Stephen H. Bach et al., “Hinge-Loss Markov Random Fields and Probabilistic Soft Logic”, Journal of Machine Learning Research, 18, 2017, 67 pages total. [cited by applicant]
Mark Chavira et al., “On probabilistic inference by weighted model counting”, Artificial Intelligence, 172, 2007, 28 pages total. [cited by applicant]
Tianshui Chen et al., “Knowledge-Embedded Routing Network for Scene Graph Generation”, Computer Vision Foundation, 2019, 9 pages total. [cited by applicant]
Bo Dai et al., “Detecting Visual Relationships with Deep Relational Networks”, Computer Vision Foundation, 2017, 11 pages total. [cited by applicant]
Wang-Zhou Dai et al., “Bridging Machine Learning and Logical Reasoning by Abductive Learning”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 12 pages total. [cited by applicant]
Artur S. d'Avila Garcez, Krysia Broda, and Dov M. Gabbay, “Neural-symbolic learning systems: foundations and applications”, Perspectives in neural computing, Springer, 2002, 1 page total. [cited by applicant]
Ivan Donadello et al., “Logic Tensor Networks for Semantic Image Interpretation”, Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI-17), 2017, 7 pages total. [cited by applicant]
Marc Fischer et al., “DL2: Training and Querying Neural Networks with Logic”, Proceedings of the 36th International Conference on Machine Learning, 2019, 11 pages total. [cited by applicant]
Alexander L. Gaunt et al., “Differentiable Programs with Neural Libraries”, Proceedings of the 34th International Conference on Machine Learning, 2017, 10 pages total. [cited by applicant]
Jiuxiang Gu et al., “Scene Graph Generation with External Knowledge and Image Reconstruction”, arXiv:1904.00560v1, Apr. 2019, 10 pages total. [cited by applicant]
Zhiting Hu et al., “Harnessing Deep Neural Networks with Logic Rules”, ACL, 2016, 13 pages total. [cited by applicant]
Zhiting Hu et al., “Deep Neural Networks with Massive Learned Knowledge”, EMNLP, 2016, 10 pages total. [cited by applicant]
Alex Kendall et al., “Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics”, Computer Vision Foundation, 2018, 10 pages total. [cited by applicant]
Ranjay Krishna et al., “Visual Genome Connecting Language and Vision Using Crowdsourced Dense Image Annotations”, arXiv:1602.07332v1, Feb. 2016, 44 pages total. [cited by applicant]
Yikang Li et al., “Scene Graph Generation from Objects, Phrases and Region Captions”, arXiv:1707.09700v2, ICCV, Sep. 2017, 10 pages total. [cited by applicant]
Yuanzhi Liang et al., “VrR-VG: Refocusing Visually-Relevant Relationships”, Computer Vision Foundation, ICCV, 2019, 10 pages total. [cited by applicant]
Ben London et al., “Collective Activity Detection using Hinge-loss Markov Random Fields”, Computer Vision Foundation, 2013, 6 pages total. [cited by applicant]
Cewu Lu et al., “Visual Relationship Detection with Language Priors”, ECCV, 2016, 19 pages total. [cited by applicant]
Robin Manhaeve et al., “DeepProbLog: Neural Probabilistic Logic Programming”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), 2018, 11 pages total. [cited by applicant]
Kenneth Marino et al., “The More You Know: Using Knowledge Graphs for Image Classification”, Computer Vision Foundation, 2017, 9 pages total. [cited by applicant]
Giuseppe Marra et al., “Relational Neural Machines”, 24th European Conference on Artificial Intelligence—ECAI, 2020, 8 pages total. [cited by applicant]
Giuseppe Marra et al., “Integrating Learning and Reasoning with Deep Logic Models”, 2019, 16 pages total. [cited by applicant]
Pasquale Minervini et al., “Adversarially Regularising Neural NLI Models to Integrate Logical Background Knowledge”, Proceedings of the 22nd Conference on Computational Natural Language Learning (CoNLL 2018), 2018, 10 p… [cited by applicant]
Mengshi Qi et al., “Attentive Relational Networks for Mapping Images to Scene Graphs”, Computer Vision Foundation, 2019, 10 pages total. [cited by applicant]
Tim Rocktäschel et al., “End-to-End Differentiable Proving”, 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, 13 pages total. [cited by applicant]
Tim Rocktäschel et al., “Injecting Logical Background Knowledge into Embeddings for Relation Extraction”, ACL, 2015, 11 pages total. [cited by applicant]
Maarten Sap et al., “ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning”, arXiv:1811.00146v3, Association for the Advancement of Artificial Intelligence, Feb. 2019, 9 pages total. [cited by applicant]
Robyn Speer et al., “ConceptNet 5.5: An Open Multilingual Graph of General Knowledge”, Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence (AAAI-17), 2017, 8 pages total. [cited by applicant]
Kaihua Tang et al., “Unbiased Scene Graph Generation from Biased Training”, Computer Vision Foundation, 2020, 10 pages total. [cited by applicant]
Kaihua Tang et al., “Learning to Compose Dynamic Tree Structures for Visual Contexts”, Computer Vision Foundation, 2019, 10 pages total. [cited by applicant]
Efthymia Tsamoura et al., “Neural-Symbolic Integration: A Compositional Perspective”, The Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21), 2021, 10 pages total. [cited by applicant]
Emile Van Krieken et al., “Semi-Supervised Learning Using Differentiable Reasoning”, arXiv:1908.04700v1, Journal of Applied Logics—IfCoLog Journal of Logics and their Applications, vol. 6. No. 4, Aug. 2019, 18 pages tot… [cited by applicant]
Emile van Krieken et al., “Analyzing Differentiable Fuzzy Implications”, Proceedings of the 17th International Conference on Principles of Knowledge Representation and Reasoning (KR 2020) Special Session on KR and Machi… [cited by applicant]
Hai Wang et al., “Deep Probabilistic Logic: A Unifying Framework for Indirect Supervision”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, 12 pages total. [cited by applicant]
Peng Wang et al., “FVQA: Fact-based Visual Question Answering”, arXiv:1606.05433v4, IEEE Transactions on Pattern Analysis and Machine Intelligence, Aug. 2017, 16 pages total. [cited by applicant]
Po-Wei Wang et al., “SATNet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver”, arXiv:1905.12149v1, Proceedings of the 36th International Conference on Machine Learning, May 2019… [cited by applicant]
Wenbin Wang et al., “Exploring Context and Visual Pattern of Relationship for Scene Graph Generation”, Computer Vision Foundation, 2019, 10 pages total. [cited by applicant]
Wenya Wang et al., “Integrating Deep Learning with Logic Fusion for Information Extraction”, The Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-20), 2020, 8 pages total. [cited by applicant]
Sanghyun Woo et al., “LinkNet: Relational Embedding for Scene Graph”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), 2018, 11 pages total. [cited by applicant]
Qi Wu et al., “Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources”, Computer Vision Foundation, 2016, 9 pages total. [cited by applicant]
Danfei Xu et al., “Scene Graph Generation by Iterative Message Passing”, Computer Vision Foundation, 2017, 10 pages total. [cited by applicant]
Jingyi Xu et al., “A Semantic Loss Function for Deep Learning with Symbolic Knowledge”, Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages total. [cited by applicant]
Jianwei Yang et al., “Graph R-CNN for Scene Graph Generation”, Computer Vision Foundation, ECCV, 2018, 16 pages total. [cited by applicant]
Zhun Yang et al., “NeurASP: Embracing Neural Networks into Answer Set Programming”, Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI-20), 2020, 8 pages total. [cited by applicant]
Guojun Yin et al., “Zoom-Net: Mining Deep Feature Interactions for Visual Relationship Recognition”, Computer Vision Foundation, ECCV, 2018, 17 pages total. [cited by applicant]
Alireza Zareian et al., “Learning Visual Commonsense for Robust Scene Graph Generation”, ECCV, 2020, 16 pages total. [cited by applicant]
Rowan Zellers et al., “Neural Motifs: Scene Graph Parsing with Global Context”, Computer Vision Foundation, 2018, 10 pages total. [cited by applicant]
Hanwang Zhang et al., “Visual Translation Embedding Network for Visual Relation Detection”, Computer Vision Foundation, 2017, 9 pages total. [cited by applicant]
Yuke Zhu et al., “Reasoning About Object Affordances in a Knowledge Base Representation”, ECCV, vol. 8690, 2014, 29 pages total. [cited by applicant]
Joseph Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection”, Computer Vision Foundation, 2016, 10 pages total. [cited by applicant]
Shaoqing Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, arXiv:1506.01497v2, NeurIPS, Sep. 2015, 10 pages total. [cited by applicant]