IP Library Granted Patent US 12,462,525
Granted Patent B2
US 12,462,525 · App. 17/969,505 · Granted Nov 4, 2025

Co-learning object and relationship detection with density aware loss

Inventors: Maksims Volkovs (Toronto, CA); Cheng Chang (Toronto, CA); Guangwei Yu (Toronto, CA); Himanshu Rai (Toronto, CA); Yichao Lu (Toronto, CA)
Assignee: The Toronto-Dominion Bank
G06V10/764G06V10/7715G06V10/774
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,525
App. No.
17/969,505
Granted
Nov 4, 2025
Kind
B2
Abstract

An object detection model and relationship prediction model are jointly trained with parameters that may be updated through a joint backbone. The offset detection model predicts object locations based on keypoint detection, such as a heatmap local peak, enabling disambiguation of objects. The relationship prediction model may predict a relationship between detected objects and be trained with a joint loss with the object detection model. The loss may include terms for object connectedness and model confidence, enabling training to focus first on highly-connected objects and later on lower-confidence items.

Claims (28)

1 . A system for image processing with object relationship detection, comprising:

a processor that executes instructions; and

a non-transitory computer-readable medium having instructions executable by the processor for:

determining a set of objects in an image based on an object detection model applied to the image, each object in the set of objects having a predicted object class;

for each object in the set of objects, determining a set of object relationship features based on a portion of a relation feature map of the image and the predicted object class, the relation feature map determined from a set of backbone layers shared with the object detection model; and

for one or more pairs of objects in the set of objects, predicting a relationship class for the pair of objects with a relationship prediction model based on the set of object relationship features of the respective objects.

2 . The system of claim 1 , wherein the object detection model identifies objects based on a central keypoint of the object.

3 . The system of claim 1 , wherein the object detection model determines an object class heatmap for an image based on a visual feature map of the image and determines the set of objects based on local peaks of each object class in the object class heatmap.

4 . The system of claim 3 , wherein the object detection model further determines an offset matrix and a size matrix based on the visual feature map and each object in the set of objects has a bounding box with coordinates determined based on the corresponding position of the local peaks in the offset matrix and the size matrix.

5 . The system of claim 1 , wherein the object detection model determines objects based on a visual feature map and the relation feature map is based on the visual feature map.

6 . The system of claim 1 , wherein for the one or more pairs of objects, the relationship prediction model predicts the relationship class with a direction between the objects.

7 . The system of claim 1 , wherein the instructions are further executable for training the object detection model features jointly with the relationship prediction model.

8 . The system of claim 7 , wherein parameters of the object detection model and relationship prediction model are continuously differentiable through joint backbone layers shared by the object detection model and the relationship prediction model.

9 . The system of claim 7 , wherein the relationship prediction model is trained with a loss function for the relationship prediction model that includes a component for the predicted relationship class, a predicted subject object class and a predicted object class.

10 . The system of claim 7 , wherein a loss function for the relationship prediction model includes a density and confidence-based loss that increases the weight for objects having a relatively high number of relationships to other objects and decreases the weight for relationship predictions having a relatively high confidence of predicted relationship class.

11 . A method for image processing with object relationship detection, comprising:

determining a set of objects in an image based on an object detection model applied to the image, each object in the set of objects having a predicted object class;

for each object in the set of objects, determining a set of object relationship features based on a portion of a relation feature map of the image and the predicted object class, the relation feature map determined from a set of backbone layers shared with the object detection model; and

for one or more pairs of objects in the set of objects, predicting a relationship class for the pair of objects with a relationship prediction model based on the set of object relationship features of the respective objects.

12 . The method of claim 11 , wherein the object detection model identifies objects based on a central keypoint of the object.

13 . The method of claim 11 , wherein the object detection model determines an object class heatmap for an image based on a visual feature map of the image and determines the set of objects based on local peaks of each object class in the object class heatmap.

14 . The method of claim 13 , wherein the object detection model further determines an offset matrix and a size matrix based on the visual feature map and each object in the set of objects has a bounding box with coordinates determined based on the corresponding position of the local peaks in the offset matrix and the size matrix.

15 . The method of claim 11 , wherein the object detection model determines objects based on a visual feature map and the relation feature map is based on the visual feature map.

16 . The method of claim 11 , wherein for the one or more pairs of objects, the relationship prediction model predicts the relationship class with a direction between the objects.

17 . The method of claim 11 , the instructions being further for training the object detection model features jointly with the relationship prediction model.

18 . The method of claim 17 , wherein parameters of the object detection model and relationship prediction model are continuously differentiable through joint backbone layers shared by the object detection model and the relationship prediction model.

19 . The method of claim 17 , wherein the relationship prediction model is trained with a loss function for the relationship prediction model that includes a component for the predicted relationship class, a predicted subject object class and a predicted object class.

20 . The method of claim 17 , wherein a loss function for the relationship prediction model includes a density and confidence-based loss that increases the weight for objects having a relatively high number of relationships to other objects and decreases the weight for relationship predictions having a relatively high confidence of predicted relationship class.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2025
From: VOLKOVS, MAKSIMS; CHANG, CHENG; YU, GUANGWEI; RAI, HIMANSHU; LU, YICHAO
To: THE TORONTO-DOMINION BANK
Reel/Frame 072242/0134 →
Continuity (2)
Provisional Application 63270416 · Oct 21, 2021
Related Publication 20230131935A1 · Apr 27, 2023
References Cited (88)
US 10262214B1 · Kim · 2019 [cited by examiner]
US 10269125B1 · Kim · 2019 [cited by examiner]
US 10789288B1 · Ranzinger · 2020 [cited by examiner]
US 11188794B2 · Yao · 2021 [cited by examiner]
US 11537811B2 · Shen · 2022 [cited by examiner]
US 11580333B2 · Bondugula · 2023 [cited by examiner]
US 11943184B2 · Back · 2024 [cited by examiner]
US 12175384B2 · Wu · 2024 [cited by examiner]
US 20150248586A1 · Gaidon · 2015 [cited by examiner]
US 20150294192A1 · Lan · 2015 [cited by examiner]
US 20160378861A1 · Eledath · 2016 [cited by examiner]
US 20170308753A1 · Wu · 2017 [cited by examiner]
US 20170330059A1 · Novotny · 2017 [cited by examiner]
US 20170344884A1 · Lin · 2017 [cited by examiner]
US 20180025249A1 · Liu · 2018 [cited by examiner]
US 20190073524A1 · Yi · 2019 [cited by examiner]
US 20190073553A1 · Yao · 2019 [cited by examiner]
US 20190095716A1 · Shrestha · 2019 [cited by examiner]
US 20190102658A1 · Wang · 2019 [cited by examiner]
US 20190279045A1 · Li · 2019 [cited by examiner]
US 20200086879A1 · Lakshmi Narayanan · 2020 [cited by examiner]
US 20200134375A1 · Zhan · 2020 [cited by examiner]
US 20200143169A1 · Vaezi Joze · 2020 [cited by examiner]
US 20200160087A1 · Redmon · 2020 [cited by examiner]
US 20200302230A1 · Chang · 2020 [cited by examiner]
US 20200356842A1 · Guo · 2020 [cited by examiner]
US 20210089841A1 · Mithun · 2021 [cited by examiner]
US 20210103776A1 · Jiang · 2021 [cited by examiner]
US 20210124993A1 · Singh · 2021 [cited by examiner]
US 20210192194A1 · Chi · 2021 [cited by examiner]
US 20210256588A1 · Moosaei · 2021 [cited by examiner]
US 20210264557A1 · Mao · 2021 [cited by examiner]
US 20210295155A1 · Vijayakumar · 2021 [cited by examiner]
US 20210319242A1 · Cholakkal · 2021 [cited by examiner]
US 20220076002A1 · Luo · 2022 [cited by examiner]
US 20220157048A1 · Ting · 2022 [cited by examiner]
US 20220157054A1 · Lin · 2022 [cited by examiner]
US 20220231979A1 · Back · 2022 [cited by examiner]
US 20220382553A1 · Li · 2022 [cited by examiner]
US 20230021551A1 · Lu · 2023 [cited by examiner]
US 20230030987A1 · Townsend · 2023 [cited by examiner]
US 20230177835A1 · Griffin · 2023 [cited by examiner]
US 20240161461A1 · Zu · 2024 [cited by examiner]
US 20240221426A1 · Xu · 2024 [cited by examiner]
CN 111462282A · 2020 [cited by applicant]
CN 113065587A · 2021 [cited by applicant]
CN 113240033A · 2021 [cited by applicant]
International Search Report and Written Opinion issued in PCT/CA2022/051546 on Jan. 1, 2023; 9 pages. [cited by applicant]
Bengio, et al., “Representation Learning: A Review and New Perspectives,” IEEE Transactions on Pattern Analysis and Machine Intelligence, arXiv preprint arXiv:1206.5538v3 [cs.LG] Apr. 23, 2014, 30 pages; https://arxiv.o… [cited by applicant]
Carion, et al., “End-to-End Object Detection with Transformers,” arXiv:2005.12872v3 [cs.CV], May 28, 2020, 26 pages; https://arxiv.org/pdf/2005.12872.pdf. [cited by applicant]
Chen, et al., “Counterfactual Critic Multi-Agent Training for Scene Graph Generation,” IEEE/CVF International Conference on Computer Vision, 2019, 11 pages; https://openaccess.thecvf.com/content_ICCV_2019/papers/Chen_Co… [cited by applicant]
Chen, et al., “Knowledge-Embedded Routing Network for Scene Graph Generation,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, 9 pages; https://openaccess.thecvf.com/content_CVPR_2019/papers/Chen_K… [cited by applicant]
Dai, et al., “Detecting Visual Relationships with Deep Relational Networks,” IEEE conference on computer vision and Pattern recognition, 2017, 11 pages; https://openaccess.thecvf.com/content_cvpr_2017/papers/Dai_Detecti… [cited by applicant]
Deng, et al., “ImageNet: A Large-Scale Hierarchical Image Database,” IEEE conference on computer vision and pattern recognition, Jun. 20, 2009, 8 pages; https://projet.liris.cnrs.fr/imagine/pub/proceedings/CVPR-2009/dat… [cited by applicant]
Everingham, et al., “The PASCAL Visual Object Classes (VOC) Challenge,” International journal of computer vision, vol. 88, No. 2, Sep. 9, 2009, 36 pages; http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.167.6629… [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition,” IEEE conference on computer vision and pattern recognition, 2016, 9 pages; https://openaccess.thecvf.com/content_cvpr_2016/papers/He_Deep_Residual_Learning_CVP… [cited by applicant]
Krishna, et al., “Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations,” International journal of computer vision, vol. 123, No., 1, Feb. 6, 2017, 42 pages; https://link.springer.com/… [cited by applicant]
Kuznetsova, et al., “The Open Images Dataset V4 Unified image classification, object detection, and visual relationship detection at scale,” International Journal of Computer Vision, vol. 128, No. 7, arXiv:1811.00982v2 … [cited by applicant]
Law, et al., “CornerNet: Detecting Objects as Paired Keypoints,” European conference on computer vision, 2018, 17 pages; https://openaccess.thecvf.com/content_ECCV_2018/papers/Hei_Law_CornerNet_Detecting_Objects_ECCV_20… [cited by applicant]
Lecun, et al., “Gradient-based learning applied to document recognition,” IEEE, 86, No. 11, Dec. 1998, 47 pages; http://lushuangning.oss-cn-beijing.aliyuncs.com/CNN%E5%AD%A6%E4%B9%A0%E7%B3%BB%E5%88%97/Gradient-Based_Lea… [cited by applicant]
Li, et al., “VIP-CNN: Visual Phrase Guided Convolutional Neural Network,” IEEE conference on computer vision and pattern recognition, 2017, 10 pages; https://openaccess.thecvf.com/content_cvpr_2017/papers/Li_ViP-CNN_Vis… [cited by applicant]
Liang, et al., “Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection,” IEEE conference on computer vision and pattern recognition, 2017, 10 pages; https://openaccess.thecvf.co… [cited by applicant]
Lin, et al., “Focal Loss for Dense Object Detection,” IEEE international conference on computer vision, 2017, 9 pages; https://openaccess.thecvf.com/content_ICCV_2017/papers/Lin_Focal_Loss_for_ICCV_2017_paper.pdf. [cited by applicant]
Lin, et al., “Microsoft COCO: Common Objects in Context,” In European conference on computer vision, Sep. 6, 2014, 16 pages; https://link.springer.com/content/pdf/10.1007/978-3-319-10602-1_48.pdf. [cited by applicant]
Liu, et al., “ConceptNet—a practical commonsense reasoning tool-kit,” BT Technology Journal, vol. 22, No. 4, Oct. 2004, 16 pages; http://eridanus.cz/id32402/jazyk/jazykove%282da/aplikovana%281_lingvistika/Ontologie/Word… [cited by applicant]
Liu, et al., “SSD: Single Shot MultiBox Detector,” European conference on computer vision, arXiv: 1512.02325v5 [cs.CV], Dec. 29, 2016, 17 pages; http://www.cs.toronto.edu/˜bonner/courses/2020s/csc2547/papers/discriminat… [cited by applicant]
Lu, et al., “Learning Effective Visual Relationship Detector on 1 GPU,” arXiv:1912.06185v1 [cs.CV], Dec. 12, 2019, 8 pages; https://arxiv.org/pdf/1912.06185.pdf. [cited by applicant]
Lu, et al., “Visual Relationship Detection with Language Priors,” European conference on computer vision, Oct. 8, 2016, 19 pages; https://www-cs.stanford.edu/people/ranjaykrishna/vrd/vrd.pdf. [cited by applicant]
Newell, et al., “Pixels to Graphs by Associative Embedding,” Advances in neural information processing systems 30, 2017, 10 pages; https://proceedings.neurips.cc/paper/2017/file/84438b7aae55a0638073ef798e50b4ef-Paper.pd… [cited by applicant]
Qi, et al., “Attentive Relational Networks for Mapping Images to Scene Graphs,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, 10 pages; https://openaccess.thecvf.com/content_CVPR_2019/papers/Qi_A… [cited by applicant]
Ren, et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Advances in neural information processing systems, vol. 28, 2015, 9 pages; https://proceedings.neurips.cc/paper/2015/file/14… [cited by applicant]
Simonyan, et al., “Very Deep Convolutional Networks For Large-Scale Image Recognition,” arXiv: 1409.1556v6 [cs.CV], Apr. 2, 2015, 14 pages; https://arxiv.org/pdf/1409.1556.pdf%E3%80%82. [cited by applicant]
Tai, et al., “Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks,” arXiv:1503.00075v3 [cs.CL], May 30, 2015, 11 pages; https:/arxiv.org/pdf/1503.00075.pdf?ref=https://codemonkey.link. [cited by applicant]
Tang, et al., “Learning to Compose Dynamic Tree Structures for Visual Contexts,” IEEE/CVF conference on computer vision and pattern recognition, 2019, 10 pages; https://openaccess.thecvf.com/content_CVPR_2019/papers/Tan… [cited by applicant]
Tang, et al., “Unbiased Scene Graph Generation from Biased Training,” IEEE/CVF conference on computer vision and pattern recognition, 2020, 10 pages; https://openaccess.thecvf.com/content_CVPR_2020/papers/Tang_Unbiased_… [cited by applicant]
Tian, et al., “FCOS: Fully Convolutional One-Stage Object Detection,” IEEE/CVF international conference on computer vision, 2019, 10 pages; https://openaccess.thecvf.com/content_ICCV_2019/papers/Tian_FCOS_Fully_Convolut… [cited by applicant]
Wang, et al., “Sketching Image Gist: Human-Mimetic Hierarchical Scene Graph Generation,” arXiv:2007.08760v1 [cs.CV], Jul. 17, 2020, 30 pages; https://arxiv.org/pdf/2007.08760.pdf. [cited by applicant]
Xu, et al., “Reasoning-RCNN: Unifying Adaptive Global Reasoning into Large-scale Object Detection,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, 10 pages; https://openaccess.thecvf.com/content_C… [cited by applicant]
Yin, et al., “Zoom-Net: Mining Deep Feature Interactions for Visual Relationship Recognition,” European Conference on Computer Vision (ECCV), 2018, 17 pages; https://openaccess.thecvf.com/content_ECCV_2018/papers/Guojun… [cited by applicant]
Yu, et al., “Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation,” IEEE International conference on computer vision, 2017, 9 pages; https://openaccess.thecvf.com/content_ICCV_2017/… [cited by applicant]
Zareian, et al., “Bridging Knowledge Graphs to Generate Scene Graphs,” arXiv:2001.02314v4 [cs.CV], Jul. 18, 2020, 29 pages; https://arxiv.org/pdf/2001.02314.pdf. [cited by applicant]
Zellers, et al., “Neural Motifs: Scene Graph Parsing with Global Context,” IEEE conference on computer vision and pattern recognition, 2018, 10 pages; https://openaccess.thecvf.com/content_cvpr_2018/papers/Zellers_Neura… [cited by applicant]
Zhang, et al., “Graphical Contrastive Losses for Scene Graph Parsing,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, 9 pages; https://openaccess.thecvf.com/content_CVPR_2019/papers/Zhang_Graphica… [cited by applicant]
Zhang, et al., “PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN,” IEEE International conference on computer vision, 2017, 9 pages; https://openaccess.thecvf.com/content_ICCV_2017/papers/… [cited by applicant]
Zhang, et al., “Visual Translation Embedding Network for Visual Relation Detection,” In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, 9 pages; https://openaccess.thecvf.com/content… [cited by applicant]
Zhao, et al., “Collaborative Training between Region Proposal Localization and Classification for Domain Adaptive Object Detection,” arXiv:2009.08119v2 [cs.CV], Sep. 18, 2020, 16 pages; https://arxiv.org/pdf/2009.08119.… [cited by applicant]
Zhou, et al., “Objects as points,” arXiv:1904.07850v2 [cs.CV], Apr. 25, 2019, 12 pages; https://arxiv.org/pdf/1904.07850.pdf%EF%BC%9B%E5%9B%9E%E5%BD%923%E4%B8%AA%E7%82%B9%EF%BC%8C%E5%B7%A6%E4%B8%8A%EF%BC%80%E5%8F%B3%E4%… [cited by applicant]
Zhuang, et al., “Towards Context-aware Interaction Recognition for Visual Relationship Detection,” IEEE International conference on computer vision, 2017, 10 pages; https://openaccess.thecvf.com/content_ICCV_2017/ paper… [cited by applicant]