IP Library › Granted Patent US 12,499,674
Granted Patent B2
US 12,499,674 · App. 18/101,133 · Granted Dec 16, 2025

Computer-implemented method, data processing apparatus and computer program for object detection

Inventor: David Nicholson Griffiths (London, GB)
Assignee: Fujitsu Limited
G06V10/82G06V10/765G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,674
App. No.
18/101,133
Granted
Dec 16, 2025
Kind
B2
Abstract

A computer-implemented method of training an object detector, the method comprising: training an embedding neural network using, as an input, cropped images from an image dataset, wherein training the embedding neural network is performed using a self-supervised learning approach and the trained embedding neural network translates input images into a lower dimensional representation; and training an object detector neural network by, for images of the image dataset, repeatedly: passing an image through the object detector neural network to obtain proposed coordinates of an object within the image, cropping the image to the proposed coordinates to obtain a cropped image, passing the cropped image through the trained embedding neural network to obtain a cropped image representation, passing an exemplar through the trained embedding neural network to obtain an exemplar representation, wherein the exemplar is a cropped manually labelled image bounding a known object, computing a distance in embedding space between the cropped image representation and the exemplar representation, computing a gradient of the cropped image representation and the exemplar representation with respect to the distance, and passing the gradient into the object detector neural network for use in backpropagation to optimise the object detector neural network.

Claims (111)

1 . A computer-implemented method of training an object detector, the method comprising:

training an embedding neural network using, as an input, cropped images from an image dataset, wherein training the embedding neural network is performed using a self-supervised learning approach and the trained embedding neural network translates input images into a lower dimensional representation; and

training an object detector neural network by, for images of the image dataset, repeatedly:

passing an image through the object detector neural network to obtain proposed coordinates of an object within the image,

cropping the image to the proposed coordinates to obtain a cropped image,

passing the cropped image through the trained embedding neural network to obtain a cropped image representation,

passing an exemplar through the trained embedding neural network to obtain an exemplar representation, wherein the exemplar is a cropped manually labelled image bounding a known object,

computing a distance in embedding space between the cropped image representation and the exemplar representation,

computing a gradient of the cropped image representation and the exemplar representation with respect to the distance using a finite difference method, wherein the finite difference method includes:

cropping the image to the proposed coordinates with a shift to obtain a shifted cropped image,

passing the shifted cropped image through the trained embedding neural network to obtain a shifted cropped image representation,

computing a second distance in embedding space between the shifted cropped image representation and the exemplar representation, and

computing the gradient as the difference between the distance and the second distance, and

passing the gradient into the object detector neural network for use in backpropagation to optimise the object detector neural network.

2 . The method of claim 1 , further comprising: optimising the object detector neural network by minimising the distance between the cropped image representation and the exemplar representation using the gradient in backpropagation.

3 . The method of claim 2 , wherein optimising the object detector neural network comprises minimising a distance-based loss function for each cropped image representation of the images and the exemplar representation.

4 . The method of claim 3 , wherein the distance-based loss function corresponds to a sum of L 1 loss and focal loss for each cropped image representation and the exemplar representation.

5 . The method of claim 1 , wherein training the object detector neural network further comprises scaling each cropped image such that all scaled cropped images are of the same size.

6 . The method of claim 5 , further comprising scaling the exemplar such that the scaled exemplar is the same size as the scaled cropped images.

7 . The method of claim 1 , wherein the method uses a plurality of exemplars for repeatedly training the object detector neural network, the method further comprising:

obtaining an exemplar representation for each exemplar, and

computing the distance and the gradient for each cropped image with respect to each exemplar representation.

8 . The method of claim 7 , wherein the method uses at least the same number of exemplars as there are classes of objects to be detected.

9 . The method of claim 1 , further comprising randomly initializing weights of a target embedding neural network, the target embedding neural network comprising the same structure as the embedding neural network, and wherein training the embedding neural network comprises, for images of the image dataset, repeatedly:

augmenting a cropped image to generate a first augmented view and a second augmented view;

passing the first augmented view through the embedding neural network to obtain a lower dimensional representation of the first augmented view;

passing the second augmented view through the target embedding network to obtain a lower dimensional representation of the second augmented view;

minimising a similarity loss between the embedding neural network and the target embedding network using stochastic gradient descent optimisation with respect to the weights of the embedding neural network.

10 . The method of claim 9 , wherein the stochastic gradient descent optimisation comprises updating the weights of the target embedding neural network as a moving average of the weights of the embedding neural network.

11 . The method of claim 9 , wherein augmenting the cropped image comprises applying at least one of the following augmentations:

colour jittering;

greyscale conversion;

Gaussian blurring;

horizontal flipping;

vertical flipping; and

random crop and resizing, optionally wherein

augmenting the cropped image comprises probabilistically applying a plurality of augmentations to the cropped image, each augmentation applied with a corresponding probability.

12 . The method of claim 1 , wherein the method is for detecting an object in an image enhancement or analysis process.

13 . The method of claim 12 , wherein the method is for detecting an object in an autonomous vehicle image analysis process.

14 . The method of claim 12 , wherein the method is for detecting an object in a railway mapping image analysis process.

15 . A computer-implemented method of object detection, the method comprising:

training an embedding neural network using, as an input, cropped images from an image dataset, wherein training the embedding neural network is performed using a self-supervised learning approach and the trained embedding neural network translates input images into a lower dimensional representation; and

training an object detector neural network by, for images of the image dataset, repeatedly:

passing an image through the object detector neural network to obtain proposed coordinates of an object within the image,

cropping the image to the proposed coordinates to obtain a cropped image,

passing the cropped image through the trained embedding neural network to obtain a cropped image representation,

passing an exemplar through the trained embedding neural network to obtain an exemplar representation, wherein the exemplar is a cropped manually labelled image bounding a known object,

computing a distance in embedding space between the cropped image representation and the exemplar representation,

computing a gradient of the cropped image representation and the exemplar representation with respect to the distance using a finite difference method, wherein the finite difference method includes:

cropping the image to the proposed coordinates with a shift to obtain a shifted cropped image,

passing the shifted cropped image through the trained embedding neural network to obtain a shifted cropped image representation,

computing a second distance in embedding space between the shifted cropped image representation and the exemplar representation, and

computing the gradient as the difference between the distance and the second distance, and

passing the gradient into the object detector neural network for use in backpropagation to optimise the object detector neural network;

receiving an input image;

passing the input image into the trained object detector neural network; and

outputting coordinates and object class of any objects detected within the input image.

16 . A data processing apparatus comprising a memory and a processor, the memory comprising instructions which, when executed by the processor:

train an embedding neural network using, as an input, cropped images from an image dataset, wherein training the embedding neural network is performed using a self-supervised learning approach and the trained embedding neural network translates input images into a lower dimensional representation; and

train an object detector neural network by, for images of the image dataset, repeatedly:

passing an image through the object detector neural network to obtain proposed coordinates of an object within the image,

cropping the image to the proposed coordinates to obtain a cropped image,

passing the cropped image through the trained embedding neural network to obtain a cropped image representation,

passing an exemplar through the trained embedding neural network to obtain an exemplar representation, wherein the exemplar is a cropped manually labelled image bounding a known object,

computing a distance in embedding space between the cropped image representation and the exemplar representation,

computing a gradient of the cropped image representation and the exemplar representation with respect to the distance using a finite difference method, wherein the finite difference method includes:

cropping the image to the proposed coordinates with a shift to obtain a shifted cropped image,

passing the shifted cropped image through the trained embedding neural network to obtain a shifted cropped image representation,

computing a second distance in embedding space between the shifted cropped image representation and the exemplar representation, and

computing the gradient as the difference between the distance and the second distance, and

passing the gradient into the object detector neural network for use in backpropagation to optimise the object detector neural network.

17 . A non-transitory computer readable medium having a computer program stored thereon, wherein the computer program, when executed by a computer, cause the computer to:

train an embedding neural network using, as an input, cropped images from an image dataset, wherein training the embedding neural network is performed using a self-supervised learning approach and the trained embedding neural network translates input images into a lower dimensional representation; and

train an object detector neural network by, for images of the image dataset, repeatedly:

passing an image through the object detector neural network to obtain proposed coordinates of an object within the image,

cropping the image to the proposed coordinates to obtain a cropped image,

passing the cropped image through the trained embedding neural network to obtain a cropped image representation,

passing an exemplar through the trained embedding neural network to obtain an exemplar representation, wherein the exemplar is a cropped manually labelled image bounding a known object,

computing a distance in embedding space between the cropped image representation and the exemplar representation,

computing a gradient of the cropped image representation and the exemplar representation with respect to the distance using a finite difference method, wherein the finite difference method includes:

cropping the image to the proposed coordinates with a shift to obtain a shifted cropped image,

passing the shifted cropped image through the trained embedding neural network to obtain a shifted cropped image representation,

computing a second distance in embedding space between the shifted cropped image representation and the exemplar representation, and

computing the gradient as the difference between the distance and the second distance, and

passing the gradient into the object detector neural network for use in backpropagation to optimise the object detector neural network.

18 . A computer-implemented method of training an object detector, the method comprising:

training an embedding neural network using, as an input, cropped images from an image dataset, wherein training the embedding neural network is performed using a self-supervised learning approach and the trained embedding neural network translates input images into a lower dimensional representation; and

training an object detector neural network by, for images of the image dataset, repeatedly:

passing an image through the object detector neural network to obtain proposed coordinates of an object within the image,

cropping the image to the proposed coordinates to obtain a cropped image,

passing the cropped image through the trained embedding neural network to obtain a cropped image representation,

passing a plurality of exemplars through the trained embedding neural network to obtain a corresponding plurality of exemplar representations, wherein each exemplar is a cropped manually labelled image bounding a known object, and wherein the method uses at least the same number of exemplars as there are classes of objects to be detected,

computing a distance in embedding space between the cropped image representation and the plurality of exemplar representations,

computing a gradient of the cropped image representation and the exemplar representations with respect to the distance, and

passing the gradient into the object detector neural network for use in backpropagation to optimise the object detector neural network.

19 . A computer-implemented method of training an object detector, the method comprising:

training an embedding neural network using, as an input, cropped images from an image dataset, wherein training the embedding neural network is performed using a self-supervised learning approach and the trained embedding neural network translates input images into a lower dimensional representation; and

training an object detector neural network by, for images of the image dataset, repeatedly:

passing an image through the object detector neural network to obtain proposed coordinates of an object within the image,

cropping the image to the proposed coordinates to obtain a cropped image,

passing the cropped image through the trained embedding neural network to obtain a cropped image representation,

passing an exemplar through the trained embedding neural network to obtain an exemplar representation, wherein the exemplar is a cropped manually labelled image bounding a known object,

computing a distance in embedding space between the cropped image representation and the exemplar representation,

computing a gradient of the cropped image representation and the exemplar representation with respect to the distance, and

passing the gradient into the object detector neural network for use in backpropagation to optimise the object detector neural network,

the method further comprising randomly initializing weights of a target embedding neural network, the target embedding neural network comprising the same structure as the embedding neural network, and wherein training the embedding neural network comprises, for images of the image dataset, repeatedly:

augmenting a cropped image to generate a first augmented view and a second augmented view;

passing the first augmented view through the embedding neural network to obtain a lower dimensional representation of the first augmented view;

passing the second augmented view through the target embedding network to obtain a lower dimensional representation of the second augmented view; and

minimising a similarity loss between the embedding neural network and the target embedding network using stochastic gradient descent optimisation with respect to the weights of the embedding neural network.

20 . The method of claim 19 , wherein the stochastic gradient descent optimisation comprises updating the weights of the target embedding neural network as a moving average of the weights of the embedding neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2023
From: GRIFFITHS, DAVID NICHOLSON
To: FUJITSU LIMITED
Reel/Frame 062480/0130 →
Priority Claims (1)
EP 22159287 · Feb 28, 2022 · regional
Continuity (1)
Related Publication 20230298335A1 · Sep 21, 2023
References Cited (35)
US 11157744B2 · Barzelay et al. · 2021 [cited by applicant]
US 20190102678A1 · Chang · 2019 [cited by examiner]
US 20200152333A1 · Tomasev · 2020 [cited by examiner]
US 20210133623A1 · Amrani et al. · 2021 [cited by applicant]
US 20210216828A1 · Ramaiah et al. · 2021 [cited by applicant]
US 20220164585A1 · Ayvaci · 2022 [cited by examiner]
US 20220175455A1 · Ungi · 2022 [cited by examiner]
CN 108985334A · 2018 [cited by applicant]
CN 112464879A · 2021 [cited by applicant]
WO 2016103651A1 · 2016 [cited by applicant]
Sun et al: “FSCE: Few-Shot Object Detection via Contrastive Proposal Encoding”, 2021 IEEE/CF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, XP034006880, Jun. 20, 2021, pp. 7352-7362. [cited by applicant]
Khosla et al., “Supervised Contrastive Learning” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), arxiv.org, Cornell University Library, XP081835152, Dec. 10, 2020, pp. 1-13. [cited by applicant]
Chen et al., “A Simple Framework for Contrastive Learning of Visual Representation”, arxiv. org, Cornell University Library, XP081632474, Feb. 13, 2020, 11 pages. [cited by applicant]
Le Jeune et al., “Experience feedback using Representation Learning for Few-Shot Object Detection on Aerial Images”, 20th IEEE International Conference on Machine Learning and Application (ICMLA), IEEE, XP034045785, Dec… [cited by applicant]
Jean-Bastien Grill et al., “Bootstrap your own latent: a new approach to self supervised learning”, arxiv.org, Cornell University Library, Jun. 14, 2020, XP081700812. [cited by applicant]
Mai Ngoc Kien “How to compute gradients in Tensorflow and Pytorch”, medium.comn/codex, XP55945624, Retrieved from the Internet: URL: https://medium.com/codex/how-to-compute-gradients-in-tensorflow-and-pytorch-59a585752f… [cited by applicant]
Paszke et al., “Automatic differentiation in PyTorch ”, NIPS 2017 Workshop Autodiff Submission, 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, pp. 1-4. [cited by applicant]
Abadi et al., “TensorFlow: A System for Large-Scale Machine Learning” In 12th (USENIX) Symposium on Operating Systems Design and Implementation (OSDI'16), Nov. 2-4, 2016, pp. 265-283. [cited by applicant]
Johnson et al., “CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2901-2910. [cited by applicant]
Lin et al., “Focal Loss for Dense Object Detection”, Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980-2988. [cited by applicant]
Qi et al., “PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space”, Computer Vision and Pattern Recognition (cs.CV), arXivpreprint, arXiv:1706.02413v1 [cs.CV], Jun. 7, 2017, pp. 1-14. [cited by applicant]
Thomas et al., “KPConv: Flexible and Deformable Convolution for Point Clouds”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 6411-6420. [cited by applicant]
Yang et al., “Instance Localization for Self-Supervised Detection Pretraining”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 3987-3996. [cited by applicant]
Chen et al., “A Simple Framework for Contrastive Learning of Visual Representations”, Proceedings of the 37th International Conference on Machine Learning, PMLR, vol. 119, Nov. 2020, pp. 1597-1607. [cited by applicant]
Chen et al., “Exploring Simple Siamese Representation Learning”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 15750-15758. [cited by applicant]
Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 779-788. [cited by applicant]
Liu et al., “SSD: Single Shot MultiBox Detector”, Springer, Part of the Lecture Notes in Computer Science book series (LNIP, vol. 9905), In European Conference on Computer Vision, Oct. 2016, pp. 21-37. [cited by applicant]
Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, Advances in Neural Information Processing Systems, vol. 28, 2015, pp. 1-9. [cited by applicant]
Girshick et al., “Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 580-587. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778. [cited by applicant]
Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, Computer Vision and Pattern Recognition (cs.CV), arXiv preprint, arXiv:1409.1556v6 [cs.CV], Apr. 10, 2015, pp. 1-14. [cited by applicant]
Extended European Search Report mailed on Sep. 12, 2022, received for EP Application 22159287.6, 11 pages. [cited by applicant]
Bsun: “FSCE/fsdet/engine/defaults.py”, Mar. 13, 2021 (Mar. 13, 2021), XP93031657, github.corn, Retrieved from the Internet: URL:https://github.corn/megvii-research/FSCE/blob/main/fsdet/engine/defaults.py [retrieved on 2… [cited by applicant]
Papers With Code: “RolPool Explained”, Sep. 8, 2023 (Sep. 8, 2023), XP93080265, Retrieved from the Internet: URL:https://paperswithcode.corn/method/roi-pooling [retrieved on Sep. 8, 2023], 5 pages. [cited by applicant]
Extended European Search Report issued Sep. 20, 2023, in corresponding European Patent Application No. 22 159 287.6, 6 pages. [cited by applicant]