IP Library › Granted Patent US 12,315,214
Granted Patent B2
US 12,315,214 · App. 17/902,025 · Granted May 27, 2025

Object detection

Inventors: Carlo Biffi (London, GB); Steven George Mcdonagh (London, GB); Ales Leonardis (London, GB); Sarah Parisot (London, GB)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06V10/454G06N20/20G06V10/764G06V10/771G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,214
App. No.
17/902,025
Granted
May 27, 2025
Kind
B2
Abstract

A device for categorising regions in images is disclosed. The device comprising: an input for receiving a first set of images, and defining one or more regions of for each image of the first set of images and a categorisation for the one or more regions, and a second set of images, and a categorisation for each image of the second set; and a processor configured to train a first machine learning algorithm to categorise features in images by: processing the images of the first and second set using the first algorithm to estimate feature regions in the images and a categorisation for each of the feature regions, and training the first algorithm in dependence on the categorisations received for the images of the first and second sets.

Claims (27)

1. A device for categorizing regions in images, comprising:

an input for receiving a first set of images, and defining one or more regions of each image of the first set of images and a categorization for the one or more regions, and a second set of images, and a categorization for each image of the second set of images; and

a processor configured to train a first machine learning algorithm to categorize features in images by performing operations comprising:

processing the images of the first and second sets of images using the first machine learning algorithm to estimate feature regions in the images and a categorization for each of the feature regions, and training the first machine learning algorithm in dependence on the categorizations received for the images of the first and second sets of images,

wherein the processor is further configured to:

train the first machine leaning algorithm to form an estimate of confidence of the categorization estimated for at least some of the images of the second set, and

train a second machine learning algorithm to categorize regions in images by selecting, for use in training the second machine learning algorithm, a subset of the images of the second set as categorized by the first machine learning algorithm such that a weight of each of the images of that subset in training the second machine learning algorithm is dependent on an estimated confidence for the respective image.

2. The device as claimed in claim 1 , wherein the processor is further configured to train the first machine learning algorithm by taking at least some of the first and second sets of images as input to the first algorithm multiple times.

3. The device as claimed in claim 1 , wherein the first machine learning algorithm operates in dependence on a set of stored weights, and the processor is further configured to train those weights in dependence on the performance of the first machine learning algorithm in classifying the images of the first and second sets.

4. The device as claimed in claim 1 , wherein the first machine learning algorithm comprises a first sub-part for estimating feature regions in images and a second sub-part for estimating a categorization for a feature region, and the processor is further configured to train the first sub-part to estimate feature regions in images of the second set of images which are categorized by the second sub-part to match the received categorization for the respective image.

5. The device as claimed in claim 4 , wherein the second machine learning algorithm is configured to train the second sub-part.

6. A method for categorizing regions in images, comprising:

receiving a first set of images, and defining one or more regions of each image of the first set of images and a categorization for the one or more regions, and a second set of images, and a categorization for each image of the second set of images;

training, by a processor, a first machine learning algorithm to categorize features in images by performing operations comprising:

processing the images of the first and second sets of images using the first machine learning algorithm to estimate feature regions in the images and a categorization for each of the feature regions, and training the first machine learning algorithm in dependence on the categorizations received for the images of the first and second sets of images;

forming, for at least some of the images of the second set, an estimate of confidence of the categorization estimated for those images; and

training a second machine learning algorithm to categorize regions in images, by selecting, for use in training the second machine learning algorithm, a subset of the images of the second set of images as categorized by the first machine learning algorithm such that a weight of each of the images of that subset in training the second machine learning algorithm is dependent on an estimated confidence for the respective image.

7. The device as claimed in claim 1 , wherein the second machine learning algorithm implements a machine learning architecture different from the first machine learning algorithm.

8. The device as claimed in claim 1 , wherein the second machine learning algorithm implements less internal feedback than the first machine learning algorithm.

9. The device as claimed in claim 1 , wherein the first and second machine learning algorithms comprise a common feature encoder.

10. The device as claimed in claim 1 , the device being configured to train the first and second machine learning algorithms simultaneously in dependence on each other's performance.

11. The device as claimed in claim 1 , wherein the first and second machine learning algorithms are end-to-end trainable.

12. The method as claimed in claim 6 , further comprising, after training the second machine learning algorithm, implementing the second machine learning algorithm, without the first machine learning algorithm in a device for categorizing features in images.

13. The method as claimed in claim 12 , wherein the first machine learning algorithm comprises a first sub-part for estimating feature regions in images and a second sub-part for estimating a categorization for a feature region, and the method further comprises training the first sub-part to estimate feature regions in images of the second set of images which are categorized by the second sub-part to match the received categorization for the respective image.

14. The method as claimed in claim 13 , wherein the second machine learning algorithm is configured to train the second sub-part.

15. The method as claimed in claim 12 , further comprising training the first machine learning algorithm by taking at least some of the first and second sets of images as input to the first algorithm multiple times.

16. The method as claimed in claim 12 , wherein the first machine learning algorithm operates in dependence on a set of stored weights, and the method further comprises training those weights in dependence on the performance of the first machine learning algorithm in classifying the images of the first and second sets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2025
From: BIFFI, CARLO; MCDONAGH, STEVEN GEORGE; LEONARDIS, ALES; PARISOT, SARAH
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 070267/0573 →
Continuity (2)
Continuation PCTEP2020055751 · Mar 4, 2020
Related Publication 20230115167A1 · Apr 13, 2023
References Cited (16)
CN 201341028Y · 2009 [cited by applicant]
CN 101982922A · 2011 [cited by applicant]
CN 109474034A · 2019 [cited by applicant]
CN 110224457A · 2019 [cited by applicant]
JP H08116604A · 1996 [cited by applicant]
WO 2017032254A1 · 2017 [cited by applicant]
EHSOD: CAM-Guided End-to-end Hybrid-Supervised Object Detection with Cascade Refinement. Fang et al. (Year: 2020). [cited by examiner]
Uijlings, J.R., Van De Sande, K.E., Gevers, T., Smeulders, A.W.: Selective search for object recognition. International journal of computer vision 104(2), 154-171 (2013). [cited by applicant]
Zitnick, C.L., Dollár, P .: Edge boxes: Locating object proposals from edges. In: European conference on computer vision. pp. 391-405. Springer (2014). [cited by applicant]
Bilen, H., Vedaldi, A.: Weakly supervised deep detection networks. In: Proceedingsof the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2846-2854 (2016). [cited by applicant]
Girshick, R.: Fast r-cnn. In: Proceedings of the IEEE international conference oncomputer vision. pp. 1440 1448 (2015). [cited by applicant]
Tang, Peng, et al. “Pcl: Proposal cluster learning for weakly supervised object detection.” IEEE transactions on pattern analysis and machine intelligence 42.1 (2018): 176-191. [cited by applicant]
Pan, Tianxiang, et al. “Low shot box correction for weakly supervised object detection.” Proceedings of the 28th International Joint Conference on Artificial Intelligence. AAAI Press, 2019. [cited by applicant]
Fang, L., Xu, H., Liu, Z., Parisot, S., Li, Z.: EHSOD: CAM-Guided End-to-EndHybrid-Supervised Object Detection with cascade refinement. In: AAAI Press 2020. [cited by applicant]
Pardo, A., Xu, M., Thabet, A., Arbelaez, P., Ghanem, B.: Baod: Budget-aware object detection. Arxiv 2019. [cited by applicant]
International Search Report and Written Opinion issued in PCT/CN2021/078659, dated Jun. 4, 2021, 9 pages. [cited by applicant]