IP Library Granted Patent US 12,373,980
Granted Patent B2
US 12,373,980 · App. 17/759,785 · Granted Jul 29, 2025

Method and system for localization with supervision

Inventors: Issam Hadj Laradji (Montreal, CA); David Vazquez Bermudez (Montreal, CA); Pau Rodriguez Lopez (Montreal, CA); Rafael Pardinas (Montreal, CA)
Assignee: ServiceNow, Inc.
G06T7/70G06V10/764G06V10/774G06V10/776G06V10/82G06T2207/20076G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,980
App. No.
17/759,785
Granted
Jul 29, 2025
Kind
B2
Abstract

A method for training a machine learning localization model to localize objects belonging to a given class within an image, the method comprising: receiving images each comprising objects of the given class; and for each image: receiving a heat map generated using the machine learning localization model; identifying proposals each corresponding to a potential object, each proposal having associated thereto an initial probability that the proposal corresponds to the potential object; for each proposal, correcting the initial probability using the heat map; selecting given ones of the proposals having a greatest corrected probability, thereby identifying object candidates; and calculating a loss for the machine learning localization model based on a location of the object candidates within the training image and the heat map; and providing the calculated loss to the machine learning localization model.

Claims (41)

1. A computer-implemented method for training a machine learning localization model to localize objects belonging to a given class within an image, the method comprising:

receiving, at a processor, a plurality of training images each comprising at least one object of the given class; and

receiving, at the processor, a heat map of a candidate training image in the plurality of training images, the heat map generated using the machine learning localization model, each pixel of the heat map having associated thereto a given probability that the pixel belongs to one of the at least one object;

identifying, at the processor within the candidate training image, a plurality of proposals each corresponding to a potential object, each proposal having associated thereto an initial probability that the proposal corresponds to the potential object;

correcting, at the processor, the initial probability using the heat map, thereby obtaining a corrected probability for each one of the plurality of proposals;

selecting given ones of the plurality of proposals having a greatest corrected probability, thereby identifying object candidates;

calculating a loss for the machine learning localization model based on a location of the object candidates within the candidate training image and the heat map; and

providing the calculated loss to the machine learning localization model to update parameters of the machine learning localization model.

2. The computer-implemented method of claim 1 , wherein a number of the object candidates is equal to a number of the at least one object of the given class contained in the candidate training image.

3. The computer-implemented method of claim 1 , wherein a number of the object candidates is less than a number of the at least one object of the given class contained in the candidate training image.

4. The computer-implemented method of claim 1 , further comprising iteratively increasing a number of object candidates and for each iteration performing said receiving the heat map, said identifying the plurality of proposals, said correcting the initial probability, said selecting the given ones of the plurality of proposals, said calculating the loss and said providing the calculated loss.

5. The computer-implemented method of claim 1 , wherein said correcting the initial probability is performed using a maximum probability of the given probability of given pixels of the heat map that corresponding to the proposal.

6. The computer-implemented method of claim 5 , wherein the corrected probability is a maximum value between the initial probability and the maximum probability.

7. The computer-implemented method of claim 1 , wherein said selecting the given ones of the plurality of proposals further comprises identifying a foreground region intersecting the given ones of the plurality of proposals and a background region intersecting none of the plurality of proposals, and said calculating the loss is performed further based on at least one of the foreground region and the background region.

8. The computer-implemented method of claim 1 , wherein the machine learning localization model comprises a fully supervised localization model.

9. The computer-implemented method of claim 8 , wherein the fully supervised localization model comprises one of a fully convolutional neural (FCN) network, a FCN with a ResNet backbone, a PSPNet, DeepLab and a Tiramisu.

10. The computer-implemented method of claim 1 , wherein said identifying the plurality of proposals is performed by one of a region proposal network, a selective search model, a sharpmask and a deepmask.

11. A system for training a machine learning localization system to localize objects belonging to a given class within an image, the system comprising:

a processor, the processor configured to:

receive a plurality of training images; and

identify, within a candidate training image in the plurality of training images, a plurality of proposals, each proposal in the plurality of proposals corresponding to a potential object, each proposal having associated thereto an initial probability that the proposal corresponds to the potential object;

receive a heat map of the candidate training image from the machine learning localization system, each pixel of the heat map having associated thereto a given probability that the pixel belongs to at least one object of a given class; and

correct the initial probability using the heat map to obtain a corrected probability for each proposal in the plurality of proposals;

select, using an object classifier model, given ones of the plurality of proposals having a greatest corrected probability to identify object candidates; and

calculating a loss for the machine learning localization system based on a location of the object candidates within the candidate training image and the heat map; and

outputting the calculated loss to update parameters of the machine learning localization system.

12. The system of claim 11 , wherein a number of the object candidates is equal to a number of the at least one object of the given class contained in the candidate training image.

13. The system of claim 11 , wherein a number of the object candidates is less than a number of the at least one object of the given class contained in the candidate training image.

14. The system of claim 11 , wherein the processor is configured to:

iteratively generate the plurality of proposals for each one of the plurality of training images;

correct the initial probability at each iteration;

increase a number of object candidates and selecting given ones of the plurality of proposals at each iteration; and

calculate the loss at each iteration.

15. The system of claim 11 , wherein the processor is configured for correcting the initial probability using a maximum probability of the given probability of given pixels of the heat map that corresponding to the proposal.

16. The system of claim 15 , wherein the corrected probability is a maximum value between the initial probability and the maximum probability.

17. The system of claim 11 , wherein the processor is further configured for:

identifying a foreground region intersecting the given ones of the plurality of proposals and a background region intersecting none of the plurality of proposals; and

calculate the loss further based on at least one of the foreground region and the background region.

18. The system of claim 11 , wherein the machine learning localization system comprises a fully supervised localization system.

19. The system of claim 18 , wherein the fully supervised localization system comprises one of a fully convolutional neural (FCN) network, a FCN with a ResNet backbone, a PSPNet, DeepLab and a Tiramisu.

20. The system of claim 11 , wherein the processor is further configured to identify the plurality of proposals based on one of a region proposal network, a selective search unit, a sharpmask and a deepmask.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2025
From: SERVICENOW CANADA INC.
To: SERVICENOW, INC.
Reel/Frame 070644/0956 →
CHANGE OF NAME Recorded Jan 22, 2025
From: ELEMENT AI INC.
To: SERVICENOW CANADA INC.
Reel/Frame 069986/0587 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2025
From: LARADJI, ISSAM HADJ; VAZQUEZ BERMUDEZ, DAVID; RODRIGUEZ LOPEZ, PAU; PARDINAS, RAFAEL
To: SERVICENOW CANADA INC.
Reel/Frame 069873/0076 →
Continuity (2)
Provisional Application 62968524 · Jan 31, 2020
Related Publication 20230082179A1 · Mar 16, 2023
References Cited (13)
US 10818017B2 · Ardö · 2020 [cited by examiner]
US 20180285682A1 · Najibi et al. · 2018 [cited by applicant]
US 20180365532A1 · Molchanov et al. · 2018 [cited by applicant]
US 20190279074A1 · Lin et al. · 2019 [cited by applicant]
US 20200050893A1 · Suresh · 2020 [cited by examiner]
US 20210192727A1 · Ward · 2021 [cited by examiner]
CN 109741347A · 2019 [cited by applicant]
Cheng, K-W. et al., “Improved Object Detection With Iterative Localization Refinement in Convolutional Neural Networks”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, No. 9, pp. 2261-2275, Sep… [cited by applicant]
Gidaris, S. et al., “Object detection via a multi-region & semantic segmentation-aware CNN model”, arXiv.org, v3, pp. 1-17, Sep. 23, 2015 (Sep. 23, 2015), retrieved from htt11s://arxiv.org/abs/1505.0I 749 on Apr. 25, 20… [cited by applicant]
Erhan, D. et al., “Scalable Object Detection using Deep Neural Networks”, arXiv.org, vl, pp. 1-8, Dec. 8, 2013 (Aug. 12, 2013), retrieved from https://arxiv.org/abs/1312.2249 on Apr. 25, 2021 (Apr. 25, 2021) *whole docu… [cited by applicant]
Gong, J. et al., “Improving Multi-stage Object Detection via Iterative Proposal Refinement”, Westwell Lab Research Group Shanghai, China, pp. 1-13, 2019, retrieved from https://bmvc2019.org/wp-content/uploads/papers/054… [cited by applicant]
Ren, S. et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, arXiv.org, v3, pp. 1-14, Jan. 6, 2016 (Jun. 1, 2016), retrieved from https://arxiv.org/abs/1506.01497 on Apr. 25, 2021 (A… [cited by applicant]
International Search Report and Written Opinion of International Application No. PCT/IB2021/050745. [cited by applicant]