IP Library › Granted Patent US 12,373,957
Granted Patent B2
US 12,373,957 · App. 18/139,941 · Granted Jul 29, 2025

Image segmentation method, network training method, electronic equipment and storage medium

Inventor: Youxian Zheng (Hangzhou, CN)
Assignee: ZHEJIANG DAHUA TECHNOLOGY CO., LTD.
G06T7/12G06T7/11G06T7/194G06V10/25G06V10/44G06V10/764G06V10/771
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,957
App. No.
18/139,941
Granted
Jul 29, 2025
Kind
B2
Abstract

Disclosed are an image segmentation method, a training method for an image segmentation network, an electronic equipment, and a storage medium. The method includes: sending an input image to an image segmentation network; obtaining a first foreground target box of the input image; obtaining a first region of interest and a first region-of-interest feature map of the input image based on the first foreground target box; dividing the first region of interest into grids, predicting a corresponding feature of each grid in the first region of interest based on the first region-of-interest feature map, obtaining a semantic feature of each pixel in the first region of interest; and obtaining an instance segmentation result based on the corresponding feature of each grid in the first region of interest, information of the first foreground target box, and the semantic feature of each pixel in the first region of interest.

Claims (79)

1. An image segmentation method, comprising:

sending an input image to an image segmentation network;

obtaining a first foreground target box of the input image by the image segmentation network;

obtaining a first region of interest and a first region-of-interest feature map of the input image based on the first foreground target box by the image segmentation network; wherein the first region of interest is a region corresponding to the first foreground target box of the input image;

dividing the first region of interest into a plurality of grids by the image segmentation network, predicting a corresponding feature of each grid in the first region of interest based on the first region-of-interest feature map by the image segmentation network, and obtaining a semantic feature of each pixel in the first region of interest by the image segmentation network; wherein the corresponding feature of each grid is a probability of an existence of a foreground target in the each grid; and

obtaining an instance segmentation result based on the corresponding feature of each grid in the first region of interest, information of the first foreground target box, and the semantic feature of each pixel in the first region of interest.

2. The method according to claim 1 , wherein the information of the first foreground target box comprises a category of the first foreground target box; the obtaining the instance segmentation result based on the corresponding feature of each grid in the first region of interest, the information of the first foreground target box, and the semantic feature of each pixel in the first region of interest comprises:

determining a region, at which a foreground target exists, in the first region of interest based on the corresponding feature of each grid in the first region of interest, and determining a category of each pixel in the first region of interest based on the semantic feature of each pixel in the first region of interest; wherein the corresponding feature of a grid corresponding to the region at which the foreground target exists is greater than a preset probability threshold; and

determining pixels that belong to the category of the first foreground target box in the region at which the foreground target exists and taking the pixels that belong to the category of the first foreground target box in the region at which the foreground target exists as the instance segmentation result, based on the category of each pixel in the first region of interest.

3. The method according to claim 1 , wherein the obtaining the first foreground target box of the input image by the image segmentation network, comprises:

obtaining a first basic feature map of the input image by the image segmentation network; and

obtaining the first foreground target box of the input image based on the first basic feature map by the image segmentation network.

4. The method according to claim 3 , wherein the information of the first foreground target box comprises a location of the first foreground target box; the obtaining the semantic feature of each pixel in the first region of interest by the image segmentation network comprises:

obtaining a corresponding first semantic segmentation feature map of each pixel of the input image by performing semantic segmentation for the first basic feature map by the image segmentation network; and

determining a corresponding feature of each pixel in the first region of interest in the first semantic segmentation feature map based on the location of the first foreground target box, and taking the corresponding feature of each pixel in the first region of interest in the first semantic segmentation feature map as the semantic feature of the each pixel in the first region of interest.

5. The method according to claim 3 , wherein the obtaining the first foreground target box of the input image based on the first basic feature map by the image segmentation network, comprises:

obtaining a first foreground candidate box of the input image based on the first basic feature map by the image segmentation network;

obtaining a second region-of-interest feature map of the input image based on the first foreground candidate box and the first basic feature map by the image segmentation network; and

obtaining the first foreground target box based on the second region-of-interest feature map by the image segmentation network.

6. The method according to claim 5 , wherein the obtaining the first foreground target box based on the second region-of-interest feature map by the image segmentation network, comprises:

obtaining a box classification feature map and a box regression feature map corresponding to the second region-of-interest feature map by the image segmentation network; wherein the box classification feature map is configured to indicate probabilities that the first foreground candidate box belongs to each category, and the first box regression feature map is configured to indicate an offset of the first foreground target box relative to the first foreground candidate box; and

obtaining the information of the first foreground target box based on the box classification feature map and the box regression feature map corresponding to the second region-of-interest feature map.

7. The method according to claim 6 , wherein the information of the first foreground target box comprises a location of the first foreground target box and a category of the first foreground target box; the obtaining the first foreground target box based on the box classification feature map corresponding to the second region-of-interest feature map and the box regression feature map corresponding to the second region-of-interest feature map, comprises:

obtaining the category of the first foreground target box by performing threshold filtering processing for the box classification feature map corresponding to the second region-of-interest feature map; and obtaining the location of the first foreground target box by performing offset transformation for the box regression feature map corresponding to the second region-of-interest feature map and a location of the first foreground candidate box.

8. The method according to claim 1 , before the obtaining the information of the first foreground target box of the input image by the image segmentation network, further comprising:

training the image segmentation network.

9. The method according to claim 8 , wherein the training the image segmentation method comprises:

sending a training image to the image segmentation network;

obtaining a second basic feature map of the training image by the image segmentation network;

obtaining a second region of interest and a third region-of-interest feature map of the training image based on the second basic feature map by the image segmentation network;

dividing the second region of interest into a plurality of grids by the image segmentation network, predicting a corresponding feature of each grid in the second region of interest based on the second region-of-interest feature map by the image segmentation network, and obtaining a second semantic segmentation feature map by performing semantic segmentation on the second basic feature map by the image segmentation network;

obtaining a first loss of the image segmentation network based on a difference between the corresponding feature of each grid in the second region of interest and a first ground truth, and obtaining a second loss of the image segmentation network based on a difference between the second semantic segmentation feature map and a second ground truth; and

adjusting weights of the image segmentation network based on the first loss and the second loss.

10. The method according claim 9 , wherein the obtaining the second region of interest and the third region-of-interest feature map in the training image based on the second basic feature map by the image segmentation network, comprises:

obtaining a second foreground candidate box of the training image based on the second basic feature map by the image segmentation network, and taking a region corresponding to the second foreground candidate box in the training image as the second region of interest; and

obtaining the third region-of-interest feature map based on the second foreground candidate box and the second basic feature map by the image segmentation network.

11. The method according to claim 10 , after the obtaining the third region-of-interest feature map based on the second foreground candidate box and the second basic feature map by the image segmentation network, further comprising:

obtaining a box classification feature map and a box regression feature map corresponding to the third region-of-interest feature map by the image segmentation network; wherein the box classification feature map corresponding to the third region-of-interest feature map is configured to represent probabilities that the second foreground candidate box belongs to each category, and the box regression feature map corresponding to the third region-of-interest feature map is configured to represent an offset of a second foreground target box relative to the second foreground candidate box;

obtaining a third loss of the image segmentation network based on a difference between the box classification feature map corresponding to the third region-of-interest feature map and a third ground truth, and obtaining a fourth loss of the image segmentation network based on a difference between the box regression feature map corresponding to the third region-of-interest feature map and a fourth ground truth; and

adjusting the weights of the image segmentation network based on the third loss and the fourth loss.

12. A training method for an image segmentation network, comprising:

sending a training image to the image segmentation network;

obtaining a second basic feature map of the training image by the image segmentation network;

obtaining a second region of interest and a third region-of-interest feature map in the training image based on the second basic feature map by the image segmentation network;

dividing the second region of interest into a plurality of grids by the image segmentation network, predicting a corresponding feature of each grid in the second region of interest based on the second region-of-interest feature map by the image segmentation network, and obtaining a second semantic segmentation feature map by performing semantic segmentation for the second basic feature map by the image segmentation network;

obtaining a first loss of the image segmentation network based on a difference between the corresponding feature of each grid in the second region of interest and a first ground truth, and obtaining a second loss of the image segmentation network based on a difference between the second semantic segmentation feature map and a second ground truth; and

adjusting weights of the image segmentation network based on the first loss and the second loss.

13. The method according to claim 12 , wherein the obtaining the second region of interest and the third region-of-interest feature map in the training image based on the second basic feature map by the image segmentation network comprises:

obtaining a second foreground candidate box of the training image based on the second basic feature map by the image segmentation network, and taking a region corresponding to the second foreground candidate box in the training image as the second region of interest; and

obtaining the third region-of-interest feature map based on the second foreground candidate box and the second basic feature map by the image segmentation network.

14. The method according to claim 13 , after the obtaining the third region-of-interest feature map based on the second foreground candidate box and the second basic feature map by the image segmentation network, further comprising:

obtaining a box classification feature map and a box regression feature map corresponding to the third region-of-interest feature map by the image segmentation network; wherein the box classification feature map corresponding to the third region-of-interest feature map is configured to represent a probabilities that the second foreground candidate box belongs to each category, and the box regression feature map corresponding to the third region-of-interest feature map is configured to represent an offset of a second foreground target box relative to the second foreground candidate box;

obtaining a third loss of the image segmentation network based on a difference between the box classification feature map corresponding to the third region-of-interest feature map and a third ground truth, and obtaining a fourth loss of the image segmentation network based on a difference between the box regression feature map corresponding to the third region-of-interest feature map and a fourth ground truth; and

adjusting the weights of the image segmentation network based on the third loss and the fourth loss.

15. An electronic device, comprising a processor and a memory connected to the processor;

wherein the memory stores a program instruction; the processor is configured to execute the program instruction stored in the memory to perform:

sending an input image to an image segmentation network;

obtaining a first foreground target box of the input image by the image segmentation network;

obtaining a first region of interest and a first region-of-interest feature map of the input image based on the first foreground target box by the image segmentation network; wherein the first region of interest is a region corresponding to the first foreground target box of the input image;

dividing the first region of interest into a plurality of grids by the image segmentation network, predicting a corresponding feature of each grid in the first region of interest based on the first region-of-interest feature map by the image segmentation network, and obtaining a semantic feature of each pixel in the first region of interest by the image segmentation network; wherein the corresponding feature of each grid is a probability of an existence of a foreground target in the each grid; and

obtaining an instance segmentation result based on the corresponding feature of each grid in the first region of interest, information of the first foreground target box, and the semantic feature of each pixel in the first region of interest.

16. The electronic device according to claim 15 , wherein the information of the first foreground target box comprises a category of the first foreground target box; the obtaining the instance segmentation result based on the corresponding feature of each grid in the first region of interest, the information of the first foreground target box, and the semantic feature of each pixel in the first region of interest comprises:

determining a region, at which a foreground target exists, in the first region of interest based on the corresponding feature of each grid in the first region of interest, and determining a category of each pixel in the first region of interest based on the semantic feature of each pixel in the first region of interest; wherein the corresponding feature of a grid corresponding to the region at which the foreground target exists is greater than a preset probability threshold; and

determining pixels that belong to the category of the first foreground target box in the region at which the foreground target exists and taking the pixels that belong to the category of the first foreground target box in the region at which the foreground target exists as the instance segmentation result, based on the category of each pixel in the first region of interest.

17. The electronic device according to claim 15 , wherein the obtaining the first foreground target box of the input image by the image segmentation network, comprises:

obtaining a first basic feature map of the input image by the image segmentation network; and

obtaining the first foreground target box of the input image based on the first basic feature map by the image segmentation network.

18. The electronic device according to claim 15 , wherein the processor is configured to execute the program instruction stored in the memory to further perform:

before the obtaining the information of the first foreground target box of the input image by the image segmentation network, training the image segmentation network.

19. The electronic device according to claim 18 , wherein the training the image segmentation method comprises:

sending a training image to the image segmentation network;

obtaining a second basic feature map of the training image by the image segmentation network;

obtaining a second region of interest and a third region-of-interest feature map of the training image based on the second basic feature map by the image segmentation network;

dividing the second region of interest into a plurality of grids by the image segmentation network, predicting a corresponding feature of each grid in the second region of interest based on the second region-of-interest feature map by the image segmentation network, and obtaining a second semantic segmentation feature map by performing semantic segmentation on the second basic feature map by the image segmentation network;

obtaining a first loss of the image segmentation network based on a difference between the corresponding feature of each grid in the second region of interest and a first ground truth, and obtaining a second loss of the image segmentation network based on a difference between the second semantic segmentation feature map and a second ground truth; and

adjusting weights of the image segmentation network based on the first loss and the second loss.

20. The electronic device according to claim 19 , wherein the obtaining the second region of interest and the third region-of-interest feature map in the training image based on the second basic feature map by the image segmentation network, comprises:

obtaining a second foreground candidate box of the training image based on the second basic feature map by the image segmentation network, and taking a region corresponding to the second foreground candidate box in the training image as the second region of interest; and

obtaining the third region-of-interest feature map based on the second foreground candidate box and the second basic feature map by the image segmentation network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2023
From: ZHENG, YOUXIAN
To: ZHEJIANG DAHUA TECHNOLOGY CO., LTD.
Reel/Frame 063455/0440 →
Priority Claims (1)
CN 202011511498.X · Dec 18, 2020 · national
Continuity (2)
Continuation PCTCN2021139223 · Dec 17, 2021
Related Publication 20240078680A1 · Mar 7, 2024
References Cited (27)
US 11238596B1 · Gadde · 2022 [cited by examiner]
US 20180108137A1 · Price · 2018 [cited by examiner]
US 20190057507A1 · El-Khamy · 2019 [cited by examiner]
US 20200005453A1 · Lin et al. · 2020 [cited by applicant]
US 20200082219A1 · Li · 2020 [cited by examiner]
US 20210056708A1 · Li · 2021 [cited by examiner]
US 20210158043A1 · Hou · 2021 [cited by examiner]
US 20210366127A1 · Gu · 2021 [cited by examiner]
US 20220051045A1 · Vandersmissen · 2022 [cited by examiner]
US 20220130141A1 · Wang · 2022 [cited by examiner]
US 20220398742A1 · Zhang · 2022 [cited by examiner]
US 20230186100A1 · Nugteren · 2023 [cited by examiner]
US 20230419648A1 · Ghafoorian · 2023 [cited by examiner]
CN 110246141A · 2019 [cited by applicant]
CN 110532954A · 2019 [cited by applicant]
CN 110599500A · 2019 [cited by applicant]
CN 111192277A · 2020 [cited by applicant]
CN 112613519A · 2021 [cited by applicant]
EP 3663982A1 · 2020 [cited by applicant]
Kaiming He et al., «Mask R-CNN», 2018. [cited by applicant]
Xinlong Wang et al., «Solo : Segmenting Objects by Locations», 2020. [cited by applicant]
International Search Report, International Application No. PCT/CN2021/139223, mailed Mar. 16, 2022 (9 pages). [cited by applicant]
Zhang Xiangyi et al: “Mask R-CNN with Feature Pyramid Attention for Instance Segmentation”, 2018 14th IEEE International Conference on Signal Processing (ICSP), IEEE, Aug. 12, 2018 (Aug. 12, 2018), pp. 1194-1197, XP0335… [cited by applicant]
He Kaiming et al: “Mask R-CNN”, IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Computer Society, USA, vol. 42, No. 2,Jun. 5, 2018 (Jun. 5, 2018), pp. 386-397, XP011765746,ISSN: 0162-8828, DOI: 10.1… [cited by applicant]
Lu Xin et al: “Grid R-CNN”, 2019 IEEE/CVF Conference On Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 15, 2019 (Jun. 15, 2019), pp. 7355-7364, XP033686690, DOI: 10.1109/CVPR.2019.00754 [retrieved on Jan. 8,… [cited by applicant]
European Search Report, European Application No. 21905836.9, mailed Feb. 13, 2024 (9 pages). [cited by applicant]
Chinese First Office Action, Chinese Application No. 202011511498.X, mailed Jun. 1, 2023 (9 pages). [cited by applicant]