IP Library › Granted Patent US 12,183,010
Granted Patent B2
US 12,183,010 · App. 17/651,243 · Granted Dec 31, 2024

Systems and methods for instance segmentation based on semantic segmentation

Inventors: Jian Tang (Beijing, CN); Chengxiang Yin (Beijing, CN); Kun Wu (Beijing, CN); Zhengping Che (Beijing, CN)
Assignee: BEIJING DIDI INFINITY TECHNOLOGY AND DEVELOPMENT CO., LTD.
G06T7/10G06N3/08G06V10/764G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,183,010
App. No.
17/651,243
Granted
Dec 31, 2024
Kind
B2
Abstract

The present disclosure relates to a system and a method for performing instance segmentation based on semantic segmentation that is capable of (1) processing HD images in real-time given semantic segmentation; 2) delivering comparable performance with Mask R-CNN in terms of accuracy when combined with a widely-used semantic segmentation method (such as DPC), while consistently outperforms a state-of-the-art real-time solution; (3) working flexibly with any semantic segmentation model for instance segmentation; (4) outperforming Mask R-CNN if the given semantic segmentation is sufficiently good; and (5) being easily extended to panoptic segmentation.

Claims (28)

1. A system for obtaining an instance segmentation or panoptic segmentation of an image based on semantic segmentation, the system comprising:

a storage medium storing a set of instructions; and

a processor in communication with the storage medium to execute the set of instructions to:

perform semantic segmentation on an input image to obtain a semantic label map having specific set of classes, using a trained semantic segmentation model;

generate a boundary map, using a trained generator, based on the obtained semantic label map concatenated with the input image; and

process the boundary map, using a post-processing step, to differentiate objects of the specific set of classes to obtain the instance segmentation or panoptic segmentation of the input image, wherein the post-processing step comprises performing Breadth-First-Search for each enclosed area of the semantic label map to get a mask for each enclosed area, a class of the mask being determined based on its semantic label map.

2. The system of claim 1 , wherein the trained semantic segmentation model is DeepLabv3+.

3. The system of claim 1 , wherein the trained semantic segmentation model is dense prediction cell (DPC).

4. The system of claim 1 , wherein the trained generator comprises a conditional Generative Adversarial Networks (GANs) coupled with deep supervision as well as a weighted fusion layer.

5. The system of claim 1 , wherein the system is able to obtain instance segmentation or panoptic segmentation in real time.

6. The system of claim 1 , wherein the set of instructions further instructs the processor to generate masks for at least one of thing classes and stuff classes.

7. The system of claim 1 , further comprising a discriminator that engages in a minimax game with a generator to form the trained generator, wherein the discriminator distinguishes between a boundary map generated by the trained generator and a corresponding boundary map of ground truth.

8. A method for obtaining an instance segmentation or panoptic segmentation of an image based on semantic segmentation, on a computing device including a storage medium storing a set of instructions, and a processor in communication with the storage medium to execute the set of instructions, the method comprising:

performing semantic segmentation on an input image to obtain a semantic label map having specific set of classes, using a trained semantic segmentation model;

generating a boundary map, using a trained generator, based on the obtained semantic label map concatenated with the input image; and

processing the boundary map, using a post-processing step, to differentiate objects of the specific set of classes to obtain the instance segmentation or panoptic segmentation of the input image, wherein the post-processing step comprises performing Breadth-First-Search for each enclosed area of the semantic label map to get a mask for each enclosed area, a class of the mask being determined based on its semantic label map.

9. The method of claim 8 , wherein the trained semantic segmentation model is DPC.

10. The method of claim 8 , wherein the trained generator comprises a conditional Generative Adversarial Networks (GANs) coupled with deep supervision as well as a weighted fusion layer.

11. The method of claim 8 , wherein the instance segmentation or panoptic segmentation is obtained in real time.

12. The method of claim 8 , further comprising generating masks for at least one of thing classes and stuff classes.

13. The method of claim 8 , further comprising using a discriminator to distinguish between a boundary map generated by the trained generator and a corresponding boundary map of ground truth to engage in a minimax game with a generator to form the trained generator.

14. A computer non-transitory readable medium, storing a set of instructions for obtaining an instance segmentation or panoptic segmentation of an image based on semantic segmentation, wherein when the set of instructions is executed by a processor of an electrical device, the electrical device performs a method comprising:

performing semantic segmentation on an input image to obtain a semantic label map having specific set of classes, using a trained semantic segmentation model;

generating a boundary map, using a trained generator, based on the obtained semantic label map concatenated with the input image; and

processing the boundary map, using a post-processing step, to differentiate objects of the specific set of classes to obtain the instance segmentation or panoptic segmentation of the input image, wherein the instance segmentation or panoptic segmentation is obtained in real time, wherein the post-processing step comprises performing Breadth-First-Search for each enclosed area of the semantic label map to get a mask for each enclosed area, a class of the mask being determined based on its semantic label map.

15. The computer non-transitory readable medium of claim 14 , wherein the trained semantic segmentation model is DPC and the trained generator comprises a conditional Generative Adversarial Networks (GANs) coupled with deep supervision as well as a weighted fusion layer.

16. The computer non-transitory readable medium of claim 14 , the method further comprising generating masks for at least one of thing classes and stuff classes.

17. The computer non-transitory readable medium of claim 14 , the method further comprising using a discriminator to distinguish between a boundary map generated by the trained generator and a corresponding boundary map of ground truth to engage in a minimax game with a generator to form the trained generator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2022
From: TANG, JIAN; YIN, CHENGXIANG; WU, KUN; CHE, ZHENGPING
To: BEIJING DIDI INFINITY TECHNOLOGY AND DEVELOPMENT CO., LTD.
Reel/Frame 061176/0899 →
Continuity (2)
Continuation PCTCN2019110539 · Oct 11, 2019
Related Publication 20220172369A1 · Jun 2, 2022
Cited By (1)
US 12,530,782