IP Library Granted Patent US 12,561,982
Granted Patent B2
US 12,561,982 · App. 18/188,701 · Granted Feb 24, 2026

Infrastructure analysis using panoptic segmentation

Inventors: Samuel Schulter (Long Island City, NY); Sparsh Garg (San Jose, CA)
Assignee: NEC Corporation
G06V20/54G06V10/26G06V10/761G06V10/774G06V10/86G06V20/70G08G1/09G08G1/16G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,982
App. No.
18/188,701
Granted
Feb 24, 2026
Kind
B2
Abstract

Methods and systems identifying road hazards include capturing an image of a road scene using a camera. The image is embedded using a segmentation model that includes an image branch having an image embedding layer that embeds images into a joint latent space and a text branch having a text embedding layer that embeds text into the joint latent space. A mask is generated for an object within the image using the segmentation model. A probability is determined that the object matches a road hazard using the segmentation mode. A signal is generated responsive to the probability to ameliorate a danger posed by the road hazard.

Claims (32)

1 . A computer-implemented method for identifying road hazards, comprising:

capturing an image of a road scene using a camera;

embedding the image using a segmentation model that includes an image branch having an image embedding layer that embeds images into a joint latent space and a text branch having a text embedding layer that embeds text into the joint latent space;

generating a mask for an object within the image using the segmentation model;

determining a probability that the object matches a road hazard using the segmentation mode; and

generating a signal responsive to the probability to ameliorate a danger posed by the road hazard.

2 . The method of claim 1 , wherein the joint latent space represents a visual object and a textual description of visual objects as vectors that are similar to one another according to a distance metric.

3 . The method of claim 2 , wherein determining the probability includes comparing masks to known road hazards using a distance metric in the joint latent space.

4 . The method of claim 3 , wherein the known road hazards are represented in the joint latent space as embeddings of textual descriptions of the road hazards.

5 . The method of claim 1 , wherein the distance metric is a cosine distance.

6 . The method of claim 1 , wherein generating the mask includes identifying a location of the object in the image and generating a label for the object.

7 . The method of claim 1 , wherein generating the signal includes transmitting the signal to a system selected from the group consisting of a vehicle, roadside signage, and a municipal authority.

8 . The method of claim 1 , wherein generating the signal includes determining that the probability exceeds a threshold value to determine that the object shows a road hazard.

9 . The method of claim 1 , further comprising ameliorating the road hazard by performing an action selected from the group consisting of displaying a warning on roadside signage, audibly announcing the road hazard, and deploying a cleanup crew to correct the road hazard.

10 . The method of claim 1 , wherein the segmentation model is trained on a plurality of training datasets that include a panoptic segmentation dataset, which includes class labels for individual image pixels, and an object detection dataset, which includes a class label for a bounding box.

11 . A system for identifying road hazards, comprising:

a hardware processor; and

a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:

capture an image of a road scene using a camera;

embed the image using a segmentation model that includes an image branch having an image embedding layer that embeds images into a joint latent space and a text branch having a text embedding layer that embeds text into the joint latent space;

generate a mask for an object within the image using the segmentation model;

determine a probability that the object matches a road hazard using the segmentation mode; and

generate a signal responsive to the probability to ameliorate a danger posed by the road hazard.

12 . The system of claim 11 , wherein the joint latent space represents a visual object and a textual description of visual objects as vectors that are similar to one another according to a distance metric.

13 . The system of claim 12 , wherein the computer program further causes the hardware processor to determine compare masks to known road hazards using a distance metric in the joint latent space.

14 . The system of claim 13 , wherein the known road hazards are represented in the joint latent space as embeddings of textual descriptions of the road hazards.

15 . The system of claim 11 , wherein the distance metric is a cosine distance.

16 . The system of claim 11 , wherein the computer program further causes the hardware processor to identify a location of the object in the image and to generate a label for the object.

17 . The system of claim 11 , wherein the computer program further causes the hardware processor to transmit the signal to a system selected from the group consisting of a vehicle, roadside signage, and a municipal authority.

18 . The system of claim 11 , wherein the computer program further causes the hardware processor to determine that the probability exceeds a threshold value to determine that the object shows a road hazard.

19 . The system of claim 11 , wherein the computer program further causes the hardware processor to ameliorate the road hazard by performing an action selected from the group consisting of displaying a warning on roadside signage, audibly announcing the road hazard, and deploying a cleanup crew to correct the road hazard.

20 . The system of claim 11 , wherein the segmentation model is trained on a plurality of training datasets that include a panoptic segmentation dataset, which includes class labels for individual image pixels, and an object detection dataset, which includes a class label for a bounding box.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2026
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 073431/0568 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2023
From: SCHULTER, SAMUEL; GARG, SPARSH
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 063076/0192 →
Continuity (4)
Continuation In Part 18178821 · Mar 6, 2023
Provisional Application 63343202 · May 18, 2022
Provisional Application 63317487 · Mar 7, 2022
Related Publication 20230281999A1 · Sep 7, 2023
References Cited (10)
US 10026020B2 · Jin · 2018 [cited by examiner]
US 11574142B2 · Lin · 2023 [cited by examiner]
US 20210271707A1 · Lin · 2021 [cited by examiner]
Li et al., “Image-text embedding learning via visual and textual semantic reasoning”, Jan. 2022 (Year: 2022). [cited by examiner]
Bevandic et al., “Multi-domain semantic segmentation with overlapping labels”, In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision 2022, Jan. 2022, pp. 2615-2624. [cited by applicant]
Cheng et al., “Masked-attention Mask Transformer for Universal Image Segmentation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022, Jun. 2022, pp. 1290-1299. [cited by applicant]
Lan et al., “Discobox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision”, In Proceedings of the IEEE/CVF International Conference on Computer Vision 2021, Oct. 2021, pp. 3406-3416. [cited by applicant]
Li et al., “Language-Driven Semantic Segmentation”, arXiv:2201.03546v2 [cs.CV], Apr. 3, 2022, pp. 1-13. [cited by applicant]
Wang et al., “FreeSOLO: Learning to Segment Objects without Annotations”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022, Jun. 2022, pp. 14176-14186. [cited by applicant]
Xu et al., “A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model”, arXiv:2112.14757v2 [cs.CV], Dec. 29, 2022, pp. 1-22. [cited by applicant]