IP Library Granted Patent US 12,205,356
Granted Patent B2
US 12,205,356 · App. 18/188,766 · Granted Jan 21, 2025

Semantic image capture fault detection

Inventors: Samuel Schulter (Long Island City, NY); Sparsh Garg (San Jose, CA); Manmohan Chandraker (Santa Clara, CA)
Assignee: NEC Corporation
G06V10/776G06T7/0002G06T7/11G06V10/761G06V10/774G06V20/70H04N17/002G06T2207/20021G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,356
App. No.
18/188,766
Granted
Jan 21, 2025
Kind
B2
Abstract

Methods and systems for detecting faults include capturing an image of a scene using a camera. The image is embedded using a segmentation model that includes an image branch having an image embedding layer that embeds images into a joint latent space and a text branch having a text embedding layer that embeds text into the joint latent space. Semantic information is generated for a region of the image corresponding to a predetermined static object using the embedded image. A fault of the camera is identified based on a discrepancy between the semantic information and semantic information of the predetermined static image. The fault of the camera is corrected.

Claims (32)

1. A computer-implemented method for detecting faults, comprising:

capturing an image of a scene using a camera;

embedding the image using a segmentation model that includes an image branch having an image embedding layer that embeds images into a joint latent space and a text branch having a text embedding layer that embeds text into the joint latent space;

generating semantic information for a region of the image corresponding to a predetermined static object using the embedded image;

identifying a fault of the camera based on a discrepancy between the semantic information and semantic information of the predetermined static image; and

correcting the fault of the camera.

2. The method of claim 1 , wherein the joint latent space represents a visual object and a textual description of visual objects as vectors that are similar to one another according to a distance metric.

3. The method of claim 2 , wherein the semantic information for the region and the semantic information of the predetermined static image each includes a textual label that is similar to the region and predetermined static object in the joint latent space.

4. The method of claim 3 , wherein identifying the fault includes determining that the textual label of the static image differs from the textual label of the region.

5. The method of claim 3 , wherein the semantic information for the region and the semantic information of the predetermined static object each includes a respective confidence score for the corresponding textual label.

6. The method of claim 5 , wherein identifying the fault includes determining that the confidence score for the textual label of the predetermined static object differs from the confidence score for the textual label of the region differs by more than a threshold amount.

7. The method of claim 2 , wherein the distance metric is a cosine distance.

8. The method of claim 1 , further comprising identifying the predetermined static object within a prior image, taken at an earlier point in time than the captured image.

9. The method of claim 1 , wherein the segmentation model is trained on a plurality of training datasets that include a panoptic segmentation dataset, which includes class labels for individual image pixels, and an object detection dataset, which includes a class label for a bounding box.

10. The method of claim 1 , wherein correcting the fault comprises performing an automatic action selected from the group consisting of performing a self-diagnostic within the camera and activating a self-cleaning function.

11. A system for detecting faults, comprising:

a hardware processor; and

a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:

capture an image of a scene using a camera;

embed the image using a segmentation model that includes an image branch having an image embedding layer that embeds images into a joint latent space and a text branch having a text embedding layer that embeds text into the joint latent space;

generate semantic information for a region of the image corresponding to a predetermined static object using the embedded image;

identify a fault of the camera based on a discrepancy between the semantic information and semantic information of the predetermined static image; and

correct the fault of the camera.

12. The system of claim 11 , wherein the joint latent space represents a visual object and a textual description of visual objects as vectors that are similar to one another according to a distance metric.

13. The system of claim 12 , wherein the semantic information for the region and the semantic information of the predetermined static image each includes a textual label that is similar to the region and predetermined static object in the joint latent space.

14. The system of claim 13 , wherein the computer program further causes the hardware processor to determine that the textual label of the static image differs from the textual label of the region.

15. The system of claim 13 , wherein the semantic information for the region and the semantic information of the predetermined static object each includes a respective confidence score for the corresponding textual label.

16. The system of claim 15 , wherein the computer program further causes the hardware processor to determine that the confidence score for the textual label of the predetermined static object differs from the confidence score for the textual label of the region differs by more than a threshold amount.

17. The system of claim 12 , wherein the distance metric is a cosine distance.

18. The system of claim 11 , wherein the computer program further causes the hardware processor to identify the predetermined static object within a prior image, taken at an earlier point in time than the captured image.

19. The system of claim 11 , wherein the segmentation model is trained on a plurality of training datasets that include a panoptic segmentation dataset, which includes class labels for individual image pixels, and an object detection dataset, which includes a class label for a bounding box.

20. The system of claim 11 , wherein the computer program further causes the hardware processor to perform an automatic action selected from the group consisting of performing a self-diagnostic within the camera and activating a self-cleaning function.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2024
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 069540/0269 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2023
From: SCHULTER, SAMUEL; GARG, SPARSH; CHANDRAKER, MANMOHAN
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 063077/0395 →
Continuity (4)
Continuation In Part 18178821 · Mar 6, 2023
Provisional Application 63343202 · May 18, 2022
Provisional Application 63317487 · Mar 7, 2022
Related Publication 20230281977A1 · Sep 7, 2023
References Cited (45)
US 5612744A · Lee · 1997 [cited by examiner]
US 6041078A · Rao · 2000 [cited by examiner]
US 6654420B1 · Snook · 2003 [cited by examiner]
US 6674904B1 · McQueen · 2004 [cited by examiner]
US 7546334B2 · Redlich · 2009 [cited by examiner]
US 8135232B2 · Kimura · 2012 [cited by examiner]
US 8402551B2 · Lee · 2013 [cited by examiner]
US 8447117B2 · Liao · 2013 [cited by examiner]
US 9736468B2 · Lee · 2017 [cited by examiner]
US 11367204B1 · Liao · 2022 [cited by examiner]
US 11688102B2 · Lin · 2023 [cited by examiner]
US 20020112181A1 · Smith · 2002 [cited by examiner]
US 20030036886A1 · Stone · 2003 [cited by examiner]
US 20040091151A1 · Jin · 2004 [cited by examiner]
US 20050138110A1 · Redlich · 2005 [cited by examiner]
US 20050193311A1 · Das · 2005 [cited by examiner]
US 20080005666A1 · Sefton · 2008 [cited by examiner]
US 20080163378A1 · Lee · 2008 [cited by examiner]
US 20090178019A1 · Bahrs · 2009 [cited by examiner]
US 20090178144A1 · Redlich · 2009 [cited by examiner]
US 20090254572A1 · Redlich · 2009 [cited by examiner]
US 20100005179A1 · Dickson · 2010 [cited by examiner]
US 20100158402A1 · Nagase · 2010 [cited by examiner]
US 20100250497A1 · Redlich · 2010 [cited by examiner]
US 20110110603A1 · Ikai · 2011 [cited by examiner]
US 20110129156A1 · Liao · 2011 [cited by examiner]
US 20110164824A1 · Kimura · 2011 [cited by examiner]
US 20120030733A1 · Andrews · 2012 [cited by examiner]
US 20120173971A1 · Sefton · 2012 [cited by examiner]
US 20120252407A1 · Poltorak · 2012 [cited by examiner]
US 20120287247A1 · Stenger · 2012 [cited by examiner]
US 20120321083A1 · Phadke · 2012 [cited by examiner]
US 20130051476A1 · Morris · 2013 [cited by examiner]
US 20130063241A1 · Simon · 2013 [cited by examiner]
US 20130091290A1 · Hirokawa · 2013 [cited by examiner]
US 20180262748A1 · Shibata · 2018 [cited by examiner]
US 20200074684A1 · Lin · 2020 [cited by examiner]
US 20200104645A1 · Ionescu · 2020 [cited by examiner]
US 20210085238A1 · Schnabel · 2021 [cited by examiner]
Bevandic et al., “Multi-domain semantic segmentation with overlapping labels”, In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision 2022, Jan. 2022, pp. 2615-2624. [cited by applicant]
Cheng et al., “Masked-attention Mask Transformer for Universal Image Segmentation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022, Jun. 2022, pp. 1290-1299. [cited by applicant]
Lan et al., “DISCOBOX: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision”, In Proceedings of the IEEE/CVF International Conference on Computer Vision 2021, Oct. 2021, pp. 3406-3416. [cited by applicant]
Li et al., “Language-Driven Semantic Segmentation”, arXiv:2201.03546v2 [cs.CV], Apr. 3, 2022, pp. 1-13. [cited by applicant]
Wang et al., “FreeSOLO: Learning to Segment Objects without Annotations”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022, Jun. 2022, pp. 14176-14186. [cited by applicant]
Xu et al., “A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model”, arXiv:2112.14757v2 [cs.CV], Dec. 29, 2022, pp. 1-22. [cited by applicant]