IP Library Patent Application 18178821
Patent Application
App. No. 18/178,821

PANOPTIC SEGMENTATION WITH MULTI-DATABASE TRAINING USING MIXED EMBEDDING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/178,821
Abstract

Methods and systems for training an image segmentation model include embedding training images, from multiple training datasets having differing label spaces, in a joint latent space to generate first features. Textual labels of the training images are embedded in the joint latent space to generate second features. A segmentation model is trained using the first features and the second features.

Claims (81)

1 . A computer-implemented method for training an image segmentation model, comprising:

embedding training images, from a plurality of training datasets having differing label spaces, in a joint latent space to generate first features;

embedding textual labels of the training images in the joint latent space to generate second features; and

training a segmentation model using the first features and the second features.

2 . The method of claim 1 , wherein the plurality of training datasets include a panoptic segmentation dataset, which includes class labels for individual image pixels, and an object detection dataset, which includes a class label for a bounding box.

3 . The method of claim 2 , wherein training the segmentation model uses a loss function that weights contributions from the panoptic segmentation dataset and the object detection dataset differently.

4 . The method of claim 3 , wherein training the segmentation model uses a pair-wise cost for the object detection dataset of:

C

ij

mb

=

1.

-

κ

in

"\[LeftBracketingBar]"

b

j

"\[RightBracketingBar]"

+

κ

out

"\[LeftBracketingBar]"

m

i

"\[RightBracketingBar]"

where κ in and κ out are a number of pixels of the mask m i inside and outside the ground truth box b j respectively.

5 . The method of claim 1 , wherein the joint latent space represents a visual object and a textual description of the visual object as vectors that are similar to one another according to a distance metric.

6 . The method of claim 1 , further comprising comparing the first features to the second features using a distance metric in the joint latent space.

7 . The method of claim 6 , wherein the distance metric is a cosine distance.

8 . The method of claim 1 , wherein the segmentation model includes an image branch having an image embedding layer embeds images into the latent space and a text branch having a text embedding layer that embeds text labels into the latent space.

9 . A computer-implemented method for image analysis, comprising:

embedding an image using a segmentation model that includes an image branch having an image embedding layer that embeds images into a joint latent space;

embedding a textual query term using the segmentation model, wherein the segmentation model further includes a text branch having a text embedding layer that embeds text into the joint latent space;

generating a mask for an object within the image using the segmentation model;

determining a probability that the object matches the textual query term using the segmentation mode; and

performing an image analysis task using the mask and the determined probability.

10 . The method of claim 9 , wherein the joint latent space represents a visual object and a textual description of the visual object as vectors that are similar to one another according to a distance metric.

11 . The method of claim 9 , wherein determining the probability includes comparing the first features to the second features using a distance metric in the joint latent space.

12 . The method of claim 11 , wherein the distance metric is a cosine distance.

13 . A system for training an image segmentation model, comprising:

a hardware processor; and

a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:

embed training images, from a plurality of training datasets having differing label spaces, in a joint latent space to generate first features;

embed textual labels of the training images in the joint latent space to generate second features; and

train a segmentation model using the first features and the second features.

14 . The system of claim 13 , wherein the plurality of training datasets include a panoptic segmentation dataset, which includes class labels for individual image pixels, and an object detection dataset, which includes a class label for a bounding box.

15 . The system of claim 14 , wherein training the segmentation model uses a loss function that weights contributions from the panoptic segmentation dataset and the object detection dataset differently.

16 . The system of claim 15 , wherein training the segmentation model uses a pair-wise cost for the object detection dataset of:

C

ij

mb

=

1.

-

κ

in

"\[LeftBracketingBar]"

b

j

"\[RightBracketingBar]"

+

κ

out

"\[LeftBracketingBar]"

m

i

"\[RightBracketingBar]"

where κ in and κ out are a number of pixels of the mask m i inside and outside the ground truth box b j respectively.

17 . The system of claim 13 , wherein the joint latent space represents a visual object and a textual description of the visual object as vectors that are similar to one another according to a distance metric.

18 . The system of claim 13 , further comprising comparing the first features to the second features using a distance metric in the joint latent space.

19 . The method of claim 18 , wherein the distance metric is a cosine distance.

20 . The system of claim 13 , wherein the segmentation model includes an image branch having an image embedding layer embeds images into the latent space and a text branch having a text embedding layer that embeds text labels into the latent space.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2023
From: SCHULTER, SAMUEL
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 062892/0137 →