IP Library › Granted Patent US 12,728,532
Granted Patent B2
US 12,728,532 · App. 18/819,791 · Granted Sep 8, 2026

System and method for object segmentation for task performance

Inventors: Anoop Cherian (Acton, MA); Siddarth Jain (Cambridge, MA); Tim Marks (Newton, MA)
Assignee: Mitsubishi Electric Research Laboratories, Inc.
B25J9/1661G06T3/02G06T7/11G06T7/168G06T7/70G06T11/00G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,728,532
App. No.
18/819,791
Granted
Sep 8, 2026
Kind
B2
Abstract

Embodiments disclosing a controller for controlling a robot to perform a task are provided. The task is performed in an environment that is represented by an input image. The controller causes segmenting of an object in the input image. A confidence level of segmentation is updated by comparing the segmented object with constrained affined transformations of a template of the object. The constrained affine transformations are based on constraints indicative of a property of the object. The property of the object and the updated confidence level of segmentation are then used for performing the task.

Claims (38)

1 . A robot for performing a task, comprising a processor causing the robot to:

segment an object in an input image to produce a segmented object and a confidence level of segmentation;

update the confidence level of segmentation, to generate an updated confidence level, by comparing the segmented object with constrained affined transformations of a template of the object, wherein a constraint limiting the affined transformations is indicative of a property of the object, wherein the property of the object is selected based on the task to be performed and limits a space of affine transformations considered during updating of the confidence level; and

perform the task based on the segmented object and the updated confidence level of the segmented object.

2 . The robot of claim 1 , wherein the property of the object includes one or a combination of a non-occlusion of the object by other objects, a pose of the object, and a distance of the robot from the object.

3 . The robot of claim 2 , wherein the property of the object is the non-occlusion of the object by other objects, and the template of the object includes only an image of a non-occluded object.

4 . The robot of claim 2 , wherein the property of the object is the pose of the object, and the constrained affine transformations are limited to transforming the template of the object into desired poses.

5 . The robot of claim 2 , wherein the property of the object is the distance of the robot from the object, and the constrained affine transformations are limited to preserving the template of the object above a predetermined size.

6 . The robot of claim 1 , wherein to perform the task based on the segmented object, the processor causes the robot to:

compare the updated confidence level of the segmented object with a confidence level threshold; and

perform the task based on the comparison.

7 . The robot of claim 6 , wherein the processor causes the robot to perform the task based on the comparison is indicative of a determination that the updated confidence level of segmentation is greater than or equal to the confidence level threshold.

8 . The robot of claim 6 , wherein the processor causes the robot to select a next segmented object based on the comparison is indicative of a determination that the updated confidence level of segmentation is lesser than the confidence level threshold.

9 . The robot of claim 1 , wherein the processor causes the robot to execute a trained neural network to update the confidence level of the segmented object based on the segmented object and the constrained affine transformations of the template of the object.

10 . The robot of claim 1 , wherein the segmented object is transmitted to an object model for generating a plurality of synthetic images for training a segmentation model, such that the segmentation model is used to segment the object in the input image to produce the segmented object.

11 . The robot of claim 10 , wherein the plurality of synthetic images are generated based on a set of affine transformations of: the segmented object and a corresponding mask of the segmented object.

12 . The robot of claim 10 , wherein the plurality of synthetic images are generated based on recursively applying each affine transformation from the set of affine transformations, on the segmented object and the corresponding mask of the segmented object.

13 . A controller for controlling a robot for performing a task, the controller comprising:

a memory to store instructions; and

a processor configured to execute the instructions to cause the controller to perform operations, the operations comprising:

segmenting an object in an input image, wherein the input image is indicative of an environment associated with the task;

updating a confidence level of segmentation by comparing the segmented object with constrained affined transformations of a template of the object with constraints indicative of a property of the object, wherein the property of the object is selected based on the task to be performed and limits a space of affine transformations considered during updating of the confidence level; and

performing the task using the property of the object based on the updated confidence level of the segmented object.

14 . The controller of claim 13 , wherein the property of the object includes one or a combination of a non-occlusion of the object by other objects, a pose of the object, and a distance of the robot from the object.

15 . The controller of claim 14 , wherein the property of the object is the non-occlusion of the object by other objects, and the template of the object includes only an image of a non-concluded object.

16 . The controller of claim 13 , wherein performing the task using the property of the object based on the updated confidence level of the segmented object comprises:

comparing the updated confidence level of the segmented object with a confidence level threshold; and

performing the task based on the comparison.

17 . The controller of claim 16 , wherein for performing the task based on the comparison, the processor is configured for:

determining that the updated confidence level of the segmented object is greater than or equal to the confidence level threshold.

18 . A non-transitory computer-readable medium having stored thereon instructions that when executed by a computer, cause the computer to perform a method for controlling a robot for performing a task, the method comprising:

segmenting an object in an input image;

updating a confidence level of segmentation by comparing the segmented object with constrained affined transformations of a template of the object with constraints indicative of a property of the object, wherein the property of the object is selected based on the task to be performed and limits a space of affine transformations considered during updating of the confidence level; and

performing the task using the property of the object based on the updated confidence level of the segmented object.

19 . The non-transitory computer-readable medium of claim 18 , wherein the method further comprising:

comparing the updated confidence level of the segmented object with a confidence level threshold such that the confidence level threshold is set according to the task; and

performing the task based on the comparison.

20 . The non-transitory computer-readable medium of claim 18 , wherein the constrained affine transformations exclude transformations inconsistent with the property of the object used for performing the task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2024
From: CHERIAN, ANOOP; JAIN, SIDDARTH; MARKS, TIM
To: MITSUBISHI ELECTRIC RESEARCH LABORATORIES, INC.
Reel/Frame 068466/0018 →
Continuity (1)
Related Publication 20260061610A1 · Mar 5, 2026
References Cited (5)
US 11961281B1 · Kim · 2024 [cited by examiner]
US 20230024736A1 · Ogura · 2023 [cited by examiner]
Yun et al., “CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features.” https://arxiv.org/pdf/1905.04899.pdf. [cited by applicant]
https://arxiv.org/pdf/1905.04899.pdf et al., “Segmenting Transparent Objects in the Wild.” https://arxiv.org/pdf/2003.13948v3.pdf. [cited by applicant]
Kalra1 et al., “Deep Polarization Cues for Transparent Object Segmentation.” CVPR 2020 , ieee xplore, open access. https://openaccess.thecvf.com/content_CVPR_2020/papers/Kalra_Deep_Polarization_Cues_for_Transparent_Obje… [cited by applicant]