IP Library Granted Patent US 12,475,649
Granted Patent B2
US 12,475,649 · App. 18/099,186 · Granted Nov 18, 2025

Method and system for facilitating generation of background replacement masks for improved labeled image dataset collection

Inventors: Matthew A. Shreve (Campbell, CA); Robert R. Price (Palo Alto, CA)
Assignee: Xerox Corporation
G06T17/205G06T15/04H04N19/17
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,649
App. No.
18/099,186
Filed
Jan 19, 2023
Granted
Nov 18, 2025
Kind
B2
Art Unit
2614
USPC
345/423
Abstract

A system captures, by a recording device, a scene with physical objects, the scene displayed as a three-dimensional (3D) mesh. The system marks 3D annotations for a physical object and identifies a mask. The mask indicates background pixels corresponding to a region behind the physical object. Each background pixel is associated with a value. The system captures a plurality of images of the scene with varying features, wherein a respective image includes: two-dimensional (2D) projections corresponding to the marked 3D annotations for the physical object; and the mask based on the associated value for each background pixel. The system updates the value of each background pixel with a new value. The system trains a machine model using the respective image as generated labeled data, thereby obtaining the generated labeled data in an automated manner based on a minimal amount of marked annotations.

Claims (98)

1 . A computer-implemented method, comprising:

capturing, by a recording device, a scene with a plurality of physical objects, wherein the scene is displayed as a three-dimensional (3D) mesh;

marking 3D annotations for a physical object in the scene;

identifying a background mask in the scene displayed as a 3D mesh, wherein the background mask indicates pixels corresponding to a region behind or partially obstructed by the physical object and each pixel in the identified background mask is associated with a value;

capturing a plurality of images of the scene with varying features, wherein a respective image includes:

two-dimensional (2D) projections corresponding to the marked 3D annotations for the physical object; and

the background mask based on the associated value for each pixel in the background mask;

updating, in the respective image, the value of each pixel in the identified background mask with a new value; and

training a machine model using the respective image as generated labeled data, thereby obtaining the generated labeled data in an automated manner based on a minimal amount of marked annotations.

2 . The method of claim 1 , wherein identifying the background mask further comprises at least one of:

inserting, by a user associated with the recording device, the background mask in the scene using tools associated with the recording device or another computing device; and

detecting, automatically by the recording device, predetermined categories of 2D surfaces or 3D surfaces or shapes.

3 . The method of claim 1 ,

wherein the background mask comprises a virtual green screen, and

wherein the value associated with each background pixel comprises a chroma key value corresponding to a shade of green.

4 . The method of claim 1 , wherein the background mask corresponds to at least one of:

a 2D surface within the 3D mesh scene that is behind or underneath the physical object relative to the recording device; and

a 3D surface or shape within the 3D mesh scene that is behind or underneath the physical object relative to the recording device.

5 . The method of claim 1 , wherein the varying features of the captured plurality of images of the scene include or are based on at least one of:

a location, pose, or angle of the recording device relative to the physical object;

a lighting condition associated with the scene; and

an occlusion factor of the physical object in the scene.

6 . The method of claim 1 , wherein the value associated with each background pixel comprises at least one of:

a chroma key value;

a red green blue (RGB) value;

a hue saturation value (HSV) value;

a hue saturation brightness (HSB) value;

a monochrome value;

a random value;

a noisy value; and

a value or flag indicating that a respective background pixel of the background mask is to be subsequently replaced by a pixel with a different value.

7 . The method of claim 1 ,

wherein a respective background pixel is of a same or a different value than a remainder of the background pixels.

8 . The method of claim 1 , wherein updating the value of each background pixel with the new value comprises at least one of:

replacing the background pixels indicated by the background mask with a natural image, wherein the natural image comprises a differing texture from the region behind the physical object; and

replacing the background pixels indicated by the background mask with pixels of a same value or a different value as each other.

9 . The method of claim 1 , further comprising:

storing an image of the scene, including the marked 3D annotations and the identified background mask, captured by the recording device;

storing the respective image, including the 2D projections and the background mask, captured by the recording device; and

storing the respective image with the updated background pixels.

10 . A computer system, comprising:

a processor; and

a storage device storing instructions that when executed by the processor cause the processor to perform a method, the method comprising:

capturing, by a recording device, a scene with a plurality of physical objects, wherein the scene is displayed as a three-dimensional (3D) mesh;

marking 3D annotations for a physical object in the scene;

identifying a background mask in the scene displayed as a 3D mesh, wherein the background mask indicates pixels corresponding to a region behind or partially obscured by the physical object and each pixel in the identified background mask is associated with a value;

capturing a plurality of images of the scene with varying features, wherein a respective image includes:

two-dimensional (2D) projections corresponding to the marked 3D annotations for the physical object; and

the background mask based on the associated value for each background pixel;

updating, in the respective image, the value of each pixel in the identified background mask with a new value; and

training a machine model using the respective image as generated labeled data, thereby obtaining the generated labeled data in an automated manner based on a minimal amount of marked annotations.

11 . The computer system of claim 10 , wherein identifying the background mask further comprises at least one of:

inserting, by a user associated with the recording device, the background mask in the scene using tools associated with the recording device or another computing device; and

detecting, automatically by the recording device, predetermined categories of 2D surfaces or 3D surfaces or shapes.

12 . The computer system of claim 10 ,

wherein the background mask comprises a virtual green screen, and

wherein the value associated with each background pixel comprises a chroma key value corresponding to a shade of green.

13 . The computer system of claim 10 , wherein the background mask corresponds to at least one of:

a 2D surface within the 3D mesh scene that is behind or underneath the physical object relative to the recording device; and

a 3D surface or shape within the 3D mesh scene that is behind or underneath the physical object relative to the recording device.

14 . The computer system of claim 10 , wherein the varying features of the captured plurality of images of the scene include or are based on at least one of:

a location, pose, or angle of the recording device relative to the physical object;

a lighting condition associated with the scene; and

an occlusion factor of the physical object in the scene.

15 . The computer system of claim 10 , wherein the value associated with each background pixel comprises at least one of:

a chroma key value;

a red green blue (RGB) value;

a hue saturation value (HSV) value;

a hue saturation brightness (HSB) value;

a monochrome value;

a random value;

a noisy value; and

a value or flag indicating that a respective background pixel of the background mask is to be subsequently replaced by a pixel with a different value.

16 . The computer system of claim 10 , wherein updating the value of each background pixel with the new value comprises at least one of:

replacing the background pixels indicated by the background mask with a natural image, wherein the natural image comprises a differing texture from the region behind the physical object; and

replacing the background pixels indicated by the background mask with pixels of a same value or a different value as each other.

17 . The computer system of claim 10 , wherein the method further comprises:

storing an image of the scene, including the marked 3D annotations and the identified background mask, captured by the recording device;

storing the respective image, including the 2D projections and the background mask, captured by the recording device; and

storing the respective image with the updated background pixels.

18 . A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method, the method comprising:

capturing, by a recording device, a scene with a plurality of physical objects, wherein the scene is displayed as a three-dimensional (3D) mesh;

marking 3D annotations for a physical object in the scene;

identifying a background mask in the scene displayed as a 3D mesh, wherein the background mask indicates pixels corresponding to a region behind or partially obstructed by the physical object and each pixel in the identified background mask is associated with a value;

capturing a plurality of images of the scene with varying features, wherein a respective image includes:

two-dimensional (2D) projections corresponding to the marked 3D annotations for the physical object; and

the background mask based on the associated value for each background pixel;

updating, in the respective image, the value of each pixel in the identified background mask with a new value; and

training a machine model using the respective image as generated labeled data, thereby obtaining the generated labeled data in an automated manner based on a minimal amount of marked annotations.

19 . The non-transitory computer-readable storage medium of claim 18 , wherein identifying the background mask further comprises at least one of:

inserting, by a user associated with the recording device, the background mask in the scene using tools associated with the recording device or another computing device; and

detecting, automatically by the recording device, predetermined categories of 2D surfaces or 3D surfaces or shapes.

20 . The non-transitory computer-readable storage medium of claim 18 , wherein the background mask corresponds to at least one of:

a 2D surface within the 3D mesh scene that is behind or underneath the physical object relative to the recording device; and

a 3D surface or shape within the 3D mesh scene that is behind or underneath the physical object relative to the recording device, and

wherein updating the value of each background pixel with the new value comprises at least one of:

replacing the background pixels indicated by the background mask with a natural image, wherein the natural image comprises a differing texture from the region behind the physical object; and

replacing the background pixels indicated by the background mask with pixels of a same value or a different value as each other.

Assignments (8)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2025
From: XEROX CORPORATION
To: GENESEE VALLEY INNOVATIONS, LLC
Reel/Frame 073225/0116 →
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVAL OF US PATENTS 9356603, 10026651, 10626048 AND INCLUSION OF US PATENT 7167871 PREVIOUSLY RECORDED ON REEL 064038 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064161/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064038/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2023
From: SHREVE, MATTHEW A.; PRICE, ROBERT R.
To: PALO ALTO RESEARCH CENTER INCORPORATED
Reel/Frame 062441/0911 →
Continuity (1)
Related Publication 20240249476A1 · Jul 25, 2024
References Cited (55)
US 8200011B2 · Eaton · 2012 [cited by examiner]
US 10147023B1 · Klaudiny · 2018 [cited by applicant]
US 10229539B2 · Shikoda · 2019 [cited by applicant]
US 10304159B2 · Aoyagi · 2019 [cited by applicant]
US 10366521B1 · Peacock · 2019 [cited by applicant]
US 11256958B1 · Subbiah · 2022 [cited by applicant]
US 11527064B1 · Croston · 2022 [cited by examiner]
US 20080189083A1 · Schell · 2008 [cited by applicant]
US 20100040272A1 · Zheng · 2010 [cited by applicant]
US 20120117072A1 · Gokturk · 2012 [cited by examiner]
US 20120200601A1 · Osterhout · 2012 [cited by applicant]
US 20130127980A1 · Haddick · 2013 [cited by applicant]
US 20150037775A1 · Ottensmeyer · 2015 [cited by applicant]
US 20160260261A1 · Hsu · 2016 [cited by applicant]
US 20160292925A1 · Montgomerie · 2016 [cited by applicant]
US 20160328887A1 · Elvezio · 2016 [cited by applicant]
US 20160335578A1 · Sagawa · 2016 [cited by applicant]
US 20160349511A1 · Meiron · 2016 [cited by applicant]
US 20170304732A1 · Velic · 2017 [cited by examiner]
US 20170366805A1 · Sevostianov · 2017 [cited by applicant]
US 20180035606A1 · Burdoucci · 2018 [cited by applicant]
US 20180204160A1 · Chehade · 2018 [cited by applicant]
US 20180315329A1 · D'Amato · 2018 [cited by applicant]
US 20180336732A1 · Schuster · 2018 [cited by applicant]
US 20190130219A1 · Shreve · 2019 [cited by examiner]
US 20190156202A1 · Falk · 2019 [cited by applicant]
US 20190196698A1 · Cohen · 2019 [cited by examiner]
US 20190362556A1 · Ben-Dor · 2019 [cited by applicant]
US 20200013219A1 · Dhua · 2020 [cited by examiner]
US 20200081249A1 · Brusnitsyn · 2020 [cited by applicant]
US 20200210780A1 · Torres · 2020 [cited by applicant]
US 20200364878A1 · Bradski · 2020 [cited by examiner]
US 20210233308A1 · Barasofsky · 2021 [cited by examiner]
US 20210243362A1 · Castillo · 2021 [cited by applicant]
US 20210342600A1 · Westmacott · 2021 [cited by examiner]
US 20210350588A1 · Tanida · 2021 [cited by applicant]
US 20220262008A1 · Kidd · 2022 [cited by examiner]
US 20220270265A1 · Dudovitch · 2022 [cited by examiner]
US 20220417532A1 · Turbell · 2022 [cited by examiner]
US 20230123749A1 · Österberg · 2023 [cited by examiner]
US 20240046568A1 · Shreve · 2024 [cited by examiner]
Kevin Lai et al., “A Large-Scale Hierarchical Multi-View RGB-D Object Dataset”, Robotics and Automation (ICRA), 2011 IEEE International Conferene on, IEEE, May 9, 2011, pp. 1817-1824. *abstract* *Section IV* *Figure 6*. [cited by applicant]
Georgios Georgakis et al., “Multiview RGB-D Dataset for Object Instance Detection”, 2016 Fourth International Conference on 3D Vision (3DV), Sep. 26, 2016, pp. 426-434. *Abstract* *Section 3*. [cited by applicant]
Aldoma Aitor et al., “Automation of “Ground Truth” Annotation for Multi-View RGB-D Object Instance Recognition Datasets”, 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems, IEEE, Sep. 14, 2014, pp… [cited by applicant]
Pat Marion et al., “LabelFusion: A Pipeline for Generating Ground Truth Labels for Real RGBD Data of Cluttered Scenes”, Jul. 15, 2017, retrieved form the Internet: URL: https://arxiv.org/ pdf/1707.04796.pdf. *abstract* … [cited by applicant]
Alhaija et al., “Augmented Reality Meets Computer Vision: Efficient Data Generation for Urban Driving Scenes”, Aug. 4, 2017. [cited by applicant]
Kaiming He, Georgia Gkioxari, Piotr Dollar Ross Girshick, “Mask R-CNN”, arXiv:1703.06870v1 [cs.CV] Mar. 20, 2017. [cited by applicant]
Umar Iqbal, Anton Milan, and Juergen Gall, “PoseTrack: Joint Multi-Person Pose Estimation and Tracking”, arXiv:1611.07727v3 [ cs.CV] Apr. 7, 2017. [cited by applicant]
Ohan Oda, etal.; “Virtual replicas for remote assistance in virtual and augmented reality”, Proceedings of the 28th Annual ACM Symposium on User Interface Software & Technology, Nov. 8-11, 2015, Charlotte, NC, pp. 405-4… [cited by applicant]
Mueller, How To Get Phone Notifications When Your Brother HL-3170CDW Printer Runs Out Of Paper, https://web.archive.org/web /20180201064332/http://www.shareyourrepair.com/2014/12/ how-to-get-push-notifications-from-brot… [cited by applicant]
Epson Printer Says Out of Paper but It Isn't, https://smartprintsupplies.com/blogs/news/ p-strong-epson-printer-says-out-of-paper-but-it-isnt-strong-p-p-p (Year: 2018). [cited by applicant]
Adepu et al. Control Behavior Integrity for Distributed Cyber-Physical Systems (Year: 2018). [cited by applicant]
Khurram Soomro “UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild” Dec. 2012. [cited by applicant]
J. L. Pech-Pacheco et al “Diatom autofocusing in brightheld microscopy: a comparative study” IEEE 2000. [cited by applicant]
Said Pertuz et al “Analysis of focus measure operators in shape-from-focus”, Research gate Nov. 2012, Article Accepted Nov. 7, 2012, Available priline Nov. 16, 2012. [cited by applicant]