IP Library Granted Patent US 10,242,449
Granted Patent B2
US 10,242,449 · App. 15/397,987 · Granted Mar 26, 2019

Automated generation of pre-labeled training data

Inventors: Rob Liston (Menlo Park, CA); John G. Apostolopoulos (Palo Alto, CA)
Assignee: Cisco Technology, Inc.
G06T7/11G06T2207/10004G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,242,449
App. No.
15/397,987
Granted
Mar 26, 2019
Kind
B2
Abstract

Presented herein are techniques for automatically generating object segmentation training data. In particular, a segmentation data generation system is configured to obtain training images derived from a scene captured by one or more image capture devices. Each training image is a still image that includes a foreground object and a background. The segmentation data generation system automatically generates a mask of the training image to delineate the object from the background and, based on the mask automatically generates a masked image. The masked image includes only the object present in the training image. The segmentation data generation system composites the masked image with an image of an environmental scene to generate a composite image that includes the masked image and the environmental scene.

Claims (58)

1. A computer implemented method comprising:

obtaining a training image derived from a scene captured by one or more image capture devices, wherein the training image includes a foreground object and a background;

automatically generating a mask of the training image to delineate the foreground object from the background;

based on the mask of the training image, automatically generating a masked image that includes only the foreground object present in the training image;

compositing the masked image with an image of an environmental scene to generate a composite image that includes the masked image and the environmental scene;

generating a region and label pair which represents the masked image in the composite image, wherein the composite image and the region and label pair form an object segmentation training data set; and

providing the object segmentation training data set to a machine learning-based image segmentation process.

2. The method of claim 1 , wherein compositing the masked image with the image of the environmental scene comprises:

performing one or more translations, rotations, scaling, or other perspective transformations to one or more of the masked image and the environmental scene.

3. The method of claim 1 , wherein compositing the masked image with the image of the environmental scene includes:

performing coherence matching of the masked image and the environmental scene to normalize the masked image and the environmental scene.

4. The method of claim 3 , wherein performing coherence matching of the masked image and the environmental scene comprises:

adjusting lighting characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene have substantially the same amount of illumination.

5. The method of claim 3 , wherein performing coherence matching of the masked image and the environmental scene comprises:

adjusting lighting characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene include substantially the same spectrum of visible light.

6. The method of claim 3 , wherein performing coherence matching of the masked image and the environmental scene comprises:

adjusting noise characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene include substantially the same amount of random noise.

7. The method of claim 1 , wherein automatically generating a mask of the training image to delineate the foreground object from the background comprises:

automatically generating a binary mask of the training image.

8. The method of claim 7 , wherein automatically generating the binary mask of the training image comprises:

using at least one of a chroma key approach or a depth camera approach to generate the binary mask.

9. An apparatus comprising:

a memory;

a network interface unit; and

a processor configured to:

obtain a training image derived from a scene captured by one or more image capture devices, wherein the training image includes a foreground object and a background;

automatically generate a mask of the training image to delineate the foreground object from the background;

based on the mask of the training image, automatically generate a masked image that includes only the foreground object present in the training image;

composite the masked image with an image of an environmental scene to generate a composite image that includes the masked image and the environmental scene;

generate a region and label pair which represents the masked image in the composite image, wherein the composite image and the region and label pair form an object segmentation training data set; and

provide the object segmentation training data set to a machine learning-based image segmentation process.

10. The apparatus of claim 9 , wherein to composite the masked image with the image of the environmental scene, the processor is configured to:

perform one or more translations, rotations, scaling, or other perspective transformations to one or more of the masked image and the environmental scene.

11. The apparatus of claim 9 , wherein to composite the masked image with the image of the environmental scene, the processor is configured to:

perform coherence matching of the masked image and the environmental scene to normalize the masked image and the environmental scene.

12. The apparatus of claim 11 , wherein to perform coherence matching of the masked image and the environmental scene, the processor is configured to:

adjust lighting characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene have substantially the same amount of illumination.

13. The apparatus of claim 11 , wherein to perform coherence matching of the masked image and the environmental scene, the processor is configured to:

adjust lighting characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene include substantially the same spectrum of visible light.

14. The apparatus of claim 11 , wherein to perform coherence matching of the masked image and the environmental scene, the processor is configured to:

adjust noise characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene include substantially the same amount of random noise.

15. One or more non-transitory computer readable storage media encoded with instructions that, when executed by a processor, cause the processor to:

obtain a training image derived from a scene captured by one or more image capture devices, wherein the training image includes a foreground object and a background;

automatically generate a mask of the training image to delineate the foreground object from the background;

based on the mask of the training image, automatically generate a masked image that includes only the foreground object present in the training image;

composite the masked image with an image of an environmental scene to generate a composite image that includes the masked image and the environmental scene;

generate a region and label pair which represents the masked image in the composite image, wherein the composite image and the region and label pair form an object segmentation training data set; and

provide the object segmentation training data set to a machine learning-based image segmentation process.

16. The non-transitory computer readable storage media of claim 15 , wherein the instructions operable to composite the masked image with the image of the environmental scene comprise instructions operable to:

perform one or more translations, rotations, scaling, or other perspective transformations to one or more of the masked image and the environmental scene.

17. The non-transitory computer readable storage media of claim 15 , wherein the instructions operable to composite the masked image with the image of the environmental scene comprise instructions operable to:

perform coherence matching of the masked image and the environmental scene to normalize the masked image and the environmental scene.

18. The non-transitory computer readable storage media of claim 17 , wherein the instructions operable to perform coherence matching of the masked image and the environmental scene comprise instructions operable to:

adjust lighting characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene have substantially the same amount of illumination.

19. The non-transitory computer readable storage media of claim 17 , wherein the instructions operable to perform coherence matching of the masked image and the environmental scene comprise instructions operable to:

adjust lighting characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene include substantially the same spectrum of visible light.

20. The non-transitory computer readable storage media of claim 17 , wherein the instructions operable to perform coherence matching of the masked image and the environmental scene comprise instructions operable to:

adjust noise characteristics of at least one of the masked image or the environmental scene such that the masked image and the environmental scene include substantially the same amount of random noise.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2017
From: LISTON, ROB; APOSTOLOPOULOS, JOHN G.
To: CISCO TECHNOLOGY, INC.
Reel/Frame 040839/0665 →
Continuity (1)
Related Publication 20180189951A1 · Jul 5, 2018
Cited By (6)
US 12,518,169 US 12,530,776 US 12,548,162 US 12,639,824 US 12,651,349 US 12,651,350