IP Library Granted Patent US 12,347,114
Granted Patent B2
US 12,347,114 · App. 18/343,591 · Granted Jul 1, 2025

Object segmentation and feature tracking

Inventors: Abhijeet Bisain (San Diego, CA); Gerhard Reitmayr (Del Mar, CA)
Assignee: QUALCOMM Incorporated
G06T7/174G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,114
App. No.
18/343,591
Granted
Jul 1, 2025
Kind
B2
Abstract

Examples are described for processing images to mask dynamic objects out of images to improve feature tracking between images. A device receives an image of an environment captured by an image sensor. The image depicts at least a static portion of the environment and a dynamic object in the environment. The device identifies a portion of the image that includes a depiction of the dynamic object. For example, the device can detect a bounding box around the dynamic object, or can detect which pixels in the image correspond to the dynamic object. The device generates a masked image at least by masking the portion of the image. The device identifies features in the masked image, and uses the features from the masked image for feature tracking from other images of the environment, masked or otherwise. The device can use this feature tracking for mapping, localization, and/or relocation.

Claims (61)

1. An apparatus for image processing, the apparatus comprising:

a memory; and

one or more processors coupled to the memory and configured to:

identify a first portion of an image of an environment captured by an image sensor, wherein the first portion of the image depicts a first object in the environment and a second portion of the image depicts a second object in the environment;

identify one or more features of the second object in a masked image, wherein the first portion of the image depicting the first object is masked in the masked image; and

track the one or more features of the second object between the masked image and one or more additional masked images of the environment, wherein the first object is masked in the one or more additional masked images.

2. The apparatus of claim 1 , wherein the one or more processors are configured to:

determine a location of a first feature of the one or more features of the second object based on tracking of the one or more features of the second object between the masked image and the one or more additional masked images of the environment; and

update a map of the environment based on the location of the first feature.

3. The apparatus of claim 2 , wherein, to update the map of the environment based on the location, the one or more processors are configured to add the location of the first feature to the map.

4. The apparatus of claim 2 , wherein, to update the map of the environment based on the location, the one or more processors are configured to modify a prior location of the first feature in the map based on the location of the first feature.

5. The apparatus of claim 1 , wherein the one or more processors are configured to:

determine a pose of the apparatus within the environment based on tracking of the one or more features of the second object between the masked image and the one or more additional masked images of the environment, wherein the pose of the apparatus within the environment includes at least one of a location of the apparatus, a pitch of the apparatus, a roll of the apparatus, or a yaw of the apparatus.

6. The apparatus of claim 1 , wherein the one or more processors are configured to:

generate a downscaled image at least by downscaling the image, wherein identifying the first portion of the image depicting the first object includes identifying a portion of the downscaled image that depicts the first object.

7. The apparatus of claim 1 , wherein the one or more processors are configured to:

generate a greyscale image at least by desaturating color in the image, wherein identifying the first portion of the image depicting the first object includes identifying a portion of the greyscale image that depicts the first object.

8. The apparatus of claim 1 , wherein, to identify the first portion of the image depicting the first object, the one or more processors are configured to:

analyze each pixel of a plurality of pixels corresponding to the image to identify a subset of the plurality of pixels that depicts at least a portion of the first object.

9. The apparatus of claim 1 , wherein, to identify the first portion of the image depicting the first object, the one or more processors are configured to identify a bounding box occupying a polygonal region of the image, wherein the first object is at least partially included within the bounding box.

10. The apparatus of claim 9 , wherein, to identify the first portion of the image depicting the first object, the one or more processors are configured to analyze each pixel of a plurality of pixels within the bounding box to identify a subset of the plurality of pixels within the bounding box that each depict a portion of the first object.

11. The apparatus of claim 10 , wherein:

to identify the bounding box, the one or more processors are configured to use at least a first trained neural network; and

to identify the subset of the plurality of pixels, the one or more processors are configured to use at least a second trained neural network.

12. The apparatus of claim 1 , wherein, to identify the first portion of the image depicting the first object, the one or more processors are configured to:

identify, using at least a first trained neural network, that the image depicts the first object; and

identify, using at least a second trained neural network in response to identification that the image depicts the first object, the first portion of the image that depicts the first object.

13. The apparatus of claim 1 , wherein the first object is a dynamic type of object and the second object is a static type of object.

14. The apparatus of claim 13 , wherein the dynamic type of object is a person, and wherein, to identify the first portion of the image depicting the first object, the one or more processors are configured to identify a depiction of a face of the person.

15. The apparatus of claim 13 , wherein the static type of object is static relative to a position of the image sensor during capture of the image, wherein the dynamic type of object moves relative to a position of the image sensor during capture of the image.

16. The apparatus of claim 1 , wherein the apparatus is one of a mobile device, a wireless communication device, a robot, a vehicle, a head-mounted display, and a camera.

17. The apparatus of claim 1 , further comprising:

the image sensor.

18. A method of image processing performed by a device, the method comprising:

identifying a first portion of an image of an environment captured by an image sensor, wherein the first portion of the image depicts a first object in the environment and a second portion of the image depicts a second object in the environment;

identifying one or more features of the second object in a masked image, wherein the first portion of the image depicting the first object is masked in the masked image; and

tracking the one or more features of the second object between the masked image and one or more additional masked images of the environment, wherein the first object is masked in the one or more additional masked images.

19. The method of claim 18 , further comprising:

determining a location of a first feature of the one or more features of the second object based on tracking of the one or more features of the second object between the masked image and the one or more additional masked images of the environment; and

updating a map of the environment based on the location of the first feature.

20. The method of claim 19 , wherein updating the map of the environment based on the location comprises adding the location of the first feature to the map.

21. The method of claim 19 , wherein updating the map of the environment based on the location comprises modifying a prior location of the first feature in the map based on the location of the first feature.

22. The method of claim 18 , further comprising:

determining a pose of the device within the environment based on tracking of the one or more features of the second object between the masked image and the one or more additional masked images of the environment, wherein the pose of the device within the environment includes at least one of a location of the device, a pitch of the device, a roll of the device, or a yaw of the device.

23. The method of claim 18 , further comprising:

generating a downscaled image at least by downscaling the image, wherein identifying the first portion of the image depicting the first object includes identifying a portion of the downscaled image that depicts the first object.

24. The method of claim 18 , further comprising:

generating a greyscale image at least by desaturating color in the image, wherein identifying the first portion of the image depicting the first object includes identifying a portion of the greyscale image that depicts the first object.

25. The method of claim 18 , wherein identifying the first portion of the image depicting the first object comprises:

analyzing each pixel of a plurality of pixels corresponding to the image to identify a subset of the plurality of pixels that depicts at least a portion of the first object.

26. The method of claim 18 , wherein identifying the first portion of the image depicting the first object comprises identifying a bounding box occupying a polygonal region of the image, wherein the first object is at least partially included within the bounding box.

27. The method of claim 26 , wherein identifying the first portion of the image depicting the first object comprises analyzing each pixel of a plurality of pixels within the bounding box to identify a subset of the plurality of pixels within the bounding box that each depict a portion of the first object.

28. The method of claim 27 , wherein:

identifying the bounding box comprises using at least a first trained neural network; and

identifying the subset of the plurality of pixels comprises using at least a second trained neural network.

29. The method of claim 18 , wherein identifying the first portion of the image depicting the first object comprises:

identifying, using at least a first trained neural network, that the image depicts the first object; and

identifying, using at least a second trained neural network in response to identification that the image depicts the first object, the first portion of the image that depicts the first object.

30. The method of claim 18 , wherein the first object is a dynamic type of object and the second object is a static type of object.

31. The apparatus of claim 1 , wherein the one or more processors are configured to:

map the environment based on the one or more features of the second object as tracked across the masked image and the one or more additional masked images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2023
From: BISAIN, ABHIJEET; REITMAYR, GERHARD
To: QUALCOMM INCORPORATED
Reel/Frame 065061/0960 →
Continuity (2)
Continuation 17127568 · Dec 18, 2020
Related Publication 20230342943A1 · Oct 26, 2023
References Cited (33)
US 9652681B1 · Tucker · 2017 [cited by examiner]
US 9904852B2 · Divakaran · 2018 [cited by examiner]
US 9965865B1 · Agrawal · 2018 [cited by examiner]
US 10860034B1 · Ziyaee · 2020 [cited by examiner]
US 11145076B1 · Horesh · 2021 [cited by examiner]
US 11727576B2 · Bisain · 2023 [cited by examiner]
US 20080226128A1 · Birtwistle · 2008 [cited by examiner]
US 20170294210A1 · Abramson · 2017 [cited by examiner]
US 20190383945A1 · Wang · 2019 [cited by examiner]
US 20200226769A1 · Das · 2020 [cited by examiner]
US 20210012503A1 · Cho · 2021 [cited by examiner]
US 20210097296A1 · Wang · 2021 [cited by examiner]
US 20210129868A1 · Nehmadi · 2021 [cited by examiner]
US 20210206312A1 · Mochizuki · 2021 [cited by examiner]
US 20210264167A1 · Chen · 2021 [cited by examiner]
US 20210383018A1 · Keskikangas · 2021 [cited by examiner]
US 20210383553A1 · Guizilini · 2021 [cited by examiner]
US 20220012916A1 · Srinivasan · 2022 [cited by examiner]
US 20220044042A1 · Kefayati · 2022 [cited by examiner]
US 20220084234A1 · Lee · 2022 [cited by examiner]
US 20220101539A1 · Lin · 2022 [cited by examiner]
US 20220156943A1 · Zhang · 2022 [cited by examiner]
US 20220189029A1 · Mequanint · 2022 [cited by examiner]
US 20220189108A1 · Shandilya · 2022 [cited by examiner]
US 20220198677A1 · Bisain · 2022 [cited by examiner]
US 20220254146A1 · Zou · 2022 [cited by examiner]
US 20220294998A1 · Ronchini Ximenes · 2022 [cited by examiner]
US 20230135137A1 · Yang · 2023 [cited by examiner]
Han S., et al., “Monocular SLAM and Obstacle Removal for Indoor Navigation”, 2018 International Conference on Machine Learning and Data Engineering (ICMLDE), IEEE, Dec. 3, 2018, pp. 67-76, XP033502183, DOI:10.1109/ICMLD… [cited by applicant]
International Search Report and Written Opinion—PCT/US2021/072551—ISA/EPO—Mar. 2, 2022. [cited by applicant]
Kaneko M., et al., “Mask-SLAM: Robust Feature-Based Monocular SLAM by Masking Using Semantic Segmentation,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Jun. 18-22, 2018, pp. 37… [cited by applicant]
Soares J.C.V., et al., “Visual SLAM in Human Populated Environments: Exploring the Trade-off between Accuracy and Speed of YOLO and Mask R-CNN,” 19th International Conference on Advanced Robotics (ICAR), Dec. 2019, pp. … [cited by applicant]
Venator M., et al., “Robust Camera Pose Estimation for Unordered Road Scene Images in Varying Viewing Conditions”, IEEE Transactions on Intelligent Vehicles, IEEE, vol. 5, No. 1, Nov. 22, 2019, pp. 165-174, XP011774774,… [cited by applicant]
Cited By (2)
US 12,400,291 US 12,548,113