IP Library Granted Patent US 10,249,044
Granted Patent B2
US 10,249,044 · App. 15/395,512 · Granted Apr 2, 2019

Image segmentation with touch interaction

Inventors: Vincent Charles Cheung (San Carlos, CA); Connie Yeewei Ho (San Jose, CA); Balmanohar Paluri (Mountain View, CA)
Assignee: Facebook, Inc.
G06T7/11G06F3/04815G06F3/04845G06F3/167G06K9/00671G06K9/3233G06F3/017G06F3/04842G06T2200/24G06T2207/10004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,249,044
App. No.
15/395,512
Granted
Apr 2, 2019
Kind
B2
Abstract

In one embodiment, a method includes detecting one or more objects in an image, generating at least one mask for each of the detected objects, wherein each of the masks is defined by a perimeter, classifying the detected objects, receiving gesture input in relation to the image, determining whether one or more locations associated with the gesture input correlate with any of the masks, and providing feedback regarding the image in response to the gesture input. Each of the masks may include data identifying the corresponding detected object, and the perimeter of each mask may correspond to a perimeter of the corresponding detected object. The perimeter of the corresponding detected object may separate the detected object from one or more portions of the image that are distinct from the detected object.

Claims (47)

1. A method comprising:

by a computing device, detecting one or more objects in an image;

by the computing device, generating at least one mask for each of the detected objects, wherein each of the masks is defined by a perimeter;

by the computing device, classifying the detected objects;

by the computing device, receiving gesture input in relation to the image;

by the computing device, determining whether one or more locations associated with the gesture input correlate with any of the masks; and

by a computing device, providing feedback regarding the image in response to the gesture input,

wherein a location associated with the gesture input corresponds to a region in which first and second masks overlap, and the feedback regarding the image comprises a description of a position of a first object that corresponds to the first mask relative to a second object that corresponds to the second mask.

2. The method of claim 1 , wherein each of the masks comprises data identifying the corresponding detected object, and the perimeter that defines each mask corresponds to a perimeter of the corresponding detected object.

3. The method of claim 2 , wherein the perimeter of the corresponding detected object separates the detected object from one or more portions of the image that are distinct from the detected object.

4. The method of claim 2 , wherein the mask comprises a matrix of data values that correspond to pixels of at least a portion of the image, pixels located inside the perimeter that defines the mask correspond to data values having a first value, and pixels located outside the perimeter that defines the mask correspond to data values having a second value.

5. The method of claim 1 , wherein the one or more objects correspond to one or more segments of the image, and the segments are detected by an image segmentation algorithm.

6. The method of claim 1 , wherein the one or more objects are classified as corresponding to a specified object type, and the feedback regarding the image comprises the specified object type.

7. The method of claim 6 , wherein the feedback is played as speech by an audio output component of the computing device.

8. The method of claim 6 , wherein the feedback is displayed as text by a display component of the computing device.

9. The method of claim 1 , wherein the one or more objects are classified as having a specified object name, and the feedback regarding the image comprises the specified object name.

10. The method of claim 1 , wherein the gesture input comprises a swipe gesture across a touch screen of the computing device or a tap gesture on the touch screen, and at least a portion of the gesture is detected at a location on the touch screen at which a portion of the image is displayed.

11. The method of claim 1 , wherein the gesture input comprises a pointing gesture that is detected as indicating a location on a display screen of the computing device at which a portion of the image is displayed.

12. The method of claim 1 , wherein determining whether one or more locations associated with the gesture input correlate with any of the masks comprises:

determining whether, for each of the detected objects, the one or more locations associated with the gesture input correspond to the mask for the detected object.

13. The method of claim 12 , wherein the one or more locations associated with the gesture input correspond to the mask for the detected object when the one or more locations are within the perimeter that defines the mask.

14. The method of claim 1 , wherein the first and second masks partially or fully overlap, the first and second objects are classified as having specified respective first and second object types, and

the feedback regarding the image further comprises the first and second object types.

15. The method of claim 1 , wherein the second object is smaller than the first object, the first and second objects are classified as having specified first and second object types, and

the feedback regarding the image further comprises an object type of the smaller of the first and second objects.

16. The method of claim 1 , wherein a location associated with the gesture input corresponds to a region in which the masks overlap, and providing feedback regarding the image in response to the gesture input comprises:

playing, by an audio component of the computing device, speech that includes an object type of the first object that corresponds to the first mask;

waiting for user input; and

in response to the user input, playing, by the audio component of the computing device, speech that includes an object type of the second object second object that corresponds to the second mask.

17. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

detect one or more objects in an image;

generate at least one mask for each of the detected objects, wherein each of the masks is defined by a perimeter;

classify the detected objects;

receive gesture input in relation to the image;

determine whether one or more locations associated with the gesture input correlate with any of the masks; and

provide feedback regarding the image in response to the gesture input,

wherein a location associated with the gesture input corresponds to a region in which first and second masks overlap, and the feedback regarding the image comprises a description of a position of a first object that corresponds to the first mask relative to a second object that corresponds to the second mask.

18. The media of claim 17 , wherein each of the masks comprises data identifying the corresponding detected object, and the perimeter that defines each mask corresponds to a perimeter of the corresponding detected object.

19. A system comprising: one or more processors; and a memory coupled to the processors comprising instructions executable by the processors, the processors being operable when executing the instructions to:

detect one or more objects in an image;

generate at least one mask for each of the detected objects, wherein each of the masks is defined by a perimeter;

classify the detected objects;

receive gesture input in relation to the image;

determine whether one or more locations associated with the gesture input correlate with any of the masks; and

provide feedback regarding the image in response to the gesture input,

wherein a location associated with the gesture input corresponds to a region in which first and second masks overlap, and the feedback regarding the image comprises a description of a position of a first object that corresponds to the first mask relative to a second object that corresponds to the second mask.

20. The system of claim 19 , wherein each of the masks comprises data identifying the corresponding detected object, and the perimeter that defines each mask corresponds to a perimeter of the corresponding detected object.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2017
From: CHEUNG, VINCENT CHARLES; HO, CONNIE YEEWEI; PALURI, BALMANOHAR
To: FACEBOOK, INC.
Reel/Frame 041240/0709 →
Continuity (1)
Related Publication 20180189598A1 · Jul 5, 2018
Cited By (15)
US 12,210,800 US 12,288,279 US 12,299,858 US 12,333,692 US 12,347,005 US 12,395,722 US 12,456,243 US 12,462,519 US 12,488,523 US 12,505,596 US 12,536,625 US 12,597,186 US 12,646,188 US 12,657,902 US 12,699,852