IP Library Granted Patent US 12664799
Granted Patent B2
US 12664799 · App. 18/360,330 · Granted Jun 23, 2026

Automatically captioning images using swipe gestures as inputs

Inventors: Bing Qin Lim (Bayan Lepas, MY); Cecilia Liaw (Bayan Lepas, MY); Ming Yeh Koh (Bayan Lepas, MY); Moh Lim Sim (Bayan Lepas, MY)
Assignee: MOTOROLA SOLUTIONS, INC.
G06V20/70G06F3/04883G06F3/14G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664799
App. No.
18/360,330
Granted
Jun 23, 2026
Kind
B2
Abstract

Devices, systems, and methods for automatically captioning images using swipe gestures as inputs. One example apparatus includes an electronic processor. The electronic processor is configured to receive an image. The electronic processor is configured to control a display to display the image. The electronic processor is configured to detect a first object in the image. The electronic processor is configured to detect a second object in the image. The electronic processor is configured to receive, from the display, a first swipe gesture. The electronic processor is configured to, responsive to receiving the first swipe gesture, determine a direction of the first swipe gesture relative to the first object and the second object. The electronic processor is configured to determine a word choice based on the first direction. The electronic processor is configured to generate a caption describing the image based on the word choice.

Claims (83)

1 . An apparatus comprising:

an electronic processor configured to:

receive an image;

control a display to display the image;

detect a first object in the image;

detect a second object in the image;

receive, from the display, a first swipe gesture;

responsive to receiving the first swipe gesture, determine a first direction of the first swipe gesture relative to the first object and the second object;

determine a word choice based on the first direction; and

generate a caption describing the image based on the word choice;

wherein:

the first direction of the first swipe gesture indicates an order in which the first object and the second object were selected by the first swipe gesture; and

the order in which the first object and the second object were selected by the first swipe gesture determines:

which of the first object and the second object is identified as the subject of the caption, and

which of the first object and the second object is identified as the predicate of the caption; and

wherein the word choice includes a linking verb that indicates which of the first object and the second object acted upon the other and how one of the first object and the second object acted upon the other.

2 . The apparatus of claim 1 , wherein the electronic processor is further configured to:

detect the first object and the second object by performing object recognition to recognize particular objects based on an incident type or a Computer Aided Dispatch (CAD) identifier;

retrieve an incident report based on the CAD identifier;

generate an initial caption for the image based on a content of the incident report; and

replace the initial caption with the caption.

3 . The apparatus of claim 1 , wherein the electronic processor is further configured to:

receive, from the display, a second swipe gesture;

responsive to receiving the second swipe gesture, determine a second direction of the second swipe gesture relative to the first object and the second object; and

determine the word choice based on the first direction and the second direction.

4 . The apparatus of claim 3 , wherein the electronic processor is further configured to:

determine an event sequence for the image based on the first swipe gesture and the second swipe gesture; and

generate the caption for the image based on the event sequence.

5 . The apparatus of claim 4 , wherein the event sequence indicates:

which of the first object and the second object performed a first action, and

which of the first object and the second object performed a second action.

6 . The apparatus of claim 4 , wherein the electronic processor is further configured to:

generate the caption describing the image by:

determining a context for the image based on the first swipe gesture, the second swipe gesture, the first direction, the second direction, the event sequence, and image analytics of the image; and

providing the context to a caption generation engine.

7 . The apparatus of claim 3 , wherein the electronic processor is further configured to:

detect the first object and the second object by performing object recognition to recognize particular objects based on an incident type or a Computer Aided Dispatch (CAD) identifier;

retrieve an incident report based on the CAD identifier;

assign one of the first object and the second object to a subject and assign the other of the first object and the second object to a predicate based on the first direction and the second direction; and

determine a verb relating the subject to the predicate based on a content of the incident report.

8 . The apparatus of claim 1 , further comprising:

a camera;

wherein the electronic processor is further configured to:

receive the image from the camera.

9 . A method for automatically captioning images using swipe gestures as inputs, the method comprising:

receiving an image;

displaying the image on a display;

detecting a first object in the image;

detecting a second object in the image;

receiving a first swipe gesture;

responsive to receiving the first swipe gesture, determining a first direction of the first swipe gesture relative to the first object and the second object;

determining a word choice based on the first direction; and

generating, with an electronic processor, a caption describing the image based on the word choice;

wherein:

the first direction of the first swipe gesture indicates an order in which the first object and the second object were selected by the first swipe gesture; and

the order in which the first object and the second object were selected by the first swipe gesture determines:

which of the first object and the second object is identified as the subject of the caption, and

which of the first object and the second object is identified as the predicate of the caption; and

wherein the word choice includes a linking verb that indicates which of the first object and the second object acted upon the other and how one of the first object and the second object acted upon the other.

10 . The method of claim 9 , further comprising:

detecting the first object and the second object by performing object recognition to recognize particular objects based on an incident type or a Computer Aided Dispatch (CAD) identifier;

retrieving an incident report based on the CAD identifier;

generating, with the electronic processor, an initial caption for the image based on a content of the incident report; and

replacing the initial caption with the caption.

11 . The method of claim 9 , further comprising:

receiving a second swipe gesture;

responsive to receiving the second swipe gesture, determining a second direction of the second swipe gesture relative to the first object and the second object; and

determining the word choice based on the first direction and the second direction.

12 . The method of claim 11 , further comprising:

determining an event sequence for the image based on the first swipe gesture and the second swipe gesture; and

generating, with the electronic processor, the caption for the image based on the event sequence.

13 . The method of claim 12 , wherein the event sequence indicates:

which of the first object and the second object performed a first action, and which of the first object and the second object performed a second action.

14 . The method of claim 12 , wherein generating the caption describing the image includes:

determining a context for the image based on the first swipe gesture, the second swipe gesture, the first direction, the second direction, the event sequence, and image analytics of the image; and

providing the context to a caption generation engine.

15 . The method of claim 11 , further comprising:

detecting the first object and the second object by performing object recognition to recognize particular objects based on an incident type or a Computer Aided Dispatch (CAD) identifier;

retrieving an incident report based on the CAD identifier;

assigning one of the first object and the second object to a subject and assigning the other of the first object and the second object to a predicate based on the first direction and the second direction; and

determining a verb relating the subject to the predicate based on a content of the incident report.

16 . The method of claim 9 , wherein receiving the image includes:

receiving the image from a camera of a portable electronic device including the electronic processor.