IP Library Granted Patent US 12703104
Granted Patent B2
US 12703104 · App. 18/705,969 · Granted Aug 11, 2026

System and method for managing a device and providing instruction from a remote location via a video display

Inventor: Rami Ayed Osaimi (Waltham, MA)
B25J9/1697B25J9/1664H04N23/661
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12703104
App. No.
18/705,969
Granted
Aug 11, 2026
Kind
B2
Abstract

A method and system for commanding a device via a video display. The device has a camera directed at the video display and is in communication with a processor. A command displayed on the video display is received by the camera. The processor interprets the command received by the camera and executes the interpreted command by instructing the device to carry out the command. The method and system enables the operation of telepresence devices without the need for a proprietary control mechanism and additional proprietary components. Specialized devices from different manufacturers are enabled to communicate with each other as described herein.

Claims (56)

1 . A method for commanding a device via a video display, the device having a camera directed at the video display and in communication with a processor, the method comprising:

receiving, by the camera directed at the video display, a command displayed on the video display;

interpreting, by the processor, the command received by the camera; and

executing, by the processor, the interpreted command by instructing the device to carry out the command;

wherein a remote user controls the device without requiring a proprietary control mechanism at the location of the remote user.

2 . The method of claim 1 , wherein the command comprises an image displayed on the video display.

3 . The method of claim 2 , wherein the command comprises text that is displayed on the video display.

4 . The method of claim 1 , wherein the command comprises a gesture that is displayed on the video display.

5 . The method of claim 1 , wherein the video display displays content that is sourced from a remote location from the camera and video display.

6 . The method of claim 5 , wherein the video display comprises a screen of a portable electronic device.

7 . The method of claim 1 , wherein interpreting the command comprises:

performing image recognition on the command displayed on the video display; and

determining when results of image recognition match a pre-determined trigger for instructing the device to perform the command.

8 . The method of claim 7 wherein determining when results of image recognition match a pre-determined trigger for instructing the device to perform the command is performed by consulting a listing of commands and triggers.

9 . The method of claim 7 , wherein the image recognition comprises gesture recognition.

10 . The method of claim 1 , wherein interpreting the command further comprises using machine learning and/or artificial intelligence to determine what command is displayed on the video display and/or the action that is to be carried out in response to the command.

11 . The method of claim 1 , wherein the device further includes a microphone and can further receive commands by the microphone that are interpreted and executed by the processor.

12 . The method of claim 1 , wherein executing, by the processor, the interpreted command by instructing the device to carry out the command comprises:

consulting a listing of commands and actions to determine the appropriate action for the received command; and

performing the determined appropriate action.

13 . The method of claim 1 , wherein the device comprises:

a telepresence robot comprising:

a body;

a mount configured to support the video display;

the camera directed at the video display on the mount;

the processor in communication with the camera; and

at least one motor in communication with the processor to motivate the body of the robot;

wherein the command received by the camera and interpreted and executed by the processor results in actuation of the at least one motor and movement of the telepresence robot in accordance with the command received.

14 . The method of claim 13 , wherein the command received is a command to rotate the body in a specified direction.

15 . The method of claim 13 , further comprising the video display supported by the mount.

16 . A telepresence robot receiving commands from a remote location via a video display, the robot comprising:

a body;

a mount supporting a video display;

a camera directed at the video display on the mount oriented to receive video images of commands displayed on the video display;

a processor in communication with the camera that interprets and executes commands received by the camera; and

at least one motor in communication with the processor that motivates at least a portion of the robot;

wherein the command received by the camera and interpreted and executed by the processor results in actuation of the motor and movement of at least a portion of the robot in accordance with the command; and

wherein a remote user controls the telepresence robot without requiring a proprietary control mechanism at the remote user's location.

17 . The robot of claim 16 , wherein the command comprises an image displayed on the video display.

18 . The robot of claim 16 , wherein the command comprises text displayed on the video display.

19 . The robot of claim 16 , wherein the command is a gesture performed by an individual displayed on the video display.

20 . The robot of claim 16 , wherein the video display displays content sourced from a remote location.

21 . The robot of claim 20 , wherein the video display comprises a screen of a portable electronic device.

22 . The robot of claim 16 , wherein the processor interprets and executes the commands by:

performing image recognition on the command displayed on the video display; and

determining when results of image recognition match a pre-determined trigger for instructing the robot to perform the command.

23 . The robot of claim 22 wherein determining when results of image recognition match a pre-determined trigger for instructing the robot to perform the command is performed by consulting a listing of commands and triggers.

24 . The robot of claim 22 , wherein processor interprets and executes the commands by:

consulting a listing of commands and actions to determine an appropriate action for the received command; and

performing the determined appropriate action.

25 . The robot of claim 22 , wherein image recognition comprises gesture recognition.

26 . The robot of claim 16 , wherein the processor interprets and executes the commands by using machine learning and/or artificial intelligence to determine what command is displayed on the video display and/or an action that is to be carried out in response to the command.

27 . The robot of claim 16 , further comprising:

a microphone, in communication with the processor, for receiving audio commands;

wherein a command received by the microphone is interpreted and executed by the processor and results in actuation of the motor and movement of at least a portion of the robot in accordance with the command.

28 . The robot of claim 16 , wherein the command received is a command to rotate the body of the robot in a specified direction.