IP Library Granted Patent US 10,671,934
Granted Patent B1
US 10,671,934 · App. 16/512,751 · Granted Jun 2, 2020

Real-time deployment of machine learning systems

Inventors: Andrew Ninh (Fountain Valley, CA); Tyler Dao (Fountain Valley, CA); Mohammad Fidaali (Chino Hills, CA)
Assignee: DOCBOT, Inc.
G06N5/04G06K9/00718G06K9/6267G06N20/00H04N19/115H04N19/186
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,671,934
App. No.
16/512,751
Granted
Jun 2, 2020
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for real-time deployment of machine learning systems. One of the operations is performed by the system receiving video data from a video image capturing device. The received video data is converted into multiple video frames. These video frames are encoded into a particular color space format. The system renders a first display output depicting imagery from the multiple encoded video frames. The system performs an inference on the video frames using a machine learning network to determine the occurrence of one or more objects in the video frames. The system renders a second display output depicting graphical information corresponding to the determined one or more objects from the multiple encoded video frames. The system then generates a composite display output including the imagery of the first display output overlaid with the graphical information of the second display output.

Claims (54)

1. A system comprising one or more processors, and a non-transitory computer-readable medium including one or more sequences of instructions that, when executed by the one or more processors, cause the system to perform operations comprising:

receiving video data, the video data having been obtained from a video image capture device;

converting the received video data into multiple video frames encoded into a particular color space format;

rendering a first display output depicting imagery from the multiple encoded video frames;

performing an inference on the multiple video frames using a machine learning network;

determining the occurrence of one or more objects in the multiple encoded video frames based on the performed inference on the multiple video frames;

in response to determining the occurrence of one or more objects, generating for a determined object, coordinates describing a bounding perimeter about the determined object;

rendering a second display output depicting graphical information in a form corresponding to the coordinates of the bounding perimeter for the determined one or more objects from the multiple encoded video frames; and

generating a composite display output, wherein the composite display output includes the imagery of the first display output overlaid with the graphical information of the second display output.

2. The system of claim 1 , wherein the first display output depicts imagery at a frame rate of 50 to 240 frames per second.

3. The system of claim 2 , wherein the second display output depicts the graphical information at a frame rate less than or equal to the frame rate of the first display output.

4. The system of claim 1 , further comprising the operations of:

generating a graphical indication around or about the one or more objects indicating the location of the identified objects in the video frame.

5. The system of claim 1 , further comprising the operations of:

determining an external environmental state of the video image capture device; and

performing the inference if the external environmental state is suitable to perform inferencing via the machine learning network.

6. The system of claim 1 , wherein the multiple encoded video frames are encoded in a color space format selected from the group consisting of NV12, I420, YV12, YUY2, YUYV, UYVY, UVYU, V308, IYU2, V408, RGB24, RGB32, V410, Y410 and Y42T.

7. The system of claim 1 , wherein the graphical information of the second display output includes graphical indications of the one or more objects disposed over a video display area of the first display output, and textual information corresponding to the one or more objects disposed over a non-video display area of the first display output.

8. A method implemented by a system comprising of one or more processors, the method comprising:

receiving video data, the video data having been obtained from a video image capture device;

converting the received video data into multiple video frames encoded into a particular color space format;

rendering a first display output depicting imagery from the multiple encoded video frames;

performing an inference on the multiple video frames using a machine learning network;

determining the occurrence of one or more objects in the multiple encoded video frames based on the performed inference on the multiple video frames;

in response to determining the occurrence of one or more objects, generating for a determined object, coordinates describing a bounding perimeter about the determined object;

rendering a second display output depicting graphical information in a form corresponding to the coordinates of the bounding perimeter for the determined one or more objects from the multiple encoded video frames; and

generating a composite display output, wherein the composite display output includes the imagery of the first display output overlaid with the graphical information of the second display output.

9. The method of claim 8 , wherein the first display output depicts imagery at a frame rate of 50 to 240 frames per second.

10. The method of claim 9 , wherein the second display output depicts the graphical information at a frame rate less than or equal to the frame rate of the first display output.

11. The method of claim 8 , further comprising the operations of:

generating a graphical indication around or about the one or more objects indicating the location of the identified objects in the video frame.

12. The method of claim 8 , further comprising the operations of:

determining an external environmental state of the video image capture device; and

performing the inference if the external environmental state is suitable to perform inferencing via the machine learning network.

13. The method of claim 8 , wherein the multiple encoded video frames are encoded in a color space format selected from the group consisting of NV12, I420, YV12, YUY2, YUYV, UYVY, UVYU, V308, IYU2, V408, RGB24, RGB32, V410, Y410 and Y42T.

14. The method of claim 8 , wherein the graphical information of the second display output includes graphical indications of the one or more objects disposed over a video display area of the first display output, and textual information corresponding to the one or more objects disposed over a non-video display area of the first display output.

15. A non-transitory computer storage medium comprising instructions that when executed by a system comprising one or more processors, cause the one or more processors to perform operations comprising:

receiving video data, the video data having been obtained from a video image capture device;

converting the received video data into multiple video frames encoded into a particular color space format;

rendering a first display output depicting imagery from the multiple encoded video frames;

performing an inference on the multiple video frames using a machine learning network;

determining the occurrence of one or more objects in the multiple encoded video frames based on the performed inference on the multiple video frames;

in response to determining the occurrence of one or more objects, generating for a determined object, coordinates describing a bounding perimeter about the determined object;

rendering a second display output depicting graphical information in a form corresponding to the coordinates of the bounding perimeter for the determined one or more objects from the multiple encoded video frames; and

generating a composite display output, wherein the composite display output includes the imagery of the first display output overlaid with the graphical information of the second display output.

16. The non-transitory computer storage medium of claim 15 , wherein the first display output depicts imagery at a frame rate of 50 to 240 frames per second.

17. The non-transitory computer storage medium of claim 16 , wherein the second display output depicts the graphical information at a frame rate less than or equal to the frame rate of the first display output.

18. The non-transitory computer storage medium of claim 15 , further comprising the operations of:

generating a graphical indication around or about the one or more objects indicating the location of the identified objects in the video frame.

19. The non-transitory computer storage medium of claim 15 , further comprising the operations of:

determining an external environmental state of the video image capture device; and

performing the inference if the external environmental state is suitable to perform inferencing via the machine learning network.

20. The non-transitory computer storage medium of claim 15 , wherein the multiple encoded video frames are encoded in a color space format selected from the group consisting of NV12, I420, YV12, YUY2, YUYV, UYVY, UVYU, V308, IYU2, V408, RGB24, RGB32, V410, Y410 and Y42T.

21. The non-transitory computer storage medium of claim 15 , wherein the graphical information of the second display output includes graphical indications of the one or more objects disposed over a video display area of the first display output, and textual information corresponding to the one or more objects disposed over a non-video display area of the first display output.

Assignments (3)
CHANGE OF NAME Recorded May 27, 2025
From: SATISFAI HEALTH INC.
To: DOVA HEALTH INTELLIGENCE INC.
Reel/Frame 071230/0699 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2022
From: DOCBOT, INC.
To: SATISFAI HEALTH INC.
Reel/Frame 061793/0798 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2020
From: NINH, ANDREW; DAO, TYLER; FIDAALI, MOHAMMAD
To: DOCBOT, INC.
Reel/Frame 051664/0584 →
Cited By (3)
US 12,315,144 US 12,342,097 US 12,574,567