IP Library Granted Patent US 11,917,268
Granted Patent B2
US 11,917,268 · App. 17/571,602 · Granted Feb 27, 2024

Prediction model training via live stream concept association

Inventors: Matthew D. Zeiler (Fort Lee, NJ); Daniel Kantor (New York, NY)
Assignee: CLARIFAI, INC.
H04N21/84G06F18/2411G06F18/41G06V10/764G06V10/7788G06V10/82H04N21/2187H04N21/23109H04N21/23418H04N21/26603H04N21/278H04N21/4223H04N21/42202H04N21/4312H04N21/44008H04N21/466H04N21/4666H04N21/482H04N21/80H04N21/8133
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,917,268
App. No.
17/571,602
Granted
Feb 27, 2024
Kind
B2
Abstract

In certain embodiments, training of a neural network or other prediction model may be facilitated via live stream concept association. In some embodiments, a live video stream may be loaded on a user interface for presentation to a user. A user selection related to a frame of the live video stream may be received via the user interface during the presentation of the live video stream on the user interface, where the user selection indicates a presence of a concept in the frame of the live video stream. In response to the user selection related to the frame, an association of at least a portion of the frame of the live video stream and the concept may be generated, and the neural network or other prediction model may be trained based on the association of at least the portion of the frame with the concept.

Claims (62)

1. A system that comprises one or more processors programmed with computer program instructions that, when executed, cause the system to:

display, on a user interface, a video stream and a plurality of predicted concepts associated with the video stream, the plurality of predicted concepts generated by a prediction model based on the video stream, and the plurality of predicted concepts include a first set of predicted concepts and a second set of predicted concepts, the first set of predicted concepts being associated with first confidence values equal to or higher than a confidence threshold and the second set of predicted concepts being associated with second confidence values less than the confidence threshold, and the first and second confidence values indicate a confidence level that a particular predicted concept is present in a particular frame of the video stream;

receive, from an input device, a user selection of a concept related to a frame of the video stream, the selected concept being from among the plurality of predicted concepts;

receive, from the input device, a selection of a portion of the frame of the video stream;

determine an association between the selected portion of the frame and the selected concept; and

cause the prediction model to be trained based on the association.

2. The system of claim 1 further comprising instructions that, when executed, cause the system to:

determine a positive association between the selected portion of the frame and the selected concept, the positive association indicating the presence of the selected concept in the selected portion of the frame; and

store the positive association in a memory;

wherein the prediction model is trained based on the stored positive association.

3. The system of claim 1 further comprising instructions that, when executed, cause the system to:

determine a negative association between the selected portion of the frame and the selected concept, the negative association indicating the absence of the selected concept in the selected portion of the frame; and

store the negative association in a memory;

wherein the prediction model is trained based on the stored negative association.

4. The system of claim 2 , further comprising instructions that, when executed, cause the system to:

receive, from the input device, a user selection of a second concept related to a second frame of the video stream, the selection of the second concept being from among the plurality of predicted concepts;

receive, from the input device, a selection of a second portion of the second frame and the selected second concept.

5. The system of claim 4 , further comprising instructions that, when executed, cause the system to:

determine a negative association between the selected second portion of the second frame and the selected second concept, the negative association indicating the absence of the selected second concept in the selected second portion of the second frame; and

store the negative association in a memory;

wherein the prediction model is trained based on the stored negative association.

6. The system of claim 5 , wherein the selected concept and the selected second concept are the same concept from the plurality of predicted concepts.

7. The system of claim 1 , wherein receiving the user selection of the concept causes computer program instructions that, when executed, to further cause the system to:

determine a pressure level applied by the user on the user interface; and

determine a confidence value associated with the concept and the frame based on the determined pressure level,

wherein the prediction is trained based on the determined confidence value.

8. The system of claim 7 , wherein the confidence value is directly proportional to the pressure level.

9. The system of claim 1 , wherein the input device is one of a touch screen, a microphone, a keyboard, and a mouse.

10. The system of claim 1 , wherein the first set of predicted concepts and the second set of predicted concepts are presented via the user interface such that the first set of predicted concepts and the second set of predicted concepts are distinguished from each other.

11. A computer-implemented method, comprising:

displaying, on a user interface, a video stream and a plurality of predicted concepts associated with the video stream, the plurality of predicted concepts generated by a prediction model based on the video stream, and the plurality of predicted concepts include a first set of predicted concepts and a second set of predicted concepts, the first set of predicted concepts being associated with first confidence values equal to or higher than a confidence threshold and the second set of predicted concepts being associated with second confidence values less than the confidence threshold, and the first and second confidence values indicate a confidence level that a particular predicted concept is present in a particular frame of the video stream;

receiving, from an input device, a user selection of a concept related to a frame of the video stream, the selected concept being from among the plurality of predicted concepts;

receiving, from the input device, a selection of a portion of the frame of the video stream;

determining an association between the selected portion of the frame and the selected concept; and

causing the prediction model to be trained based on the association.

12. The computer-implemented method of claim 11 , further comprising:

determining a positive association between the selected portion of the frame and the selected concept, the positive association indicating the presence of the selected concept in the selected portion of the frame; and

storing the positive association in a memory;

wherein the prediction model is trained based on the stored positive association.

13. The computer-implemented method of claim 11 , further comprising:

determining a negative association between the selected portion of the frame and the selected concept, the negative association indicating the absence of the selected concept in the selected portion of the frame; and

storing the negative association in a memory;

wherein the prediction model is trained based on the stored negative association.

14. The computer-implemented method of claim 12 , further comprising:

receiving, from the input device, a user selection of a second concept related to a second frame of the video stream, the selection of the second concept being from among the plurality of predicted concepts;

receiving, from the input device, a selection of a second portion of the second frame and the selected second concept.

15. The computer-implemented method of claim 14 , further comprising:

determining a negative association between the selected second portion of the second frame and the selected second concept, the negative association indicating the absence of the selected second concept in the selected second portion of the second frame; and

storing the negative association in a memory;

wherein the prediction model is trained based on the stored negative association.

16. The computer-implemented method of claim 15 , wherein the selected concept and the selected second concept are the same concept from the plurality of predicted concepts.

17. The computer-implemented method of claim 11 , wherein receiving the user selection of the concept further comprises:

determining a pressure level applied by the user on the user interface; and

determining a confidence value associated with the concept and the frame based on the determined pressure level,

wherein the prediction is trained based on the determined confidence value.

18. A non-transitory computer-readable medium comprising instructions which when executed by a processor cause a computer to perform a method comprising:

displaying, on a user interface, a video stream and a plurality of predicted concepts associated with the video stream, the plurality of predicted concepts generated by a prediction model based on the video stream, and the plurality of predicted concepts include a first set of predicted concepts and a second set of predicted concepts, the first set of predicted concepts being associated with first confidence values equal to or higher than a confidence threshold and the second set of predicted concepts being associated with second confidence values less than the confidence threshold, and the first and second confidence values indicate a confidence level that a particular predicted concept is present in a particular frame of the video stream;

receiving, from an input device, a user selection of a concept related to a frame of the video stream, the selected concept being from among the plurality of predicted concepts;

receiving, from the input device, a selection of a portion of the frame of the video stream;

determining an association between the selected portion of the frame and the selected concept; and

causing the prediction model to be trained based on the association.

19. The non-transitory computer-readable medium of claim 18 , wherein the input device is one of a touch screen, a microphone, a keyboard, and a mouse.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2026
From: CLARIFAI, INC.
To: NEBIUS BV
Reel/Frame 075712/0109 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2026
From: ZEILER, MATTHEW D.; KANTOR, DANIEL
To: CLARIFAI, INC.
Reel/Frame 074809/0024 →
Continuity (5)
Continuation 17016457 · Sep 10, 2020
Continuation 15986239 · May 22, 2018
Continuation 15717114 · Sep 27, 2017
Provisional Application 62400538 · Sep 27, 2016
Related Publication 20220132222A1 · Apr 28, 2022