IP Library Granted Patent US 11,245,968
Granted Patent B2
US 11,245,968 · App. 17/016,457 · Granted Feb 8, 2022

Prediction model training via live stream concept association

Inventors: Matthew D. Zeiler (New York, NY); Daniel Kantor (New York, NY)
Assignee: CLARIFAI, INC.
H04N21/84G06K9/6254G06K9/6269H04N21/2187H04N21/23109H04N21/23418H04N21/26603H04N21/278H04N21/4223H04N21/42202H04N21/4312H04N21/44008H04N21/466H04N21/4666H04N21/482H04N21/80H04N21/8133
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,245,968
App. No.
17/016,457
Granted
Feb 8, 2022
Kind
B2
Abstract

In certain embodiments, training of a neural network or other prediction model may be facilitated via live stream concept association. In some embodiments, a live video stream may be loaded on a user interface for presentation to a user. A user selection related to a frame of the live video stream may be received via the user interface during the presentation of the live video stream on the user interface, where the user selection indicates a presence of a concept in the frame of the live video stream. In response to the user selection related to the frame, an association of at least a portion of the frame of the live video stream and the concept may be generated, and the neural network or other prediction model may be trained based on the association of at least the portion of the frame with the concept.

Claims (69)

1. A system for training a prediction model via association of a concept with a video stream, the system comprising:

a computer system that comprises one or more processors programmed with computer program instructions that, when executed, cause the computer system to:

obtain a video stream;

process, via a prediction model, the video stream to generate a plurality of predicted concepts relating to the video stream;

cause the plurality of predicted concepts to be presented via a user interface during presentation of the video stream via the user interface;

obtain a selection of a concept related to a frame of the video stream, the selection of the concept being from among the plurality of predicted concepts and including moving the concept over a portion of the frame of the video stream presented via the user interface;

determine an association between the portion of the frame of the video stream and the concept; and

train the prediction model based on the association between the portion of the frame of the video stream and the concept.

2. The system of claim 1 , wherein the computer system is caused to:

determine a pressure level applied by a user on the user interface during the selection related to the frame of the video stream;

determine a confidence value for a presence of the concept in the frame of the video stream based on the determined pressure level; and

train the prediction model based on the determined confidence value.

3. The system of claim 2 , wherein the confidence value is directly proportional to the pressure level.

4. The system of claim 1 , wherein the selection is a user selection and the user selection is based on at least one of a voice instruction from a user, a visual instruction from the user, or a touch instruction from the user.

5. The system of claim 1 , wherein the computer system is caused to:

obtain metadata related to another frame of the video stream, the metadata describing another concept in the another frame of the video stream;

determine another association between the another frame and the metadata; and

train the prediction model based on the another association between the another frame and the metadata.

6. The system of claim 1 ,

wherein the plurality of predicted concepts include a first set of predicted concepts and a second set of predicted concepts,

wherein the first set of predicted concepts are associated with first confidence values equal to or higher than a confidence threshold and the second set of predicted concepts are associated with second confidence values less than the confidence threshold, and

wherein the first and second confidence values indicate a confidence level that a particular predicted concept is present in a particular frame of the video stream.

7. The system of claim 6 , wherein the first set of predicted concepts and the second set of predicted concepts are presented via the user interface such that the first set of predicted concepts and the second set of predicted concepts are distinguished from each other.

8. A method comprising:

obtaining a video stream;

processing, via a prediction model, the video stream to generate a plurality of predicted concepts relating to the video stream;

causing the plurality of predicted concepts to be presented via a user interface during presentation of the video stream via the user interface;

obtaining a selection of a concept related to a frame of the video stream, the selection of the concept being from among the plurality of predicted concepts and including moving the concept over a portion of the frame of the video stream presented via the user interface;

determining an association between the portion of the frame of the video stream and the concept; and

training the prediction model based on the association between the portion of the frame of the video stream and the concept.

9. The method of claim 8 , further comprising:

determining a pressure level applied by a user on the user interface during the selection related to the frame of the video stream;

determining a confidence value for a presence of the concept in the frame of the video stream based on the determined pressure level; and

training the prediction model based on the determined confidence value.

10. The method of claim 9 , wherein the confidence value is directly proportional to the pressure level.

11. The method of claim 8 , wherein the selection is a user selection and the user selection is based on at least one of a voice instruction from a user, a visual instruction from the user, or a touch instruction from the user.

12. The method of claim 8 , further comprising:

obtaining metadata related to another frame of the video stream, the metadata describing another concept in the another frame of the video stream;

determining another association between the another frame and the metadata; and

training the prediction model based on the another association between the another frame

and the metadata.

13. The method of claim 8 ,

wherein the plurality of predicted concepts include a first set of predicted concepts and a second set of predicted concepts,

wherein the first set of predicted concepts are associated with first confidence values equal to or higher than a confidence threshold and the second set of predicted concepts are associated with second confidence values less than the confidence threshold, and

wherein the first and second confidence values indicate a confidence level that a particular predicted concept is present in a particular frame of the video stream.

14. The method of claim 13 , wherein the first set of predicted concepts and the second set of predicted concepts are presented via the user interface such that the first set of predicted concepts and the second set of predicted concepts are distinguished from each other.

15. One or more non-transitory, computer-readable media storing instructions that, when executed by one or more processors, effectuate operations comprising:

obtaining a video stream;

processing, via a prediction model, the video stream to generate a plurality of predicted concepts relating to the video stream;

causing the plurality of predicted concepts to be presented via a user interface during presentation of the video stream via the user interface;

obtaining a selection of a concept related to a frame of the video stream, the selection of the concept being from among the plurality of predicted concepts and including moving the concept over a portion of the frame of the video stream presented via the user interface;

determining an association between the portion of the frame of the video stream and the concept; and

training the prediction model based on the association between the portion of the frame of the video stream and the concept.

16. The non-transitory, computer-readable media of claim 15 , further comprising:

determining a pressure level applied by a user on the user interface during the selection related to the frame of the video stream;

determining a confidence value for a presence of the concept in the frame of the video stream based on the determined pressure level; and

training the prediction model based on the determined confidence value.

17. The non-transitory, computer-readable media of claim 16 , wherein the confidence value is directly proportional to the pressure level.

18. The non-transitory, computer-readable media of claim 15 , wherein the selection is a user selection and the user selection is based on at least one of a voice instruction from a user, a visual instruction from the user, or a touch instruction from the user.

19. The non-transitory, computer-readable media of claim 15 , further comprising: obtaining metadata related to another frame of the video stream, the metadata describing

another concept in the another frame of the video stream;

determining another association between the another frame and the metadata; and

training the prediction model based on the another association between the another frame

and the metadata.

20. The non-transitory, computer-readable media of claim 15 ,

wherein the plurality of predicted concepts include a first set of predicted concepts and a second set of predicted concepts,

wherein the first set of predicted concepts are associated with first confidence values equal to or higher than a confidence threshold and the second set of predicted concepts are associated with second confidence values less than the confidence threshold,

wherein the first and second confidence values indicate a confidence level that a particular predicted concept is present in a particular frame of the video stream, and

wherein the first set of predicted concepts and the second set of predicted concepts are presented via the user interface such that the first set of predicted concepts and the second set of predicted concepts are distinguished from each other.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2026
From: CLARIFAI, INC.
To: NEBIUS BV
Reel/Frame 075712/0109 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2026
From: ZEILER, MATTHEW D.; KANTOR, DANIEL
To: CLARIFAI, INC.
Reel/Frame 074806/0350 →
Continuity (4)
Continuation 15986239 · May 22, 2018
Continuation 15717114 · Sep 27, 2017
Provisional Application 62400538 · Sep 27, 2016
Related Publication 20200413156A1 · Dec 31, 2020