IP Library Granted Patent US 10,921,957
Granted Patent B1
US 10,921,957 · App. 16/251,317 · Granted Feb 16, 2021

User interface for context labeling of multimedia items

Inventors: Matthew D. Zeiler (New York, NY); Adam L. Berenzweig (New York, NY)
Assignee: Clarifai, Inc.
G06F3/0482G06F3/04842G06F7/08G06F16/44G06F40/166G06F40/35G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,921,957
App. No.
16/251,317
Granted
Feb 16, 2021
Kind
B1
Abstract

In certain embodiments, a neural network may be trained to associated context information with multimedia items. In some embodiments, context predictions for multimedia items may be obtained via a neural network. A first multimedia item and a first task related to a first context prediction for the first multimedia item may be presented on a user interface. A user response to the first task may be obtained via the user interface. Based on the user response to the first task, prediction feedback related to the first context prediction or the first multimedia item may be provided to the neural network to cause the neural network to be updated based on the prediction feedback.

Claims (56)

1. A method being implemented by one or more processors executing computer program instructions that, when executed, perform the method, the method comprising:

obtaining a label that is predicted by a prediction model to be a corresponding label for multimedia items;

assigning the multimedia items to a group based on the prediction model predicting the label as a corresponding label for the multimedia items;

generating, based on the predicted label, a task related to the predicted label, the task soliciting a user to indicate one or more of the multimedia items that are relevant to the predicted label;

causing the multimedia items and the task to be presented together on a user interface at a same time based on the assignment of the multimedia items to the group;

obtaining, via the user interface, a user response to the task, the user response comprising (i) a first user indication of one of the multimedia items as being relevant to the predicted label and (ii) a second user indication of another one of the multimedia items as being relevant or not relevant to the predicted label; and

providing, to the prediction model, the first and second user indications to cause the prediction model to be updated based on the first and second user indications.

2. The method of claim 1 , further comprising:

providing, to one or more other prediction models, the first and second user indications to cause the one or more other prediction models to be updated based on the first and second user indications.

3. The method of claim 1 , further comprising:

obtaining the user response to the task from a user;

obtaining one or more other user responses to the task from one or more other users; and

assigning the task to one or more additional users based on the user response and the one or more other user responses collectively not satisfying a consensus threshold.

4. The method of claim 3 , further comprising:

obtaining one or more additional user responses to the task from the one or more additional users; and

removing the task from a set of pending tasks based on the user response, the one or more other user responses, and the one or more additional user responses collectively satisfying the consensus threshold.

5. The method of claim 1 , wherein each of the multimedia items comprises at least one of an image, an audio, or a video, and the prediction model comprises a machine learning model.

6. The method of claim 1 , wherein the prediction model comprises a neural network.

7. A system comprising:

a computer system comprising one or more processors programmed with computer program instructions that, when executed, cause the computer system to:

obtain a label that is predicted by a prediction model to be a corresponding label for multimedia items;

assign the multimedia items to a group based on the prediction model predicting the label as a corresponding label for the multimedia items;

generate, based on the predicted label, a task related to the predicted label, the task soliciting a user to indicate one or more of the multimedia items that are relevant to the predicted label;

cause the multimedia items and the task to be presented together on a user interface at a same time based on the assignment of the multimedia items to the group;

obtain, via the user interface, a user response to the task, the user response comprising (i) a first user indication of one of the multimedia items as being relevant to the predicted label and (ii) a second user indication of another one of the multimedia items as being relevant or not relevant to the predicted label; and

provide, to the prediction model, the first and second user indications to cause the prediction model to be updated based on the first and second user indications.

8. The system of claim 7 , wherein the computer system is caused to:

provide, to one or more other prediction models, the first and second user indications to cause the one or more other prediction models to be updated based on the first and second user indications.

9. The system of claim 7 , wherein the computer system is caused to:

obtain the user response to the task from a user;

obtain one or more other user responses to the task from one or more other users; and

assign the task to one or more additional users based on the user response and the one or more other user responses collectively not satisfying a consensus threshold.

10. The system of claim 9 , wherein the computer system is caused to:

obtain one or more additional user responses to the task from the one or more additional users; and

remove the task from a set of pending tasks based on the user response, the one or more other user responses, and the one or more additional user responses collectively satisfying the consensus threshold.

11. The system of claim 7 , wherein each of the multimedia items comprises at least one of an image, an audio, or a video, and the prediction model comprises a machine learning model.

12. The system of claim 7 , wherein the prediction model comprises a neural network.

13. A non-transitory computer-readable media comprising instructions that, when executed by one or more processors, cause operations comprising:

obtaining a label that is predicted by a prediction model to be a corresponding label for multimedia items;

assigning the multimedia items to a group based on the prediction model predicting the label as a corresponding label for the multimedia items;

generating, based on the predicted label, a task related to the predicted label, the task soliciting a user to indicate one or more of the multimedia items that are relevant to the predicted label;

causing the multimedia items and the task to be presented together on a user interface at a same time based on the assignment of the multimedia items to the group;

obtaining, via the user interface, a user response to the task, the user response comprising one or more user indications related to the predicted label; and

providing, to the prediction model, the one or more user indications to cause the prediction model to be updated based on the one or more user indications.

14. The media of claim 13 , the operations further comprising:

providing, to one or more other prediction models, the one or more user indications to cause the one or more other prediction models to be updated based on the one or more user indications.

15. The media of claim 13 , the operations further comprising:

obtaining the user response to the task from a user;

obtaining one or more other user responses to the task from one or more other users; and

assigning the task to one or more additional users based on the user response and the one or more other user responses collectively not satisfying a consensus threshold.

16. The media of claim 15 , the operations further comprising:

obtaining one or more additional user responses to the task from the one or more additional users; and

removing the task from a set of pending tasks based on the user response, the one or more other user responses, and the one or more additional user responses collectively satisfying the consensus threshold.

17. The media of claim 13 , wherein the one or more user indications comprises (i) a first user indication of one of the multimedia items as being relevant to the predicted label and (ii) a second user indication of another one of the multimedia items as being relevant or not relevant to the predicted label.

18. The media of claim 13 , wherein each of the multimedia items comprises at least one of an image, an audio, or a video, and the prediction model comprises a machine learning model.

19. The media of claim 13 , wherein the prediction model comprises a neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2026
From: CLARIFAI, INC.
To: NEBIUS BV
Reel/Frame 075712/0109 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2019
From: ZEILER, MATTHEW D.; BERENZWEIG, ADAM L.
To: CLARIFAI, INC.
Reel/Frame 048056/0459 →
Continuity (2)
Continuation 15002248 · Jan 20, 2016
Provisional Application 62106648 · Jan 22, 2015