IP Library Granted Patent US 11,645,579
Granted Patent B2
US 11,645,579 · App. 16/723,623 · Granted May 9, 2023

Automated machine learning tagging and optimization of review procedures

Inventors: Miquel Angel Farré Guiu (Bern, CH); Monica Alfaro Vendrell (Barcelona, ES); Marc Junyent Martin (Barcelona, ES); Anthony M. Accardo (Los Angeles, CA)
Assignee: Disney Enterprises, Inc.
G06N20/00G06F18/2185G06V10/7784G06V20/41G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,645,579
App. No.
16/723,623
Granted
May 9, 2023
Kind
B2
Abstract

Techniques for machine learning optimization are provided. A video comprising a plurality of segments is received, and a first segment of the plurality of segments is processed with a machine learning (ML) model to generate a plurality of tags, where each of the plurality of tags indicates presence of an element in the first segment. A respective accuracy value is determined for each respective tag of the plurality of tags, where the respective accuracy value is based at least in part on a maturity score for the ML model. The first segment is classified as accurate, based on determining that an aggregate accuracy of tags corresponding to the first segment exceeds a predefined threshold. Upon classifying the first segment as accurate, the first segment is bypassed during a review process.

Claims (85)

1. A method, comprising:

receiving a video comprising a plurality of segments;

processing a first segment of the plurality of segments with a machine learning (ML) model to generate a plurality of tags, wherein each of the plurality of tags indicates presence of an element in the first segment;

determining, for each respective tag of the plurality of tags, a respective accuracy value, wherein the respective accuracy value is based at least in part on a maturity score generated for the ML model based on aggregating a plurality of model-specific scores for a plurality of versions of the ML model, wherein a respective weight assigned to each respective version of the plurality of versions is inversely proportional to a respective age of the respective version;

classifying the first segment as accurate, based on determining that an aggregate accuracy of tags corresponding to the first segment exceeds a predefined threshold; and

upon classifying the first segment as accurate, bypassing the first segment during a review process.

2. The method of claim 1 , wherein the review process comprises:

outputting a second segment from the plurality of segments via a graphical user interface (GUI);

outputting an indication of corresponding tags associated with the second segment;

upon receiving feedback on the corresponding tags, identifying a third segment from the plurality of segments via the GUI; and

upon determining that the third segment is classified as accurate:

bypassing the third segment; and

outputting a fourth segment from the plurality of segments.

3. The method of claim 2 , wherein bypassing the third segment is further based on:

outputting, via the GUI, an indication that the third segment is accurate; and

receiving input specifying to skip the third segment.

4. The method of claim 1 , wherein the plurality of tags include a plurality of unknown tags, wherein each unknown tag corresponds to an element that could not be identified by the ML model, the method further comprising:

grouping unknown tags of the plurality of unknown tags into one or more clusters based on similarity between the unknown tags.

5. The method of claim 4 , the method further comprising:

upon receiving an identification for a first unknown tag assigned to a first cluster of the one or more clusters, assigning the identification to each other unknown tag in the first cluster.

6. A method, comprising:

receiving a video comprising a plurality of segments;

processing a first segment of the plurality of segments with a machine learning (ML) model to generate a plurality of tags, wherein each of the plurality of tags indicates presence of an element in the first segment;

determining, for each respective tag of the plurality of tags, a respective accuracy value, wherein the respective accuracy value is based at least in part on a maturity score generated for the ML model based on how many times the ML model has correctly identified the element of the respective tag as being present in video segments previously processed with the ML model, compared to how many times the element has actually been present in the video segments previously processed with the ML model;

classifying the first segment as accurate, based on determining that an aggregate accuracy of tags corresponding to the first segment exceeds a predefined threshold; and

upon classifying the first segment as accurate, bypassing the first segment during a review process.

7. The method of claim 6 , wherein the maturity score comprises a plurality of element-specific scores, such that a first element is associated with a first element-specific score and a second element is associated with a second element-specific score.

8. A non-transitory computer-readable medium containing computer program code that, when executed by operation of one or more computer processors, performs an operation comprising:

receiving a video comprising a plurality of segments;

processing a first segment of the plurality of segments with a machine learning (ML) model to generate a plurality of tags, wherein each of the plurality of tags indicates presence of an element in the first segment;

determining, for each respective tag of the plurality of tags, a respective accuracy value, wherein the respective accuracy value is based at least in part on a maturity score generated for the ML model based on aggregating a plurality of model-specific scores for a plurality of versions of the ML model, wherein a respective weight assigned to each respective version of the plurality of versions is inversely proportional to a respective age of the respective version;

classifying the first segment as accurate, based on determining that an aggregate accuracy of tags corresponding to the first segment exceeds a predefined threshold; and

upon classifying the first segment as accurate, bypassing the first segment during a review process.

9. The computer-readable medium of claim 8 , wherein the review process comprises:

outputting a second segment from the plurality of segments via a graphical user interface (GUI);

outputting an indication of corresponding tags associated with the second segment;

upon receiving feedback on the corresponding tags, identifying a third segment from the plurality of segments via the GUI; and

upon determining that the third segment is classified as accurate:

bypassing the third segment; and

outputting a fourth segment from the plurality of segments.

10. The computer-readable medium of claim 9 , wherein bypassing the third segment is further based on:

outputting, via the GUI, an indication that the third segment is accurate; and

receiving input specifying to skip the third segment.

11. The computer-readable medium of claim 8 , wherein the plurality of tags include a plurality of unknown tags, wherein each unknown tag corresponds to an element that could not be identified by the ML model, the operation further comprising:

grouping unknown tags of the plurality of unknown tags into one or more clusters based on similarity between the unknown tags.

12. The computer-readable medium of claim 11 , the operation further comprising:

upon receiving an identification for a first unknown tag assigned to a first cluster of the one or more clusters, assigning the identification to each other unknown tag in the first cluster.

13. A system, comprising:

one or more computer processors; and

a memory containing a program which when executed by the one or more computer processors performs an operation, the operation comprising:

receiving a video comprising a plurality of segments;

processing a first segment of the plurality of segments with a machine learning (ML) model to generate a plurality of tags, wherein each of the plurality of tags indicates presence of an element in the first segment;

determining, for each respective tag of the plurality of tags, a respective accuracy value, wherein the respective accuracy value is based at least in part on a maturity score generated for the ML model based on aggregating a plurality of model-specific scores for a plurality of versions of the ML model, wherein a respective weight assigned to each respective version of the plurality of versions is inversely proportional to a respective age of the respective version;

classifying the first segment as accurate, based on determining that an aggregate accuracy of tags corresponding to the first segment exceeds a predefined threshold; and

upon classifying the first segment as accurate, bypassing the first segment during a review process.

14. The system of claim 13 , wherein the review process comprises:

outputting a second segment from the plurality of segments via a graphical user interface (GUI);

outputting an indication of corresponding tags associated with the second segment;

upon receiving feedback on the corresponding tags, identifying a third segment from the plurality of segments via the GUI; and

upon determining that the third segment is classified as accurate:

bypassing the third segment; and

outputting a fourth segment from the plurality of segments.

15. The system of claim 14 , wherein bypassing the third segment is further based on:

outputting, via the GUI, an indication that the third segment is accurate; and

receiving input specifying to skip the third segment.

16. The system of claim 13 , wherein the plurality of tags include a plurality of unknown tags, wherein each unknown tag corresponds to an element that could not be identified by the ML model, the operation further comprising:

grouping unknown tags of the plurality of unknown tags into one or more clusters based on similarity between the unknown tags.

17. The system of claim 16 , the operation further comprising:

upon receiving an identification for a first unknown tag assigned to a first cluster of the one or more clusters, assigning the identification to each other unknown tag in the first cluster.

18. A non-transitory computer-readable medium containing computer program code that, when executed by operation of one or more computer processors, performs an operation comprising:

receiving a video comprising a plurality of segments;

processing a first segment of the plurality of segments with a machine learning (ML) model to generate a plurality of tags, wherein each of the plurality of tags indicates presence of an element in the first segment;

determining, for each respective tag of the plurality of tags, a respective accuracy value, wherein the respective accuracy value is based at least in part on a maturity score generated for the ML model based on how many times the ML model has correctly identified the element of the respective tag as being present in video segments previously processed with the ML model, compared to how many times the element has actually been present in the video segments previously processed with the ML model;

classifying the first segment as accurate, based on determining that an aggregate accuracy of tags corresponding to the first segment exceeds a predefined threshold; and

upon classifying the first segment as accurate, bypassing the first segment during a review process.

19. The non-transitory computer-readable medium containing computer program code of claim 18 , wherein the maturity score comprises a plurality of element-specific scores, such that a first element is associated with a first element-specific score and a second element is associated with a second element-specific score.

20. A system, comprising:

one or more computer processors; and

a memory containing a program which when executed by the one or more computer processors performs an operation, the operation comprising:

receiving a video comprising a plurality of segments;

processing a first segment of the plurality of segments with a machine learning (ML) model to generate a plurality of tags, wherein each of the plurality of tags indicates presence of an element in the first segment;

determining, for each respective tag of the plurality of tags, a respective accuracy value, wherein the respective accuracy value is based at least in part on a maturity score generated for the ML model based on how many times the ML model has correctly identified the element of the respective tag as being present in video segments previously processed with the ML model, compared to how many times the element has actually been present in the video segments previously processed with the ML model;

classifying the first segment as accurate, based on determining that an aggregate accuracy of tags corresponding to the first segment exceeds a predefined threshold; and

upon classifying the first segment as accurate, bypassing the first segment during a review process.

21. The system of claim 20 , wherein the maturity score comprises a plurality of element-specific scores, such that a first element is associated with a first element-specific score and a second element is associated with a second element-specific score.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: FARRÉ GUIU, MIQUEL ANGEL; ALFARO VENDRELL, MONICA; JUNYENT MARTIN, MARC
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 051348/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 051348/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: ACCARDO, ANTHONY M.
To: DISNEY ENTERPRISES, INC.
Reel/Frame 051348/0919 →
Continuity (1)
Related Publication 20210192385A1 · Jun 24, 2021
Cited By (1)
US 12,639,816