IP Library Granted Patent US 11,017,780
Granted Patent B2
US 11,017,780 · App. 16/287,892 · Granted May 25, 2021

System and methods for neural network orchestration

Inventors: Chad Steelberg (Newport Beach, CA); Peter Nguyen (Costa Mesa, CA); David Kettler (Bellevue, WA); Karl Schwamb (Mission Viejo, CA); Yu Zhao (Irvine, CA)
Assignee: VERITONE, INC.
G10L15/32G06F40/20G06F40/253G06F40/284G06F40/30G06N3/0445G06N3/0454G06N3/08G06N5/003G06N20/00G10L15/02G10L15/16G10L15/1815G10L15/30G06N3/0472
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,017,780
App. No.
16/287,892
Granted
May 25, 2021
Kind
B2
Abstract

Methods and systems for classifying a multimedia file using interclass data is disclosed. One of the methods can use classification results from one or more engines of different classes to select a different engine for the original classification task. For example, given an audio segment with associated metadata and image data, the disclosed interclass method can use the classification results from a topic classification of metadata and/or an image classification result of the image data as inputs for selecting a new transcription engine to transcribe the audio segment.

Claims (27)

1. A method for classifying a first media segment of a first data type having a corresponding media segment of a second data type, the method comprising:

extracting a first set of media features of the first media segment of the first data type;

generating, using an engine prediction neural network, a best candidate neural network based at least on the first set of media features, wherein the best candidate neural network comprises a neural network having a highest predicted value of accuracy;

determining whether a predicted value of accuracy of the best candidate neural network is above a predetermined accuracy threshold;

when the predicted value of accuracy of the best candidate neural network is below the predetermined accuracy threshold, classifying the corresponding media segment of a second data type using a second classification neural network; and

selecting, based at least on results of the classification of the corresponding media segment of a second data type, a third classification neural network to classify the first media segment of the first data type, wherein the first and second data types are different, and wherein the third classification neural network and the best candidate neural network are different.

2. The method of claim 1 , wherein extracting the first set of media features of the first data type comprises extracting audio features of the first media segment using outputs of one or more layers of a speech-to-text classification neural network, wherein the first data type comprises audio data and the corresponding media segment of second data type comprises image data, transcription data, or metadata.

3. The method of claim 2 , wherein the corresponding media segment spans a duration of 30 seconds before and after from when the first media segment occurs within a multimedia file.

4. The method of claim 1 , wherein extracting the first set of media features of the first data type comprises extracting image features of the first media segment using outputs of one or more layers of an image classification neural network, wherein the first data type comprises image data and the corresponding media segment of second data type comprises audio data, transcription data, or metadata.

5. The method of claim 1 , further classifying the corresponding media segment of a second data type using the second classification neural network comprising:

extracting a second set of media features of the corresponding media segment of the second data type; and

generating, using the engine prediction neural network, a best candidate neural network based at least on the second set of media features, wherein the second classification neural network comprises the best candidate neural network.

6. The method of claim 1 , wherein extracting the first set of media features of the first data type comprises extracting the first set of media features of the first media segment using outputs of one or more hidden layers of a fourth classification neural network trained to classify data of the first data type, wherein the first data type comprises audio data, image data, transcription data, or metadata.

7. The method of claim 6 , wherein the engine prediction neural network is trained to associate outputs of the one or more hidden layers of the fourth classification neural network with predicted classification performances of a plurality of neural network.

8. The method of claim 6 , wherein the first set of media features is extracted from a last hidden layer of the fourth classification neural network.

9. The method of claim 6 , wherein the first set of media features is extracted from a first and last hidden layer of the fourth classification neural network.

10. The method of claim 1 , wherein selecting the third image classification neural network comprises:

determining context of the corresponding media segment based at least on the classification result of the corresponding media segment received from the second classification neural network; and

selecting the third classification neural network based on the determined context.

11. A system for classifying a first media segment of a first data type having a corresponding media segment of a second data type, the system comprising:

a memory;

one or more processors coupled to the memory, the one or more processors configured to:

extract a first set of media features of the first media segment of the first data type;

generate, using an engine prediction neural network, a best candidate neural network based at least on the first set of media features, wherein the best candidate neural network comprises a neural network having a highest predicted value of accuracy;

determine whether a predicted value of accuracy of the best candidate neural network is above a predetermined accuracy threshold;

when the predicted value of accuracy of the best candidate neural network is below the predetermined accuracy threshold, classify the corresponding media segment of a second data type using a second classification neural network; and

select, based at least on results of the classification of the corresponding media segment of a second data type, a third classification neural network to classify the first media segment of the first data type, wherein the first and second data types are different, and wherein the third classification neural network and the best candidate neural network are different.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 19, 2025
From: WILMINGTON SAVINGS FUND SOCIETY, FSB, AS COLLATERAL AGENT
To: VERITONE, INC.
Reel/Frame 073634/0333 →
SECURITY INTEREST Recorded Dec 13, 2023
From: VERITONE, INC.
To: WILMINGTON SAVINGS FUND SOCIETY, FSB, AS COLLATERAL AGENT
Reel/Frame 066140/0513 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2020
From: STEELBERG, CHAD; NGUYEN, PETER; KETTLER, DAVID; SCHWAMB, KARL; ZHAO, YU
To: VERITONE, INC.
Reel/Frame 051957/0449 →
Continuity (9)
Continuation In Part 16243033 · Jan 8, 2019
Continuation In Part 16109516 · Aug 22, 2018
Continuation 16052459 · Aug 1, 2018
Provisional Application 62713937 · Aug 2, 2018
Provisional Application 62735769 · Sep 24, 2018
Provisional Application 62638745 · Mar 5, 2018
Provisional Application 62633023 · Feb 20, 2018
Provisional Application 62540508 · Aug 2, 2017
Related Publication 20200066278A1 · Feb 27, 2020
Cited By (5)
US 12,255,749 US 12,339,960 US 12,405,934 US 12,499,314 US 12,505,296