IP Library Granted Patent US 9,183,296
Granted Patent B1
US 9,183,296 · App. 14/461,076 · Granted Nov 10, 2015

Large scale video event classification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,183,296
App. No.
14/461,076
Granted
Nov 10, 2015
Kind
B1
Abstract

Systems and methods are provided herein relating to video classification. A text mining component is disclosed that automatically generates a plurality of video event categories. Part-of-Speech (POS) analysis can be applied to video titles and descriptions, further using a lexical hierarchy to filter potential classifications. Classification performance can be further improved by extracting content-based features from a video sample. Using the content based features a set of classifier scores can be generated. A hyper classifier can use both the classifier scores and the content-based features of the video to classify the video sample.

Claims (32)

1. A non-transitory computer-readable storage medium storing computer-executable instructions that, in response to execution, cause a device comprising a processor to perform operations, comprising:

text mining metadata of a word combination that is at least one of a noun-verb combination or a verb-noun combination and that is associated with a video sample for the word combination;

storing in the memory the at least one of the noun-verb combination or the verb-noun combination a video category label among a set of video category labels; and

filtering the set of video category labels containing the at least one of the noun-verb combination or the verb-noun combination stored in the memory based upon a lexical hierarchy.

2. The non-transitory computer-readable storage medium of claim 1 , wherein a noun and a verb of the at least one noun-verb combination or verb-noun combination are non-adjacent to each other in the at least one description or title.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the lexical hierarchy constrains a noun of the at least one of the noun-verb combination or the verb-noun combination to be within a hierarchy of physical entity.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the lexical hierarchy constrains a verb of the at least one of the noun-verb combination or the verb-noun combination to be within hierarchies of act, move; act, human action, human activity; or happening, occurrence, natural event.

5. The non-transitory computer-readable storage medium of claim 1 , wherein the operations further comprise:

extracting a plurality of features from the video sample.

6. The non-transitory computer-readable storage medium of claim 5 , wherein the plurality of features comprise at least one of a histogram of local features, a color histogram, edge features, a histogram of textons, face features, color motion, shot boundary features or audio features.

7. The non-transitory computer-readable storage medium of claim 5 , wherein the operations further comprise:

employing a plurality of models based on the plurality of features to generate a plurality of classification scores; and

associating the video sample with the video category label based on the plurality of features and the plurality of classification scores.

8. The non-transitory computer-readable storage medium of claim 7 , wherein the associating the video sample with the video category label is performed employing a hyper classifier.

9. The non-transitory computer-readable storage medium of claim 7 , wherein a model of the plurality of models is employed based on a determination that an accuracy score of the model is greater than a defined threshold.

10. A video classification method, comprising:

text mining, by a device comprising a processor, metadata of a word combination that is a noun-verb combination or a verb-noun combination and that is associated with a video sample for the word combination;

storing in a memory of the device, the at least one of the noun-verb combination or the verb-noun combination as a video category label among a set of category labels; and

filtering the set of video category labels containing the at least one of the noun-verb combination or the verb-noun combination stored in the memory based upon a lexical hierarchy.

11. The method of claim 10 , wherein a noun and a verb of the at least one noun-verb combination or verb-noun combination are non-adjacent to each other in the at least one description or title.

12. The method of claim 10 , wherein the lexical hierarchy constrains a noun of the at least one of the noun-verb combination or the verb-noun combination to be within a hierarchy of physical entity.

13. The method of claim 10 , wherein the lexical hierarchy constrains a verb of the at least one of the noun-verb combination or the verb-noun combination to be within hierarchies of act, move; act, human action, human activity; or happening, occurrence, natural event.

14. The method of claim 10 , further comprising:

extracting a plurality of features from the video sample.

15. The method of claim 14 , wherein the plurality of features comprise at least one of a histogram of local features, a color histogram, edge features, a histogram of textons, face features, color motion, shot boundary features or audio features.

16. The method of claim 14 , further comprising:

generating a plurality of classification scores based on the plurality of features determined employing a plurality of models; and

associating the video sample with the video category label based on the plurality of features and the plurality of classification scores.

17. The method of claim 16 , wherein the associating the video sample with the video category label comprises using a hyper classifier that is automatically trained based on the plurality of features and the plurality of classification scores.

18. The method of claim 16 , wherein the employing the plurality of models comprises using a plurality of binary classifiers, wherein one or more of the plurality of binary classifiers is trained based upon an associated video category label.

19. The method of claim 18 , wherein the binary classifier is a classifier.

20. The method of claim 18 , wherein one or more of the plurality of models is employed based on an accuracy score of the one or more of the plurality of models being greater than a defined threshold.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044334/0466 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2014
From: SONG, YANG; ZHAO, MING; NI, BINGBING
To: GOOGLE INC.
Reel/Frame 033548/0075 →