IP Library › Granted Patent US 8,682,654
Granted Patent B2
US 8,682,654 · App. 11/411,016 · Granted Mar 25, 2014

Systems and methods for classifying sports video

Inventors: Ming-Jun Chen (Tai Nan, TW); Jiun-Fu Chen (Hemei Township, Changhua County, TW); Shih-Min Tang (Jiali Township, Tainan County, TW); Ho-Chao Huang (Shindian, TW)
Assignee: Cyberlink Corp.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,682,654
App. No.
11/411,016
Granted
Mar 25, 2014
Kind
B2
Abstract

Disclosed are systems, methods, and computer readable media having programs for classifying sports video. In one embodiment, a method includes: extracting, from an audio stream of a video clip, a plurality of key audio components contained therein; and classifying, using at least one of the plurality of key audio components, a sport type contained in the video clip. In one embodiment, a computer readable medium having a computer program for classifying ports video includes: logic configured to extract a plurality of key audio components from a video clip; and logic configured to classify a sport type corresponding to the video clip.

Claims (35)

1. A system for classifying sports video, comprising:

a processor;

a memory storing instructions that when execute cause the system to:

collect a plurality of key audio samples from a plurality of types of sports;

extract a plurality of sample audio features from a plurality of frames within each of the plurality of key audio samples;

generate a plurality of patterns corresponding to the plurality of key audio samples;

extract a plurality of audio features from the plurality of frames within an audio stream of a video clip;

compare the plurality of sample audio features in the plurality of patterns with the a plurality of audio features extracted from the audio stream; and

classify the video clip into a sport type from a plurality of sport types based on the location, the distribution, and the frequency of the key audio components, wherein the classification based on the distribution compares a first distribution of the plurality of key audio components with a second distribution of key audio components which corresponds to one of the plurality of sport types, wherein the first and second distributions are related to a regularity of the occurrences of the respective key audio components.

2. The system of claim 1 , further comprising means for editing the video clip by using the location and distribution of the plurality of key audio components within the video clip.

3. The system of claim 1 , wherein a portion of the plurality of audio features are selected from the group consisting of: mel-frequency cepstrum coefficients, pitch, LPC coefficients, LSP coefficients, audio energy, zero-crossing-rate, and noise frame ratio.

4. A method for classifying sports video, comprising:

extracting, using an instruction execution system, a plurality of key audio components from an audio stream of a video clip; and

classifying, using the instruction execution system, the video clip into a sport type from a plurality of sport types, based on a comparison of a first distribution of one or more of the plurality of key audio components extracted from the audio stream as compared to a second distribution of key audio components which corresponds to one of the plurality of sport types, wherein the first and second distributions are related to the regularity of the occurrences of the respective key audio components.

5. The method of claim 4 , further comprising generating, using the instruction execution system, a plurality of patterns corresponding to a plurality of key audio components.

6. The method of claim 5 , wherein the generating further comprises collecting the plurality of key audio components that correspond to a plurality of sports.

7. The method of claim 6 , wherein the generating further comprises determining a plurality of features for each of the plurality of key audio components.

8. The method of claim 7 , wherein a portion of the plurality of features are selected from the group consisting of: mel-frequency cepstrum coefficients, pitch, LPC coefficients, LSP coefficients, audio energy, zero-crossing-rate, and noise frame ratio.

9. The method of claim 7 , wherein the generating further comprises determining a plurality of patterns of the plurality of features, wherein each pattern of the plurality of patterns corresponds to one of the plurality of key audio components.

10. The method of claim 9 , wherein the generating further comprises training a model for each of the plurality of key audio components using the corresponding pattern of features.

11. The method of claim 4 , wherein the extracting comprises determining a plurality of features from an audio stream of a video clip, wherein each of the features corresponds to one of a plurality of frames in the video clip.

12. The method of claim 11 , wherein a portion of the plurality of features are selected from the group consisting of: mel-frequency cepstrum coefficients, pitch, LPC coefficients, LSP coefficients, audio energy, zero-crossing rate, and noise frame ratio.

13. The method of claim 4 , wherein the extracting comprises comparing the plurality of features in the audio stream with the plurality of patterns.

14. The method of claim 13 , wherein the comparing comprises determining which of the plurality of key audio components is present in the video clip.

15. The method of claim 13 , wherein the comparing comprises determining a distribution of the plurality of key audio components that occur in the video clip.

16. The method of claim 4 , wherein an event that generates one of the plurality of key audio components is selected from the group consisting of: a ball strike, a whistle sound, a car engine sound, a water splash, a referee sound, a commentator sound, and a bell ringing.

17. A non-transitory computer readable medium storing instructions for classifying sports video, wherein the instructions when executed by an instruction execution system cause the instruction execution system to at least:

extract a plurality of key audio components from a video clip; and

classify the video clip into a sport type in a plurality of sport types, based on a comparison of a first distribution of one or more of the plurality of key audio components extracted from the audio stream as compared to a second distribution of key audio components which corresponds to one of a plurality of sport types, wherein the first and second distributions are related to the regularity of the occurrences of the respective key audio components.

18. The non-transitory computer readable medium of claim 17 , wherein the instructions when executed by the instruction execution system further cause the instruction execution system to at least generate a plurality of key audio patterns corresponding to a plurality of key audio sample components.

19. The non-transitory computer readable medium of claim 17 , wherein the instructions when executed by the instruction execution system further cause the instruction execution system to at least determine a plurality of audio characteristics from the video clip.

20. The non-transitory computer readable medium of claim 19 , wherein a portion of the plurality of audio characteristics are selected from the group consisting of: mel-frequency cepstrum coefficients, pitch, LPC coefficients, LSP coefficients, audio energy, zero-crossing-rate, and noise frame ratio.

21. The non-transitory computer readable medium of claim 17 , wherein the instructions when executed by the instruction execution system further cause the instruction execution system to at least compare the plurality of audio characteristics within the video clip to the plurality of key audio patterns.

22. The non-transitory computer readable medium of claim 17 , wherein the instructions to classify the sport type further cause the instruction execution system to determine the sport type based on which of the key audio components is in the video clip.

23. The non-transitory computer readable medium of claim 17 , wherein the instructions to classify the sport type further cause the instruction execution system to determine the sport type based on a distribution of the key audio components within the video clip.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2006
From: CHEN, MING-JUN; CHEN, JIUN-FU; TANG, SHIH-MIN; HUANG, HO-CHAO
To: CYBERLINK CORP.
Reel/Frame 017828/0250 →
Continuity (1)
Related Publication 20070250777A1 · Oct 25, 2007