IP Library Granted Patent US 7,184,959
Granted Patent B2
US 7,184,959 · App. 10/686,459 · Granted Feb 27, 2007

System and method for automated multimedia content indexing and retrieval

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,184,959
App. No.
10/686,459
Granted
Feb 27, 2007
Kind
B2
Abstract

The invention provides a system and method for automatically indexing and retrieving multimedia content. The method may include separating a multimedia data stream into audio, visual and text components, segmenting the audio, visual and text components based on semantic differences, identifying at least one target speaker using the audio and visual components, identifying a topic of the multimedia event using the segmented text and topic category models, generating a summary of the multimedia event based on the audio, visual and text components, the identified topic and the identified target speaker, and generating a multimedia description of the multimedia event based on the identified target speaker, the identified topic, and the generated summary.

Claims (50)

1. A method for automatically indexing and retrieving a multimedia event, comprising:

separating a multimedia data stream into audio, visual and text components;

segmenting the audio, visual and text components of the multimedia data stream based on semantic differences, wherein frame-level features are extracted from the segmented audio component in a plurality of subbands;

identifying at least one target speaker using the audio and visual components;

identifying semantic boundaries of text for at least one of the identified target speakers to generate semantically coherent text blocks;

generating a summary of multimedia content based on the audio, visual and text components, the semantically coherent text blocks and the identified target speaker;

deriving a topic for each of the semantically coherent text blocks based on a set of topic category models; and

generating a multimedia description of the multimedia event based on the identified target speaker, the semantically coherent text blocks, the topic, and the generated summary.

2. The method of claim 1 , further comprising:

automatically identifying a hierarchy of multimedia content types.

3. The method of claim 2 , wherein the multimedia content types include at least one of speakers, anchors, interviews, correspondence reports, multimedia content segments, general news stories, topical news stories, news summaries, and commercials.

4. The method of claim 1 , further comprising:

converting the multimedia data stream from an analog multimedia data stream to a digital multimedia data stream; and

compressing the digital multimedia data stream.

5. The method of claim 1 , wherein the extracted audio features from the audio component further comprise clip level features.

6. The method of claim 1 , wherein the multimedia event includes a news broadcast and the target speakers include news anchorpersons.

7. The method of claim 1 , wherein the step of identifying at least one speaker includes the process of identifying using Gaussian Mixture Models.

8. The method of claim 1 , wherein the generated multimedia description is represented by at least one of a text description, a video description and a story icon.

9. The method of claim 1 , further comprising:

storing the generated multimedia descriptions in a database.

10. The method of claim 1 , further comprising:

presenting the generated multimedia description to a user.

11. The method of claim 10 , further comprising:

playing back the segment of the multimedia event corresponding to the generated multimedia description to the user.

12. The method of claim 1 , wherein the plurality of subbands comprises three subbands.

13. The method of claim 12 , wherein the frame level features in the three subbands are at least one of volume, zero crossing rate, pitch period, frequency centroid, frequency bandwidth and energy ratios.

14. A terminal that displays the multimedia descriptions generated by the multimedia description generator of claim 1 .

15. A system that automatically indexes and retrieves a multimedia event, comprising:

a multimedia data stream separation unit that separates a multimedia data stream into audio, visual and text components;

a data stream component segmentation unit that segments the audio, visual and text components of the multimedia data stream based on semantic differences;

a feature extraction unit that extracts audio features from the audio component and the audio features comprising a frame-level feature in a plurality of subbands;

a target speaker detection unit that identifies at least one target speaker using the audio and visual components;

a content segmentation unit that identifies semantic boundaries of text for at least one of the identified target speakers, to generate semantically coherent text blocks;

a summary generator that generates a summary of multimedia content based on the audio, visual and text components, the semantically coherent text blocks and the identified target speaker;

a topic categorization unit that derives a topic for each of the semantically coherent text blocks based on a set of topic category models; and

a multimedia description generator that generates a multimedia description of the multimedia event based on the identified target speaker, the semantically coherent text blocks, the topic and the generated summary.

16. The system of claim 15 , wherein the multimedia description generator automatically identifies a hierarchy of multimedia content types.

17. The system of claim 16 , wherein the multimedia content types include at least one of speakers, anchors, interviews, correspondence reports, multimedia content segments, general news stories, topical news stories, news summaries, and commercials.

18. The system of claim 15 , further comprising:

an analog-to-digital converter that converts the multimedia data stream from an analog multimedia data stream to a digital multimedia data stream; and

a compression unit that compresses the digital multimedia data stream.

19. The system of claim 15 , wherein the multimedia event includes a news broadcast and the target speakers include news anchorpersons.

20. The system of claim 15 , wherein the target speaker detection unit identifies at least one target speaker using Gaussian Mixture Models.

21. The system of claim 15 , wherein the multimedia description generator generates one or more multimedia description that are represented by at least one of a text description, a video description and a story icon.

22. The system of claim 15 , further comprising:

a database that stores the generated multimedia descriptions.

23. The system of claim 15 , wherein the generated multimedia descriptions are retrieved from the database and presented to a user.

24. The system of claim 23 , further comprising:

a playback device that plays back the segment of the multimedia event corresponding to the generated multimedia description to the user.

25. The system of claim 15 , wherein the plurality of subbands comprises three sub-bands.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041498/0316 →
CORRECTIVE ASSIGNMENT TO CORRECT THE WRONG INVENTOR ASSIGNMENT SUBMITTED PREVIOUSLY RECORDED AT REEL: 038959 FRAME: 0712. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 9, 2016
From: GIBBON, DAVID CRAWFORD; HUANG, QIAN; LIU, ZHU; ROSENBERG, AARON EDWARD; SHAHRARAY, BEHZAD
To: AT&T CORP.
Reel/Frame 039635/0470 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2016
From: HUANG, QIAN; MAGRIN-CHAGNOLLEAU, IVAN; PARTHASARATHY, SARANGARAJAN; ROSENBERG, AARON EDWARD
To: AT&T CORP.
Reel/Frame 038959/0712 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038961/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038961/0431 →