IP Library Granted Patent US 10,210,184
Granted Patent B2
US 10,210,184 · App. 14/670,561 · Granted Feb 19, 2019

Methods and systems for enhancing metadata

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,210,184
App. No.
14/670,561
Granted
Feb 19, 2019
Kind
B2
Abstract

A method and system for utilizing metadata to search for media, such as multimedia and streaming media, includes searching for the media, receiving results, extracting metadata associated with the media, enhancing the extracted metadata, and grouping the search results in accordance with attributes of the enhanced metadata. Enhancing and grouping include adding related metadata to the database of metadata, iteratively using metadata to search for more media related data, removing duplicate URLs, collapsing URLs that are variants of each other, and masking out superfluous terms from URLs. The resultant metadata and media files are available to users and search engines.

Claims (68)

1. A system for enhancing metadata, the system comprising:

a metadata obtaining section that obtains metadata of a file that includes media; and

a digital storage storing processing instructions;

one or more processing sections that execute the processing instructions so as to implement functions of an extraction agent that:

separates noisy items of the metadata into keywords,

performs, using separated noisy items of the metadata, a full-text query against the potential ground truth database,

calculates a score quantifying a degree of similarity between the separated noisy items of metadata and potential ground truth data, and

qualifies the potential ground truth database, based on a comparison of the calculated score to a threshold score, as a ground truth database;

identifies, in the ground truth database, valid metadata that at least partially matches the obtained metadata; and

modifies the obtained metadata using at least a portion of the valid metadata so as to generate enhanced metadata, by

comparing contents in the obtained metadata with corresponding contents in the valid metadata,

identifying contents in the valid metadata that are not in the obtained metadata, and

adding the identified contents in the valid metadata to the obtained metadata.

2. The system of claim 1 , wherein the metadata obtaining section obtains the obtained metadata from a source other than the ground truth database.

3. The system of claim 1 , wherein the metadata obtaining section obtains at least some of the obtained metadata using a spider.

4. The system of claim 1 , wherein the metadata obtaining section parses and indexes into metadata fields the obtained metadata, and wherein the extraction agent identifies valid metadata via a field-by-field comparison of the metadata fields.

5. The system of claim 1 , wherein the extraction agent identifies the valid metadata by:

calculating a score based on a degree of similarity between the obtained metadata and the valid metadata; and

determining that the valid metadata at least partially matches the obtained metadata based at least in part on the calculated score.

6. A method of enhancing metadata, the method comprising:

separating noisy items of the metadata into keywords,

performing, using separated noisy items of the metadata, a full-text query against the potential ground truth database,

calculating a score quantifying a degree of similarity between the separated noisy items of metadata and potential ground truth data, and

qualifying the potential ground truth database, based on a comparison of the calculated score to a threshold score, as a ground truth database;

identifying, in the ground truth database, valid metadata that at least partially matches obtained metadata of a file that includes media; and

modifying the obtained metadata using at least a portion of the valid metadata so as to generate enhanced metadata, by

comparing contents in the obtained metadata with corresponding contents in the valid metadata,

identifying contents in the valid metadata that are not in the obtained metadata, and

adding the identified contents in the valid metadata to the obtained metadata.

7. The method of claim 6 , wherein the obtained metadata is obtained from a source other than the ground truth database.

8. The method of claim 6 , wherein the obtained metadata is at least partially obtained using a spider.

9. The method of claim 6 , further comprising parsing and indexing into metadata fields the obtained metadata, and wherein the identifying the valid metadata includes a field-by-field comparison of the metadata fields.

10. The system of method 6 , wherein the identifying valid metadata includes:

calculating a score based on a degree of similarity between the obtained metadata and the valid metadata; and

determining whether the valid metadata at least partially matches the obtained metadata based at least in part on the calculated score.

11. A device for enhancing metadata, the device comprising:

a metadata obtainer that obtains metadata of a file that includes media;

a digital storage that stores processing instructions; and

one or more hardware processors that execute the processing instructions so as to provide functions of an extractor that:

separates noisy items of the metadata into keywords,

performs, using separated noisy items of the metadata, a full-text query against the potential ground truth database,

calculates a score quantifying a degree of similarity between the separated noisy items of metadata and potential ground truth data, and

qualifies the potential ground truth database, based on a comparison of the calculated score to a threshold score, as a ground truth database;

identifies, in the ground truth database, valid metadata that at least partially matches the obtained metadata; and

enhances the obtained metadata using at least a portion of the valid metadata so as to generate enhanced metadata, by

comparing contents in the obtained metadata with corresponding contents in the valid metadata,

identifying contents in the valid metadata that are not in the obtained metadata, and

adding the identified contents in the valid metadata to the obtained metadata.

12. The system of claim 11 , wherein the metadata obtainer obtains the metadata of a file from a source other than the ground truth database.

13. The system of claim 11 , wherein the metadata obtainer obtains at least some of the metadata of the file using a spider.

14. The system of claim 11 , wherein the metadata obtainer parses and indexes into metadata fields the metadata of the file that includes media, and wherein the extraction agent identifies the valid metadata via a field-by-field comparison of the metadata fields.

15. The system of claim 11 , wherein the extractor identifies the valid metadata by:

calculating a score based on a degree of similarity between the metadata of the file that includes media and the valid metadata; and

determining that the valid metadata at least partially matches the metadata of the file that includes media based at least in part on the calculated score.

16. A device for enhancing metadata, the device comprising:

a metadata obtainer that obtains metadata of a file that includes media;

a digital storage that stores processing instructions;

one or more hardware processors that execute the processing instructions so as to provide functions of a potential ground truth database qualifier that

separates noisy items of the metadata into keywords,

performing, using separated noisy items of the metadata, a full-text query against the potential ground truth database,

calculating a score quantifying a degree of similarity between the separated noisy items of metadata and potential ground truth data, and

qualifying the potential ground truth database, based on a comparison of the calculated score to a threshold score, as a ground truth database; and

one or more hardware processors that execute the processing instructions so as to provide functions of an extractor that:

identifies, in the ground truth database, valid metadata that at least partially matches the obtained metadata; and

enhances the obtained metadata using at least a portion of the valid metadata so as to generate enhanced metadata, by

comparing contents in the obtained metadata with corresponding contents in the valid metadata,

identifying contents in the valid metadata that do not match contents in the obtained metadata, and

replacing the non-matching contents in the obtained metadata with the identified contents in the valid metadata.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2017
From: ABAJIAN, ARAM CHRISTIAN; ALEXANDER, ROBIN ANDREW; LEE, SCOTT CHAO-CHUEH; DAHL, AUSTIN DAVID; DEROSA, JOHN ANTHONY; PORTER, CHARLES A.; REHM, ERIC CARL; KOLAR, JENNIFER LYNN; SUDANAGUNTA, SRINIVASAN
To: THOMSON LICENSING S.A.
Reel/Frame 041725/0148 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2017
From: THOMSON LICENSING S.A.
To: AMERICA ONLINE, INC.
Reel/Frame 041725/0229 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2017
From: AOL INC.
To: MICROSOFT CORPORATION
Reel/Frame 041727/0105 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2017
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 041727/0254 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2017
From: AOL LLC
To: AOL INC.
Reel/Frame 042086/0001 →
CHANGE OF NAME Recorded Mar 24, 2017
From: AMERICA ONLINE, INC.
To: AOL LLC
Reel/Frame 042087/0438 →