IP Library Granted Patent US 8,843,490
Granted Patent B2
US 8,843,490 · App. 13/191,369 · Granted Sep 23, 2014

Method and system for automatically extracting data from web sites

Inventors: Bora C. Gazen (Huntington Beach, CA); Steven N. Minton (El Segundo, CA)
Assignee: Connotate, Inc.
G06F17/3071G06F17/30861
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,843,490
App. No.
13/191,369
Granted
Sep 23, 2014
Kind
B2
Abstract

In accordance with an embodiment, data may be automatically extracted from semi-structured web sites. Unsupervised learning may be used to analyze web sites and discover their structure. One method utilizes a set of heterogeneous “experts,” each expert being capable of identifying certain types of generic structure. Each expert represents its discoveries as “hints.” Based on these hints, the system may cluster the pages and text segments and identify semi-structured data that can be extracted. To identify a good clustering, a probabilistic model of the hint-generation process may be used.

Claims (42)

1. A method for automatically identifying semi-structured data from a semi-structured web site, the method comprising:

analyzing links and pages on the semi-structured web site using a set of heterogeneous experts, each of the experts focusing on a respective type of structure included in the semi-structured web site;

identifying, by the set of experts, similarities and dissimilarities between the analyzed links and pages;

clustering pages and text segments based on the similarities and dissimilarities identified by at least two experts in the set of heterogeneous experts,

wherein each of the at least two experts produces hints indicating whether two items should be together in a cluster, the hints containing respective levels of confidence; and

wherein the clustering text segments comprises:

finding page clusters;

determining a set of text segments for each of the found pare clusters; and

clustering text segments of the set of text segments;

identifying, based on the clustering of pages and text segments, at least some of the semi-structured data to be extracted from the semi-structured web site;

extracting the at least some of the identified semi-structured data; and

transforming the extracted semi-structured data into a relational structured form.

2. The method of claim 1 , further comprising evaluating a probability of a clustering based on the hints to determine a quality of the clustering.

3. The method of claim 1 , wherein the clustering of pages and text segments provides at least two alternative clusterings.

4. The method of claim 3 , further comprising employing probabilistic models to rate the alternative clusterings.

5. The method of claim 1 , further comprising employing a generative probabilistic model to enable assignment of probabilities to the hints in view of a clustering.

6. The method of claim 5 wherein all hints are assigned the probabilities.

7. The method of claim 6 , wherein probabilities of page hints are determined from page clusters.

8. The method of claim 1 , further comprising adding to the hints a binary hint that indicates that a particular pair of items are in the same cluster.

9. The method of claim 8 , further comprising extending a constraint language for constraint clustering, wherein constraints for the constraint clustering are defined in a form of must-link or cannot-link pairs.

10. The method of claim 9 , further comprising extending the constraint language so that the constraints are assigned confidence scores.

11. A system for automatically identifying semi-structured data from a semi-structured web site by executing instructions stored in a computer-readable memory by a computer processor, the system comprising:

means for analyzing links and pages on the semi-structured web site using a set of heterogeneous experts, each of the experts focusing on a respective type of structure included in the semi-structured web site;

means for identifying, by the set of experts, similarities and dissimilarities between the analyzed links and pages;

means for clustering pages and text segments based on the similarities and dissimilarities identified by at least two experts in the set of heterogeneous experts,

wherein each of the at least two experts produces hints indicating whether two items should be together in a cluster, the hints containing respective levels of confidence; and

wherein the clustering of text segments comprises:

finding page clusters;

determining a set of text segments for each of the found pare clusters; and

clustering text segments of the set of text segments;

means for identifying, based on the clustering of pages and text segments, at least some of the semi-structured data to be extracted from the semi-structured web site;

means for extracting the at least some of the identified semi-structured data; and

means for transforming the extracted semi-structured data into a relational structured form.

12. The system of claim 11 , further comprising means for evaluating a probability of a clustering based on the hints to determine a quality of the clustering.

13. The system of claim 11 , wherein the clustering means provides at least two alternative clusterings.

14. The system of claim 13 further comprises means for employing probabilistic models to rate the alternative clusterings.

15. The system of claim 11 , further comprising means for employing a generative probabilistic model to enable assignment of probabilities to the hints in view of a clustering.

16. The system of claim 15 , wherein all hints are assigned the probabilities.

17. The system of claim 16 , wherein probabilities of page hints are determined from page clusters.

18. The system of claim 11 , further comprising means for adding to the hints a binary hint that indicates that a particular pair of items are in the same cluster.

19. The system of claim 18 , further comprising means for extending a constraint language for constraint clustering, wherein constraints for the constraint clustering are defined in a form of must-link or cannot-link pairs.

20. The system of claim 19 further comprising means for extending the constraint language so that the constraints are assigned confidence scores.

Assignments (7)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2022
From: IMPORT.IO GLOBAL, INC.
To: IMPORT.IO CORPORATION
Reel/Frame 061550/0909 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2019
From: CONNOTATE, INC.
To: IMPORT.IO GLOBAL INC.
Reel/Frame 048888/0452 →
RELEASE OF SECURITY INTEREST Recorded Feb 14, 2019
From: PACIFIC WESTERN BANK
To: CONNOTATE, INC.
Reel/Frame 048329/0116 →
SECURITY AGREEMENT Recorded Oct 10, 2012
From: CONNOTATE, INC.
To: SQUARE 1 BANK
Reel/Frame 029102/0293 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2012
From: FETCH TECHNOLOGIES, INC.
To: CONNOTATE, INC.
Reel/Frame 028411/0237 →
CHANGE OF ADDRESS Recorded Sep 6, 2011
From: FETCH TECHNOLOGIES, INC.
To: FETCH TECHNOLOGIES, INC.
Reel/Frame 026862/0286 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2011
From: GAZEN, BORA C.; MINTON, STEVEN N.
To: FETCH TECHNOLOGIES, INC.
Reel/Frame 026668/0453 →
Continuity (4)
Continuation 12014532 · Jan 15, 2008
Continuation PCTUS2006027335 · Jul 14, 2006
Provisional Application 60699519 · Jul 15, 2005
Related Publication 20110282877A1 · Nov 17, 2011