IP Library Granted Patent US 10,318,564
Granted Patent B2
US 10,318,564 · App. 14/867,620 · Granted Jun 11, 2019

Domain-specific unstructured text retrieval

Inventors: Achraf Abdel Moneim Tawfik Chalabi (Cairo, EG); Eslam Kamal Abdel-Aal Abdel-Reheem (Cairo, EG); Sayed Hassan Sayed Abdelaziz (Redmond, WA); Yuval Yehezkel Marton (Seattle, WA); Michel Naim Naguib Gerguis (Cairo, EG)
Assignee: Microsoft Technology Licensing, LLC
G06F16/334G06F16/35G06F16/951G06F16/958G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,318,564
App. No.
14/867,620
Granted
Jun 11, 2019
Kind
B2
Abstract

Retrieving from the Internet unstructured text related to a specified domain is described. Training data is accessed; the training data comprises unstructured text related to the specified domain. A first classifier is trained using features of the training data. It is used to classify unstructured text having plurality of features, to obtain unstructured text examples related to the domain. The unstructured text examples are used to retrieve from the Internet similar examples which do not have at least some of the plurality of features. Optionally, a second classifier is trained using the similar examples. Additional unstructured text is retrieved from the Internet and the second classifier is used to label the additional unstructured text for domain relevance.

Claims (37)

1. An apparatus for retrieving unstructured text from the Internet related to a specified domain, the apparatus comprising:

one or more processors; and

a memory having instructions stored therein, the instructions executable by the one or more processors to perform operations as

a first classifier having been trained using training data comprising unstructured text related to the specified domain, the training data having a plurality of features, the unstructured text being separated from structured data and semi-structured data;

a similar web page retriever configured to retrieve, from the Internet, only web pages that include text that is unstructured and do not have at least some of the plurality of features of the training data, and where the retrieved web pages are similar to web pages classified by the first classifier; and

a second classifier having been trained using unstructured text examples which do not have at least one of the plurality of features;

wherein the second classifier is configured to label web pages retrieved by the similar web page retriever to select web pages which are relevant to the specified domain.

2. The apparatus of claim 1 wherein the similar web page retriever is configured to assign confidence values to the similar web pages to indicate likelihood of being relevant to the specified domain and wherein the second classifier has been trained using unstructured text examples from web pages retrieved by the similar web page retriever and selected according to confidence values.

3. The apparatus of claim 1 wherein the similar web page retriever identifies the similar web pages from inbound links of web pages classified by the first classifier.

4. The apparatus of claim 1 wherein the similar web page retriever identifies first similar web pages from inbound links of web pages classified by the first classifier and from inbound links of first similar web pages.

5. The apparatus of claim 1 wherein the similar web page retriever accesses an index of web pages, the index comprising inbound link data of the indexed web pages.

6. The apparatus of claim 1 wherein the similar web page retriever accesses a click log comprising a record of web pages observed as having been selected by different users in connection with a same query.

7. The apparatus of claim 1 wherein the similar web page retriever accesses an impression log comprising a record of web pages occurring in results lists returned by an information retrieval system in response to a same query.

8. The apparatus of claim 1 wherein the similar web page retriever is configured to assign a confidence value to a similar web page on the basis of a number of inbound links of the similar web page.

9. The apparatus of claim 1 wherein the first classifier is configured to use features comprising one or more of: a category of a web page, a title of a web page, metatags of a web page, an information box of a web page.

10. The apparatus of claim 1 wherein the first classifier is configured to classify web pages of a public online encyclopedia as being relevant to the specified domain or not.

11. The apparatus of claim 1 wherein the first classifier has been trained using training data retrieved from a source known to comprise web pages having the plurality of features and using queries comprising seed examples.

12. The apparatus of claim 1 comprising a communications interface configured to enable the apparatus to be accessed as a web service.

13. The apparatus of claim 1 further comprising a feature extractor configured to extract sentences from the web pages retained by the second classifier, where the extracted sentences are likely to comprise facts.

14. The apparatus of claim 13 further comprising a clustering component configured. to cluster the extracted facts into relation clusters and assign confidence values to the clusters' facts.

15. The apparatus of claim 14 further comprising a mapping component configured to map the relation clusters of extracted facts to an ontology of a knowledge store.

16. A computer-implemented method of retrieving unstructured text from the Internet related to a specified domain, the method comprising:

accessing training data comprising unstructured text related to the specified domain, the training data having a plurality of features, the unstructured text being separated from structured data and semi-structured data;

training a first classifier using the training data including the unstructured text related to the specified domain and the plurality of features;

using the trained first classifier to classify unstructured text having the plurality of features, to obtain unstructured text examples related to the domain;

using the unstructured text examples to retrieve, from the Internet, only similar examples that include text that is unstructured and do not have at least one of the plurality of features of the training data;

training a second classifier using at least some of the similar examples,

retrieving additional unstructured text from the Internet; and

using the second classifier to classify the additional unstructured text as being related to the specified domain or not.

17. The method of claim 16 comprising assigning confidence values to the similar web pages to indicate likelihood of being relevant to the specified domain and training the second classifier using web pages selected according to the confidence values.

18. The method of claim 16 comprising identifying the similar web pages from inbound links of web pages classified by the first classifier.

19. The method of claim 16 comprising identifying the similar web pages from a cascade of inbound links of web pages classified by both the first and the second classifiers.

20. An apparatus for retrieving unstructured text from the Internet related to a specified domain, the apparatus comprising:

one or more processors; and

a memory having instructions stored therein, the instructions executable by the one or more processors to perform operations as

a first classifier having been trained using training data comprising unstructured text related to the specified domain, the training data having a plurality of features, the unstructured text being separated from structured data and semi-structured data; and

a similar web page retriever configured to retrieve from the Internet, only web pages that include text that is unstructured and do not have at least some of the plurality of features of the training data, and where the retrieved web pages are similar to web pages classified by the first classifier, the similar web page retriever being configured to assign confidence values to the similar web pages to indicate likelihood of being relevant to the specified domain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2015
From: CHALABI, ACHRAF ABDEL MONEIM TAWFIK; ABDEL-REHEEM, ESLAM KAMAL ABDEL-AAL; ABDELAZIZ, SAYED HASSAN SAYED; MARTON, YUVAL YEHEZKEL; GERGUIS, MICHEL NAIM NAGUIB
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 036671/0316 →
Continuity (1)
Related Publication 20170091313A1 · Mar 30, 2017
Cited By (1)
US 12,705,273