IP Library Granted Patent US 8,484,014
Granted Patent B2
US 8,484,014 · App. 12/362,428 · Granted Jul 9, 2013

Retrieval using a generalized sentence collocation

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,484,014
App. No.
12/362,428
Granted
Jul 9, 2013
Kind
B2
Abstract

A method and system for identifying documents relevant to a query that specifies a part of speech is provided. A retrieval system receives from a user an input query that includes a word and a part of speech. Upon receiving an input query that includes a word and a part of speech, the retrieval system identifies documents with a sentence that includes that word collocated with a word that is used as that part of speech. The retrieval system displays to the user an indication of the identified documents.

Claims (41)

1. A method in a computing device for searching for sentences, the method comprising:

providing a collection of sentences;

for each sentence in the collection,

identifying by the computing device pairs of collocated words of the sentence;

for each identified pair of collocated words,

identifying by the computing device a part of speech of each word of the pair;

generating a first part of speech and word pair that includes the identified part of speech of the first word and the second word and a second part of speech and word pair that includes the first word and the identified part of speech of the second word; and

generating a mapping from the first part of speech and word pair and the second part of speech and word pair for the sentence;

receiving an input query that includes a first word, a part of speech, and a second word;

identifying from the mapping sentences that include the first word of the input query collocated with a word with the part of speech of the input query and a word with the part of speech collocated with the second word of the input query; and

displaying to a user the identified sentences.

2. The method of claim 1 wherein the identifying of sentences includes identifying a sentence such that the same word with the part of speech of input query is collocated with both the first word and the second word of the input query.

3. The method of claim 1 wherein the mapping includes an entry for each unique part of speech and word pair that identifies sentences that contain collocated words with that part of speech and word pair.

4. The method of claim 3 wherein each entry of the mapping identifies a position of the part of speech and a position of the word within each sentence that contains collocated words.

5. The method of claim 1 wherein the identifying of a part of speech of a word within a sentence includes invoking a natural language parser.

6. The method of claim 1 including ranking the identified sentences based on readability and displaying the sentences in an order base on their ranking.

7. The method of claim 1 wherein when the input query includes multiple parts of speech identifying sentences that include each of the multiple parts of speech collocated with a word of the input query.

8. A computer-readable storage medium that is not a signal containing instructions for controlling a computing device to identify documents relevant to a query, by a method comprising:

receiving from a user an input query that includes a word and a part of speech, the part of speech representing a wildcard for any word that is that part of speech;

identifying documents with a sentence that includes the word collocated with any word of that part of speech, the document being identified based on mappings from part of speech and word pairs to documents, the mappings generated by:

identifying collocated words of sentences of the documents; and

for each identified pair of collocated words of a sentence,

identifying a part of speech of each word of the pair;

generating a first part of speech and word pair that includes the identified part of speech of the first word and the second word and a second part of speech and word pair that includes the first word and the identified part of speech of the second word; and

generating a mapping from the first part of speech and word pair and the second part of speech and word pair to the document that contains the sentence;

ranking the identified documents; and

displaying to the user the identified documents in order of their rankings.

9. The computer-readable storage medium of claim 8 wherein the documents are web pages identified by crawling web sites.

10. The computer-readable storage medium of claim 9 wherein the ranking of the documents includes ranking web pages based on the relevance of sentences to the input query.

11. A computing device for identifying sentences having words of a designated part of speech, comprising:

a component that inputs from a user a query having a word and a part of speech, the part of speech representing a wildcard for any word that is that part of speech;

a component that identifies sentences that have the word of the query collocated with any word used as the part of speech of the query, the sentences being identified based on mappings from part of speech and word pairs to sentences, the mappings having been generated by a component that:

identifies collocated words of sentences; and

for each identified pair of collocated words of a sentence,

identifies a part of speech of each word of the pair;

generates a first part of speech and word pair that includes the identified part of speech of the first word and the second word and a second part of speech and word pair that includes the first word and the identified part of speech of the second word; and

generates a mapping from the first part of speech and word pair and the second part of speech and word pair to the sentence; and

a component that displays to the user the identified sentences.

12. The computing of claim 11 wherein the word of the query and the word used as the part of speech are different words.

13. The computing device of claim 11 including a component that ranks the sentences and wherein the sentences are displayed in an order based on their ranking.

14. The computing device of claim 11 including a component that generates a mapping of collocated words and their parts of speech for sentences that contain them.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034564/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2009
From: LIU, XIAOHUA; ZHOU, MING; WEI, HAO; ZHAO, JING; SCOTT, MATTHEW R.; JIANG, LONG; CHEN, GANG
To: MICROSOFT CORPORATION
Reel/Frame 022204/0322 →