IP Library › Granted Patent US 8,521,761
Granted Patent B2
US 8,521,761 · App. 12/503,806 · Granted Aug 27, 2013

Transliteration for query expansion

Inventors: Lalitesh Katragadda (Hyderabad, IN); Vineet Gupta (Bangalore, IN); Piyush Prahladka (Palo Alto, CA)
Assignee: Google Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,521,761
App. No.
12/503,806
Filed
Jul 15, 2009
Granted
Aug 27, 2013
Kind
B2
Art Unit
2166
USPC
707/4
Abstract

Methods, systems, and apparatus, including computer program products, for identifying candidate synonyms of transliterated terms for query expansion. In one aspect, a method includes identifying multiple transliterated terms in a target language. For each transliterated term of the multiple transliterated terms in the target language, the transliterated term is mapped to one or more terms in a source language. For a first transliterated term of the multiple transliterated terms in the target language, one or more second transliterated terms of the multiple transliterated terms in the target language are identified as candidate synonyms of the first transliterated term, where each of the one or more second transliterated terms is mapped to at least one term in the source language that is also mapped from the first transliterated term.

Claims (69)

1. A computer-implemented method comprising:

identifying, using one or more computers, a plurality of transliterated terms in a target language, where identifying the plurality of transliterated terms in the target language comprises:

identifying terms containing only characters of the target language,

computing a statistic for each identified term of the terms containing only characters of the target language, where the statistic for each said identified term is based on a ratio of a probability of occurrence of the identified term in resources associated with one or more locales where the source language is spoken to a probability of occurrence of the identified term in resources associated with any locale,

comparing the statistic for each said identified term to a specified threshold, and

including a particular said identified term in the plurality of transliterated terms in the target language if the statistic for the identified term satisfies the specified threshold;

for each transliterated term of the plurality of transliterated terms in the target language, mapping the transliterated term to one or more terms in a source language; and

for a first transliterated term of the plurality of transliterated terms in the target language, identifying one or more second transliterated terms of the plurality of transliterated terms in the target language as candidate synonyms of the first transliterated term, where each of the one or more second transliterated terms is mapped to at least one term in the source language that is also mapped from the first transliterated term.

2. The method of claim 1 , where the ratio for the statistic for each said identified term is a probability of occurrence of the identified term in web resources of a top-level domain associated with one or more locales where the source language is spoken to a probability of occurrence of the identified term in web resources of a top-level domain associated with any locale.

3. The method of claim 1 , where the resources associated with one or more locales where the source language is spoken are determined by a top-level domain of the resources.

4. The method of claim 1 , where mapping the transliterated term to one or more terms in the source language further comprises:

transliterating the transliterated term in the target language to the one or more terms in the source language.

5. The method of claim 4 , where each of the one or more second transliterated terms identified as candidate synonyms of the first transliterated term has a confidence value with respect to the first transliterated term that is above a specified threshold.

6. The method of claim 4 , where the confidence value of a second transliterated term is a function of a number of terms in the source language that are mapped from both the first transliterated term and the second transliterated term.

7. The method of claim 4 , where transliterating the transliterated term in the target language to a term in the source language further comprises:

generating a transliteration score for the transliteration of the transliterated term in the target language to the term in the source language.

8. The method of claim 7 , where the confidence value of a second transliterated term is a function of one or more of a probability of occurrence of the second transliterated term in web resources, the transliteration score for the transliteration of the second transliterated term to a term in the source language that is also mapped from the first transliterated term, and the transliteration score for the transliteration of the first transliterated term to the term in the source language.

9. The method of claim 1 , further comprising:

for the first transliterated term of the plurality of transliterated terms in the target language, identifying one or more terms in the source language that are mapped from the first transliterated term and from at least one of the one or more second transliterated terms as candidate synonyms of the first transliterated term.

10. The method of claim 1 , further comprising:

receiving a query including the first transliterated term;

expanding the query with one or more of the candidate synonyms of the first transliterated term;

providing the expanded query to a search engine; and

receiving search results for the expanded query.

11. The method of claim 1 , further comprising:

receiving a query including the first transliterated term; and

providing one or more expanded queries for selection by a user, each expanded query including the query and one or more of the candidate synonyms of the first transliterated term.

12. The method of claim 1 , further comprising:

receiving a query including the first transliterated term;

providing the query to a search engine, where the search engine identifies as a possible search result for the query a web resource that includes at least one of the candidate synonyms of the first transliterated term but does not include any term in the query; and

modifying a score associated with the web resource, the score for use in ranking possible search results for the query.

13. The method of claim 1 , further comprising:

receiving a query including the first transliterated term;

providing the query to a search engine, where the search engine identifies as a possible search result for the query a web resource that includes at least one of the terms in the source language that is mapped from the first transliterated term and from at least one of the one or more second transliterated terms but does not include any term in the query; and

modifying an information retrieval score associated with the web resource, the information retrieval score for use in ranking possible search results for the query.

14. A system comprising:

one or more computers configured to perform operations including:

identifying a plurality of transliterated terms in a target language, where identifying the plurality of transliterated terms in the target language comprises:

identifying terms containing only characters of the target language,

computing a statistic for each identified term of the terms containing only characters of the target language, where the statistic for each said identified term is based on a ratio of a probability of occurrence of the identified term in resources associated with one or more locales where the source language is spoken to a probability of occurrence of the identified term in resources associated with any locale,

comparing the statistic for each identified term to a specified threshold, and

including a particular identified term in the plurality of transliterated terms in the target language if the statistic for the particular identified term satisfies the specified threshold;

for each transliterated term of the plurality of transliterated terms in the target language, mapping the transliterated term to one or more terms in a source language; and

for a first transliterated term of the plurality of transliterated terms in the target language, identifying one or more second transliterated terms of the plurality of transliterated terms in the target language as candidate synonyms of the first transliterated term, where each of the one or more second transliterated terms is mapped to at least one term in the source language that is also mapped from the first transliterated term.

15. A computer-implemented method comprising:

identifying, using one or more computers, a plurality of transliterated terms in a target language, where identifying the plurality of transliterated terms in the target language comprises:

identifying terms containing only characters of the target language,

computing a statistic for each identified term of the terms containing only characters of the target language, where the statistic for each said identified term is based on a ratio of a probability of occurrence of the identified term in resources associated with one or more locales where the source language is spoken to a probability of occurrence of the identified term in resources associated with any locale,

comparing the statistic for each identified term to a specified threshold, and

including a particular identified term in the plurality of transliterated terms in the target language if the statistic for the particular identified term satisfies the specified threshold;

for a first transliterated term of the plurality of transliterated terms in the target language, identifying one or more second transliterated terms of the plurality of transliterated terms in the target language as candidate synonyms of the first transliterated term; and

using the candidate synonyms of the first transliterated term to expand queries including the first transliterated term.

16. A system comprising:

one or more computers configured to perform operations including:

identifying, using one or more computers, a plurality of transliterated terms in a target language, where identifying the plurality of transliterated terms in the target language comprises:

identifying terms containing only characters of the target language,

computing a statistic for each identified term of the terms containing only characters of the target language, where the statistic for each said identified term is based on a ratio of a probability of occurrence of the identified term in resources associated with one or more locales where the source language is spoken to a probability of occurrence of the identified term in resources associated with any locale,

comparing the statistic for each identified term to a specified threshold, and

including a particular identified term in the plurality of transliterated terms in the target language if the statistic for the particular identified term satisfies the specified threshold;

for a first transliterated term of the plurality of transliterated terms in the target language, identifying one or more second transliterated terms of the plurality of transliterated terms in the target language as candidate synonyms of the first transliterated term; and

using the candidate synonyms of the first transliterated term to expand queries including the first transliterated term.

17. A non-transitory computer readable storage medium storing computer instructions executable by a processor to perform a method comprising:

identifying, using one or more computers, a plurality of transliterated terms in a target language, where identifying the plurality of transliterated terms in the target language comprises:

identifying terms containing only characters of the target language,

computing a statistic for each identified term of the terms containing only characters of the target language, where the statistic for each said identified term is based on a ratio of a probability of occurrence of the identified term in resources associated with one or more locales where the source language is spoken to a probability of occurrence of the identified term in resources associated with any locale,

comparing the statistic for each identified term to a specified threshold, and

including a particular identified term in the plurality of transliterated terms in the target language if the statistic for the particular identified term satisfies the specified threshold;

for each transliterated term of the plurality of transliterated terms in the target language, mapping the transliterated term to one or more terms in a source language; and

for a first transliterated term of the plurality of transliterated terms in the target language, identifying one or more second transliterated terms of the plurality of transliterated terms in the target language as candidate synonyms of the first transliterated term, where each of the one or more second transliterated terms is mapped to at least one term in the source language that is also mapped from the first transliterated term.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044101/0299 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2009
From: KATRAGADDA, LALITESH; PRAHLADKA, PIYUSH; GUPTA, VINEET
To: GOOGLE INC.
Reel/Frame 023033/0519 →
Continuity (2)
Provisional Application 61082165 · Jul 18, 2008
Related Publication 20100017382A1 · Jan 21, 2010