IP Library Granted Patent US 9,239,823
Granted Patent B1
US 9,239,823 · App. 13/893,462 · Granted Jan 19, 2016

Identifying common co-occurring elements in lists

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,239,823
App. No.
13/893,462
Granted
Jan 19, 2016
Kind
B1
Abstract

One embodiment of the present invention provides a system for detecting correlations between terms. During operation, the system identifies one or more lists contained in one or more documents and identifies two terms co-occurring in the lists. The system further determines a correlation between the co-occurring terms, and places the co-occurring terms in a correlated-pair list based on the correlation.

Claims (30)

1. A computer-implemented method comprising:

obtaining, at one or more computers, a pair of terms in a first language, the pair of terms being commonly co-occurring non-synonyms in a corpus of documents, the corpus of documents being in the first language;

determining a set of variations for each term in the pair of terms;

generating a set of known related input pairs based on the sets of variations for each term in the pair of terms;

for each input pair of terms in the set of known related input pairs, translating, by an automatic translation system, each term in the pair of terms into a second language plurality of languages to generate a set of translated terms;

adding, at the one or more computers, the set of translated terms to a blacklist of known non-synonym pairs for at least one of the plurality of languages; and

determining, based on the blacklist of known non-synonym pairs, whether a pair of candidate terms in at least one of the plurality of languages are synonyms.

2. The method of claim 1 , wherein determining variations for each term in the pair of terms comprises determining one or more normalized versions of each term in the pair of terms.

3. The method of claim 1 , wherein generating a set of known related input pairs based on the sets of variations for each term in the pair comprises calculating a cross-product between the sets of variations for each term in the pair.

4. The method of claim 1 , further comprising generating one or more normalized versions of one or more terms in the set of translated terms.

5. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining, at one or more computers, a pair of terms in a first language, the pair of terms being commonly co-occurring non-synonyms in a corpus of documents, the corpus of documents being in the first language;

determining a set of variations for each term in the pair of terms;

generating a set of known related input pairs based on the sets of variations for each term in the pair of terms;

for each input pair of terms in the set of known related input pairs, translating, by an automatic translation system, each term in the pair of terms into a plurality of languages to generate a set of translated terms;

adding, at the one or more computers, the set of translated terms to a blacklist of known non-synonym pairs for at least one of the plurality of languages; and

determining, based on the blacklist of known non-synonym pairs, whether a pair of candidate terms in at least one of the plurality of languages are synonyms.

6. The system of claim 5 , wherein determining variations for each term in the pair of terms comprises determining one or more normalized versions of each term in the pair of terms.

7. The system of claim 5 , wherein generating a set of known related input pairs based on the sets of variations for each term in the pair comprises calculating a cross-product between the sets of variations for each term in the pair.

8. The system of claim 5 , further comprising generating one or more normalized versions of one or more terms in the set of translated terms.

9. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

obtaining, at one or more computers, a pair of terms in a first language, the pair of terms being commonly co-occurring non-synonyms in a corpus of documents, the corpus of documents being in the first language;

determining a set of variations for each term in the pair of terms;

generating a set of known related input pairs based on the sets of variations for each term in the pair of terms;

for each input pair of terms in the set of known related input pairs, translating, by an automatic translation system, each term in the pair of terms into a plurality of languages to generate a set of translated terms;

adding, at the one or more computers, the set of translated terms to a blacklist of known non-synonym pairs for at least one of the plurality of languages; and

determining, based on the blacklist of known non-synonym pairs, whether a pair of candidate terms in at least one of the plurality of languages are synonyms.

10. The computer-readable medium of claim 9 , wherein determining variations for each term in the pair of terms comprises determining one or more normalized versions of each term in the pair of terms.

11. The computer-readable medium of claim 9 , wherein generating a set of known related input pairs based on the sets of variations for each term in the pair comprises calculating a cross-product between the sets of variations for each term in the pair.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044566/0657 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2013
From: UPSTILL, TRYSTAN G.; BAKER, STEVEN D.
To: GOOGLE INC.
Reel/Frame 030668/0803 →