IP Library Granted Patent US 9,298,757
Granted Patent B1
US 9,298,757 · App. 13/801,278 · Granted Mar 29, 2016

Determining similarity of linguistic objects

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,298,757
App. No.
13/801,278
Granted
Mar 29, 2016
Kind
B1
Abstract

A computer-implemented system for searching includes a data store accessible via a network for storing a data set; an indexing system coupled to the network and indexing the data set, the indexing system configured to generate content vectors for terms in the data set; generate index vectors for terms in the data set; and generate a bitset signature from the index vector. The system further includes a search module coupled to the network and configured to receive a search query and perform a search on one or more terms in the search query by accessing a bitset signature and content vector corresponding to the term; retrieving bitset signatures that are within a predetermined closeness to the bitset signature; selecting content vectors corresponding to retrieved bitset signatures; and selecting content vectors that are within a predetermined similarity to the term content vector; and return the terms corresponding to the content vectors.

Claims (55)

1. A computer-implemented system for searching, comprising:

a data store accessible via a network for storing a data set;

an indexing system coupled to the network and indexing the data set, the indexing system including a processor configured to:

generate content vectors for terms in the data set, wherein the content vectors define a similarity metric;

generate index vectors for the terms in the data set from the content vectors to access the terms in the data set; and

generate bitset signatures from the index vectors to determine similarity with the terms in the data set, wherein the bitset signatures include a first section for positive magnitude values and a second section for negative magnitude values, and generating the bitset signatures from the index vectors comprises:

for each of a predetermined number of highest magnitude positive values in the corresponding index vector, setting a corresponding bitset signature value at a corresponding dimension in the first section to a predetermined value; and

for each of a predetermined number of highest magnitude negative values in the corresponding index vector, setting a corresponding bitset signature value at a corresponding dimension in the second section to the predetermined value; and

a search module coupled to the network and including a processor configured to receive a search query and perform a search on each of one or more terms in the search query by:

accessing a bitset signature and content vector corresponding to a term in the search query;

retrieving bitset signatures that are within a predetermined closeness to the accessed bitset signature;

selecting content vectors corresponding to the retrieved bitset signatures;

identifying the selected content vectors that are within a predetermined similarity to the accessed content vector corresponding to the term in the search query; and

returning the terms of the data set corresponding to the identified content vectors.

2. A system according to claim 1 , wherein the index vectors and the content vectors for the terms in the data set are generated using random indexing.

3. A system according to claim 1 , wherein the index vectors and the content vectors for the term s in the data set are generated using hashing.

4. A system according to claim 1 , wherein the generated index vectors are smaller in dimension than the generated content vectors.

5. A system according to claim 1 , wherein the bitset signatures are twice a length of the generated index vectors.

6. A system according to claim 1 , wherein the one or more terms in the search query and the terms corresponding to the identified content vectors are used to search the data set.

7. A computer program product comprising one or more non-transitory computer readable storage media storing instructions translatable by one or more processors to perform:

generating content vectors for terms in a data set, wherein the content vectors define a similarity metric;

generating index vectors for the terms in the data set from the content vectors to access the terms in the data set; and

generating bitset signatures from the index vectors to determine similarity with the terms in the data set, wherein the bitset signatures include a first section for positive magnitude values and a second section for negative magnitude values, and generating the bitset signatures from the index vectors comprises:

for each of a predetermined number of highest magnitude positive values in the corresponding index vector, setting a corresponding bitset signature value at a corresponding dimension in the first section to a predetermined value; and

for each of a predetermined number of highest magnitude negative values in the corresponding index vector, setting a corresponding bitset signature value at a corresponding dimension in the second section to the predetermined value;

storing the bitset signatures and content vectors in an index; and

performing a search on each of one or more terms in a search query by:

accessing a bitset signature and content vector from the index corresponding to a term in the search query;

retrieving bitset signatures from the index that are within a predetermined closeness to the accessed bitset signature;

selecting content vectors from the index corresponding to the retrieved bitset signatures;

identifying the selected content vectors that are within a predetermined similarity to the accessed content vector corresponding to the term in the search query; and

returning the terms of the data set corresponding to the identified content vectors.

8. A computer program product according to claim 7 , wherein the index vectors and the content vectors for the terms in the data set are generated using random indexing.

9. A computer program product according to claim 7 , wherein the index vectors and the content vectors for the terms in the data set are generated using hashing.

10. A computer program product according to claim 7 , wherein the generated index vectors are smaller in dimension than the generated content vectors.

11. A computer program product according to claim 7 , wherein the bitset signatures are twice a length of the generated index vectors.

12. A computer program product according to claim 7 , wherein the one or more terms in the search query and the terms corresponding to the identified content vectors are used to search the data set.

13. A computer-implemented method, comprising:

generating content vectors for terms in a data set stored in a data store, wherein the content vectors define a similarity metric;

generating index vectors for the terms in the data set from the content vectors to access the terms in the data set;

generating bitset signatures from the index vectors to determine similarity with the terms in the data set, wherein the bitset signatures include a first section for positive magnitude values and a second section for negative magnitude values, and generating the bitset signatures from the index vectors comprises:

for each of a predetermined number of highest magnitude positive values in the corresponding index vector, setting a corresponding bitset signature value at a corresponding dimension in the first section to a predetermined value; and

for each of a predetermined number of highest magnitude negative values in the corresponding index vector, setting a corresponding bitset signature value at a corresponding dimension in the second section to the predetermined value;

storing the bitset signatures and content vectors in an index; and

performing a search on each of one or more terms in a search query received via a computer by:

accessing a bitset signature and content vector from the index corresponding to a term in the search query;

retrieving bitset signatures from the index that are within a predetermined closeness to the accessed bitset signature;

selecting content vectors from the index corresponding to the retrieved bitset signatures;

identifying the selected content vectors that are within a predetermined similarity to the accessed content vector corresponding to the term in the search query; and

returning the terms of the data set corresponding to the identified content vectors from the data store.

14. A method according to claim 13 , wherein the index vectors and the content vectors for the terms in the data set are generated using random indexing.

15. A method according to claim 13 , wherein the index vectors and the content vectors for the terms in the data set are generated using hashing.

16. A method according to claim 13 , wherein the generated index vectors are smaller in dimension than the generated content vectors.

17. A method according to claim 13 , wherein the bitset signatures are twice a length of the generated index vectors.

18. A method according to claim 13 , wherein the one or more terms in the search query and the terms corresponding to the identified content vectors are used to search the data set.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2024
From: BREAKWATER SOLUTIONS, LLC; REPARIO DATA, LLC
To: JETTEE, INC.
Reel/Frame 069124/0227 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 11, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: BREAKWATER SOLUTIONS LLC
Reel/Frame 058616/0384 →
NUNC PRO TUNC ASSIGNMENT Recorded Apr 2, 2014
From: STOREDIQ, INC.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 032584/0044 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2013
From: PONVERT, ELIAS; TRAN, MICHAEL TUYEN
To: STOREDIQ, INC.
Reel/Frame 030056/0896 →