IP Library Granted Patent US 8,612,457
Granted Patent B2
US 8,612,457 · App. 13/073,836 · Granted Dec 17, 2013

Method and system for comparing documents based on different document-similarity calculation methods using adaptive weighting

Inventors: Oliver Brdiczka (Mountain View, CA); Petro Hizalev (Palo Alto, CA)
Assignee: Palo Alto Research Center Incorporated
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,612,457
App. No.
13/073,836
Filed
Mar 28, 2011
Granted
Dec 17, 2013
Kind
B2
Art Unit
2168
USPC
707/728
Abstract

One embodiment provides a system for comparing documents based on different document-similarity calculation methods using adaptive weighting. During operation, the system receives at least two document-similarity values associated with two documents, wherein the document-similarity values are calculated by different document-similarity calculation methods. The system then determines the weight of a respective document-similarity calculation method for each of the two documents, as well as a weight-combination function for calculating a combined weight of the respective document-similarity calculation method associated with the two documents. Next, the system generates a combined similarity value based on the document-similarity values and the weight-combination function.

Claims (131)

1. A computer-executable method for comparing documents, the method comprising:

receiving, by a computer, a number of document-similarity values s i (a,b) associated with two documents, the document-similarity values being calculated by different document-similarity calculation methods, which are indexed by i;

determining weights (α i , β i ) of a respective document-similarity calculation method i for each of the two documents a and b respectively, wherein α i corresponds to document a and β i corresponds to document b, wherein determining the weights of a document-similarity calculation method involves:

initializing the weights; and

updating the weights based on user feedback, wherein updating the weights of the document-similarity calculation method for each document involves applying a learning algorithm to the weights based on user feedback;

determining a weight-combination function ƒ(α i ,β i ) for calculating a combined weight of the respective document-similarity calculation method i associated with the two documents, wherein the weight-combination function accounts for the direction of comparison between the documents a and b; and

generating, by the computer, a combined similarity value S(a,b) based on the document-similarity values calculated by k different document-similarity calculation methods and the weight-combination function, given by:

S

(

a

,

b

)

=

i

=

1

k

f

(

α

i

,

β

i

)

·

s

i

(

a

,

b

)

.

2. The method of claim 1 , wherein the document-similarity calculation methods include one or more of: text-based, visual-based, usage-based, and social network-based document-similarity calculation methods.

3. The method of claim 1 , wherein the weight of the document-similarity calculation method for each document is initialized based at least on one of: document type, document location, document structure, document usage, and weights of related documents.

4. The method of claim 1 , wherein the weight-combination function is a real-valued function that calculates an average, a minimum, or a maximum of the weights.

5. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for comparing documents, the method comprising:

receiving a number of document-similarity values s i (a,b) associated with two documents, the document-similarity values being calculated by different document-similarity calculation methods, which are indexed by i;

determining weights (α i ,β i ) of a respective document-similarity calculation method i for each of the two documents a and b respectively, wherein α i corresponds to document a and β i corresponds to document b, wherein determining the weights of a document-similarity calculation method involves:

initializing the weights; and

updating the weights based on user feedback, wherein updating the weights of the document-similarity calculation method for each document involves applying a learning algorithm to the weights based on user feedback;

determining a weight-combination function ƒ(α i ,β i ) for calculating a combined weight of the respective document-similarity calculation method i associated with the two documents, wherein the weight-combination function accounts for the direction of comparison between the documents a and b; and

generating a combined similarity value S(a,b) based on the document-similarity values calculated by k different document-similarity calculation methods and the weight-combination function, given by:

S

(

a

,

b

)

=

i

=

1

k

f

(

α

i

,

β

i

)

·

s

i

(

a

,

b

)

.

6. The computer-readable storage medium of claim 5 , wherein the document-similarity calculation methods include one or more of: text-based, visual-based, usage-based, and social network-based document-similarity calculation methods.

7. The computer-readable storage medium of claim 5 , wherein the weight of the document-similarity calculation method for each document is initialized based at least on one of: document type, document location, document structure, document usage, and weights of related documents.

8. The computer-readable storage medium of claim 5 , wherein the weight-combination function is a real-valued function that calculates an average, a minimum, or a maximum of the weights.

9. A system for estimating a similarity level between documents, comprising:

at least one hardware processor couple to at least one memory for estimating the similarity level between documents;

a receiving mechanism configured to receive a number of document-similarity values s i (a,b) associated with two documents, the document-similarity values being calculated by different document-similarity calculation methods, which are indexed by i;

a determination mechanism configured to determine:

weights (α i ,β i ) of a respective document-similarity calculation method i for each of the two documents a and b respectively, wherein α i corresponds to document a and β i corresponds to document b, wherein determining the weights of a document-similarity calculation method involves:

initializing the weights; and

updating the weights based on user feedback, wherein updating the weights of the document-similarity calculation method for each document involves applying a learning algorithm to the weights based on user feedback; and

a weight-combination function ƒ(α i ,β i ) for calculating a combined weight of the respective document-similarity calculation method i associated with the two documents, wherein the weight-combination function accounts for the direction of comparison between the documents a and b; and

generating a combined similarity value S(a,b) based on the document-similarity values calculated by k different document-similarity calculation methods and the weight-combination function, given by:

S

(

a

,

b

)

=

i

=

1

k

f

(

α

i

,

β

i

)

·

s

i

(

a

,

b

)

.

10. The system of claim 9 , wherein the document-similarity calculation methods include one or more of: text-based, visual-based, usage-based, and social network-based document-similarity calculation methods.

11. The system of claim 9 , wherein the weight of the document-similarity calculation method for each document is initialized based at least on one of: document type, document location, document structure, document usage, and weights of related documents.

12. The system of claim 9 , wherein the weight-combination function is a real-valued function that calculates an average, a minimum, or a maximum of the weights.

Assignments (9)
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVAL OF US PATENTS 9356603, 10026651, 10626048 AND INCLUSION OF US PATENT 7167871 PREVIOUSLY RECORDED ON REEL 064038 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064161/0001 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064038/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2011
From: BRDICZKA, OLIVER; HIZALEV, PETRO
To: PALO ALTO RESEARCH CENTER INCORPORATED
Reel/Frame 026043/0187 →
Continuity (1)
Related Publication 20120254165A1 · Oct 4, 2012