IP Library Granted Patent US 8,538,896
Granted Patent B2
US 8,538,896 · App. 12/872,105 · Granted Sep 17, 2013

Retrieval systems and methods employing probabilistic cross-media relevance feedback

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,538,896
App. No.
12/872,105
Granted
Sep 17, 2013
Kind
B2
Abstract

In a retrieval application, a document relevance scoring function comprises a weighted combination of scoring components including at least one of a pseudo-relevance scoring component and a cross-media relevance scoring component. Weights of the document relevance scoring function are optimized to generate a trained document relevance scoring function. The optimizing is respective to a set of training documents including at least some multimedia training documents and a set of training queries and corresponding training document relevance annotations. A retrieval operation is performed for an input query respective to a database using the trained document relevance scoring function to retrieve one or more documents from the database.

Claims (112)

1. A non-transitory storage medium storing instructions executable by a digital processor to perform a method comprising:

optimizing weights of a document relevance scoring function to generate a trained document relevance scoring function ƒ(q,d) where q denotes a query and d denotes a document, wherein the document relevance scoring function comprises a weighted combination of scoring components including at least one pseudo-relevance scoring component and at least one cross-media relevance scoring component, and the optimizing is respective to a set of training documents including at least some multimedia training documents and a set of training queries and corresponding training document relevance annotations, the optimizing comprising optimizing a distribution-matchinq objective function respective to matching between:

a distribution

p

q

(

d

)

=

exp

(

f

(

d

,

q

)

)

d

D

exp

(

f

(

d

,

q

)

 of document relevance computed using the document relevance scoring function ƒ(q,d) for the set of training queries and the set of training documents D, and

a distribution

p

q

*

(

d

)

=

{

R

q

-

1

:

d

R

q

0

:

d

R

q

 of the training document relevance annotations corresponding to the set of training queries where R q is the set of documents relevant to training query q; and

performing a retrieval operation for an input query respective to a database using the trained document relevance scoring function to retrieve one or more documents from the database.

2. The non-transitory storage medium as set forth in claim 1 , wherein the document relevance scoring function ƒ(q,d) comprises a weighted linear combination of scoring components including at least one pseudo-relevance scoring component, at least one cross-media relevance scoring component, and at least one direct relevance scoring component.

3. The non-transitory storage medium as set forth in claim 1 , wherein:

the at least one pseudo-relevance scoring component of the document relevance scoring function ƒ(q,d) includes at least one of a pseudo-relevance textual scoring component and a pseudo-relevance image scoring component; and

the at least one cross-media relevance scoring component of the document relevance scoring function ƒ(q,d) includes at least one of a cross-media relevance scoring component having a text query modality and an image feedback modality and a cross-media relevance scoring component having an image query modality and a text feedback modality.

4. The method non-transitory storage medium as set forth in claim 1 , wherein the optimizing comprises:

training a classifier employing the document relevance scoring function ƒ(q,d) to predict document relevance for an input query, the training using the set of training documents and the set of training queries and corresponding training document relevance annotations.

5. The non-transitory storage medium as set forth in claim 1 , wherein the optimizing comprises:

training a classifier employing the document relevance scoring function ƒ(q,d) to predict a set of most relevant documents of the set of training documents for an input query, the training using the set of training documents and the set of training queries and corresponding training document relevance annotations.

6. The non-transitory storage medium as set forth in claim 1 , wherein the optimizing comprises:

for each training query of the set of training queries, computing training document relevance values for the training documents using the document relevance scoring function ƒ(q,d); and

scaling the computed training document relevance values using training query-dependent scaling factors.

7. A method comprising:

optimizing weights of a document relevance scoring function to generate a trained document relevance scoring function ƒ(q,d) where q denotes a query and d denotes a document, wherein the document relevance scoring function comprises a weighted combination of scoring components including at least one pseudo-relevance scoring component and at least one cross-media relevance scoring component, and the optimizing is respective to a set of training documents including at least some multimedia training documents and a set of training queries and corresponding training document relevance annotations, the optimizing including optimizing a distribution-matching objective function respective to matching between:

a distribution p q (d) of document relevance computed using the document relevance scoring function ƒ(q,d) for the set of training queries and the set of training documents D, and

a distribution p q *(d) of the training document relevance annotations corresponding to the set of training queries wherein p q *(d) is uniform over a set of documents R q that are relevant to training query q and zero for all other training documents, the optimizing further including:

for each training query of the set of training queries, computing training document relevance values for the training documents using the document relevance scoring function ƒ(q,d); and

scaling the computed training document relevance values using training query-dependent scaling factors, wherein the optimizing also optimizes the training query-dependent scaling factors; and

performing a retrieval operation for an input query respective to a database using the trained document relevance scoring function to retrieve one or more documents from the database;

wherein the optimizing and the performing are performed by a digital processor.

8. The non-transitory storage medium as set forth in claim 6 , wherein the training query-dependent scaling factors include a linear scaling factor α q for each training query.

9. The non-transitory storage medium as set forth in claim 8 , wherein the training query-dependent scaling factors further include an offset scaling factor β q for each training query.

10. The non-transitory storage medium as set forth in claim 6 , wherein the trained document relevance scoring function does not include the training query-dependent scaling factors.

11. The non-transitory storage medium as set forth in claim 6 , wherein each of the at least one pseudo-relevance scoring component and at least one cross-media relevance scoring component comprises a weighted linear combination of an ordered list of top k most similar feedback scoring sub-components weighted by a feedback ordinal position weighting, and the optimizing further comprises:

optimizing at least one parameter controlling the feedback ordinal position weighting.

12. The non-transitory storage medium as set forth in claim 1 , wherein each of the at least one pseudo-relevance scoring component and at least one cross-media relevance scoring component comprises a weighted linear combination of an ordered list of top k most similar feedback scoring sub-components weighted by a feedback ordinal position weighting, and the optimizing further comprises:

optimizing at least one parameter controlling the feedback ordinal position weighting.

13. The non-transitory storage medium as set forth in claim 1 , wherein the document relevance scoring function ƒ(q,d) comprises a weighted linear combination of scoring components including at least one pseudo-relevance scoring component and at least one cross-media relevance scoring component.

14. The non-transitory storage medium as set forth in claim 1 , wherein the database includes annotated images, the retrieval operation is performed for an input query image and the method further comprises:

constructing an annotation for the input query image based on annotations of annotated images retrieved from the database by the retrieval operation.

15. An apparatus comprising:

a digital processor configured to train a document relevance scoring function to generate a trained document relevance scoring function ƒ(q,d) where q denotes a query and d denotes a document, wherein the document relevance scoring function comprises a weighted linear combination of scoring components including at least one pseudo-relevance scoring component and at least one cross-media relevance scoring component, the training adjusts weights of the weighted linear combination of scoring components, and the training is respective to a set of training documents including at least some multimedia training documents and a set of training queries and corresponding training document relevance annotations, wherein the digital processor is configured to train the document relevance scoring function by a process including:

for each training query of the set of training queries, computing training document relevance values for the training documents using the document relevance scoring function ƒ(q,d);

scaling the computed training document relevance values using training query-dependent scaling factors; and

adjusting (i) weights of the weighted linear combination of scoring components and (ii) the training query-dependent scaling factors to optimize a distribution-matching objective function measuring an aggregate similarity between the computed training document relevance values and the corresponding training document relevance annotations, wherein the distribution-matching objective function is respective to matching between (1) a distribution p q (d) of document relevance computed using the document relevance scoring function ƒ(q,d) for the set of training queries and the set of training documents D, and (2) a distribution p q *(d) of the training document relevance annotations corresponding to the set of training queries wherein p q *(d) is uniform over a set of documents R q that are relevant to training query q and zero for all other training documents.

16. The apparatus as set forth in claim 15 , wherein:

the at least one pseudo-relevance scoring component of the document relevance scoring function includes at least one of a pseudo-relevance textual scoring component and a pseudo-relevance image scoring component;

the at least one cross-media relevance scoring component of the document relevance scoring function includes at least one of a cross-media relevance scoring component having a text query modality and an image feedback modality and a cross-media relevance scoring component having an image query modality and a text feedback modality; and

the linear combination of scoring components further includes at least one of a direct relevance text scoring component and a direct relevance image scoring component.

17. The apparatus as set forth in claim 15 , wherein the training query-dependent scaling factors include a linear scaling factor for each training query of the set of training queries.

18. The apparatus as set forth in claim 15 , wherein trained document relevance scoring function does not include the training query-dependent scaling factors, and the digital processor is further configured to perform a retrieval operation using the trained document relevance scoring function.

19. The method as set forth in claim 7 , wherein the training query-dependent scaling factors include a linear scaling factor α q for each training query.

20. The method as set forth in claim 19 , wherein the training query-dependent scaling factors further include an offset scaling factor β q for each training query.

21. The method as set forth in claim 7 , wherein the trained document relevance scoring function does not include the training query-dependent scaling factors.

22. The method as set forth in claim 7 , wherein each of the at least one pseudo-relevance scoring component and at least one cross-media relevance scoring component comprises a weighted linear combination of an ordered list of top k most similar feedback scoring sub-components weighted by a feedback ordinal position weighting, and the optimizing further comprises:

optimizing at least one parameter controlling the feedback ordinal position weighting.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2025
From: XEROX CORPORATION
To: GENESEE VALLEY INNOVATIONS, LLC
Reel/Frame 073842/0479 →
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
RELEASE OF SECURITY INTEREST IN PATENTS AT R/F 062740/0214 Recorded May 18, 2023
From: CITIBANK, N.A., AS AGENT
To: XEROX CORPORATION
Reel/Frame 063694/0122 →
SECURITY INTEREST Recorded Nov 10, 2022
From: XEROX CORPORATION
To: CITIBANK, N.A., AS AGENT
Reel/Frame 062740/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2011
From: MENSINK, THOMAS; CSURKA, GABRIELA; VERBEEK, JAKOB
To: XEROX CORPORATION
Reel/Frame 025793/0314 →