IP Library › Granted Patent US 11,487,827
Granted Patent B2
US 11,487,827 · App. 16/234,537 · Granted Nov 1, 2022

Extended query performance prediction framework utilizing passage-level information

Inventor: Haggai Roitman (Yokneam Illit, IL)
Assignee: International Business Machines Corporation
G06F16/93G06F16/90335
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,487,827
App. No.
16/234,537
Granted
Nov 1, 2022
Kind
B2
Abstract

An illustrative embodiment includes a method for post-retrieval query performance prediction using hybrid document-passage information. The method includes: obtaining a set of documents responsive to a specific query; extracting document-level information regarding respective documents within the set; extracting passage-level information regarding respective passages of documents within the set; and estimating a likelihood that the set of documents includes relevant information to the specific query using both the document-level information and the passage-level information.

Claims (56)

1. A method for post-retrieval query performance prediction using hybrid document-passage information, the method comprising:

obtaining a set of documents of a corpus of documents;

extracting document-level information regarding respective documents within the set;

extracting passage-level information regarding respective passages of a proper subset of documents within the document corpus, based on a probability of the proper subset of documents including relevant information independent of a specific query; and

estimating a likelihood that the proper subset of documents within the document corpus includes relevant information to the specific query by using the passage-level information for the proper subset of documents of the document corpus, the proper subset of documents retrieved using the document-level information, the proper subset of documents being the top-k documents relative to the specific query.

2. A method for post-retrieval query performance prediction using hybrid document-passage information, the method comprising:

with a computerized information retrieval system, obtaining a set of documents responsive to a specific query;

with the computerized information retrieval system, extracting document-level information regarding respective documents within the set;

with the computerized information retrieval system, extracting passage-level information regarding respective passages of documents within the set;

with the computerized information retrieval system, estimating a likelihood that the set of documents includes relevant information to the specific query using both the document-level information and the passage-level information;

with the computerized information retrieval system, estimating one or more score calibration signals using the passage-level information; and

with the computerized information retrieval system, estimating the likelihood that the set of documents includes relevant information to the specific query based at least in part on the one or more score-calibration signals estimated using the passage-level information.

3. The method of claim 2 , further comprising:

generating a WPM2 weighted product model using the document-level information; and

estimating the likelihood that the set of documents includes relevant information to the specific query based on the WPM2 weighted product model using the document-level information and the one or more score-calibration signals using the passage-level information.

4. The method of claim 2 , wherein a first of the score-calibration signals estimated using the passage-level information denotes a likelihood, estimated using a representative passage for a given document, that the given document within the set includes relevant information regardless of the specific query.

5. The method of claim 4 , wherein the representative passage for the given document is a single passage within the given document having a highest retrieval score for the specific query.

6. The method of claim 4 , wherein the representative passage is extracted from the given document using a first window size when the specific query has a first length, and wherein the representative passage is extracted from the given document using a second window size when the specific query has a second length, the first query length being less than the second length, and the first window size being greater than the second window size.

7. The method of claim 4 , wherein estimating the likelihood that the given document within the set includes relevant information regardless of the specific query comprises:

estimating a likelihood that the representative passage includes the relevant information; and

estimating a relationship between the representative passage and the given document.

8. The method of claim 7 , wherein the likelihood that the representative passage includes the relevant information is estimated as a combination of:

a language model entropy of the representative passage; and

a position of the representative passage within the given document.

9. The method of claim 7 , wherein the likelihood that the representative passage includes the relevant information is estimated so as to prefer that the representative passage be more diverse and be located earlier within the given document.

10. The method of claim 4 , wherein the likelihood that a set of documents includes relevant information to a specific query is estimated further based at least in part on a normalization term estimated based on a length of the specific query.

11. The method of claim 10 , wherein estimating the likelihood that the set of documents includes relevant information to the specific query based at least in part on the estimated one or more score-calibration signals using passage-level information comprises calibrating respective weights assigned at least to the first score calibration signal and to the normalization term.

12. The method of claim 4 , wherein a second of the score-calibration signals is estimated using the representative passage for the given document and captures a relationship between the given document and the set of documents.

13. The method of claim 12 , wherein estimating the second score-calibration signal comprises:

estimating a similarity of the representative passage to the given document; and

estimating a similarity of the representative passage to a relevance model for the set of documents.

14. The method of claim 12 , wherein a third of the score-calibration signals is estimated using at least one representative passage for the set of documents and denotes a likelihood that the set of documents includes relevant information regardless of the specific query.

15. The method of claim 14 :

wherein the representative passage for the given document is a single passage within the given document having a highest retrieval score for the specific query, and

wherein the representative passage for the set of documents is a single passage within the set of documents having a highest retrieval score for the specific query.

16. The method of claim 14 , wherein the third of the score-calibration signals is estimated using an average or a standard deviation for a set of passages within the set of documents selected using respective retrieval scores for the set of passages.

17. The method of claim 14 , wherein estimating the likelihood that the set of documents within includes relevant information regardless of the specific query comprises:

estimating a likelihood that the representative passage for the set of documents includes the relevant information; and

estimating a similarity between the representative passage for the set of documents and a centroid language model for the set of documents.

18. The method of claim 14 , wherein estimating the likelihood that the set of documents includes relevant information to the specific query based at least in part on the estimated one or more score-calibration signals using passage-level information comprises calibrating respective weights assigned at least to the first score calibration signal, the second score calibration signal, and the third score calibration signal.

19. An apparatus for post-retrieval query performance prediction using hybrid document-passage information, the apparatus comprising:

a memory; and

at least one processor coupled to the memory, the processor being operative:

to obtain a set of documents responsive to a specific query;

to extract document-level information regarding respective documents within the set;

to extract passage-level information regarding respective passages of documents within the set;

to estimate a likelihood that the set of documents includes relevant information to the specific query using both the document-level information and the passage-level information;

to estimate one or more score calibration signals using the passage-level information; and

to estimate the likelihood that the set of documents includes relevant information to the specific query based at least in part on the one or more score-calibration signals estimated using the passage-level information.

20. A computer program product comprising a non-transitory machine-readable storage medium having machine-readable program code embodied therewith, said machine-readable program code comprising machine-readable program code configured:

to obtain a set of documents responsive to a specific query;

to extract document-level information regarding respective documents within the set;

to extract passage-level information regarding respective passages of documents within the set;

to estimate a likelihood that the set of documents includes relevant information to the specific query using both the document-level information and the passage-level information;

to estimate one or more score calibration signals using the passage-level information; and

to estimate the likelihood that the set of documents includes relevant information to the specific query based at least in part on the one or more score-calibration signals estimated using the passage-level information.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2026
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: ANTHROPIC, PBC
Reel/Frame 075991/0666 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2018
From: ROITMAN, HAGGAI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047864/0020 →
Continuity (1)
Related Publication 20200210489A1 · Jul 2, 2020