Extended query performance prediction framework utilizing passage-level information
View Patent ↗An illustrative embodiment includes a method for post-retrieval query performance prediction using hybrid document-passage information. The method includes: obtaining a set of documents responsive to a specific query; extracting document-level information regarding respective documents within the set; extracting passage-level information regarding respective passages of documents within the set; and estimating a likelihood that the set of documents includes relevant information to the specific query using both the document-level information and the passage-level information.
1. An apparatus for post-retrieval query performance prediction using hybrid document-passage information, the apparatus comprising:
a memory; and
at least one processor coupled to the memory, the processor being operative:
to obtain a set of documents of a corpus of documents responsive to a specific query;
to extract document-level information regarding respective documents within the set;
to extract passage-level information regarding respective passages of a proper subset of documents within the document corpus, based on a probability of the proper subset of documents including relevant information independent of a specific query;
to estimate a likelihood that the proper subset of documents within the document corpus set of documents includes relevant information to the specific query by using both the document-level information and the passage-level information for the proper subset of documents of the document corpus, the proper subset of documents retrieved using the document-level information, the proper subset of documents being the top-k documents relative to the specific query.
2. A computer program product comprising a non-transitory machine-readable storage medium having machine-readable program code embodied therewith, said machine-readable program code comprising machine-readable program code configured:
to obtain a set of documents of a corpus of documents responsive to a specific query;
to extract document-level information regarding respective documents within the set;
to extract passage-level information regarding respective passages of a proper subset of documents within the document corpus, based on a probability of the proper subset of documents including relevant information independent of a specific query;
to estimate a likelihood that the proper subset of documents within the document corpus set of documents includes relevant information to the specific query by using both the document-level information and the passage-level information for the proper subset of documents of the document corpus, the proper subset of documents retrieved using the document-level information, the proper subset of documents being the top-k documents relative to the specific query.