Enhanced context aware content retrieval using logical modeling
Context aware content retrieval can be performed and enhanced. Document manager (DM) can determine respective full, summary, and/or partial versions of respective electronic documents, and respective VI scores and respective token sizes associated therewith, based on analysis of the documents and initial queries. In connection with a received query, from the group of respective versions of respective electronic documents, DM can determine a group of respective candidate versions of respective electronic documents based on respective VI scores and token sizes, first VI score criterion, and first threshold token size. From the candidate group, DM can determine a subgroup of respective versions of respective electronic documents based on respective VI scores and token sizes, second VI score criterion, and second threshold token size that can correspond to a maximum token size limit that can be input into an AI-based model. Updates can be performed based on feedback information.
1 . A method, comprising:
with regard to a group of respective versions of respective electronic documents, determining, by a system comprising at least one processor, respective value information scores associated with the respective versions of respective electronic documents based on a result of a first analysis of the respective versions of respective electronic documents, wherein the group of respective versions of respective electronic documents comprises respective first versions of respective electronic documents and respective second versions of respective electronic documents that are variations of the respective first versions of respective electronic documents; and
in connection with a query, and from the group of respective versions of respective electronic documents, determining, by the system, a subgroup of respective versions of respective electronic documents to input to an artificial intelligence-based model for a second analysis based on a threshold token size associated with the artificial intelligence-based model and the respective value information scores and respective token sizes associated with the respective versions of respective electronic documents, wherein the subgroup of respective versions of respective electronic documents is input into the artificial intelligence-based model for the second analysis to facilitate determining a response to the query,
wherein the threshold token size indicates a total token size of tokens that are able to be input to the artificial intelligence-based model in connection with the query,
wherein the group of respective versions of respective electronic documents comprises the subgroup of respective versions of respective electronic documents and other respective subgroups of respective versions of respective electronic documents that are determined to have respective total token sizes that satisfy the threshold token size, and
wherein the subgroup of respective versions of respective electronic documents has a total value information score that is higher than other respective total value information scores of the other respective subgroups of respective versions of respective electronic documents.
2 . The method of claim 1 , wherein the respective first versions of respective electronic documents are respective full versions of respective electronic documents, and wherein the respective second versions of respective electronic documents comprise respective partial versions of respective electronic documents and respective summary versions of respective electronic documents that are respective variations of, and correspond to, the respective full versions of respective electronic documents.
3 . The method of claim 2 , wherein the result is a first result, and wherein the method further comprises:
generating, by the system, the respective partial versions of respective electronic documents and the respective summary versions of respective electronic documents based on a second result of a third analysis of the respective full versions of respective electronic documents.
4 . The method of claim 3 , further comprising:
determining, by the system, respective initial queries, relating to the respective full versions of respective electronic documents, that have been determined to satisfy a defined threshold likelihood of being received by the system,
wherein the generating of the respective partial versions of respective electronic documents and the respective summary versions of respective electronic documents comprises generating the respective partial versions of respective electronic documents and the respective summary versions of respective electronic documents based on the second result of the third analysis of the respective full versions of respective electronic documents or the respective initial queries, and
wherein the determining of the respective value information scores comprises determining the respective value information scores associated with the respective full versions of respective electronic documents, the respective partial versions of respective electronic documents, and the respective summary versions of respective electronic documents based on respective contexts, respective intents, or respective keywords associated with the respective initial queries.
5 . The method of claim 2 , wherein the result is a first result, and wherein the method further comprises:
determining, by the system, respective first token sizes associated with the respective full versions of respective electronic documents, respective second token sizes associated with the respective partial versions of respective electronic documents, and respective third token sizes associated with the respective summary versions of respective electronic documents based on a second result of a third analysis of the respective full versions of respective electronic documents, the respective partial versions of respective electronic documents, and the respective summary versions of respective electronic documents.
6 . The method of claim 2 , wherein the threshold token size is a second threshold token size, wherein a first threshold token size is greater than the second threshold token size, wherein the subgroup of respective versions of respective electronic documents is a second subgroup of respective versions of respective electronic documents, and wherein the method further comprises:
determining, by the system, that the group of respective versions of respective electronic documents is potentially responsive to the query based on a determination of a context, an intent, or a group of keywords of the query,
wherein the determining of the second subgroup of respective versions of respective electronic documents from the group of respective versions of respective electronic documents comprises: from the group of respective versions of respective electronic documents, determining respective first subgroups of respective versions of respective electronic documents that satisfy the first threshold token size and are associated with respective total value information scores that satisfy a defined value information score criterion based on the respective token sizes and the respective value information scores associated with the respective versions of respective electronic documents of the group of respective versions of respective electronic documents,
wherein the respective first subgroups of respective versions of respective electronic documents comprise the second subgroup of respective versions of respective electronic documents and the other subgroups of respective versions of respective electronic documents, and
wherein the respective first subgroups of respective versions of respective electronic documents comprise respective combinations of the respective full versions of respective electronic documents, the respective partial versions of respective electronic documents, or the respective summary versions of respective electronic documents.
7 . The method of claim 6 , wherein the respective first subgroups of respective versions of respective electronic documents are associated with the respective total value information scores that are determined based on the respective value information scores associated with the respective versions of respective electronic documents of the respective first subgroups,
wherein the respective first subgroups of respective versions of respective electronic documents are associated with the respective total token sizes that are determined based on the respective token sizes associated with the respective versions of respective electronic documents of the respective first subgroups, and
wherein the determining of the second subgroup of respective versions of respective electronic documents comprises: from the respective first subgroups of respective versions of respective electronic documents, determining the second subgroup of respective versions of respective electronic documents based on a combined token size of a token size associated with the query and the total token size associated with the second subgroup being determined to satisfy the second threshold token size, and based on the total value information score associated with the second subgroup being determined to be higher than the other respective total value information scores associated with the other respective first subgroups of respective versions of respective electronic documents.
8 . The method of claim 1 , wherein the respective value information scores comprise a value information score associated with a version of an electronic document of the respective versions of the respective electronic documents, and wherein the method further comprises:
determining, by the system, the value information score associated with the version of the electronic document based on a similarity evaluation score associated with the version of the electronic document, a first score weight associated with the similarity evaluation score, a semantic score associated with the version of the electronic document, and a second score weight associated with the semantic score, wherein the similarity evaluation score relates to a textual similarity between the version of the electronic document and a reference electronic document.
9 . The method of claim 8 , wherein the similarity evaluation score is a recall-oriented understudy for gisting evaluation score, a bidirectional-encoder-representations-from-transformers score, or a metric-for-evaluation-of-translation-with-explicit-ordering score.
10 . The method of claim 1 , further comprising:
inputting, by the system, the subgroup of respective versions of respective electronic documents into the artificial intelligence-based model for the second analysis of the subgroup of respective versions of respective electronic documents using the artificial intelligence-based model; and
presenting, by the system, the response to the query to a device or a user, wherein the response to the query is determined based on the second analysis of the subgroup of respective versions of respective electronic documents using the artificial intelligence-based model.
11 . The method of claim 10 , wherein the query is a first query, wherein the respective versions of respective electronic documents comprise a version of an electronic document, wherein the version of the electronic document is a full version of the electronic document, a first partial version of the electronic document, or a first summary version of the electronic document, and wherein the method further comprises:
receiving, by the system, feedback information, relating to the response to the query, from the device or the user; and
based on the feedback information, at least one of:
generating, by the system, a second query relating to at least some of the respective versions of respective electronic documents;
modifying, by the system, an initial query relating to at least some of the respective versions of respective electronic documents, wherein the initial query had been generated by the system;
modifying, by the system, a value information score associated with the version of the respective electronic document;
modifying, by the system, the partial version of the electronic document or the summary version of the electronic document; or
from the full version of the electronic document, generating, by the system, a second partial version of the electronic document or a second summary version of the electronic document.
12 . A system, comprising:
at least one memory that stores computer executable components; and
at least one processor that executes computer executable components stored in the at least one memory, wherein the computer executable components comprise:
a value information determinator that, with regard to a group of respective forms of respective electronic documents, determines respective value information scores associated with the respective forms of respective electronic documents based on a result of a first analysis of the respective forms of respective electronic documents, wherein the group of respective forms of respective electronic documents comprises respective first forms of respective electronic documents and respective second forms of respective electronic documents that are derived from the respective first forms of respective electronic documents; and
a document selector, wherein, in response to receiving a query, and from the group of respective forms of respective electronic documents, the document selector determines a portion of the respective forms of respective electronic documents to input to an artificial intelligence-based model for a second analysis based on a threshold token size associated with the artificial intelligence-based model and the respective value information scores and respective token sizes associated with the respective forms of respective electronic documents, wherein the portion of the respective forms of respective electronic documents is input into the artificial intelligence-based model for the second analysis to facilitate a determination of a response to the query,
wherein the threshold token size indicates an overall token size of tokens that are able to be input to the artificial intelligence-based model with respect to the query,
wherein the group of respective forms of respective electronic documents comprises the portion of respective forms of respective electronic documents and other respective portions of respective forms of respective electronic documents that are determined to have respective overall token sizes that satisfy the threshold token size, and
wherein the portion of respective forms of respective electronic documents has an overall value information score that is determined to be greater than other respective overall value information scores of the other respective portions of respective versions of respective electronic documents.
13 . The system of claim 12 , wherein the respective first forms of respective electronic documents are respective full forms of respective electronic documents, and wherein the respective second forms of respective electronic documents comprise respective partial forms of respective electronic documents or respective summary forms of respective electronic documents that are derived from and correspond to the respective full forms of respective electronic documents.
14 . The system of claim 13 , wherein the result is a first result, and wherein the computer executable components further comprise:
a document processor that determines and generates the respective partial forms of respective electronic documents or the respective summary forms of respective electronic documents based on a second result of a third analysis of the respective full forms of respective electronic documents.
15 . The system of claim 14 , wherein the computer executable components further comprise:
a query generator that determines respective initial queries, relating to the respective full forms of respective electronic documents, that at least satisfy a defined likelihood of being received by the system,
wherein the document processor generates the respective partial forms of respective electronic documents or the respective summary forms of respective electronic documents based on the second result of the third analysis of the respective full forms of respective electronic documents or the respective initial queries, and
wherein the value information determinator determines the respective value information scores associated with the respective full forms of respective electronic documents, the respective partial forms of respective electronic documents, or the respective summary forms of respective electronic documents based on respective contexts, respective intents, or respective keywords associated with the respective initial queries.
16 . The system of claim 13 , wherein the result is a first result, and wherein the computer executable components further comprise:
a token size determinator that determines respective first token sizes associated with the respective full forms of respective electronic documents, respective second token sizes associated with the respective partial forms of respective electronic documents, or respective third token sizes associated with the respective summary forms of respective electronic documents based on a second result of a third analysis of the respective full forms of respective electronic documents, the respective partial forms of respective electronic documents, or the respective summary forms of respective electronic documents.
17 . The system of claim 13 , wherein the threshold token size is a second threshold token size, wherein a first threshold token size is larger than the second threshold token size, wherein the portion of respective forms of respective electronic documents is a second portion of respective forms of respective electronic documents,
wherein the document selector determines that the group of respective forms of respective electronic documents is potentially responsive to the query based on a determination of a context, an intent, or a group of keywords of the query,
wherein, from the group of respective forms of respective electronic documents, the document selector determines respective first portions of respective forms of respective electronic documents that satisfy the first threshold token size and are associated with the respective overall value information scores that satisfy a defined value information score criterion based on the respective token sizes and the respective value information scores associated with the respective forms of respective electronic documents of the group of respective forms of respective electronic documents,
wherein the respective first portions of respective forms of respective electronic documents comprise the second portion of respective forms of respective electronic documents and the other portions of respective forms of respective electronic documents, and
wherein the respective first portions of respective forms of respective electronic documents comprise respective combinations of the respective full forms of respective electronic documents, the respective partial forms of respective electronic documents, or the respective summary forms of respective electronic documents.
18 . The system of claim 17 , wherein the respective first portions of respective forms of respective electronic documents are associated with the respective overall value information scores that are determined based on the respective value information scores associated with the respective forms of respective electronic documents of the respective first portions,
wherein the respective first portions of respective forms of respective electronic documents are associated with the respective overall token sizes that are determined based on the respective token sizes associated with the respective forms of respective electronic documents of the respective first subgroups, and
wherein, from the respective first portions of respective forms of respective electronic documents, the document selector determines and selects the second s portion of respective forms of respective electronic documents based on a total of a token size associated with the query and the overall token size associated with the second portion being determined to satisfy the second threshold token size, and based on the overall value information score associated with the second s portion being determined to be greater than the other respective overall value information scores associated with the other respective first portions of respective forms of respective electronic documents.
19 . A non-transitory machine-readable medium, comprising executable instructions that, when executed by at least one processor, facilitate performance of operations, comprising:
with regard to a group of respective versions of respective electronic documents, generating respective value scores associated with the respective versions of respective electronic documents of the group based on a result of a first analysis of the respective versions of respective electronic documents, wherein the group of respective versions of respective electronic documents comprises respective first versions of respective electronic documents and respective second versions of respective electronic documents that are derived based on the respective first versions of respective electronic documents; and
in response to receiving a query, and from the group of respective versions of respective electronic documents, selecting a subgroup of respective versions of respective electronic documents for a second analysis using an artificial intelligence-based model based on a threshold token size associated with the artificial intelligence-based model and based on the respective value scores and respective token sizes associated with the respective versions of respective electronic documents, wherein the subgroup of respective versions of respective electronic documents is input into the artificial intelligence-based model for the second analysis to facilitate generation of a response to the query,
wherein the threshold token size indicates a total token size of tokens that are able to be input to the artificial intelligence-based model with respect to the query,
wherein the group of respective versions of respective electronic documents comprises the subgroup of respective versions of respective electronic documents and other respective subgroups of respective versions of respective electronic documents that are determined to have respective total token sizes that satisfy the threshold token size, and
wherein the subgroup of respective versions of respective electronic documents has a total value score that is determined to be higher than other respective total value scores of the other respective subgroups of respective versions of respective electronic documents.
20 . The non-transitory machine-readable medium of claim 19 , wherein the result is a first result, and wherein the operations further comprise:
determining respective initial queries, relating to respective full versions of respective electronic documents, that potentially will be received; and
generating respective partial versions of respective electronic documents or respective summary versions of respective electronic documents based on a second result of a third analysis of the respective full versions of respective electronic documents or the respective initial queries, wherein the respective first versions of respective electronic documents are the respective full versions of respective electronic documents, wherein the respective second versions of respective electronic documents comprise the respective partial versions of respective electronic documents or the respective summary versions of respective electronic documents,
wherein the generating of the respective value scores comprises determining the respective value scores associated with the respective full versions of respective electronic documents, the respective partial versions of respective electronic documents, or the respective summary versions of respective electronic documents based on respective contexts, respective intents, or respective keywords associated with the respective initial queries,
wherein the selecting comprises: from the respective full versions of respective electronic documents, the respective partial versions of respective electronic documents, or the respective summary versions of respective electronic documents, selecting the subgroup of respective versions of respective electronic documents for the second analysis using the artificial intelligence-based model based on the threshold token size, based on the respective value scores and the respective token sizes, and based on a token size associated with the query, wherein the subgroup of respective versions of respective electronic documents is determined to have the total value score that is determined to be higher than the other total value scores associated with the other subgroups of respective versions of respective electronic documents, and wherein a combined token size associated with the query and the subgroup of respective versions of respective electronic documents is determined to satisfy the threshold token size.