Constructing a search index based on an automatically-selected set of document properties
A facility is described for automatically selecting document properties to be included in an index on a corpus of documents. In some cases, the facility performs this automatic selection by subjecting the documents of the corpus to an inference model. In some cases, the facility performs this automatic selection by submitting to a large language model a prompt automatically constructed to include, for each of one or more example documents, (1) the example document's contents, and (2) a list of properties of the sample document that should be indexed.
1 . One or more memory devices collectively having contents configured to cause a computing system to perform a method, the method comprising:
accessing one or more training documents, each of the training documents having contents;
accessing, for each of the training documents, a list of properties of the training document that should be indexed;
for each of the accessed training documents, constructing a training example comprising (1) the training document's contents and (2) the list of properties of the training document that should be indexed;
accessing a corpus of documents to be indexed;
constructing and submitting a textual prompt to a large language model, wherein the textual prompt specifies (1) the constructed training examples, (2) the accessed corpus of documents, (3) an indication that the training examples should be used as a basis to identify in the corpus of documents properties that should be used to index documents of the corpus, and (4) a request for a response that specifies at least the identified properties that should be used to index documents of the corpus;
receiving from the large language model the response to the prompt that identifies, for each of the documents of the corpus, a list of properties that are automatically selected by the large language model and proposed to be a basis for an index should be included in an index on the corpus for the document of the corpus; and
selecting the properties identified in the response received from the large language model to construct, at least in part, an index on the documents of the corpus.
2 . The one or more memory devices of claim 1 wherein the accessed lists of properties include properties that would be common to search on or filter on for the document of the corpus.
3 . The one or more memory devices of claim 1 , the method further comprising:
constructing the index on the documents of the corpus that is based on the selected properties.
4 . The one or more memory devices of claim 3 , the method further comprising:
before the constructing:
causing the properties identified in the response received from the large language model to be presented to a user; and
receiving input originated by the user:
selecting one or more unselected document properties,
unselecting one or more selected document properties, or
selecting one or more unselected document properties and unselecting one or more selected document properties; and
adjusting selection among the document properties in accordance with the received input.
5 . The one or more memory devices of claim 3 wherein the constructed index is a tree structure mapping from the selected properties to the documents of the corpus having those properties.
6 . The one or more memory devices of claim 3 wherein the constructing comprises:
for each document of the corpus,
concatenating the document's contents of the selected properties into a representation of the document;
encoding the representation of the document into a vector wherein the constructed index maps from the vectors produced by the encoding to the corresponding document of the corpus.
7 . The one or more memory devices of claim 3 , the method further comprising:
using the constructed index to conduct a search of the corpus.
8 . The one or more memory devices of claim 3 , the method further comprising:
using the constructed index to service a chatbot interaction with respect to the corpus.
9 . A computer-implemented method comprising:
accessing one or more training documents, each of the training documents having contents;
accessing, for each of the training documents, a list of properties of the training document that should be indexed;
for each of the accessed training documents, constructing a training example comprising (1) the training document's contents and (2) the list of properties of the training document that should be indexed;
accessing a corpus of documents to be indexed;
constructing and submitting a textual prompt to a large language model, wherein the textual prompt specifies (1) the constructed training examples, (2) the accessed corpus of documents, (3) an indication that the training examples should be used as a basis to identify in the corpus of documents properties that should be used to index documents of the corpus, and (4) a request for a response that specifies at least the identified properties that should be used to index documents of the corpus;
receiving from the large language model the response to the prompt that identifies, for each of the documents of the corpus, a list of properties that are automatically selected by the large language model and proposed to be a basis for an index should be included in an index on the corpus for the document of the corpus; and
selecting the properties identified in the response received from the large language model to construct, at least in part, an index on the documents of the corpus.
10 . The method of claim 9 wherein the accessed lists of properties include properties that would be common to search on or filter on for the document of the corpus.
11 . The method of claim 9 , further comprising:
constructing the index on the documents of the corpus that is based on the selected properties.
12 . The method of claim 11 , further comprising:
before the constructing:
causing the properties identified in the response received from the large language model to be presented to a user; and
receiving input originated by the user:
selecting one or more unselected document properties,
unselecting one or more selected document properties, or
selecting one or more unselected document properties and unselecting one or more selected document properties; and
adjusting selection among the document properties in accordance with the received input.
13 . The method of claim 11 wherein the constructed index is a tree structure mapping from the selected properties to the documents of the corpus having those properties.
14 . The method of claim 11 wherein the constructing comprises:
for each document of the corpus,
concatenating the document's contents of the selected properties into a representation of the document;
encoding the representation of the document into a vector wherein the constructed index maps from the vectors produced by the encoding to the corresponding document of the corpus.
15 . The method of claim 11 , further comprising:
using the constructed index to conduct a search of the corpus.
16 . The method of claim 11 , further comprising:
using the constructed index to service a chatbot interaction with respect to the corpus.
17 . A system, comprising:
one or more processors; and
memory storing content that, when executed by the one or more processors, cause the system to perform a method comprising:
accessing one or more training documents, each of the training documents having contents;
accessing, for each of the training documents, a list of properties of the training document that should be indexed;
for each of the accessed training documents, constructing a training example comprising (1) the training document's contents and (2) the list of properties of the training document that should be indexed;
accessing a corpus of documents to be indexed;
constructing and submitting a textual prompt to a large language model, wherein the textual prompt specifies (1) the constructed training examples, (2) the accessed corpus of documents, (3) an indication that the training examples should be used as a basis to identify in the corpus of documents properties that should be used to index documents of the corpus, and (4) a request for a response that specifies at least the identified properties that should be used to index documents of the corpus;
receiving from the large language model the response to the prompt that identifies, for each of the documents of the corpus, a list of properties that are automatically selected by the large language model and proposed to be a basis for an index should be included in an index on the corpus for the document of the corpus; and
selecting the properties identified in the response received from the large language model to construct, at least in part, an index on the documents of the corpus.
18 . The system of claim 17 wherein the accessed lists of properties include properties that would be common to search on or filter on for the document of the corpus.
19 . The system of claim 17 , the method further comprising:
constructing the index on the documents of the corpus that is based on the selected properties.
20 . The system of claim 19 , the method further comprising:
before the constructing:
causing the properties identified in the response received from the large language model to be presented to a user; and
receiving input originated by the user:
selecting one or more unselected document properties,
unselecting one or more selected document properties, or
selecting one or more unselected document properties and unselecting one or more selected document properties; and
adjusting selection among the document properties in accordance with the received input.
21 . The system of claim 19 wherein the constructed index is a tree structure mapping from the selected properties to the documents of the corpus having those properties.
22 . The system of claim 19 wherein the constructing comprises:
for each document of the corpus,
concatenating the document's contents of the selected properties into a representation of the document;
encoding the representation of the document into a vector wherein the constructed index maps from the vectors produced by the encoding to the corresponding document of the corpus.
23 . The system of claim 19 , the method further comprising:
using the constructed index to conduct a search of the corpus.
24 . The system of claim 19 , the method further comprising:
using the constructed index to service a chatbot interaction with respect to the corpus.