Fine tuning large language models
A method, a system, and a computer program product for tuning a large language model. One or more first electronic documents are sampled to generate one or more sampled electronic documents. One or more portions of one or more second electronic documents are identified and extracted from the second electronic documents. The sampled electronic documents are sent to a generative artificial intelligence (AI) model to generate one or more first labels. One or more portions of the second electronic documents are sent to the generative AI model to generate one or more second labels. A large language model is trained using one or more first labels and one or more second labels to generate a trained large language model.
1 . A computer-implemented method, comprising:
sampling, using at least one processor, one or more first electronic documents to generate one or more sampled electronic documents, the one or more first electronic documents are stored in a first storage location using a first storage arrangement, wherein sampling includes
searching the one or more first electronic documents using one or more context-specific portions and identifying the one or more context-specific portions in the one or more first electronic documents; and
selecting, using the identified one or more context-specific portions, the one or more sampled electronic documents containing the one or more context-specific portions:
selecting, using the at least one processor, one or more portions in one or more second electronic documents contextually related to the identified one or more context-specific portions in the one or more first electronic documents, and extracting the one or more portions from the one or more second electronic documents, wherein the one or more second electronic documents are stored in a second storage location using a second storage arrangement being different from the first storage arrangement;
sending, using the at least one processor, the one or more sampled electronic documents and the one or more context-specific portions to a generative artificial intelligence (AI) model to generate one or more first labels for the one or more sampled electronic documents, and sending the one or more portions of the one or more second electronic documents to the generative AI model to generate one or more second labels;
training, using the at least one processor, a large language model using the one or more first labels and the one or more second labels; and
generating, using the at least one processor, a trained large language model.
2 . The method of claim 1 , wherein the sampling includes
identifying one or more context-based portions of the one or more first electronic documents;
assigning one or more identifiers to the one or more context-based portions; and
generating the one or more sampled electronic documents using the one or more assigned identifiers, wherein at least one sampled electronic document in the one or more sampled electronic documents corresponds to at least one context-based portion in the one or more context-based portions.
3 . The method of claim 1 , wherein the one or more first electronic documents are received from one or more electronic data sources.
4 . The method of claim 3 , wherein the one or more electronic data sources include at least one of the following: one or more public databases, one or more non-public databases, one or more government databases, one or more internet database, and any combination thereof.
5 . The method of claim 1 , wherein the one or more first electronic documents and the one or more second electronic documents are received from different electronic data sources.
6 . The method of claim 1 , wherein the one or more portions of the one or more second electronic documents are identified based on one or more tokens associated with at least one portion in the one or more portions.
7 . The method of claim 6 , wherein the one or more tokens are determined based on a content of the one or more portions.
8 . The method of claim 1 , wherein the training includes training, based on the or more first labels and the one or more second labels, the large language model using low-rank adaptation.
9 . A system, comprising:
at least one processor; and
at least one non-transitory storage media storing instructions, that when executed by the at least one processor, cause the at least one processor to
retrieve one or more first electronic documents from a plurality of electronic data sources;
generate one or more sampled electronic documents based on the one or more first electronic documents, the one or more first electronic documents are stored in a first storage location using a first storage arrangement, wherein generating one or more sampled documents includes
searching the one or more first electronic documents using one or more context-specific portions and identifying one or more context-specific portions in the one or more first electronic documents; and
selecting, using the identified one or more context-specific portions, the one or more sampled electronic documents containing the one or more context-specific portions:
select and extract one or more portions from one or more second electronic documents contextually related to the identified one or more context-specific portions in the one or more first electronic documents, wherein the one or more second electronic documents are stored in a second storage location using a second storage arrangement being different from the first storage arrangement;
send the one or more sampled electronic documents and the one or more context-specific portions to a generative artificial intelligence (AI) model to generate one or more first labels for the one or more sampled electronic documents, and send the one or more portions of the one or more second electronic documents to the generative AI model to generate one or more second labels;
train a large language model using the one or more first labels and the one or more second labels; and
generate a trained large language model.
10 . The system of claim 9 , wherein the at least one processor is configured to
identify one or more context-based portions of the one or more first electronic documents;
assign one or more identifiers to the one or more context-based portions; and
generate the one or more sampled electronic documents using the one or more identifiers, wherein at least one sampled electronic document in the one or more sampled electronic documents corresponds to at least one context-based portion in the one or more context-based portions.
11 . The system of claim 9 , wherein the plurality of electronic data sources includes at least one of the following: one or more public databases, one or more non-public databases, one or more government databases, one or more internet database, and any combination thereof.
12 . The system of claim 9 , wherein the one or more first electronic documents and the one or more second electronic documents are received from different electronic data sources.
13 . The system of claim 9 , wherein the one or more portions of the one or more second electronic documents are identified based on one or more tokens associated with at least one portion in the one or more portions.
14 . The system of claim 13 , wherein the one or more tokens are determined based on content of the one or more portions.
15 . The system of claim 9 , wherein the training includes training, based on the or more first labels and the one or more second labels, the large language model using low-rank adaptation.
16 . A computer program product comprising a non-transitory machine-readable medium storing instructions that, when executed by at least one programmable processor, cause the at least one programmable processor to:
train a large language model using one or more labels, wherein the one or more labels include at least one of: one or more first labels and one or more second labels,
the one or more first labels are generated based on one or more first electronic documents, the one or more first electronic documents are stored in a first storage location using a first storage arrangement, wherein the one or more first labels are generated using one or more sampled electronic documents, wherein generation of the one or more sampled electronic documents includes
searching the one or more first electronic documents using one or more context-specific portions and identifying one or more context-specific portions in the one or more first electronic documents; and
selecting, using the identified one or more context-specific portions, the one or more sampled electronic documents containing the one or more context-specific portions;
and
the one or more second labels are generated based on one or more portions selected and extracted from one or more second electronic documents and contextually related to the identified one or more context-based portions in the one or more first electronic documents, wherein the one or more second electronic documents are stored in a second storage location using a second storage arrangement being different from the first storage arrangement; and
generate a trained large language model.
17 . The computer program product of claim 16 , wherein the one or more labels are generated using a generative artificial intelligence (AI).
18 . The computer program product of claim 17 , wherein the at least one processor is configured to
identify the one or more context-based portions of the one or more first electronic documents;
assign one or more identifiers to the one or more context-based portions; and
generate the one or more sampled electronic documents using the one or more identifiers, wherein at least one sampled electronic document in the one or more sampled electronic documents corresponds to at least one context-based portion in the one or more context-based portions.
19 . The computer program product of claim 17 , wherein the one or more first electronic documents are retrieved from a plurality of electronic data sources, wherein the plurality of electronic data sources includes at least one of the following: one or more public databases, one or more non-public databases, one or more government databases, one or more internet database, and any combination thereof.
20 . The computer program product of claim 19 , wherein the one or more first electronic documents and the one or more second electronic documents are received from different electronic data sources.