Portable personalized large language models
One example method includes transmitting, by a client device, a request for a reduced large language model (“LLM”) to a remote server; receiving, by the client device from the remote server, and storing the reduced LLM, the reduced LLM based on a trained general LLM; receiving, by the client device, a request to generate content using the reduced LLM; providing the request to the reduced LLM; and receiving generated content from the reduced LLM based on the request.
1 . A method comprising:
transmitting, by a client device, a request for a reduced large language model (“LLM”) to a remote server;
receiving, by the client device from the remote server, and storing the reduced LLM, the reduced LLM based on a trained general LLM;
training the reduced LLM based on one or more user-generated content items to generate a personalized reduced LLM;
receiving, by the client device, a request to generate content using the personalized reduced LLM;
providing the request to the personalized reduced LLM; and
receiving generated content from the personalized reduced LLM based on the request.
2 . The method of claim 1 , further comprising:
receiving a request to train the reduced LLM; and
receiving an identification of one or more user-generated content items.
3 . The method of claim 2 , wherein the personalized reduced LLM is trained based on a selected type of user-generated content items.
4 . The method of claim 1 , wherein the personalized reduced LLM is trained based on the trained general LLM.
5 . The method of claim 1 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters having fewer parameters than the first set of parameters.
6 . The method of claim 1 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters comprises one or more parameters having different numerical representations than corresponding parameters in the first set of parameters.
7 . The method of claim 6 , wherein the second set of parameters comprises one or more parameters having a floating-point representation using fewer bits that a floating-point representation of the corresponding parameters in the trained LLM.
8 . A system comprising:
a non-transitory computer-readable medium; and
one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:
transmit a request for a reduced large language model (“LLM”) to a remote server;
receive, from the remote server, and store a reduced LLM, the reduced LLM based on a trained general LLM;
train the reduced LLM based on one or more user-generated content items to generate a personalized reduced LLM;
receive a request to generate content using the personalized reduced LLM;
provide the request to the personalized reduced LLM; and
receive generated content from the personalized reduced LLM based on the request.
9 . The system of claim 8 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:
receive a request to train the reduced LLM; and
receive an identification of one or more user-generated content items.
10 . The system of claim 9 , wherein the personalized reduced LLM is trained based on a selected type of user-generated content items.
11 . The system of claim 8 , wherein the personalized reduced LLM is trained based on the trained general LLM.
12 . The system of claim 8 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters having fewer parameters than the first set of parameters.
13 . The system of claim 8 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters comprises one or more parameters having different numerical representations than corresponding parameters in the first set of parameters.
14 . The system of claim 13 , wherein the second set of parameters comprises one or more parameters having a floating-point representation using fewer bits that a floating-point representation of the corresponding parameters in the trained LLM.
15 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:
transmit a request for a reduced large language model (“LLM”) to a remote server;
receive, from the remote server, and store a reduced LLM, the reduced LLM based on a trained general LLM;
train the reduced LLM based on one or more user-generated content items to generate a personalized reduced LLM;
receive a request to generate content using the personalized reduced LLM;
provide the request to the personalized reduced LLM; and
receive generated content from the personalized reduced LLM based on the request.
16 . The non-transitory computer-readable medium of claim 15 , further comprising processor-executable instructions configured to cause the one or more processors to:
receive a request to train the reduced LLM; and
receive an identification of one or more user-generated content items.
17 . The non-transitory computer-readable medium of claim 16 , wherein the personalized reduced LLM is trained based on a selected type of user-generated content items.
18 . The non-transitory computer-readable medium of claim 15 , wherein the personalized reduced LLM is trained based on the trained general LLM.
19 . The non-transitory computer-readable medium of claim 15 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters having fewer parameters than the first set of parameters.
20 . The non-transitory computer-readable medium of claim 15 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters comprises one or more parameters having different numerical representations than corresponding parameters in the first set of parameters.