IP Library › Granted Patent US 12,676,147
Granted Patent B2
US 12,676,147 · App. 18/385,158 · Granted Jul 7, 2026

Portable personalized large language models

Inventors: Bilung Lee (Irvine, CA); Vijay Venkataswamy Parthasarathy (San Jose, CA); Renjie Tao (Santa Clara, CA); Zheng Yuan (Saratoga, CA); Bing Zhao (San Jose, CA)
Assignee: Zoom Communications, Inc.
G10L15/183G10L15/063G10L15/22G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,676,147
App. No.
18/385,158
Filed
Oct 30, 2023
Granted
Jul 7, 2026
Kind
B2
Examiner
HANG, VU B
Art Unit
2654
USPC
704/243
Abstract

One example method includes transmitting, by a client device, a request for a reduced large language model (“LLM”) to a remote server; receiving, by the client device from the remote server, and storing the reduced LLM, the reduced LLM based on a trained general LLM; receiving, by the client device, a request to generate content using the reduced LLM; providing the request to the reduced LLM; and receiving generated content from the reduced LLM based on the request.

Claims (46)

1 . A method comprising:

transmitting, by a client device, a request for a reduced large language model (“LLM”) to a remote server;

receiving, by the client device from the remote server, and storing the reduced LLM, the reduced LLM based on a trained general LLM;

training the reduced LLM based on one or more user-generated content items to generate a personalized reduced LLM;

receiving, by the client device, a request to generate content using the personalized reduced LLM;

providing the request to the personalized reduced LLM; and

receiving generated content from the personalized reduced LLM based on the request.

2 . The method of claim 1 , further comprising:

receiving a request to train the reduced LLM; and

receiving an identification of one or more user-generated content items.

3 . The method of claim 2 , wherein the personalized reduced LLM is trained based on a selected type of user-generated content items.

4 . The method of claim 1 , wherein the personalized reduced LLM is trained based on the trained general LLM.

5 . The method of claim 1 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters having fewer parameters than the first set of parameters.

6 . The method of claim 1 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters comprises one or more parameters having different numerical representations than corresponding parameters in the first set of parameters.

7 . The method of claim 6 , wherein the second set of parameters comprises one or more parameters having a floating-point representation using fewer bits that a floating-point representation of the corresponding parameters in the trained LLM.

8 . A system comprising:

a non-transitory computer-readable medium; and

one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

transmit a request for a reduced large language model (“LLM”) to a remote server;

receive, from the remote server, and store a reduced LLM, the reduced LLM based on a trained general LLM;

train the reduced LLM based on one or more user-generated content items to generate a personalized reduced LLM;

receive a request to generate content using the personalized reduced LLM;

provide the request to the personalized reduced LLM; and

receive generated content from the personalized reduced LLM based on the request.

9 . The system of claim 8 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

receive a request to train the reduced LLM; and

receive an identification of one or more user-generated content items.

10 . The system of claim 9 , wherein the personalized reduced LLM is trained based on a selected type of user-generated content items.

11 . The system of claim 8 , wherein the personalized reduced LLM is trained based on the trained general LLM.

12 . The system of claim 8 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters having fewer parameters than the first set of parameters.

13 . The system of claim 8 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters comprises one or more parameters having different numerical representations than corresponding parameters in the first set of parameters.

14 . The system of claim 13 , wherein the second set of parameters comprises one or more parameters having a floating-point representation using fewer bits that a floating-point representation of the corresponding parameters in the trained LLM.

15 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:

transmit a request for a reduced large language model (“LLM”) to a remote server;

receive, from the remote server, and store a reduced LLM, the reduced LLM based on a trained general LLM;

train the reduced LLM based on one or more user-generated content items to generate a personalized reduced LLM;

receive a request to generate content using the personalized reduced LLM;

provide the request to the personalized reduced LLM; and

receive generated content from the personalized reduced LLM based on the request.

16 . The non-transitory computer-readable medium of claim 15 , further comprising processor-executable instructions configured to cause the one or more processors to:

receive a request to train the reduced LLM; and

receive an identification of one or more user-generated content items.

17 . The non-transitory computer-readable medium of claim 16 , wherein the personalized reduced LLM is trained based on a selected type of user-generated content items.

18 . The non-transitory computer-readable medium of claim 15 , wherein the personalized reduced LLM is trained based on the trained general LLM.

19 . The non-transitory computer-readable medium of claim 15 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters having fewer parameters than the first set of parameters.

20 . The non-transitory computer-readable medium of claim 15 , wherein the trained general LLM has a first set of parameters and the reduced LLM has a second set of parameters, the second set of parameters comprises one or more parameters having different numerical representations than corresponding parameters in the first set of parameters.

Assignments (2)
CHANGE OF NAME Recorded Jun 8, 2026
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 075699/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2026
From: LEE, BILUNG; PARTHASARATHY, VIJAY VENKATASWAMY; TAO, RENJIE; YUAN, ZHENG; ZHAO, BING
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 074876/0988 →
Continuity (1)
Related Publication 20250140245A1 · May 1, 2025
References Cited (15)
US 11176934B1 · Venkatesh Raman · 2021 [cited by examiner]
US 20210019616A1 · Chen · 2021 [cited by examiner]
US 20240411798A1 · Gerard et al. · 2024 [cited by applicant]
US 20250078484A1 · Nguyen · 2025 [cited by examiner]
International Search Report and Written Opinion for PCT/US2024/046853 mailed Jan. 28, 2025. [cited by applicant]
Xu et al., “Compress, Then Prompt: Improving Accuracy -Efficiency Trade-off of LLM Inference with Transferable Prompt”, ARXIV.org, Cornell University Library, Ithaca, New York, May 17, 2023; pp. 1-20. [cited by applicant]
Lin et al., “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration”, ARXIV.org, Cornell University Library, Ithaca, New York, Jun. 1, 2023; pp. 1-15. [cited by applicant]
Dettmers et al. “QLoRA: Efficient Finetuning of Quantized LLMs”, ARXIV.org, Cornell University Library, Ithaca, NY, May 23, 2023; pp. 1-26. [cited by applicant]
U.S. Appl. No. 18/385,141 , Non-Final Office Action, Mailed On Jul. 15, 2025, 16 pages. [cited by applicant]
Du , “Summarising Your Meeting With Chatgpt and Langchain”, Medium Available Online at: https://dxiaochuan.medium.com/summarising-your-meeting-with-chatgpt-and-langchain-8eb646cfcdd, Jun. 8, 2023, 7 pages. [cited by applicant]
Du , “Summarizing your Meeting with ChatGPT and LangChain”, Available online at: https://dxiaochuan.medium.com/summarising-your-meeting-with-chatgpt-and-langchain-8eb646cfcdd1, Jun. 8, 2023, 14 pages. [cited by applicant]
Li et al., “DQ-BART: Efficient Sequence-to-Sequence Model via Joint Distillation and Quantization”, Available online at: https://arxiv.org/pdf/2203.11239, Mar. 21, 2022, 9 pages. [cited by applicant]
Application No. PCT/US2024/046584 , International Search Report and Written Opinion, Mailed On Jan. 17, 2025, 12 pages. [cited by applicant]
Zhang et al., “Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models”, Available online at: https://arxiv.org/pdf/2305.12356, May 21, 2023, pp. 1-11. [cited by applicant]
U.S. Appl. No. 18/385,141, filed Oct. 30, 2023. [cited by applicant]