IP Library › Granted Patent US 12,499,312
Granted Patent B2
US 12,499,312 · App. 18/335,898 · Granted Dec 16, 2025

Systems and methods for training a neural network model using knowledge from pre-trained large language models

Inventors: Shiva Kumar Pentyala (Mountain View, CA); Prafulla Kumar Choubey (San Jose, CA); Shashank Harinath (Mountain View, CA); Sitaram Asur (Newark, CA); Chien-Sheng Jason Wu (Mountain View, CA); Zachary Alexander (Berkeley, CA); Caiming Xiong (Menlo Park, CA)
Assignee: Salesforce, Inc.
G06F40/284G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,312
App. No.
18/335,898
Granted
Dec 16, 2025
Kind
B2
Abstract

Embodiments described herein provide a training framework for generative NLP models. Specifically, the training input, e.g., in the form of a sequence of tokens representing a user-agent dialogue, may be randomly masked for a few spans, which can be one or more tokens, one or more words, one or more sentences, or one or more paragraphs. These masked spans are replaced with their embeddings generated from pre-trained large language models are then used for training the NLP model.

Claims (40)

1 . A method for pretraining a neural network based natural language processing (NLP) model, comprising:

receiving, at a server and via a communication interface over a network, a training input including tokens of a prior user-agent dialogue;

randomly replacing one or more spans in the training input with one or more corresponding span embeddings obtained by embedding the one or more spans via a pretrained large language model (LLM), thereby resulting in a masked training input;

training the neural network based NLP model implemented on one or more hardware processors at the server using a pre-processed dataset including the masked training input; and

generating, by the trained neural network based NLP model, an agent response in response to a user utterance.

2 . The method of claim 1 , wherein at least one span of the one or more spans includes a sentence or a paragraph.

3 . The method of claim 1 , wherein the one or more corresponding span embeddings have a same dimension of the neural network based NLP model.

4 . The method of claim 1 , wherein the one or more corresponding span embeddings are precomputed by the pretrained LLM prior to runtime.

5 . The method of claim 1 , wherein the prior user-agent dialogue comprises a prior user utterance and at least one prior system response generated in response to the prior user utterance based on at least one or more knowledge documents.

6 . The method of claim 1 , wherein the one or more corresponding span embeddings are retrieved via an application programming interface (API) at the server, which connects the server to an external server that hosts the pretrained LLM.

7 . The method of claim 1 , wherein the training the neural network based NLP model using the pre-processed dataset including the masked training input further comprises:

predicting, by the neural network based NLP model, a next token subsequent to a subset of tokens in response to the masked training input;

computing a causal language modeling loss based on the predicted next token and a ground-truth token from tokens of the prior user-agent dialogue; and

updating parameters of the neural network based NLP model based on the causal language modeling loss.

8 . A system for pretraining a neural network based natural language processing (NLP) model, the system comprising:

a memory that stores the neural network based NLP model and a plurality of processor executable instructions;

a communication interface that receives a training input including tokens of a prior user-agent dialogue; and

one or more hardware processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:

randomly replacing one or more spans in the training input with one or more corresponding span embeddings obtained by embedding the one or more spans via a pretrained large language model (LLM), thereby resulting in a masked training input;

training the neural network based NLP model implemented on one or more hardware processors at the server using a pre-processed dataset including the masked training input; and

generating, by the trained neural network based NLP model, an agent response in response to a user utterance.

9 . The system of claim 8 , wherein at least one span of the one or more spans includes a sentence or a paragraph.

10 . The system of claim 8 , wherein the one or more corresponding span embeddings have a same dimension of the neural network based NLP model.

11 . The system of claim 8 , wherein the one or more corresponding span embeddings are precomputed by the pretrained LLM prior to runtime.

12 . The system of claim 8 , wherein the prior user-agent dialogue comprises a prior user utterance and at least one prior system response generated in response to the prior user utterance based on at least one or more knowledge documents.

13 . The system of claim 8 , wherein the one or more corresponding span embeddings are retrieved via an application programming interface (API) at the server, which connects the server to an external server that hosts the pretrained LLM.

14 . The system of claim 8 , wherein the operation of training the neural network based NLP model using the pre-processed dataset including the masked training input further comprises:

predicting, by the neural network based NLP model, a next token subsequent to a subset of tokens in response to the masked training input;

computing a causal language modeling loss based on the predicted next token and a ground-truth token from tokens of the prior user-agent dialogue; and

updating parameters of the neural network based NLP model based on the causal language modeling loss.

15 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:

receiving, at a server and via a communication interface over a network, a training input including tokens of a prior user-agent dialogue;

randomly replacing one or more spans in the training input with one or more corresponding span embeddings obtained by embedding the one or more spans via a pretrained large language model (LLM), thereby resulting in a masked training input; and

training the neural network based NLP model implemented on one or more hardware processors at the server using a pre-processed dataset including the masked training input; and

generating, by the trained neural network based NLP model, an agent response in response to a user utterance.

16 . The non-transitory machine-readable medium of claim 15 , wherein at least one span of the one or more spans includes a sentence or a paragraph.

17 . The non-transitory machine-readable medium of claim 15 , wherein the one or more corresponding span embeddings have a same dimension of the neural network based NLP model.

18 . The non-transitory machine-readable medium of claim 15 , wherein the one or more corresponding span embeddings are precomputed by the pretrained LLM prior to runtime.

19 . The non-transitory machine-readable medium of claim 15 , wherein the prior user-agent dialogue comprises a prior user utterance and at least one prior system response generated in response to the prior user utterance based on at least one or more knowledge documents.

20 . The non-transitory machine-readable medium of claim 15 , wherein the one or more corresponding span embeddings are retrieved via an application programming interface (API) at the server, which connects the server to an external server that hosts the pretrained LLM.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: PENTYALA, SHIVA KUMAR; ALEXANDER, ZACHARY; WU, CHIEN-SHENG; ASUR, SITARAM; HARINATH, SHASHANK; XIONG, CAIMING; CHOUBEY, PRAFULLA KUMAR
To: SALESFORCE, INC.
Reel/Frame 066633/0907 →
Continuity (2)
Continuation 18330216 · Jun 6, 2023
Related Publication 20240411992A1 · Dec 12, 2024
References Cited (17)
US 20190188326A1 · Daianu · 2019 [cited by examiner]
US 20220101052A1 · Canim · 2022 [cited by examiner]
US 20230154221A1 · Gu · 2023 [cited by examiner]
US 20240070394A1 · Peng · 2024 [cited by examiner]
US 20240143927A1 · Yoon · 2024 [cited by examiner]
US 20240176958A1 · Raimondo · 2024 [cited by examiner]
US 20240185001A1 · Nagaraju · 2024 [cited by examiner]
US 20240256906A1 · Yadav · 2024 [cited by examiner]
US 20240281472A1 · LaRhette · 2024 [cited by examiner]
Wen, Haoyang, et al. “Sequence-to-sequence learning for task-oriented dialogue with dialogue state representation.” arXiv preprint arXiv:1806.04441 (2018). (Year: 2018). [cited by examiner]
Izacard, Gautier, et al. “Few-shot learning with retrieval augmented language models.” arXiv preprint arXiv:2208.03299 1.2 (2022): 4. (Year: 2022). [cited by examiner]
Dua, Dheeru, et al. “To adapt or to annotate: Challenges and interventions for domain adaptation in open-domain question answering.” arXiv preprint arXiv:2212.10381 (2022). (Year: 2022). [cited by examiner]
Hudeček, Vojtěch, and Ondřej Dušek. “Are LLMs all you need for task-oriented dialogue?.” arXiv preprint arXiv:2304.06556 (2023). (Year: 2023). [cited by examiner]
Dua, Dheeru, et al. “To adapt or to annotate: Challenges and interventions for domain adaptation in open-domain question answering.” arXiv preprint arXiv:2212.10381, Dec. 20, 2022, 13 pages. [cited by applicant]
Hudecek, Vojtech, and Ondrej Dusek. “Are LLMs all you need for task-oriented dialogue?.” arXiv preprint arXiv:2304.06556 , Apr. 13, 2023, 11 pages. [cited by applicant]
Izacard, Gautier, et al. “Few-shot learning with retrieval augmented language models.” arXiv preprint arXiv:2208.03299 1.2, Nov. 16, 2022, 33 pages. [cited by applicant]
Wen, Haoyang, et al. “Sequence-to-sequence learning for task-oriented dialogue with dialogue state representation.” arXiv preprint arXiv:1806.04441 , Jun. 12, 2018, 12 pages. [cited by applicant]