IP Library › Granted Patent US 12,456,013
Granted Patent B2
US 12,456,013 · App. 18/330,216 · Granted Oct 28, 2025

Systems and methods for training a neural network model using knowledge from pre-trained large language models

Inventors: Shiva Kumar Pentyala (Mountain View, CA); Prafulla Kumar Choubey (San Jose, CA); Shashank Harinath (Mountain View, CA); Sitaram Asur (Newark, CA); Chien-Sheng Jason Wu (Mountain View, CA); Zachary Alexander (Berkeley, CA); Caiming Xiong (Menlo Park, CA)
Assignee: Salesforce, Inc.
G06F40/284G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,013
App. No.
18/330,216
Granted
Oct 28, 2025
Kind
B2
Abstract

Embodiments described herein provide a training framework for generative NLP models that operate on previously learnt knowledge from pretrained large language models. Specifically, to train an NLP model to generate a response to a user utterance (e.g., “resolve login issue”), document embeddings of support IT documents encoded by a pretrained LLM are fed to an NLP decoder together with a training dialogue (e.g., a dialogue between the chat agent on how to “resolve login issue”). The NLP decoder can thus be trained by a causal language modeling loss computed based on the predicted next token and the ground-truth token from the training dialogue.

Claims (62)

1 . A method for training a neural network based natural language processing (NLP) model, comprising:

receiving, at a server and via a communication interface over a network, a training input including tokens of a prior user-agent dialogue;

retrieving, at the server, one or more precomputed document embeddings of one or more knowledge documents from a pretrained large language model (LLM);

generating an augmented training input by combining the one or more precomputed document embeddings and token embeddings corresponding to at least a subset of tokens from the training input in an embedding space, wherein the generating comprises inserting a first token indicating a first knowledge document from the one or more knowledge documents at a start of a first prior system response in the prior user-agent dialogue;

training the neural network based NLP model implemented on one or more hardware processors at the server using the augmented training input, wherein the training comprises:

predicting, by the neural network based NLP model, a document identification in response to the augmented training input,

computing a factual alignment loss by comparing the predicted document identification and the first token, and

updating parameters of the neural network based NLP model based on the factual alignment loss; and

generating, by the trained neural network based NLP model, an agent response that is based on the one or more knowledge documents in response to a user utterance.

2 . The method of claim 1 , wherein the prior user-agent dialogue comprises a prior user utterance and at least one prior system response generated in response to the prior user utterance based on at least one of the one or more knowledge documents.

3 . The method of claim 1 , wherein the one or more precomputed document embeddings are retrieved via an application programming interface (API) at the server, which connects the server to an external server that hosts the pretrained LLM, and

wherein the one or more precomputed document embeddings are generated by the pretrained LLM and stored at the external server prior to training the neural network based natural language processing (NLP) model.

4 . The method of claim 1 , wherein the training the neural network based NLP model using the augmented training input further comprises:

predicting, by the neural network based NLP model, a next token subsequent to the subset of tokens in response to the augmented training input;

computing a causal language modeling loss based on the predicted next token and a ground-truth token from the tokens of a prior user-agent dialogue; and

updating parameters of the neural network based NLP model based on the causal language modeling loss.

5 . The method of claim 1 , further comprising:

generating, by the trained neural network based NLP model, a document identification token at a start of the agent response in response to an user utterance, wherein the document identification token indicates which knowledge document the agent response is based on.

6 . The method of claim 1 , further comprising:

pretraining the neural network based NLP model by:

randomly replacing one or more spans in the training input with one or more corresponding embeddings, thereby resulting in a masked training input,

wherein the one or more corresponding embeddings have a same dimension of the neural network based NLP model; and

training the neural network NLP model using a pre-processed dataset including the masked training input.

7 . The method of claim 6 , wherein at least one span includes a sentence or a paragraph, and wherein the one or more corresponding embeddings are generated by the pretrained LLM.

8 . The method of claim 1 , further comprising:

transforming multiple randomly selected training samples into corresponding embeddings generated by the pretrained LLM; and

training the neural network based NLP model using the corresponding embeddings as in-context examples.

9 . A system for training a neural network based natural language processing (NLP) model, comprising:

a memory that stores the neural network based NLP model and a plurality of processor executable instructions;

a communication interface that receives a training input including tokens of a prior user-agent dialogue; and

one or more hardware processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:

retrieving, at the server, one or more precomputed document embeddings of one or more knowledge documents from a pretrained large language model (LLM);

generating an augmented training input by combining the one or more precomputed document embeddings and token embeddings corresponding to at least a subset of tokens from the training input in an embedding space, wherein the generating comprises inserting a first token indicating a first knowledge document from the one or more knowledge documents at a start of a first prior system response in the prior user-agent dialogue;

training the neural network based NLP model implemented on one or more hardware processors at the server using the augmented training input, wherein the training comprises:

predicting, by the neural network based NLP model, a document identification in response to the augmented training input,

computing a factual alignment loss by comparing the predicted document identification and the first token, and

updating parameters of the neural network based NLP model based on the factual alignment loss; and

generating, by the trained neural network based NLP model, an agent response that is based on the one or more knowledge documents in response to a user utterance.

10 . The system of claim 9 , wherein the prior user-agent dialogue comprises a prior user utterance and at least one prior system response generated in response to the prior user utterance based on at least one of the one or more knowledge documents.

11 . The system of claim 9 , wherein the one or more precomputed document embeddings are retrieved via an application programming interface (API) at the server, which connects the server to an external server that hosts the pretrained LLM, and

wherein the one or more precomputed document embeddings are generated by the pretrained LLM and stored at the external server prior to training the neural network based natural language processing (NLP) model.

12 . The system of claim 9 , wherein the operation of training the neural network based NLP model using the augmented training input further comprises:

predicting, by the neural network based NLP model, a next token subsequent to the subset of tokens in response to the augmented training input;

computing a causal language modeling loss based on the predicted next token and a ground-truth token from the tokens of a prior user-agent dialogue; and

updating parameters of the neural network based NLP model based on the causal language modeling loss.

13 . The system of claim 9 , wherein the operations further comprise:

generating, by the trained neural network based NLP model, a document identification token at a start of the agent response in response to an user utterance, wherein the document identification token indicates which knowledge document the agent response is based on.

14 . The system of claim 9 , wherein the operations further comprise:

pretraining the neural network based NLP model by:

randomly replacing one or more spans in the training input with one or more corresponding embeddings, thereby resulting in a masked training input,

wherein the one or more corresponding embeddings have a same dimension of the neural network based NLP model; and

training the neural network NLP model using a pre-processed dataset including the masked training input.

15 . The system of claim 14 , wherein at least one span includes a sentence or a paragraph, and wherein the one or more corresponding embeddings are generated by the pretrained LLM.

16 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:

receiving, at a server and via a communication interface over a network, a training input including tokens of a prior user-agent dialogue;

retrieving, at the server, one or more precomputed document embeddings of one or more knowledge documents from a pretrained large language model (LLM);

generating an augmented training input by combining the one or more precomputed document embeddings and token embeddings corresponding to at least a subset of tokens from the training input in an embedding space, wherein the generating comprises inserting a first token indicating a first knowledge document from the one or more knowledge documents at a start of a first prior system response in the prior user-agent dialogue;

training the neural network based NLP model implemented on one or more hardware processors at the server using the augmented training input, wherein the training comprises:

predicting, by the neural network based NLP model, a document identification in response to the augmented training input,

computing a factual alignment loss by comparing the predicted document identification and the first token, and

updating parameters of the neural network based NLP model based on the factual alignment loss; and

generating, by the trained neural network based NLP model, an agent response that is based on the one or more knowledge documents in response to a user utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2024
From: PENTYALA, SHIVA KUMAR; CHOUBEY, PRAFULLA KUMAR; HARINATH, SHASHANK; ASUR, SITARAM; WU, CHIEN-SHENG; ALEXANDER, ZACHARY; XIONG, CAIMING
To: SALESFORCE, INC.
Reel/Frame 066614/0836 →
Continuity (1)
Related Publication 20240411991A1 · Dec 12, 2024
References Cited (13)
US 20190188326A1 · Daianu · 2019 [cited by examiner]
US 20220101052A1 · Canim · 2022 [cited by examiner]
US 20230154221A1 · Gu · 2023 [cited by examiner]
US 20240070394A1 · Peng · 2024 [cited by examiner]
US 20240143927A1 · Yoon · 2024 [cited by examiner]
US 20240176958A1 · Raimondo · 2024 [cited by examiner]
US 20240185001A1 · Nagaraju · 2024 [cited by examiner]
US 20240256906A1 · Yadav · 2024 [cited by examiner]
US 20240281472A1 · LaRhette · 2024 [cited by examiner]
Hudeček, Vojtěch, and Ondřej Dušek. “Are LLMs all you need for task-oriented dialogue?. ” arXiv preprint arXiv:2304.06556 (2023). (Year: 2023). [cited by examiner]
Wen, Haoyang, et al. “Sequence-to-sequence learning for task-oriented dialogue with dialogue state representation.” arXiv preprint arXiv:1806.04441 (2018). (Year: 2018). [cited by examiner]
Dua, Dheeru, et al. “To adapt or to annotate: Challenges and interventions for domain adaptation in open-domain question answering.” arXiv preprint arXiv:2212.10381 (2022). (Year: 2022). [cited by examiner]
Izacard, Gautier, et al. “Few-shot learning with retrieval augmented language models.” arXiv preprint arXiv:2208.03299 1.2 (2022): 4. (Year: 2022). [cited by examiner]