IP Library › Patent Application 18898509
Patent Application
App. No. 18/898,509

COPILOT IMPLEMENTATION: TRAINING AN EXPANSION MACHINE LEARNING TOOL

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/898,509
Abstract

Apparatus and methods are disclosed for implementing a copilot as a network of microservices including specialized large language models (LLMs) or other trained machine learning (ML) tools. This architecture supports flexible, customizable, or dynamically determinable dataflow. Compared to much larger competing LLMs, comparable or superior performance is achieved for certain tasks, while significantly reducing computation time and hardware requirements, even to a single compute node with a single GPU. An expansion microservice converts client input tokens into associated tokens to diversify targeted tasks presented to a core microservice. An ML tool in the expansion microservice is pretrained in multiple stages for varying tasks using varying general and target-specific corpora. Synthesized training data derived from a knowledge graph can also be used. The expansion ML tool is subsequently fine-tuned in one or multiple stages. Variations and additional techniques are disclosed.

Claims (73)

1 . A computer-implemented method of customizing a copilot for a target deployment, wherein the copilot comprises a client interface, a core microservice, and an expansion microservice in turn comprising a trained expansion machine learning (ML) tool, the method comprising:

pretraining the expansion ML tool in a plurality of pretraining stages;

wherein the expansion microservice is coupled between the client interface and the core microservice, and is configured to augment first tokens received from the client interface with additional second tokens related to the received tokens, and to forward the first and second tokens toward the core microservice; and

wherein the pretraining stages comprise two or more of:

a first pretraining stage in which the expansion ML tool is pretrained on a first general corpus to optimize performance for a first task;

a second pretraining stage in which the expansion ML tool is pretrained on a second general corpus to optimize performance for a second task distinct from the first task;

a third pretraining stage in which the expansion ML tool is pretrained on synthesized training data derived from a pruned knowledge graph; or

a fourth pretraining stage in which the expansion ML tool is pretrained on a target corpus specific to the target deployment; and

subsequent to the pretraining, fine-tuning the expansion ML tool with a training dataset derived from the target corpus, the training dataset comprising a plurality of training records, each training record comprising a training input and a desired response.

2 . The computer-implemented method of claim 1 , wherein the plurality of pretraining stages comprise the first, second, and fourth pretraining stages.

3 . The computer-implemented method of claim 1 , wherein the first or second task is a masked language replacement task.

4 . The computer-implemented method of claim 1 , wherein the plurality of pretraining stages comprises the first and fourth pretraining stages, and the fourth pretraining stage optimizes performance on the first task.

5 . The computer-implemented method of claim 1 , wherein the plurality of pretraining stages comprises the fourth pretraining stage, the third pretraining stage optimizes performance on a fourth task, and the pretraining stages further comprise one or more fifth pretraining stages in which the expansion ML tool is pretrained on the target corpus to optimize performance for respective fifth tasks distinct from the fourth task.

6 . The computer-implemented method of claim 1 , further comprising:

incorporating, into the training dataset, a first set of the training records for which the training inputs and the desired responses are created by a first human having expert knowledge of the target corpus; and

at least one of:

incorporating, into the training dataset, a second set of the training records for which the training inputs are created by a second human lacking expert knowledge of the target corpus, and the desired responses are created by the first human; or

incorporating, into the training dataset, a third set of the training records synthesized based on the first set of the training records.

7 . The computer-implemented method of claim 1 , wherein the training records comprise:

one or more first training records for each of which the respective desired response is an answer to the respective training input;

one or more second training records for each of which the respective desired response comprises a clarification directed toward the client interface; and

one or more third training records for each of which the respective desired response comprises an output directed toward the core microservice.

8 . The computer-implemented method of claim 1 , wherein the expansion ML tool is a large language model (LLM), large multimodal model (LMM), or deep neural network (DNN).

9 . The computer-implemented method of claim 1 , wherein the expansion ML tool is configured as an encoder-decoder transformer neural network.

10 . One or more computer-readable media storing instructions which, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations for customizing an expansion machine learning (ML) tool of a copilot for a target deployment comprising:

pretraining the expansion ML tool in a plurality of pretraining stages;

wherein the copilot comprises a client interface, a core microservice, and an expansion ML tool coupled therebetween;

wherein the expansion ML tool is configured to augment first tokens received from the client interface with additional second tokens related to the received tokens, and to forward the first and second tokens toward the core microservice; and

wherein the pretraining stages comprise two or more of:

a first pretraining stage in which the expansion ML tool is pretrained on a first general corpus to optimize performance for a first task;

a second pretraining stage in which the expansion ML tool is pretrained on a second general corpus to optimize performance for a second task distinct from the first task;

a third pretraining stage in which the expansion ML tool is pretrained on synthesized training data derived from a pruned knowledge graph; or

a fourth pretraining stage in which the expansion ML tool is pretrained on a target corpus specific to the target deployment; and

subsequent to the pretraining, fine-tuning the expansion ML tool with a training dataset derived from the target corpus, the training dataset comprising a plurality of training records, each training record comprising a training input and a desired response.

11 . The one or more computer-readable media of claim 10 , wherein the plurality of pretraining stages comprise the first, second, and fourth pretraining stages.

12 . The one or more computer-readable media of claim 10 , wherein the expansion ML tool is a large language model (LLM), large multimodal model (LMM), or deep neural network (DNN).

13 . The one or more computer-readable media of claim 10 , wherein the expansion ML tool is configured as an encoder-decoder transformer neural network.

14 . A system comprising:

one or more hardware processors, with memory coupled thereto; and

one or more computer readable media storing first and second instructions, the first instructions comprising a plurality of modules which, when executed by the one or more hardware processors, implement respective microservices, the microservices forming a weakly connected network of microservices configured as a copilot for one or more first client applications, and wherein:

each of the microservices is configured to:

receive input from (i) a respective first group comprising one or more others of the microservices or (ii) one or more second client applications; and

transmit output to (i) a second group comprising one or more of the microservices or (ii) one or more third client applications;

a plurality of the microservices incorporate respective trained machine learning tools; and

the network of microservices comprises at least an expansion microservice, comprising an expansion machine learning (ML) tool and coupled to a client interface, and a core microservice;

output of the expansion microservice is directed toward the core microservice; and

wherein the second instructions, upon execution by the one or more hardware processors, cause the one or more hardware processors to perform operations for customizing the copilot for a target deployment, the operations comprising:

pretraining the expansion ML tool in a plurality of pretraining stages;

wherein the expansion microservice is configured to augment first tokens received from the client interface with additional second tokens related to the received tokens, and to forward the first and second tokens toward the core microservice; and

wherein the pretraining stages comprise two or more of:

a first pretraining stage in which the expansion ML tool is pretrained on a first general corpus to optimize performance for a first task;

a second pretraining stage in which the expansion ML tool is pretrained on a second general corpus to optimize performance for a second task distinct from the first task;

a third pretraining stage in which the expansion ML tool is pretrained on synthesized training data derived from a pruned knowledge graph; or

a fourth pretraining stage in which the expansion ML tool is pretrained on a target corpus specific to the target deployment; and

subsequent to the pretraining, fine-tuning the expansion ML tool with a training dataset derived from the target corpus, the training dataset comprising a plurality of training records, each training record comprising a training input and a desired response.

15 . The system of claim 14 , wherein the plurality of pretraining stages comprise the first, second, and fourth pretraining stages.

16 . The system of claim 14 , wherein the first or second task is a masked language replacement task.

17 . The system of claim 14 , wherein the plurality of pretraining stages comprises the first and fourth pretraining stages, and the fourth pretraining stage optimizes performance on the first task.

18 . The system of claim 14 , wherein the training records comprise:

one or more first training records for each of which the respective desired response is an answer to the respective training input;

one or more second training records for each of which the respective desired response comprises a clarification directed toward the client interface; and

one or more third training records for each of which the respective desired response comprises an output directed toward the core microservice.

19 . The system of claim 14 , wherein the expansion ML tool is a large language model (LLM), large multimodal model (LMM), or deep neural network (DNN).

20 . The system of claim 14 , wherein the expansion ML tool is configured as an encoder-decoder transformer neural network.

21 . The system of claim 14 , wherein:

the fine-tuning operations are one of a plurality of fine-tuning stages performed subsequent to the pretraining;

an earlier one of the fine-tuning stages is performed with a first training dataset;

a later one of the fine-tuning stages is performed after the earlier one of the fine-tuning stages, with a second training dataset;

the second training dataset is specific to the target deployment; and

the first training dataset is more general than the second training dataset.

22 . The system of claim 14 , wherein the core microservice is configured to:

generate, based at least partly on one or more of the second tokens, a response to a first input, comprising the first tokens, received at the client interface; and

transmit the response toward the client interface.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2024
From: EDUWORKS CORPORATION
To: THIA ST CO.
Reel/Frame 069643/0557 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2024
From: KELSEY, ELAINE; ROBSON, ELLIOT NICHOLAS; NASIR, SAZZAD MAHMUD; YARBRO, JEFFREY THOMAS; ROBSON, ROBERT OSCAR; EGERTON, LAUREN ELIZABETH; WARD, SPENCER THOMAS; KELLY, BRENDAN MICHAEL
To: EDUWORKS CORPORATION
Reel/Frame 069093/0097 →