IP Library Granted Patent US 12,367,426
Granted Patent B1
US 12,367,426 · App. 19/005,712 · Granted Jul 22, 2025

Customization of machine learning tools with occupation training

Inventors: Elaine Kelsey (Corvallis, OR); Sazzad Mahmud Nasir (Muncie, IN); Jeffrey Thomas Yarbro (Memphis, TN); Lauren Elizabeth Egerton (New York, NY); Elliot Nicholas Robson (Seoul, KR); Brendan Michael Kelly (Somerville, MA); Robert Oscar Robson (Corvallis, OR); Spencer Thomas Ward (Kent, WA)
Assignee: THIA ST Co.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,426
App. No.
19/005,712
Granted
Jul 22, 2025
Kind
B1
Abstract

Methods and apparatus are disclosed for customizing a copilot or other machine learning tool. Document records related to copilot objectives or tasks are obtained and used to identify corresponding data sources. Data sources can be integrated into data producers or data repositories, to be used by a retrieval microservice. Data producers and other microservices are individually fine-tuned for the custom application, before or after integration into the copilot. The integrated copilot is tested end-to-end, and can be further refined. Disclosed techniques range from fully automated to human-in-the-loop (e.g. guided by an expert) to fully interactive. Some techniques produce tools which can automate certain customization operations. Training data, tasks, or document records blend expert-generated, developer-generated, or synthesized items.

Claims (49)

1. A method comprising:

(a) producing a language-trained machine learning tool ML 2 from a machine learning tool ML 1 by training the machine learning tool ML 1 to perform general language tasks with at least a first predetermined performance level;

(b) producing an occupation-trained machine learning tool ML 3 from the language-trained machine learning tool ML 2 by training the machine learning tool ML 2 to perform occupation-related tasks of a given occupation with at least a second predetermined performance level; and

(c) producing a custom-trained machine learning tool ML 4 from the occupation-trained machine learning tool ML 3 by training the machine learning tool ML 3 to perform custom tasks of the given occupation with at least a third predetermined performance level.

2. The method of claim 1 , wherein the custom tasks of the given occupation or the occupation-related tasks comprise identifying one or more pertinent document records from a data corpus in response to a given task description, and the method further comprises:

prompting the machine learning tool ML 3 or ML 4 with a task description within scope of a copilot objective; and

receiving the pertinent document record(s) from the machine learning tool ML 3 or ML 4 .

3. The method of claim 2 , further comprising:

validating a subset of the received document record(s).

4. The method of claim 1 , wherein the custom tasks of the given occupation or the occupation-related tasks comprise determining, in response to a given document record, one or more data sources which support the given document record, and the method further comprises:

prompting the machine learning tool ML 3 or ML 4 with a document record pertinent to a copilot objective; and

receiving the supporting data source(s) from the machine learning tool ML 3 or ML 4 .

5. The method of claim 4 , further comprising:

validating a subset of the received data source(s).

6. The method of claim 1 , wherein the given occupation comprises interviewing.

7. The method of claim 1 , wherein the given occupation comprises annotating interviews or annotating recorded work sessions.

8. The method of claim 1 , wherein the custom-trained machine learning tool ML 4 is a first custom-trained machine learning tool, and the method further comprises:

(d) producing a second custom-trained machine learning tool ML 5 from the occupation-trained machine learning tool ML 3 by training the machine learning tool ML 3 to perform second custom tasks of the given occupation with at least a fourth predetermined performance level.

9. The method of claim 8 , further comprising:

(e) deploying machine learning tools ML 4 and ML 5 at distinct first and second organizations respectively;

wherein the training at act (c) uses proprietary data of the first organization without using proprietary data of the second organization, and the training at act (d) uses proprietary data of the second organization without using proprietary data of the first organization.

10. The method of claim 1 , further comprising:

(d) deploying machine learning tool ML 4 within a first microservice of a copilot, the copilot comprising a weakly connected network of microservices including the first microservice.

11. One or more computer-readable media storing instructions which, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising:

(a) producing a language-trained machine learning tool ML 2 from a machine learning tool ML 1 by training the machine learning tool ML 1 to perform general language tasks with at least a first predetermined performance level;

(b) producing an occupation-trained machine learning tool ML 3 from the language-trained machine learning tool ML 2 by training the machine learning tool ML 2 to perform occupation-related tasks of a given occupation with at least a second predetermined performance level; and

(c) producing a custom-trained machine learning tool ML 4 from the occupation-trained machine learning tool ML 3 by training the machine learning tool ML 3 to perform custom tasks of the given occupation with at least a third predetermined performance level.

12. The one or more computer-readable media of claim 11 , wherein the custom tasks of the given occupation or the occupation-related tasks comprise identifying one or more pertinent document records from a data corpus in response to a given task description, and the operations further comprise:

prompting the machine learning tool ML 3 or ML 4 with a task description within scope of a copilot objective; and

receiving the pertinent document record(s) from the machine learning tool ML 3 or ML 4 .

13. The one or more computer-readable media of claim 12 , wherein the operations further comprise:

validating a subset of the received document record(s).

14. The one or more computer-readable media of claim 11 , wherein the custom-trained machine learning tool ML 4 is a first custom-trained machine learning tool, and the operations further comprise:

(d) producing a second custom-trained machine learning tool ML 5 from the occupation-trained machine learning tool ML 3 by training the machine learning tool ML 3 to perform second custom tasks of the given occupation with at least a fourth predetermined performance level.

15. A system, comprising:

one or more hardware processors with memory coupled thereto; and

computer-readable media storing instructions which, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising:

(a) producing a language-trained machine learning tool ML 2 from a machine learning tool ML 1 by training the machine learning tool ML 1 to perform general language tasks with at least a first predetermined performance level;

(b) producing an occupation-trained machine learning tool ML 3 from the language-trained machine learning tool ML 2 by training the machine learning tool ML 2 to perform occupation-related tasks of a given occupation with at least a second predetermined performance level; and

(c) producing a custom-trained machine learning tool ML 4 from the occupation-trained machine learning tool ML 3 by training the machine learning tool ML 3 to perform custom tasks of the given occupation with at least a third predetermined performance level.

16. The system of claim 15 , wherein the custom tasks of the given occupation or the occupation-related tasks comprise determining, in response to a given document record, one or more data sources which support the given document record, and the operations further comprise:

prompting the machine learning tool ML 3 or ML 4 with a document record pertinent to a copilot objective; and

receiving the supporting data source(s) from the machine learning tool ML 3 or ML 4 .

17. The system of claim 16 , wherein the operations further comprise:

validating a subset of the received data source(s).

18. The system of claim 15 , wherein the custom-trained machine learning tool ML 4 is a first custom-trained machine learning tool, and the operations further comprise:

(d) producing a second custom-trained machine learning tool ML 5 from the occupation-trained machine learning tool ML 3 by training the machine learning tool ML 3 to perform second custom tasks of the given occupation with at least a fourth predetermined performance level.

19. The system of claim 15 , wherein the given occupation comprises interviewing.

20. The system of claim 15 , wherein the given occupation comprises annotating interviews or annotating recorded work sessions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2025
From: KELSEY, ELAINE; NASIR, SAZZAD MAHMUD; YARBRO, JEFFREY THOMAS; EGERTON, LAUREN ELIZABETH; ROBSON, ELLIOT NICHOLAS; KELLY, BRENDAN MICHAEL; ROBSON, ROBERT OSCAR; WARD, SPENCER THOMAS
To: THIA ST CO.
Reel/Frame 069871/0410 →
Continuity (7)
Continuation PCTUS2024061934 · Dec 26, 2024
Continuation In Part 18898502 · Sep 26, 2024
Provisional Application 63717151 · Nov 6, 2024
Provisional Application 63709258 · Oct 18, 2024
Provisional Application 63646613 · May 13, 2024
Provisional Application 63561654 · Mar 5, 2024
Provisional Application 63620329 · Jan 12, 2024
References Cited (62)
US 8442940B1 · Faletti et al. · 2013 [cited by applicant]
US 10802488B1 · Abeloe · 2020 [cited by examiner]
US 10803127B2 · Alexander · 2020 [cited by examiner]
US 11379715B2 · Tang · 2022 [cited by examiner]
US 11443102B1 · Wilson et al. · 2022 [cited by applicant]
US 11922121B2 · Wang · 2024 [cited by examiner]
US 12093658B1 · Silver et al. · 2024 [cited by applicant]
US 12136043B1 · Davis · 2024 [cited by examiner]
US 12210973B2 · Johnson · 2025 [cited by examiner]
US 20150199646A1 · Taylor et al. · 2015 [cited by applicant]
US 20160246824A1 · Furuhashi et al. · 2016 [cited by applicant]
US 20170262811A1 · Ovadya · 2017 [cited by applicant]
US 20170262949A1 · Jay · 2017 [cited by applicant]
US 20180089593A1 · Patel et al. · 2018 [cited by applicant]
US 20190325353A1 · Aftab et al. · 2019 [cited by applicant]
US 20200050946A1 · Lecue et al. · 2020 [cited by applicant]
US 20200057946A1 · Singaraju et al. · 2020 [cited by applicant]
US 20200069208A1 · Keane · 2020 [cited by applicant]
US 20200241944A1 · Derdak et al. · 2020 [cited by applicant]
US 20200257733A1 · Mei et al. · 2020 [cited by applicant]
US 20200380254A1 · Zeng et al. · 2020 [cited by applicant]
US 20210035047A1 · Mossoba · 2021 [cited by examiner]
US 20210174347A1 · Rose · 2021 [cited by applicant]
US 20210233030A1 · Preuss et al. · 2021 [cited by applicant]
US 20210286831A1 · Girardi et al. · 2021 [cited by applicant]
US 20210295238A1 · Poon · 2021 [cited by examiner]
US 20210312399A1 · Asokan et al. · 2021 [cited by applicant]
US 20210357705A1 · Sung · 2021 [cited by examiner]
US 20210365782A1 · Huang · 2021 [cited by examiner]
US 20220067665A1 · Westerheide · 2022 [cited by examiner]
US 20220232085A1 · Panikkar et al. · 2022 [cited by applicant]
US 20220237892A1 · Anderton-Yang · 2022 [cited by applicant]
US 20220284312A1 · Brecque · 2022 [cited by applicant]
US 20220336060A1 · Koop et al. · 2022 [cited by applicant]
US 20230133373A1 · McGonnell et al. · 2023 [cited by applicant]
US 20230214192A1 · Makhija et al. · 2023 [cited by applicant]
US 20230214238A1 · Yitzhaki et al. · 2023 [cited by applicant]
US 20230289691A1 · Sabourin · 2023 [cited by applicant]
US 20240012842A1 · Kislal et al. · 2024 [cited by applicant]
US 20240037949A1 · Johnston et al. · 2024 [cited by applicant]
US 20240087743A1 · Grimm et al. · 2024 [cited by applicant]
US 20240095679A1 · Nowak et al. · 2024 [cited by applicant]
US 20240119383A1 · Bowers et al. · 2024 [cited by applicant]
US 20240176629A1 · Reddy · 2024 [cited by applicant]
US 20240202177A1 · Bandlamudi et al. · 2024 [cited by applicant]
US 20240303569A1 · Yu et al. · 2024 [cited by applicant]
US 20240362503A1 · Prasad et al. · 2024 [cited by applicant]
US 20250005288A1 · Amatriain-Rubio · 2025 [cited by examiner]
US 20250069308A1 · Cameron · 2025 [cited by examiner]
Morisot, “Add a SideNet to your MainNet,” ArXiv:2007.13512v1, 15 pages (Jul. 2020). [cited by applicant]
Agarwal et al., “Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training,” ArXiv:2010.12688 v2, 12 pages, (Mar. 13, 2021). [cited by applicant]
Bodor et al., “From Development to Deployment: An Approach to MLOps Monitoring for Machine Learning Model,” 14 [cited by applicant]
Dorsch et al., “GraphGuard: Enhancing Data Quality in Knowledge Graph Pipelines,” SEMIIM, 14 pages (Nov. 2023). [cited by applicant]
Feng et al., “Knowledge Card: Filling LLMs' Knowledge Gaps With Plug-in Specialized Language Models,” ArXiv: 2305.09955v2, pp. 1-24 (Oct. 2023). [cited by applicant]
Jeong, “A Study on the Implementation of Generative AI Services Using an Enterprise Data-Based LLM Application Architecture,” ArXiv 2309.01105vl, pp. 1-26 (Sep. 2023). [cited by applicant]
Kaddour et al., “Challenges and Applications of Large Language Models,” ArXiv 2307.10169v1, pp. 1-72 (Jul. 2023). [cited by applicant]
Liang et al. “Modular Retrieval for Generalization and Interpretation,” ArXiv:2303.13419v1, pp. 1-15 (Mar. 2023). [cited by applicant]
Liu et al., “Multi-stage pre-training over simplified multimodal pre-training models,” ArXiv:2107.14596, 10 pages (Jul. 22, 2021). [cited by applicant]
PCT/US2024/061299 Invitation to Pay Additional Fees, including Partial Search Report and Provisional Opinion, 15 pages (Apr. 9, 2025). [cited by applicant]
PCT/US2024/061934 Invitation to Pay Additional Fees, including Partial Search Report and Provisional Opinion, 20 pages (Apr. 4, 2025). [cited by applicant]
Roca et al., “Microservice chatbot architecture for chronic patient support,” Journal of Biomedical Informatics, pp. 1-19 (Oct. 2019). [cited by applicant]
Shao et al., “Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy,” ArXiv: 2305.15294v1, pp. 1-12 (May 2023). [cited by applicant]