IP Library › Granted Patent US 12,210,828
Granted Patent B2
US 12,210,828 · App. 18/630,990 · Granted Jan 28, 2025

Methods and systems for generating mobile enabled extraction models

Inventors: Dominic Miguel Rossi (San Diego, CA); Hui Fang Lee (Mountain View, CA); Tharathorn Rimchala (San Francisco, CA)
Assignee: INTUIT INC.
G06F40/284G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,828
App. No.
18/630,990
Granted
Jan 28, 2025
Kind
B2
Abstract

A computing system generates a plurality of training data sets for generating the NLP model. The computing system trains a teacher network to extract and classify tokens from a document. The training includes a pre-training stage where the teacher network is trained to classify generic data in the plurality of training data sets and a fine-tuning stage where the teacher network is trained to classify targeted data in the plurality of training data sets. The computing system trains a student network to extract and classify tokens from a document by distilling knowledge learned by the teacher network during the fine-tuning stage from the teacher network to the student network. The computing system outputs the NLP model based on the training. The computing system causes the NLP model to be deployed in a remote computing environment.

Claims (51)

1. A computer-implemented method comprising:

generating a plurality of training data sets for generating a machine learning model;

training a teacher network, the training comprising a first stage where the teacher network is trained to classify a first data in the plurality of training data sets and a second stage where the teacher network is trained to classify a second data in the plurality of training data sets;

training a student network by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network, wherein distilling knowledge comprises implementing a distillation loss using weights that are same; and

outputting the machine learning model based on the training of the student network.

2. The computer-implemented method of claim 1 , wherein the teacher network is a bidirectional encoder representations for transformer (BERT) model and wherein the student network is a MobileBERT model.

3. The computer-implemented method of claim 1 , wherein generating the plurality of training data sets for generating the machine learning model, comprises:

generating a first plurality of training data sets using a text corpus for a generic training process; and

generating a second plurality of training data sets using targeted data for which the machine learning model is optimized.

4. The computer-implemented method of claim 1 , wherein training the student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

training the student network to approximate an original function learned by the teacher network during training of the teacher network.

5. The computer-implemented method of claim 1 , wherein training the student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

classifying tokens for subsequent mapping of tokens to one or more of a plurality of input data fields.

6. The computer-implemented method of claim 1 , wherein training a student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

inputting soft labels generated by the teacher network during training of the teacher network into the student network.

7. The computer-implemented method of claim 1 , wherein outputting the machine learning model based on the training, comprises:

compressing a trained student model into the machine learning model using dynamic quantization.

8. A non-transitory computer readable medium having one or more sequences of instructions, which, when executed by a processor, causes operations comprising:

generating a plurality of training data sets for generating a machine learning model;

training a teacher network, the training comprising a first stage where the teacher network is trained to classify a first data in the plurality of training data sets and a second stage where the teacher network is trained to classify a second data in the plurality of training data sets;

training a student network by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network, wherein distilling knowledge comprises implementing a distillation loss using weights that are same; and

outputting the machine learning model based on the training of the student network.

9. The non-transitory computer readable medium of claim 8 , wherein the teacher network is a bidirectional encoder representations for transformer (BERT) model and wherein the student network is a MobileBERT model.

10. The non-transitory computer readable medium of claim 8 , wherein generating the plurality of training data sets for generating the machine learning model, comprises:

generating a first plurality of training data sets using a text corpus for a generic training process; and

generating a second plurality of training data sets using targeted data for which the machine learning model is optimized.

11. The non-transitory computer readable medium of claim 8 , wherein training the student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

training the student network to approximate an original function learned by the teacher network during training of the teacher network.

12. The non-transitory computer readable medium of claim 8 , wherein training the student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

classifying tokens for subsequent mapping of tokens to one or more of a plurality of input data fields.

13. The non-transitory computer readable medium of claim 8 , wherein training a student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

inputting soft labels generated by the teacher network during training of the teacher network into the student network.

14. The non-transitory computer readable medium of claim 8 , wherein outputting the machine learning model based on the training, comprises:

compressing a trained student model into the machine learning model using dynamic quantization.

15. A system comprising:

a processor; and

a memory having one or more instructions stored thereon, which, when executed by the processor, causes the system to perform operations comprising:

generating a plurality of training data sets for generating a machine learning model;

training a teacher network, the training comprising a first stage where the teacher network is trained to classify a first data in the plurality of training data sets and a second stage where the teacher network is trained to classify a second data in the plurality of training data sets;

training a student network by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network, wherein distilling knowledge comprises implementing a distillation loss using weights that are same; and

outputting the machine learning model based on the training of the student network.

16. The system of claim 15 , wherein the teacher network is a bidirectional encoder representations for transformer (BERT) model and wherein the student network is a MobileBERT model.

17. The system of claim 15 , wherein generating the plurality of training data sets for generating the machine learning model, comprises:

generating a first plurality of training data sets using a text corpus for a generic training process; and

generating a second plurality of training data sets using targeted data for which the machine learning model is optimized.

18. The system of claim 15 , wherein training the student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

training the student network to approximate an original function learned by the teacher network during training of the teacher network.

19. The system of claim 15 , wherein training the student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

classifying tokens for subsequent mapping of tokens to one or more of a plurality of input data fields.

20. The system of claim 15 , wherein training a student network from a document by distilling knowledge learned by the teacher network during the second stage from the teacher network to the student network comprises:

inputting soft labels generated by the teacher network during training of the teacher network into the student network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2024
From: ROSSI, DOMINIC MIGUEL; LEE, HUI FANG; RIMCHALA, THARATHORN
To: INTUIT INC.
Reel/Frame 067067/0436 →
Continuity (2)
Continuation 17246277 · Apr 30, 2021
Related Publication 20240256775A1 · Aug 1, 2024
References Cited (12)
US 10559299B1 · Arel · 2020 [cited by applicant]
US 11461415B2 · Lu et al. · 2022 [cited by applicant]
US 20220067274A1 · Wang · 2022 [cited by applicant]
US 20220093253A1 · Misrilall · 2022 [cited by examiner]
US 20230229912A1 · Zhang · 2023 [cited by applicant]
WO WO2021060899 · 2021 [cited by applicant]
Houlsby et al., “Parameter-Efficient Transfer Learning for NLP”, arXiv:1902.00751v2 [cs.LG], Jun. 13, 2019, 13 pages. [cited by applicant]
Liu et al., “Multi-Task Deep Neural Networks for Natural Language Understanding”, arXiv:1901.11504v2 [cs.CL], May 30, 2019, 10 pages. [cited by applicant]
Sanh et al., “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”, arXiv:1910.01108v4 [cs.CL], Mar. 1, 2020, 5 pages. [cited by applicant]
Sun et al., “MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices”, arXiv:2004.02984v2 [cs.CL], Apr. 14, 2020, 13 pages. [cited by applicant]
Search Report dated Jun. 28, 2022 issued in International Application No. PCT/US2022/021703. [cited by applicant]
Written Opinion dated Jun. 28, 2022 issued in International Application No. PCT/US2022/021703. [cited by applicant]