IP Library › Granted Patent US 12,572,852
Granted Patent B2
US 12,572,852 · App. 18/087,647 · Granted Mar 10, 2026

Lexical dropout for natural language processing

Inventors: Tuyen Quang Pham (Melbourne, AU); Cong Duy Vu Hoang (Melbourne, AU); Thanh Tien Vu (Brisbane, AU); Mark Edward Johnson (Sydney, AU); Thanh Long Duong (Melbourne, AU)
Assignee: Oracle International Corporation
G06N20/00G06F40/253G06F40/284G06F40/295G06F40/35G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,852
App. No.
18/087,647
Granted
Mar 10, 2026
Kind
B2
Abstract

Techniques are provided for improved training of a machine learning model using lexical dropout. A machine learning model and a training data set are accessed. The training data set can include sample utterances and corresponding labels. A dropout parameter is identified. The dropout parameter can indicate a likelihood for dropping out one or more feature vectors for tokens associated with respective entities during training of the machine learning model. The dropout parameter is applied to feature vectors for tokens associated with respective entities. The machine learning model is trained using the training data set and the dropout parameter to generate a trained machine learning model. The use of the trained the machine learning model is facilitated.

Claims (51)

1 . A computer-implemented method for training a machine learning model to process audio or textual language input, the method comprising:

accessing a machine learning model;

accessing a training data set that includes sample utterances and corresponding labels;

identifying a dropout parameter that indicates a likelihood for dropping out one or more feature vectors for tokens labeled as respective entities during training of the machine learning model;

training the machine learning model using the training data set and the dropout parameter to generate a trained machine learning model, including selectively applying the dropout parameter to the one or more feature vectors for the tokens labeled as the respective entities, wherein, when the dropout parameter is applied, the machine learning model is caused to learn contextual information of the training data set, and wherein the contextual information includes a subset of the training data that is not labeled as an entity, and wherein the contextual information includes at least one or more verbs or at least one or more prepositions;

and

facilitating use of the trained the machine learning model to identify a named entity based upon an input utterance.

2 . The method of claim 1 , wherein the dropout parameter is a first dropout parameter, the method further comprising:

processing the training data set using a first model to generate a first feature vector; and

processing the training data set using a second model to generate a second feature vector,

wherein selectively applying the dropout parameter comprises applying the first dropout parameter to the first feature vector and applying a second dropout parameter to the second feature vector.

3 . The method of claim 1 , wherein the dropout parameter is a hyperparameter of the machine learning model, the method further comprising performing hypertuning to identify the dropout parameter.

4 . The method of claim 1 , wherein selectively applying the dropout parameter to the one or more feature vectors includes selectively applying the dropout parameter to one or more named entity tokens generated from the one or more feature vectors.

5 . The method of claim 1 , wherein selectively applying the dropout parameter to the one or more feature vectors includes selectively applying the dropout parameter to one or more tokens generated from the one or more feature vectors and matched via a gazetteer.

6 . The method of claim 1 , wherein:

the dropout parameter is a dropout rate;

applying the dropout parameter comprises dropping out the one or more feature vectors for the tokens labeled as the respective entities according to the dropout rate; and

the machine learning model includes a plurality of self-attention layers.

7 . A system comprising:

one or more processors; and

a non-transitory computer-readable memory coupled to the one or more processors, the memory comprising a plurality of instructions executable by the one or more processors to cause the one or more processors to perform operations comprising:

accessing a machine learning model;

accessing a training data set that includes sample utterances and corresponding labels;

identifying a dropout parameter that indicates a likelihood for dropping out one or more feature vectors for tokens labeled as respective entities during training of the machine learning model;

training the machine learning model using the training data set and the dropout parameter to generate a trained machine learning model, including selectively applying the dropout parameter to the one or more feature vectors for the tokens labeled as the respective entities, wherein, when the dropout parameter is applied, the machine learning model is caused to learn contextual information of the training data set, and wherein the contextual information includes a subset of the training data that is not labeled as an entity, and wherein the contextual information includes at least one or more verbs or at least one or more prepositions;

and

facilitating use of the trained the machine learning model to identify a named entity based upon an input utterance.

8 . The system of claim 7 , wherein the dropout parameter is a first dropout parameter, and wherein the operations further comprise:

processing the training data set using a first model to generate a first feature vector; and

processing the training data set using a second model to generate a second feature vector,

wherein selectively applying the dropout parameter comprises applying the first dropout parameter to the first feature vector and applying a second dropout parameter to the second feature vector.

9 . The system of claim 7 , wherein the dropout parameter is a hyperparameter of the machine learning model, and wherein the operations further comprise performing hypertuning to identify the dropout parameter.

10 . The system of claim 7 , wherein the operation of selectively applying the dropout parameter to the one or more feature vectors includes selectively applying the dropout parameter to one or more named entity tokens generated from the one or more feature vectors.

11 . The system of claim 7 , wherein the operation of selectively applying the dropout parameter to the one or more feature vectors includes selectively applying the dropout parameter to one or more tokens generated from the one or more feature vectors and matched via a gazetteer.

12 . The system of claim 7 , wherein:

the dropout parameter is a dropout rate;

the operation of applying the dropout parameter comprises dropping out the one or more feature vectors for the tokens labeled as the respective entities according to the dropout rate; and

the machine learning model includes a plurality of self-attention layers.

13 . A non-transitory computer-readable memory comprising a plurality of instructions executable by one or more processors to cause the one or more processors to perform operations comprising:

accessing a machine learning model and a training data set that includes sample utterances and corresponding labels;

identifying a dropout parameter that indicates a likelihood for dropping out one or more feature vectors for tokens labeled as respective entities during training of the machine learning model;

training the machine learning model using the training data set and the dropout parameter to generate a trained machine learning model, including selectively applying the dropout parameter to the one or more feature vectors for the tokens labeled as the respective entities, wherein, when the dropout parameter is applied, the machine learning model is caused to learn contextual information of the training data set, and wherein the contextual information includes a subset of the training data that is not labeled as an entity, and wherein the contextual information includes at least one or more verbs or at least one or more prepositions;

and

facilitating use of the trained the machine learning model to identify a named entity based upon an input utterance.

14 . The non-transitory computer-readable memory of claim 13 , wherein the dropout parameter is a first dropout parameter, and wherein the operations further comprise:

processing the training data set using a first model to generate a first feature vector; and

processing the training data set using a second model to generate a second feature vector,

wherein selectively applying the dropout parameter comprises applying the first dropout parameter to the first feature vector and applying a second dropout parameter to the second feature vector.

15 . The non-transitory computer-readable memory of claim 13 , wherein the dropout parameter is a hyperparameter of the machine learning model, and wherein the operations further comprise performing hypertuning to identify the dropout parameter.

16 . The non-transitory computer-readable memory of claim 13 , wherein the operation of selectively applying the dropout parameter to the one or more feature vectors includes selectively applying the dropout parameter to one or more named entity tokens generated from the one or more feature vectors.

17 . The non-transitory computer-readable memory of claim 13 , wherein the operation of selectively applying the dropout parameter to the one or more feature vectors includes selectively applying the dropout parameter to one or more tokens generated from the one or more feature vectors and matched via a gazetteer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2022
From: PHAM, TUYEN QUANG; HOANG, CONG DUY VU; VU, THANH TIEN; JOHNSON, MARK EDWARD; DUONG, THANH LONG
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 062214/0808 →
Continuity (2)
Provisional Application 63293441 · Dec 23, 2021
Related Publication 20230206125A1 · Jun 29, 2023
References Cited (23)
US 11604925B1 · Lee et al. · 2023 [cited by applicant]
US 20190213470A1 · Schmidt · 2019 [cited by applicant]
US 20190377790A1 · Redmond · 2019 [cited by examiner]
US 20230154455A1 · Vu et al. · 2023 [cited by applicant]
Meng, Yu, et al. “Distantly-supervised named entity recognition with noise-robust learning and language model augmented self-training.” arXiv preprint arXiv:2109.05003 (2021). (Year: 2021). [cited by examiner]
Eduardo Maldonado-Cruz, Michael J. Pyrcz, Tuning machine learning dropout for subsurface uncertainty model accuracy, Journal of Petroleum Science and Engineering, vol. 205, 2021, 108975, ISSN 0920-4105, https://doi.org/… [cited by examiner]
Peshterliev, Stanislav, Christophe Dupuy, and Imre Kiss. “Self-attention gazetteer embeddings for named-entity recognition.” arXiv preprint arXiv:2004.04060 (2020). (Year: 2020). [cited by examiner]
Xue, Xia, et al. “Convolutional recurrent neural networks with a self-attention mechanism for personnel performance prediction.” Entropy 21.12 (2019): 1227. (Year: 2019). [cited by examiner]
Emelyanov et al., “Multilingual Named Entity Recognition Using Pretrained Embeddings, Attention Mechanism and NCRF”, Proceedings of the 7th Workshop on Balto-Slavic Natural Language Processing, Aug. 2, 2019, pp. 94-99. [cited by applicant]
Liu et al., “Towards Improving Neural Named Entity Recognition with Gazetteers”, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28-Aug. 2, 2019, pp. 5301-5307. [cited by applicant]
Magnolini et al., “How to Use Gazetteers for Entity Recognition with Neural Models”, Proceedings of the 5th Workshop on Semantic Deep Learning (SemDeep-5), Available Online at: https://aclanthology.org/W19-5807.pdf, Aug… [cited by applicant]
Peshterliev et al., “Self-Attention Gazetteer Embeddings for Named-Entity Recognition”, Available Online at: https://arxiv.org/pdf/2004.04060.pdf, Apr. 18, 2020, 6 pages. [cited by applicant]
Ratinov et al., “Design Challenges and Misconceptions in Named Entity Recognition”, Proceedings of the Thirteenth Conference on Computational Natural Language Learning (CoNLL), Jun. 2009, pp. 147-155. [cited by applicant]
Song et al., “Gazetteer Generation for Neural Named Entity Recognition”, The Thirty-Third International Flairs Conference (Flairs-33), May 2020, pp. 298-301. [cited by applicant]
Song et al., “Improving Neural Named Entity Recognition with Gazetteers”, Available Online at: https://arxiv.org/pdf/2003.03072.pdf, Mar. 2020, 8 pages. [cited by applicant]
Srivastava et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”, Journal of Machine Learning Research, vol. 15, No. 1, Jan. 2014, pp. 1929-1958. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Jun. 2017, 11 pages. [cited by applicant]
Yang et al., “Drop-Out Conditional Random Fields for Twitter with Huge Mined Gazetteer”, Proceedings of NAACL-HLT, Jun. 12-17, 2016, pp. 282-288. [cited by applicant]
Zhang et al., “Token Drop Mechanism for Neural Machine Translation”, Proceedings of the 28th International Conference on Computational Linguistics, Dec. 8-13, 2020, pp. 4298-4303. [cited by applicant]
U.S. Appl. No. 18/087,629, Non-Final Office Action, mailed on Feb. 27, 2025, 17 pages. [cited by applicant]
Fetahu et al., “Gazetteer Enhanced Named Entity Recognition for Code-Mixed Web Queries”, SIGIR '21: Proceedings of the 44th International Association for Computing Machinery SIGIR Conference on Research and Development … [cited by applicant]
Lin et al., “Gazetteer-Enhanced Attentive Neural Networks for Named Entity Recognition”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference … [cited by applicant]
U.S. Appl. No. 18/087,629, “Non-Final Office Action”, mailed on May 27, 2025, 17 pages. [cited by applicant]