IP Library › Granted Patent US 12,417,348
Granted Patent B2
US 12,417,348 · App. 18/185,675 · Granted Sep 16, 2025

Training data augmentation using gazetteers and perturbations to facilitate training named entity recognition models

Inventors: Omid Mohamad Nezami (Sydney, AU); Shivashankar Subramanian (Melbourne, AU); Thanh Tien Vu (Herston, AU); Tuyen Quang Pham (Springvale, AU); Budhaditya Saha (Sydney, AU); Aashna Devang Kanuga (Foster City, CA); Shubham Pawankumar Shah (Foster City, CA)
Assignee: Oracle International Corporation
G06F40/295G06N3/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,348
App. No.
18/185,675
Granted
Sep 16, 2025
Kind
B2
Abstract

Techniques are provided for augmenting training data using gazetteers and perturbations to facilitate training named entity recognition models. The training data can be augmented by generating additional utterances from original utterances in the training data and combining the generated additional utterances with the original utterances to form the augmented training data. The additional utterances can be generated by replacing the named entities in the original utterances with different named entities and/or perturbed versions of the named entities in the original utterances selected from a gazetteer. Gazetteers of named entities can be generated from the training data and expanded by searching a knowledge base and/or perturbing the named entities therein. The named entity recognition model can be trained using the augmented training data.

Claims (42)

1. A computer-implemented method comprising:

accessing training data comprising a plurality of original utterances, wherein each original utterance of the plurality of original utterances comprises at least one named entity corresponding to a named entity category of a plurality of named entity categories;

accessing one or more gazetteers, wherein each gazetteer of the one or more gazetteers comprises a plurality of named entities extracted from the plurality of original utterances and a plurality of perturbed named entities derived from one or more named entities of the plurality of named entities;

generating a plurality of template utterances, wherein each template utterance of the plurality of template utterances comprises information from an original utterance of the plurality of original utterances and at least one placeholder identifier representing a named entity in the original utterance of the plurality of original utterances;

generating a plurality of additional utterances, wherein each additional utterance of the plurality of additional utterances comprises a template utterance of the plurality of template utterances populated with at least one named entity selected from a gazetteer of the one or more gazetteers;

augmenting the training data by adding the plurality of additional utterances to the plurality of original utterances; and

training a named entity recognition (NER) model with the augmented training data.

2. The computer-implemented method of claim 1 , wherein each gazetteer of the one or more gazetteers corresponds to a different named entity category of the plurality of named entity categories.

3. The computer-implemented method of claim 1 , wherein at least one gazetteer of the one or more gazetteers comprises a plurality of named entities retrieved from a source other than the training data.

4. The computer-implemented method of claim 3 , wherein the plurality of named entities retrieved from the source other than the training data comprises named entities retrieved using at least one of a pre-trained model and a query-based search.

5. The computer-implemented method of claim 4 , wherein the source other than the training data comprises at least one knowledge base.

6. The computer-implemented method of claim 1 , wherein the plurality of perturbed named entities derived from one or more named entities of the plurality of named entities comprises perturbed versions of named entities of the plurality of named entities.

7. The computer-implemented method of claim 6 , wherein a particular perturbed version of the perturbed versions of named entities of the plurality of named entities comprises a named entity having at least one typographical error.

8. The computer-implemented method of claim 1 , further comprising:

providing the trained NER model to a system, wherein the providing the trained NER model includes detecting and classifying named entities in utterances received by the system from a user.

9. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions which, when executed by the one or more processors, cause the system to perform operations comprising:

accessing training data comprising a plurality of original utterances, wherein each original utterance of the plurality of original utterances comprises at least one named entity corresponding to a named entity category of a plurality of named entity categories;

accessing one or more gazetteers, wherein each gazetteer of the one or more gazetteers comprises a plurality of named entities extracted from the plurality of original utterances and a plurality of perturbed named entities derived from one or more named entities of the plurality of named entities;

generating a plurality of template utterances, wherein each template utterance of the plurality of template utterances comprises information from an original utterance of the plurality of original utterances and at least one placeholder identifier representing a named entity in the original utterance of the plurality of original utterances;

generating a plurality of additional utterances, wherein each additional utterance of the plurality of additional utterances comprises a template utterance of the plurality of template utterances populated with at least one named entity selected from a gazetteer of the one or more gazetteers;

augmenting the training data by adding the plurality of additional utterances to the plurality of original utterances; and

training a named entity recognition (NER) model with the augmented training data.

10. The system of claim 9 , wherein each gazetteer of the one or more gazetteers corresponds to a different named entity category of the plurality of named entity categories.

11. The system of claim 9 , wherein at least one gazetteer of the one or more gazetteers comprises a plurality of named entities retrieved from a particular source other than the training data.

12. The system of claim 11 , wherein the plurality of named entities retrieved from the particular source other than the training data comprises named entities retrieved using at least one of a pre-trained model and a query-based search.

13. The system of claim 12 , wherein the particular source other than the training data comprises at least one knowledge base.

14. The system of claim 9 , wherein the plurality of perturbed named entities derived from one or more named entities of the plurality of named entities comprises perturbed versions of named entities of the plurality of named entities.

15. The system of claim 14 , wherein a particular perturbed version of the perturbed versions of named entities of the plurality of named entities comprises a named entity having at least one typographical error.

16. The system of claim 9 , the operations further comprising:

providing the trained NER model to a system, wherein providing the trained NER model includes detecting and classifying named entities in utterances received by the system in which the trained NER model is provided to from a user.

17. One or more non-transitory computer-readable media storing instructions which, when executed by one or more processors, cause a system to perform operations comprising:

accessing training data comprising a plurality of original utterances, wherein each original utterance of the plurality of original utterances comprises at least one named entity corresponding to a named entity category of a plurality of named entity categories;

accessing one or more gazetteers, wherein each gazetteer of the one or more gazetteers comprises a plurality of named entities extracted from the plurality of original utterances and a plurality of perturbed named entities derived from one or more named entities of the plurality of named entities;

generating a plurality of template utterances, wherein each template utterance of the plurality of template utterances comprises information from an original utterance of the plurality of original utterances and at least one placeholder identifier representing a named entity in the original utterance of the plurality of original utterances;

generating a plurality of additional utterances, wherein each additional utterance of the plurality of additional utterances comprises a template utterance of the plurality of template utterances populated with at least one named entity selected from a gazetteer of the one or more gazetteers;

augmenting the training data by adding the plurality of additional utterances to the plurality of original utterances; and

training a named entity recognition (NER) model with the augmented training data.

18. The one or more non-transitory computer-readable media of claim 17 , wherein each gazetteer of the one or more gazetteers corresponds to a different named entity category of the plurality of named entity categories.

19. The one or more non-transitory computer-readable media of claim 17 , wherein at least one gazetteer of the one or more gazetteers comprises a plurality of named entities retrieved from a particular source other than the training data.

20. The one or more non-transitory computer-readable media of claim 17 , wherein the plurality of perturbed named entities derived from one or more named entities of the plurality of named entities comprises perturbed versions of named entities of the plurality of named entities.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2023
From: VU, THANH TIEN; SUBRAMANIAN, SHIVASHANKAR; NEZAMI, OMID MOHAMAD; PHAM, TUYEN QUANG; SAHA, BUDHADITYA; KANUGA, AASHNA DEVANG; SHAH, SHUBHAM PAWANKUMAR
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 063030/0209 →
Continuity (3)
Provisional Application 63362233 · Mar 31, 2022
Provisional Application 63362234 · Mar 31, 2022
Related Publication 20230325599A1 · Oct 12, 2023
References Cited (8)
US 20220222489A1 · Liu · 2022 [cited by examiner]
CN 111554278A · 2020 [cited by examiner]
Lara-Clares et al., LSI2 Uned at eHealth-KD Challenge 2019 a Few-shot Learning Model for Knowledge Discovery from eHealth Documents, CEUR Workshop Proceedings, Sep. 24, 2019, pp. 60-66. [cited by applicant]
Peshterliev et al., Self-Attention Gazetteer Embeddings for Named-Entity Recognition, Computer Science, Available Online at: https://arxiv.org/pdf/2004.04060.pdf, Apr. 18, 2020, 6 pages. [cited by applicant]
Runge et al., Exploring Neural Entity Representations for Semantic Information, Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, Nov. 17, 2020, 13 pages. [cited by applicant]
Sang et al., Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition, Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, Jun. 12, 2003, 6 pages. [cited by applicant]
Song et al., Gazetteer Generation for Neural Named Entity Recognition, The Thirty-Third International FLAIRS Conference (FLAIRS-33), May 2020, pp. 298-301. [cited by applicant]
Yamada et al., Wikipedia2Vec: An Efficient Toolkit for Learning and Visualizing the Embeddings of Words and Entities from Wikipedia, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing… [cited by applicant]