IP Library › Granted Patent US 12,682,257
Granted Patent B2
US 12,682,257 · App. 18/163,235 · Granted Jul 14, 2026

Techniques for negative entity aware augmentation

Inventors: Ahmed Ataallah Ataallah Abobakr (Geelong, AU); Shivashankar Subramanian (Melbourne, AU); Ying Xu (Albion, AU); Vladislav Blinov (Melbourne, AU); Umanga Bista (Southbank, AU); Tuyen Quang Pham (Springvale, AU); Thanh Long Duong (Seabrook, AU); Mark Edward Johnson (Sydney, AU); Elias Luqman Jalaluddin (Seattle, WA); Vanshika Sridharan (San Mateo, CA); Xin Xu (San Jose, CA); Srinivasa Phani Kumar Gadde (Fremont, CA); Vishal Vishnoi (Redwood City, CA)
Assignee: Oracle International Corporation
G06N5/022G06F40/247G06F40/295G06F40/56G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,257
App. No.
18/163,235
Filed
Feb 1, 2023
Granted
Jul 14, 2026
Kind
B2
Art Unit
2654
USPC
704/9
Abstract

Novel techniques are described for negative entity-aware augmentation using a two-stage augmentation to improve the stability of the model to entity value changes for intent prediction. In some embodiments, a method comprises accessing a first set of training data for an intent prediction model, the first set of training data comprising utterances and intent labels; applying one or more negative entity-aware data augmentation techniques to the first set of training data, depending on the tuning requirements for hyper-parameters, to result in a second set of training data, where the one or more negative entity-aware data augmentation techniques comprise Keyword Augmentation Technique (“KAT”) plus entity without context technique and KAT plus entity in random context as OOD technique; combining the first set of training data and the second set of training data to generate expanded training data; and training the intent prediction model using the expanded training data.

Claims (43)

1 . A method, comprising:

accessing a first set of training data, the first set of training data comprising utterances;

extracting named entities from the first set of training data using a trained Named Entity Recognition (NER) model;

applying one or more negative entity-aware data augmentation techniques to the first set of training data, depending on tuning requirements for hyper-parameters, to result in a second set of training data;

wherein the one or more negative entity-aware data augmentation techniques comprise Keyword Augmentation Technique (“KAT”) plus entity without context technique and KAT plus entity in random context as OOD technique;

combining the first set of training data and the second set of training data to generate expanded training data;

training a machine learning model using the expanded training data to result in a trained machined learning model; and

deploying the trained machine learning model.

2 . The method of claim 1 , further comprising applying Stop Word Augmentation Technique (“SWAT”) by replacing non-stop words in the first set of training data with random stop words while preserving existing stop words to result in the second set of training data.

3 . The method of claim 1 , wherein the KAT plus entity without context technique comprises removing in-domain context information around the named entities within the first set of training data resulting in a modified data as an out-of-domain training data.

4 . The method of claim 3 , wherein vector distance between the out-of-domain training data and any in-domain training data is above a threshold value.

5 . The method of claim 1 , wherein the KAT plus entity in random context as OOD technique comprises replacing context around entity values within the first set of training data with random out-of-domain context.

6 . The method of claim 1 , wherein the one or more negative entity-aware data augmentation techniques train the machine learning model to focus on overall context rather than changes to an entity value of the utterances.

7 . The method of claim 6 , wherein the entity value comprises one or more words representing an individual named entity within a named entity category.

8 . A non-transitory computer-readable medium storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

accessing a first set of training data, the first set of training data comprising utterances;

extracting named entities from the first set of training data using a trained NER model;

applying one or more negative entity-aware data augmentation techniques to the first set of training data, depending on tuning requirements for hyper-parameters, to result in a second set of training data;

wherein the one or more negative entity-aware data augmentation techniques comprise Keyword Augmentation Technique (“KAT”) plus entity without context technique and KAT plus entity in random context as OOD technique;

combining the first set of training data and the second set of training data to generate expanded training data;

training a machine learning model using the expanded training data to result in a trained machined learning model; and

deploying the trained machine learning model.

9 . The non-transitory computer-readable medium of claim 8 , further comprising applying Stop Word Augmentation Technique (“SWAT”) by replacing non-stop words in the first set of training data with random stop words while preserving existing stop words to result in the second set of training data.

10 . The non-transitory computer-readable medium of claim 8 , wherein the KAT plus entity without context technique comprises removing in-domain context information around the named entities within the first set of training data resulting in a modified data as an out-of-domain training data.

11 . The non-transitory computer-readable medium of claim 10 , wherein vector distance between the out-of-domain training data and any in-domain training data is above a threshold value.

12 . The non-transitory computer-readable medium of claim 8 , wherein the KAT plus entity in random context as OOD technique comprises replacing context around entity values within the first set of training data with random out-of-domain context.

13 . The non-transitory computer-readable medium of claim 8 , wherein the one or more negative entity-aware data augmentation techniques train the machine learning model to focus on overall context rather than changes to an entity value of the utterances.

14 . The non-transitory computer-readable medium of claim 13 , wherein the entity value comprises one or more words representing an individual named entity within a named entity category.

15 . A system, comprising:

one or more processors; and

one or more non-transitory computer readable media storing computer-executable instructions that, when executed by the one or more processors, cause the system to perform:

accessing a first set of training data, the first set of training data comprising utterances;

extracting named entities from the first set of training data using a trained NER model;

applying one or more negative entity-aware data augmentation techniques to the first set of training data, depending on tuning requirements for hyper-parameters, to result in a second set of training data;

wherein the one or more negative entity-aware data augmentation techniques comprise Keyword Augmentation Technique (“KAT”) plus entity without context technique and KAT plus entity in random context as OOD technique;

combining the first set of training data and the second set of training data to generate expanded training data;

training a machine learning model using the expanded training data to result in a trained machine learning model; and

deploying the trained machine learning model.

16 . The system of claim 15 , further comprising applying Stop Word Augmentation Technique (“SWAT”) by replacing non-stop words in the first set of training data with random stop words while preserving existing stop words to result in the second set of training data.

17 . The system of claim 15 , wherein the KAT plus entity without context technique comprises removing in-domain context information around the named entities within the first set of training data resulting in a modified data as an out-of-domain training data.

18 . The system of claim 17 , wherein vector distance between the out-of-domain training data and any in-domain training data is above a threshold value.

19 . The system of claim 15 , wherein the KAT plus entity in random context as OOD technique comprises replacing context around entity values within the first set of training data with random out-of-domain context.

20 . The system of claim 3 , wherein the one or more negative entity-aware data augmentation techniques train the machine learning model to focus on overall context rather than changes to an entity value of the utterances.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 062833 FRAME: 0393. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 2, 2023
From: ABOBAKR, AHMED ATAALLAH ATAALLAH; SUBRAMANIAN, SHIVASHANKAR; XU, YING; BLINOV, VLADISLAV; BISTA, UMANGA; PHAM, TUYEN QUANG; DUONG, THANH LONG; JOHNSON, MARK EDWARD; JALALUDDIN, ELIAS LUQMAN; SRIDHARAN, VANSHIKA; XU, XIN; GADDE, SRINIVASA PHANI KUMAR; VISHNOI, VISHAL
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 062917/0703 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2023
From: ABOB, AHMED ATAALLAH ATAALLAH; SUBRAMANIAN, SHIVASHANKAR; XU, YING; BLINOV, VLADISLAV; BISTA, UMANGA; PHAM, TUYEN QUANG; DUONG, THANH LONG; JOHNSON, MARK EDWARD; JALALUDDIN, ELIAS LUQMAN; SRIDHARAN, VANSHIKA; XU, XIN; GADDE, SRINIVASA PHANI KUMAR; VISHNOI, VISHAL
To: CORPORATION, ORACLE I
Reel/Frame 062833/0393 →
Continuity (2)
Provisional Application 63354675 · Jun 22, 2022
Related Publication 20230419127A1 · Dec 28, 2023
References Cited (61)
US 8321220B1 · Chotimongkol et al. · 2012 [cited by applicant]
US 10607042B1 · Dasgupta et al. · 2020 [cited by applicant]
US 11158311B1 · Zhang · 2021 [cited by applicant]
US 11281857B1 · Craft et al. · 2022 [cited by applicant]
US 11538457B2 · Jalaluddin et al. · 2022 [cited by applicant]
US 12340792B2 · Qu · 2025 [cited by examiner]
US 20160055240A1 · Tur et al. · 2016 [cited by applicant]
US 20160350288A1 · Wick et al. · 2016 [cited by applicant]
US 20170069310A1 · Hakkani-Tur et al. · 2017 [cited by applicant]
US 20180314689A1 · Wang · 2018 [cited by examiner]
US 20180358000A1 · Amid et al. · 2018 [cited by applicant]
US 20180358001A1 · Amid et al. · 2018 [cited by applicant]
US 20190180196A1 · Terry · 2019 [cited by examiner]
US 20200226212A1 · Tan et al. · 2020 [cited by applicant]
US 20200257857A1 · Peper et al. · 2020 [cited by applicant]
US 20200320351A1 · Nikolenko et al. · 2020 [cited by applicant]
US 20220129644A1 · Kang et al. · 2022 [cited by applicant]
US 20220171930A1 · Jalaluddin et al. · 2022 [cited by applicant]
US 20220222441A1 · Liu · 2022 [cited by examiner]
US 20220366893A1 · Qu · 2022 [cited by examiner]
US 20230103728A1 · Liu et al. · 2023 [cited by applicant]
US 20230419040A1 · Abobakr · 2023 [cited by examiner]
US 20230419052A1 · Abobakr · 2023 [cited by examiner]
US 20230419127A1 · Abobakr · 2023 [cited by examiner]
US 20240062264A1 · Trikha · 2024 [cited by applicant]
US 20240096155A1 · Sachdeva et al. · 2024 [cited by applicant]
CN 107515857A · 2017 [cited by applicant]
EP 3183728A1 · 2017 [cited by applicant]
WO 2016028946A1 · 2016 [cited by applicant]
WO 2016055240A1 · 2016 [cited by applicant]
Botframework—Training None Intent in LUIS—Stack Overflow, Available Online at: https://stackoverflow.com/questions/46116260/training-none-intent-in-luis, Sep. 8, 2017, pp. 1-2. [cited by applicant]
U.S. Appl. No. 17/016,117, First Action Interview Pilot Program Pre-Interview Communication mailed on Jun. 9, 2022, 4 pages. [cited by applicant]
U.S. Appl. No. 17/016,117, Notice of Allowance mailed on Aug. 10, 2022, 13 pages. [cited by applicant]
U.S. Appl. No. 17/016,122, Final Office Action mailed on Sep. 30, 2022, 47 pages. [cited by applicant]
U.S. Appl. No. 17/016,122, Non-Final Office Action mailed on Apr. 26, 2022, 21 pages. [cited by applicant]
U.S. Appl. No. 17/345,288, Non-Final Office Action mailed on Dec. 22, 2022, 12 pages. [cited by applicant]
Abulaish et al., A Text Data Augmentation Approach for Improving the Performance of CNN, 2019 11th International Conference on Communication Systems & Networks (COMSNETS), Jan. 7, 2019, 6 pages. [cited by applicant]
Anaby-Tavor et al., Do Not Have Enough Data? Deep Learning to the Rescue!, Available Online at: arXiv:1911.03118v2, Cornell University Library, Nov. 27, 2019, 9 pages. [cited by applicant]
Bird et al., Chatbot Interaction with Artificial Intelligence: Human Data Augmentation with TS and Language Transformer Ensemble for Text Classification, Available Online at: arXiv:2010.05990v2, Cornell University Libra… [cited by applicant]
Coulombe, Text Data Augmentation Made Simple by Leveraging NLP Cloud APIs, Olin Library Cornell University, Ithaca, Available on internet at: https://arxiv.org/ftp/arxiv/papers/1812/1812.04718.pdf, Dec. 5, 2018, 33 page… [cited by applicant]
Dai et al., An Analysis of Simple Data Augmentation for Named Entity Recognition, Arxiv.Org, Cornell University Library, 201, Oct. 22, 2020, 7 pages. [cited by applicant]
Jalalvand et al., Automatic Data Expansion for Customer-Care Spoken Language Understanding, Available Online at: arXiv:1810.00670v1, Cornell University Library, Sep. 27, 2018, 10 pages. [cited by applicant]
Jurafsky et al., Regular Expressions, Text Normalization, Edit Distance, Speech and Language Processing, Available Online at: https://web.archive.org/web/20180219015352if_/http://web.stanford.edu:80/~jurafsky/slp3/2.pdf… [cited by applicant]
Li et al., An Unsupervised Learning Approach for NER Based on Online Encyclopedia, Advances in Databases And Information Systems; [Lecture Notes In Computer Science], Springer International Publishing, Jul. 18, 2019, 16… [cited by applicant]
Mintz et al., Distant Supervision for Relation Extraction Without Labeled Data, Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Pr… [cited by applicant]
International Application No. PCT/US2020/050342, International Search Report and Written Opinion mailed on Dec. 14, 2020, 9 pages. [cited by applicant]
International Application No. PCT/US2020/05040, International Preliminary Report on Patentability mailed on Mar. 31, 2022, 8 pages. [cited by applicant]
International Application No. PCT/US2020/050407, International Search Report and Written Opinion mailed on Dec. 14, 2020, 10 pages. [cited by applicant]
International Application No. PCT/US2021/036939, International Search Report and Written Opinion mailed on Sep. 24, 2021, 11 pages. [cited by applicant]
International Application No. PCT/US2021/060953, International Search Report and Written Opinion mailed on Mar. 15, 2022, 14 pages. [cited by applicant]
International Application No. PCT/US2021/060956, International Search Report and the Written Opinion mailed on May 24, 2022, 9 pages. [cited by applicant]
Wei et al., EDA: Easy Data Augmentation Techniques for Boosting performance on Text Classification Tasks, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International … [cited by applicant]
Xie et al., Data Noising as Smoothing in Neural Network Language Models, 5th International Conference on Learning Representations, Mar. 7, 2017, 12 pages. [cited by applicant]
International Application No. PCT/US2023/018351, International Search Report and Written Opinion mailed on Jul. 18, 2023, 10 pages. [cited by applicant]
Issifu et al., A Simple Data Augmentation Method to Improve the Performance of Named Entity Recognition Models in Medical Domain, IEEE Xplore, 6th International Conference on Computer Science and Engineering (UBMK), ret… [cited by applicant]
U.S. Appl. No. 18/163,230, Non-Final Office Action mailed on Feb. 21, 2025, 44 pages. [cited by applicant]
U.S. Appl. No. 18/163,231, Non-Final Office Action mailed on Feb. 19, 2025, 27 pages. [cited by applicant]
Kang et al., “UMLS-Based Data Augmentation for Natural Language Processing of Clinical Research Literature”, Journal of the American Medical Informatics Association, vol. 28, No. 4, Mar. 18, 2021, pp. 812-823. [cited by applicant]
Kobayashi, “Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations”, Available online at: https://arxiv.org/abs/1805.06201, May 16, 2018, 6 pages. [cited by applicant]
Application No. PCT/US2023/018351, International Preliminary Report on Patentability mailed On Jan. 2, 2025, 7 pages. [cited by applicant]
Sinha et al., “Negative Data Augmentation”, Cornell University, Available Online at: https://arxiv.org/pdf/2102.05113v1.pdf, Feb. 9, 2021, 17 pages. [cited by applicant]