IP Library Granted Patent US 11,538,457
Granted Patent B2
US 11,538,457 · App. 17/016,117 · Granted Dec 27, 2022

Noise data augmentation for natural language processing

Inventors: Elias Luqman Jalaluddin (Seattle, WA); Vishal Vishnoi (Redwood City, CA); Mark Edward Johnson (Castle Cove, AU); Thanh Long Duong (Seabrook, AU); Yu-Heng Hong (Carlton, AU); Balakota Srinivas Vinnakota (Sunnyvale, CA)
Assignee: ORACLE INTERNATIONAL CORPORATION
G10L15/063G10L15/05G10L15/18G10L15/22G10L15/26G10L2015/0633G10L2015/0638G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,538,457
App. No.
17/016,117
Granted
Dec 27, 2022
Kind
B2
Abstract

Techniques for noise data augmentation for training chatbot systems in natural language processing. In one particular aspect, a method is provided that includes receiving a training set of utterances for training an intent classifier to identify one or more intents for one or more utterances; augmenting the training set of utterances with noise text to generate an augmented training set of utterances; and training the intent classifier using the augmented training set of utterances. The augmenting includes: obtaining the noise text from a list of words, a text corpus, a publication, a dictionary, or any combination thereof irrelevant of original text within the utterances of the training set of utterances, and incorporating the noise text within the utterances relative to the original text in the utterances of the training set of utterances at a predefined augmentation ratio to generate augmented utterances.

Claims (20)

1. A method for training an intent classifier, the method comprising:

receiving, at a data processing system, a training set of utterances for training the intent classifier to identify one or more intents for one or more utterances;

augmenting, by the data processing system, the training set of utterances with noise text to generate an augmented training set of utterances, wherein the augmenting comprises:

obtaining the noise text from a list of words, a text corpus, a publication, a dictionary, or any combination thereof irrelevant of original text within utterances of the training set of utterances, wherein the noise text is random strings of text or sentences of text automatically generated or copied from the list of words, the text corpus, the publication, the dictionary, or any combination with or without consideration of frequencies of words or characters selected for the random strings of text or sentences of text, and

incorporating, using one or more noise augmentation operations, the noise text within the utterances positionally relative to the original text in the utterances of the training set of utterances at a predefined augmentation ratio of 1:0.5 to 1:5 to artificially generate augmented utterances, wherein the one or more noise augmentation operations cause the noise text to be incorporated: (i) in front of the original text within the utterances, (ii) after the original text of the utterances, (iii) flanking the original text of the utterances, (iv) integrated within the original text of the utterances, (v) or a combination thereof; and

training, by the data processing system, the intent classifier using the augmented training set of utterances.

2. A system comprising:

one or more data processors; and

a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform actions including:

receiving a training set of utterances for training an intent classifier to identify one or more intents for one or more utterances;

augmenting the training set of utterances with noise text to generate an augmented training set of utterances, wherein the augmenting comprises:

obtaining the noise text from a list of words, a text corpus, a publication, a dictionary, or any combination thereof irrelevant of original text within utterances of the training set of utterances, wherein the noise text is random strings of text or sentences of text automatically generated or copied from the list of words, the text corpus, the publication, the dictionary, or any combination with or without consideration of frequencies of words or characters selected for the random strings of text or sentences of text, and

incorporating, using one or more noise augmentation operations, the noise text within the utterances positionally relative to the original text in the utterances of the training set of utterances at a predefined augmentation ratio of 1:0.5 to 1:5 to artificially generate augmented utterances, wherein the one or more noise augmentation operations cause the noise text to be incorporated: (i) in front of the original text within the utterances, (ii) after the original text of the utterances, (iii) flanking the original text of the utterances, (iv) integrated within the original text of the utterances, (v) or a combination thereof; and

training the intent classifier using the augmented training set of utterances.

3. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform actions including:

receiving a training set of utterances for training an intent classifier to identify one or more intents for one or more utterances;

augmenting the training set of utterances with noise text to generate an augmented training set of utterances, wherein the augmenting comprises:

obtaining the noise text from a list of words, a text corpus, a publication, a dictionary, or any combination thereof irrelevant of original text within utterances of the training set of utterances, wherein the noise text is random strings of text or sentences of text automatically generated or copied from the list of words, the text corpus, the publication, the dictionary, or any combination with or without consideration of frequencies of words or characters selected for the random strings of text or sentences of text, and

incorporating, using one or more noise augmentation operations, the noise text within the utterances positionally relative to the original text in the utterances of the training set of utterances at a predefined augmentation ratio of 1:0.5 to 1:5 to artificially generate augmented utterances, wherein the one or more noise augmentation operations cause the noise text to be incorporated: (i) in front of the original text within the utterances, (ii) after the original text of the utterances, (iii) flanking the original text of the utterances, (iv) integrated within the original text of the utterances, (v) or a combination thereof; and

training the intent classifier using the augmented training set of utterances.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2020
From: JALALUDDIN, ELIAS LUQMAN; VISHNOI, VISHAL; JOHNSON, MARK EDWARD; DUONG, THANH LONG; HONG, YU-HENG; VINNAKOTA, BALAKOTA SRINIVAS
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 053728/0285 →
Continuity (2)
Provisional Application 63002066 · Mar 30, 2020
Related Publication 20210304733A1 · Sep 30, 2021
Cited By (3)
US 12,579,447 US 12,579,448 US 12,682,257