IP Library Granted Patent US 11,442,804
Granted Patent B2
US 11,442,804 · App. 16/728,102 · Granted Sep 13, 2022

Anomaly detection in data object text using natural language processing (NLP)

Inventor: Dmitry Martyanov (Cupertino, CA)
Assignee: PAYPAL, INC.
G06F11/079G06F11/3476G06F40/20G06F40/279G06K9/6267G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,442,804
App. No.
16/728,102
Granted
Sep 13, 2022
Kind
B2
Abstract

Systems and methods are disclosed for detecting anomalies in text content of data objects even when a format of the data and/or data object is unknown. These may include receiving a first data object that corresponds to a first application service and that includes first text content. An anomaly classifier may be trained based on an artificial neural network by using a natural language processing algorithm on respective text content of at least a portion of each of a plurality of data objects corresponding to the first computing service. Each of the plurality of data objects may be labeled as belonging a category. The trained anomaly classifier may identify one or more text character sequences in the first text content of the first data object as anomalous and output identifying information indicating the one or more anomalous text character sequences in the first text content of the first data object.

Claims (53)

1. A system, comprising:

a non-transitory memory storing instructions; and

one or more hardware processors coupled to the non-transitory memory and configured to read the instructions from the non-transitory memory to cause the system to perform operations comprising:

accessing at least a portion of a first plurality of data objects corresponding to a computing service, wherein each of the first plurality of data objects includes text content;

training an anomaly classifier based on an artificial neural network by using a natural language processing algorithm on the text content of each of at least the portion of the first plurality of data objects, wherein each of the first plurality of data objects is labeled as having a first condition or as having a second condition, and wherein a complete structural format of the text content of the data objects is not available to the system during the training; and

based on the training, producing a trained anomaly classifier that can identify one or more text character sequences in particular text content of a particular data object having the second condition as anomalous.

2. The system of claim 1 , wherein the operations further comprise:

receiving a first data object that corresponds to the computing service and that includes first text content;

identifying, using the trained anomaly classifier, one or more text character sequences in the first text content of the first data object as anomalous; and

outputting identifying information indicating the one or more anomalous text character sequences in the first text content of the first data object.

3. The system of claim 2 , wherein the first data object is labeled as having the first condition, and wherein the first condition is a no error condition.

4. The system of claim 1 , wherein the second condition is an error condition.

5. The system of claim 1 , wherein the operations further comprise:

receiving a second plurality of data objects that each corresponds to the computing service and that each includes text content;

identifying, using the trained anomaly classifier, one or more first common anomalous text character sequences in first text content of a first set of the second plurality of data objects;

identifying, using the trained anomaly classifier, one or more second common anomalous text character sequences in second text content of a second set of the second plurality of data objects; and

outputting, in response to determining that the one or more first common anomalous text character sequences satisfies a predetermined condition, identifying information indicating the one or more first common anomalous text character sequences in the first set of the second plurality of data objects.

6. The system of claim 5 , wherein the operations further comprise:

omitting, in response to determining that the one or more second common anomalous text character sequences does not satisfy a predetermined condition, identifying information indicating the one or more second common anomalous text character sequences in the second set of the second plurality of data objects.

7. The system of claim 1 , wherein the training the trained anomaly classifier includes:

using an unsupervised neural network to extract features from the text content; and

training, using the extracted features and the labels of each of the first plurality of data objects, a supervised neural network to determine one or more text character sequences in the text content in the first plurality of data objects as anomalous.

8. A method, comprising

receiving a first data object that corresponds to a first computing service and that includes first text content;

accessing a trained anomaly classifier, wherein the trained anomaly classifier was trained based on an artificial neural network by using a natural language processing algorithm on respective text content of at least a portion of each of a plurality of data objects corresponding to the first computing service, and wherein each of the plurality of data objects is labeled as belonging to one of a plurality of categories;

identifying, using the trained anomaly classifier, one or more text character sequences in the first text content of the first data object as anomalous; and

outputting identifying information indicating the one or more anomalous text character sequences in the first text content of the first data object.

9. The method of claim 8 , wherein the first data object is a log file that is generated based on communications between the first computing service and a second computing service.

10. The method of claim 8 , further comprising:

determining that the first data object is included in a first category of the plurality of categories, wherein the accessing the trained anomaly classifier is in response to the determining that the first data object is included in the first category of the plurality of categories.

11. The method of claim 10 , wherein the first category of the plurality of categories is an error condition.

12. The method of claim 8 , further comprising:

sending the first data object to a second computing service; and

receiving an error message from the second computing service that the first data object resulted in an error, wherein the identifying is performed in response to receiving the error message.

13. The method of claim 8 , wherein outputting identifying information indicating the one or more anomalous text character sequences includes causing the one or more anomalous text character sequences to be visually augmented on a user interface to appear different than other text content of the first text content.

14. The method of claim 8 , further comprising:

updating the trained anomaly classifier by at least one of penalizing the artificial neural network for incorrectly identifying one or more anomalous text character sequences or rewarding the artificial neural network for correctly identifying one or more anomalous text character sequences.

15. A non-transitory machine-readable medium having stored thereon machine-readable instructions executable to cause a machine to perform operations comprising:

receiving a first data object that corresponds to a first computing service and that includes first text content;

accessing a trained anomaly classifier, wherein the trained anomaly classifier was trained based on an artificial neural network by using a natural language processing algorithm on respective text content of at least a portion of each of a plurality of data objects corresponding to the first computing service, and wherein each of the plurality of data objects is labeled an error condition or a no error condition;

identifying, using the trained anomaly classifier, one or more text character sequences in the first text content of the first data object as anomalous; and

outputting identifying information indicating the one or more anomalous text character sequences in the first text content of the first data object.

16. The non-transitory machine-readable medium of claim 15 , wherein the first data object is a log file that is generated from communications between the first computing service and a second computing service.

17. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:

determining that the first data object is associated with an error condition, wherein the accessing the trained anomaly classifier is in response to the determining that the first data object is associated with the error condition.

18. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:

sending the first data object to a second computing service; and

receiving an error message from the second computing service that the first data object resulted in an error, wherein the identifying is performed in response to receiving the error message.

19. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:

receiving a second data object that corresponds to the first computing service and that includes second text content;

identifying, using the trained anomaly classifier, one or more text character sequences in the second text content of the second data object as non-anomalous; and

outputting identifying information indicating the one or more non-anomalous text character sequences in the second text content of the second data object that correspond with the one or more anomalous text character sequences in the first text content of the first data object.

20. The non-transitory machine-readable medium of claim 19 , wherein the identifying information indicating the one or more non-anomalous text character sequences is visually augmented differently than the identifying information indicating the one or more anomalous text character sequences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2019
From: MARTYANOV, DMITRY
To: PAYPAL, INC.
Reel/Frame 051395/0156 →
Continuity (1)
Related Publication 20210200612A1 · Jul 1, 2021
Cited By (1)
US 12,277,223