IP Library Granted Patent US 11,507,742
Granted Patent B1
US 11,507,742 · App. 16/454,220 · Granted Nov 22, 2022

Log parsing using language processing

Inventor: Wah-Kwan Lin (Melrose, MA)
Assignee: Rapid7, Inc.
G06F40/205G06N3/04G06N3/08H04L63/1425G06F40/295
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,507,742
App. No.
16/454,220
Granted
Nov 22, 2022
Kind
B1
Abstract

Methods and systems for parsing log records. A method involves receiving a log record including data regarding a network device's operation and providing the log record to a natural language processing model. The natural language processing model may analyze the log record to identify items in the log record and relationships between items in the log record.

Claims (40)

1. A method comprising:

receiving at an interface a first log record in a first format from a first network device;

providing the first log record to a language processing model implemented by a processor executing instructions stored on a memory;

performing, by the language processing model:

tokenizing the first log record to identify a plurality of items;

identifying respective item classes of the items in the first log record, wherein the language processing model is trained to recognize item classes including devices, users, actions, timestamps, and time durations; and

identifying relationships between the items identified in the first log record, wherein the language processing model is trained to recognize relationships between the item classes including an action is performed by a user, an action is performed on a device, a user is associated with a device, an action is performed at a timestamp, and an action is performed for a time duration; and

receiving from the model a report comprising the items, the item classes, and the relationships between the items identified in the first log record.

2. The method of claim 1 further comprising providing the report to a threat detection module to detect malicious activity associated with the first log record.

3. The method of claim 1 further comprising:

receiving at the interface a second log record in a second format from a second network device; and

providing the second log record to the language processing model,

wherein the report comprises the items included in the second log record and relationships between the items in the second log record.

4. The method of claim 1 wherein the language processing model is trained on a plurality of log records from different sources.

5. The method of claim 1 wherein the items identified in the first log record include at least one of a byte count, a port, and an IP address.

6. The method of claim 1 wherein the report includes a probabilistic assessment of the identified items or a probabilistic assessment of the relationships between the identified items.

7. The method of claim 1 wherein the language processing model is based on a convolutional neural network.

8. The method of claim 1 further comprising receiving feedback to the language processing model to revise the language processing model, wherein the feedback includes user input that annotates items and relationships in log records.

9. The method of claim 1 further comprising:

receiving a plurality of training log records from different sources;

annotating the plurality of training log records; and

providing the plurality of annotated training log records to train the language processing model.

10. The method of claim 1 further comprising providing the report to a log searching tool configured to conduct searches on log records.

11. A system comprising:

an interface for receiving at least a first log record in a first format from a first network device;

a memory; and

a processor configured to execute instructions stored on the memory to implement a language processing model and use the language processing model to:

tokenize the first log record to identify a plurality of items;

identify respective item classes of the items in the first log record, wherein the language processing model is trained to recognize item classes including devices, users, actions, timestamps, and time durations;

identify relationships between the items identified in the first log record, wherein the language processing model is trained to recognize relationships between the item classes including an action is performed by a user, an action is performed on a device, a user is associated with a device, an action is performed at a timestamp, and an action is performed for a time duration; and

generate a report comprising the items and item classes identified in the first log record and relationships between the items identified in the first log record.

12. The system of claim 11 further comprising a threat detection module configured to analyze the report to detect malicious activity associated with the first log record.

13. The system of claim 11 wherein the interface is further configured to receive a second log record in a second format from a second network device, and the report comprises the items identified in the second log record and relationships between the items identified in the second log record.

14. The system of claim 11 wherein the language processing model is trained on a plurality of log records from different sources.

15. The system of claim 11 wherein the items identified in the first log record include at least one of a byte count, a port, and an IP address.

16. The system of claim 11 wherein the report includes a probabilistic assessment of the identified items or a probabilistic assessment of the relationships between the identified items.

17. The system of claim 11 wherein the language processing model is based on a convolutional neural network.

18. The system of claim 11 wherein the interface is further configured to receive feedback regarding the report to revise the language processing model, wherein the feedback includes user input that annotates items and relationships in log records.

19. The system of claim 11 wherein the interface is further configured to receive a plurality of annotated training log records from different sources and provide the plurality of annotated training log records to train the language processing model.

20. The system of claim 11 further comprising a log searching tool configured to conduct searches on log records.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2025
From: RAPID7, INC.; RAPID7 LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 071743/0537 →
RELEASE OF SECURITY INTEREST Recorded Dec 27, 2024
From: KEYBANK NATIONAL ASSOCIATION, AS ADMINISTRATIVE AGENT
To: RAPID7, INC.
Reel/Frame 069785/0328 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 24, 2020
From: RAPID7, INC.
To: KEYBANK NATIONAL ASSOCIATION
Reel/Frame 052489/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2019
From: LIN, WAH-KWAN
To: RAPID7, INC.
Reel/Frame 051260/0514 →