IP Library Patent Application 18315516
Patent Application
App. No. 18/315,516

DATA PREPROCESSING SYSTEM FOR CLEANING SMALL MOLECULE COMPOUND AND METHOD THEREOF

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/315,516
Abstract

The present invention provides a data preprocessing method for cleaning a small molecule compound, the data preprocessing method comprising: an S 1 text preprocessing step including: preprocessing an original SMILES text of a small molecule compound into a standardized SMILES text of the small molecule compound; and an S 2 chemical graph formatting step including: splitting in a format each text element of the standardized SMILES text of the small molecule compound of S 1 to obtain chemical graph information of the small molecule compound. The present invention also provides a data preprocessing system for cleaning a small molecule compound. The present invention enables the cleaning, deduplication, and standardization of global datasets, providing an efficient, fast, accurate integration method for the cleaning of end-to-end small molecule compounds.

Claims (46)

1 . A data preprocessing method for cleaning a small molecule compound, characterized in that the data preprocessing method comprising:

a step S 1 , text preprocessing step, including: preprocessing an original SMILES text of the small molecule compound into a standardized SMILES text of the small molecule compound according to predetermined text processing rules; and

a step S 2 , chemical graph formatting step, including: splitting in a format each text element of the standardized SMILES text of the small molecule compound of S 1 according to the predetermined text processing rules to obtain a digitized graph structure of chemical information of the small molecule compound.

2 . The data preprocessing method for cleaning a small molecule compound of claim 1 , further comprising:

a step S 3 , wherein the digitized graph structure of the chemical information of the small molecule compound of step S 2 is used for the construction of an artificial intelligence model.

3 . The data preprocessing method for cleaning a small molecule compound of claim 1 , wherein when the original SMILES text of the small molecule compound is preprocessed into the standardized SMILES text of the small molecule compound in the S 1 text preprocessing step, the predetermined text processing rules comprises:

step S 1 - 1 , optional structural normalization, wherein the data of the small molecule compound is processed into the original SMILES text;

step S 1 - 2 , respondent to the original SMILES text comprises heavy metal components and organic compound components, removing the heavy metal components from and retaining the organic compound components in the original SMILES text;

step S 1 - 3 , respondent to the original SMILES text comprises multimer components, removing the multimer components from and retaining a longest component in the original SMILES text;

step S 1 - 4 , respondent to the original SMILES text comprises a charge, adding or subtracting a hydrogen atom in the original SMILES text to remove the charge;

step S 1 - 5 , removing special SMILES text information; and

step S 1 - 6 , exporting normalized sequences to obtain the normalized SMILES text for the small molecule compound.

4 . The data preprocessing method for cleaning a small molecule compound of claim 1 , further comprising:

respondent to each text element of the standardized SMILES text of the small molecule compound of S 1 is split in a format in the S 2 chemical graph formatting step, the predetermined text processing rules comprises:

step S 2 - 1 , splitting the standardized SMILES text of the small molecule compound of S 1 into text elements of each core to obtain text elements of the small molecule compound;

step S 2 - 2 , performing text processing and identification on the properties of the text elements of the small molecule compound of step S 2 - 1 , and identifying and completing simplified chemical information to obtain a chemical information graph of the small molecule compound;

step S 2 - 3 , according to the chemical information graph of the small molecule compound in step S 2 - 2 , establishing a coordinate system with an atomic element as a node, and constructing a digital coordinate system of the chemical information graph of the small molecule compound; and

step S 2 - 4 , according to the digital coordinate system of the chemical information graph of the small molecule compound in step S 2 - 3 , and adding element attributes of nodes and edges to obtain a digitized graph structure of the chemical information of the small molecule compound.

5 . The data preprocessing method for cleaning a small molecule compound of claim 4 , further comprising:

step S 2 - 5 , complementing the hydrogen atom information of the digitized graph structure of the chemical information, if necessary.

6 . A data preprocessing system for cleaning a small molecule compound adapted for a data preprocessing method for cleaning a small molecule compound, characterized in that the data preprocessing method comprising:

a step S 1 , text preprocessing step, including: preprocessing an original SMILES text of the small molecule compound into a standardized SMILES text of the small molecule compound according to predetermined text processing rules; and

a step S 2 , chemical graph formatting step, including: splitting in a format each text element of the standardized SMILES text of the small molecule compound of S 1 according to the predetermined text processing rules to obtain a digitized graph structure of chemical information of the small molecule compound;

wherein the system comprises:

an S 1 text preprocessing unit configured to include preprocessing original SMILES data of the small molecule compound into a standardized SMILES text of the small molecule compound according to predetermined text processing rules; and

an S 2 chemical graph formatting unit configured to include: splitting in a format each text element of the standardized SMILES text of the small molecule compound of S 1 according to the predetermined text processing rules to obtain a digitized graph structure of chemical information of the small molecule compound.

7 . The data preprocessing system for cleaning a small molecule compound of claim 6 , further comprising an S 3 unit configured such that a digitized graph structure of the chemical information of the small molecule compound of S 2 is used in the construction of an artificial intelligence model.

8 . The data preprocessing system for cleaning a small molecule compound of claim 6 , when the original SMILES text of the small molecule compound is preprocessed into the standardized SMILES text of the small molecule compound in the S 1 text preprocessing unit, the predetermined text processing rule comprises:

an S 1 - 1 unit configured for optional structural normalization, wherein the data of the small molecule compound is processed into the original SMILES text;

an S 1 - 2 unit configured for, if the original SMILES text comprises heavy metal components and organic compound components, removing the heavy metal components from and retaining the organic compound components in the original SMILES text;

an S 1 - 3 unit configured for, if the original SMILES text comprises multimer components, removing the multimer components from and retaining a longest component in the original SMILES text;

an S 1 - 4 unit configured for, if the original SMILES text comprises a charge, adding or subtracting a hydrogen atom in the original SMILES text to remove the charge;

an S 1 - 5 unit configured for removing special SMILES text information; and

an S 1 - 6 unit configured for exporting normalized sequences to obtain the normalized SMILES text for the small molecule compound.

9 . The data preprocessing system for cleaning a small molecule compound of claim 6 , when each text element of the standardized SMILES text of the small molecule compound of S 1 is split in a format in the S 2 chemical graph formatting unit, the predetermined text processing rules comprises:

an S 2 - 1 unit configured for splitting the standardized SMILES text of the small molecule compound of S 1 into text elements of each core to obtain text elements of the small molecule compound;

an S 2 - 2 unit configured for performing text processing and identification on the properties of the text elements of the small molecule compound of the S 2 - 1 unit, and identifying and completing simplified chemical information to obtain a chemical information graph of the small molecule compound;

an S 2 - 3 unit configured for, according to the chemical information graph of the small molecule compound in the S 2 - 2 unit, establishing a coordinate system with an atomic element as a node, and constructing a digital coordinate system of the chemical information graph of the small molecule compound; and

an S 2 - 4 unit configured for, according to the digital coordinate system of the chemical information graph of the small molecule compound in the S 2 - 3 unit, adding element attributes of nodes and edges to obtain a digitized graph structure of the chemical information of the small molecule compound.

10 . The data preprocessing system for cleaning a small molecule compound of claim 9 , further comprising:

an S 2 - 5 unit configured for, if necessary, complementing the hydrogen atom information of the digitized graph structure of the chemical information.

11 . An electronic device comprising:

a memory; and

a processor, wherein the memory is configured to store one or more computer instructions which, when executed by the processor, implements a data preprocessing method for cleaning a small molecule compound, wherein the data preprocessing method comprises:

a step S 1 , text preprocessing step, including: preprocessing an original SMILES text of the small molecule compound into a standardized SMILES text of the small molecule compound according to predetermined text processing rules; and

a step S 2 , chemical graph formatting step, including: splitting in a format each text element of the standardized SMILES text of the small molecule compound of S 1 according to the predetermined text processing rules to obtain a digitized graph structure of chemical information of the small molecule compound.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2026
From: PAN, LURONG, DR.
To: AINNOCENCE LLC
Reel/Frame 073754/0650 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2023
From: AINNOCENCE INC.
To: PAN, LURONG, DR.
Reel/Frame 065741/0445 →
NUNC PRO TUNC ASSIGNMENT Recorded Nov 14, 2023
From: AINNOCENCE INC.
To: PAN, LURONG, DR.
Reel/Frame 065549/0701 →