IP Library › Granted Patent US 11,829,889
Granted Patent B2
US 11,829,889 · App. 17/831,450 · Granted Nov 28, 2023

Processing method and device for data of well site test based on knowledge graph

Inventors: Fei Tian (Beijing, CN); Qingyun Di (Beijing, CN); Wenhao Zheng (Beijing, CN); Zhongxing Wang (Beijing, CN); Yongyou Yang (Beijing, CN); Wenxiu Zhang (Beijing, CN); Renzhong Pei (Beijing, CN)
Assignee: Institute of Geology and Geophysics, Chinese Academy of Sciences
G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,829,889
App. No.
17/831,450
Granted
Nov 28, 2023
Kind
B2
Abstract

The present invention provides a processing method and device for data of a well site test based on a knowledge graph. The processing method for the data of the well site test based on the knowledge graph comprises: carrying out format identification on received historical data of the well site test to generate format identification results; establishing a mind map according to the format identification results; generating the knowledge graph of the data of the well site test according to the mind map; and processing the historical data of the well site test and new data of the well site test according to the knowledge graph.

Claims (184)

1. A processing method for data of a well site test based on a knowledge graph, comprising:

carrying out format identification on received historical data of a well site test, so as to generate format identification results;

establishing a mind map according to the format identification results;

generating a knowledge graph of the data of the well site test according to the mind map; and

processing the historical data of the well site test and new data of the well site test according to the knowledge graph;

the carrying out format identification on received historical data of the well site test to generate format identification results, comprising:

receiving and resolving binary streams and attached operation commands of files sent by a user through a network request; extracting information of file names, categories of operation objects and file formats from command parameters; scanning a target directory to determine whether folders corresponding to the extracted file formats exist or not; if a folder corresponding to the extracted file format is determined to be nonexistent, creating a new folder corresponding to the file format and writing a corresponding file under a newly created folder to generate a heterogeneous integrated database;

the establishing the mind map according to the format identification results, comprising:

determining keywords of the historical data of the well site test according to the format identification results;

establishing a data storage bank with a multi-level relationship according to the multiple keywords and a preset term dictionary of the well site test;

establishing the mind map according to the data storage bank;

the format identification results comprising: structured data, semi-structured data and unstructured data; the unstructured data comprising: a technical file, a picture/an audio/a video, an instrument and equipment ledger, actual drilling data;

the determining keywords of the historical data of the well site test according to the format identification results comprising:

carrying out grammar analysis on the structured data to determine keywords of the structured data;

calibrating labels of the semi-structured data and the unstructured data to determine keywords of the semi-structured data and the unstructured data;

the carrying out grammar analysis on the structured data to determine keywords of the structured data, specifically comprising: carrying out term extraction on the structured data according to the term dictionary of the well site test; selecting terms, a number of appearing frequencies of which is more than a preset number of times, from extraction results; generating feature vectors of the structured data according to the terms, the number of appearing frequencies of which is more than the preset number of times; and generating the keywords of the structured data according to the feature vectors;

the calibrating labels of the semi-structured data and the unstructured data to determine keywords of the semi-structured data and the unstructured data, comprising: calculating literal text similarities of the semi-structured data/the unstructured data and the term dictionary of the well site test; and selecting part of data from the semi-structured data and the unstructured data and calibrating the part of data according to the literal text similarities;

wherein a to-be-matched word B with the highest literal text similarity is selected as a label corresponding to data to be calibrated by the following formula, and the label is determined as the keyword to achieve extraction for the data of the semi-structured/unstructured data;

sim

=

60

×

(

xsword

ctrlword

+

xsword

keyword

)

/

2

+

40

×

dp

×

(

∑

c_xsword

⁢

(

i

)

ctrlword

⁡

(

i

)

+

∑

k_xsword

⁢

(

i

)

keyword

(

i

)

)

/

2

wherein sim represents the literal text similarity; xsword represents the number of same characters contained in a to-be-matched word A and a to-be-matched word B; ctrlword represents the total number of characters contained in the to-be-matched word A; keyword represents the total number of characters contained in the to-be-matched word B; dp represents a position coefficient representing a ratio of the total characters of the to-be-matched word A and the to-be-matched word B;

∑

c_xsword

⁢

(

i

)

∑

ctrlword

⁡

(

i

)

represents the sum of weight of the positions of the same characters contained in the to-be-matched word A and the to-be-matched word B in the to-be-matched word A;

∑

k_xsword

⁢

(

i

)

∑

keyword

(

i

)

represents the sum of weight of the positions of the same characters contained in the to-be-matched word A and the to-be-matched word B in the to-be-matched word B;

the generating a knowledge graph of the historical data of the well site test according to the mind map specifically comprising: carrying out granularity entity identification on the format identification results according to the mind map to generate identification results; establishing a knowledge level of the historical data of the well site test according to the identification results; extracting entity data of the historical data of the well site test according to the identification results; and generating the knowledge graph according to the knowledge level and the entity data.

2. An electronic equipment, comprising a memory, a processor and a computer program which is stored in the memory and can operate in the processor, wherein the processor is configured to achieve the steps of the processing method for data of a well site test based on a knowledge graph of claim 1 when executing the program.

3. A non-transitory computer readable storage medium, having stored thereon a computer program, wherein the computer program when executed by a processor achieves the steps of the processing method for data of a well site test based on a knowledge graph of claim 1 .

4. A processing device for data of a well site test based on a knowledge graph, comprising:

an identification result generation module, used for carrying out format identification on received historical data of the well site test to generate format identification results;

a mind map establishing module, used for establishing a mind map according to the format identification results;

a knowledge graph generation module, used for generating a knowledge graph of the historical data of the well site test according to the mind map; and

a data processing module, used for processing the historical data of the well site test and new data of the well site test according to the knowledge graph;

the carrying out format identification on received historical data of the well site test to generate format identification results, comprising:

receiving and resolving binary streams and attached operation commands of files sent by a user through a network request extracting information of file names, categories of operation objects and file formats from command parameters; scanning a target directory to determine whether folders corresponding to the extracted file formats exist or not if a folder corresponding to the extracted file format is determined to be nonexistent, creating a new folder corresponding to the file format and writing a corresponding file under a newly created folder to generate a heterogeneous integrated database;

the mind map establishing module comprising:

a keyword determination unit, used for determining keywords of the historical data of the well site test according to the format identification results;

a data storage bank establishing unit, used for establishing a data storage bank with a multi-level relationship according to the multiple keywords and a preset term dictionary of the well site test; and

a mind map establishing unit, used for establishing the mind map according to the data storage bank;

the format identification results comprising: structured data, semi-structured data and unstructured data; the unstructured data comprising: a technical file, a picture/an audio/a video, an instrument and equipment ledger, actual drilling data;

the keyword determination unit comprising:

a data grammar analysis unit, used for carrying out grammar analysis on the structured data to determine keywords of the structured data; and

a label calibrating unit, used for calibrating labels of the semi-structured data and the unstructured data to determine keywords of the semi-structured data and the unstructured data;

the data grammar analysis unit being specifically used for:

carrying out term extraction on the structured data according to the term dictionary of the well site test selecting terms, a number of appearing frequencies of which is more than a preset number of times, from extraction results; generating feature vectors of the structured data according to the terms, the number of appearing frequencies of which is more than the preset number of times; and

generating the keywords of the structured data according to the feature vectors;

the label calibrating unit being specifically used for:

calculating literal text similarities of the semi-structured data/the unstructured data and the term dictionary of the well site test and selecting part of data from the semi-structured data and the unstructured data and calibrating the part of data according to the literal text similarities;

wherein a to-be-matched word B with the highest literal text similarity is selected as a label corresponding to data to be calibrated by the following formula, and the label is determined as the keyword to achieve extraction for the data of the semi-structured/unstructured data;

sim

=

60

×

(

xsword

ctrlword

+

xsword

keyword

)

/

2

+

40

×

dp

×

(

∑

c_xsword

⁢

(

i

)

ctrlword

⁡

(

i

)

+

∑

k_xsword

⁢

(

i

)

keyword

(

i

)

)

/

2

wherein sim represents the literal text similarity; xsword represents the number of same characters contained in a to-be-matched word A and a to-be-matched word B; ctrlword represents the total number of characters contained in the to-be-matched word A; keyword represents the total number of characters contained in the to-be-matched word B; dp represents a position coefficient representing a ratio of the total characters of the to-be-matched word A and the to-be-matched word B;

∑

c_xsword

⁢

(

i

)

∑

ctrlword

⁡

(

i

)

represents the sum of weight of the positions of the same characters contained in the to-be-matched word A and the to-be-matched word B in the to-be-matched word A;

∑

k_xsword

⁢

(

i

)

∑

keyword

(

i

)

represents the sum of weight of the positions of the same characters contained in the to-be-matched word A and the to-be-matched word B in the to-be-matched word B;

the knowledge graph generation module being specifically used for:

carrying out granularity entity identification on the format identification results according to the mind map to generate identification results; establishing a knowledge level of the historical data of the well site test according to the identification results; extracting entity data of the historical data of the well site test according to the identification results; and generating the knowledge graph according to the knowledge level and the entity data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2022
From: TIAN, FEI; DI, QINGYUN; ZHENG, WENHAO; WANG, ZHONGXING; YANG, YONGYOU; ZHANG, WENXIU; PEI, RENZHONG
To: INSTITUTE OF GEOLOGY AND GEOPHYSICS, CHINESE ACADEMY OF SCIENCES
Reel/Frame 060907/0567 →
Priority Claims (1)
CN 202110719605.6 · Jun 28, 2021 · national
Continuity (1)
Related Publication 20220414488A1 · Dec 29, 2022