IP Library › Granted Patent US 12,639,349
Granted Patent B2
US 12,639,349 · App. 18/744,771 · Granted May 26, 2026

Method and system for metadata extraction for document identification

Inventors: Hydar A. Aqeel (Dhahran, SA); Hasan M. Asfoor (Dhahran, SA); Mario A. Dourado (Dhahran, SA); Dalal A. Alharbi (Dhahran, SA); Parwez P. Kothwaal Sheik (Dhahran, SA)
Assignee: SAUDI ARABIAN OIL COMPANY
G06F16/3334G06F16/383G06F16/387
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,349
App. No.
18/744,771
Granted
May 26, 2026
Kind
B2
Abstract

A method includes obtaining a document comprising earth property data regarding a geological region of interest, preprocessing the document to form at least one preprocessed document and determining, using a set of trained machine-learned models processing the at least one preprocessed document, a category of the document and a type of the document. The method further includes determining, using a natural language processing algorithm, metadata attributes of the document and a title of the document, and updating a database storing the document with the title, the category, the type and the metadata attributes. The method further includes identifying, by a planning module processing a query, the document from the database based on at least one of the title, the category, the type and the metadata attributes and planning a wellbore path in the geological region of interest using the earth property data comprised in the document.

Claims (72)

1 . A method, comprising:

obtaining a document comprising earth property data regarding a geological region of interest;

preprocessing the document to form at least one preprocessed document;

determining, using a set of trained machine-learned models processing the at least one preprocessed document, a category of the document and a type of the document, wherein the category is a class of subsurface exploration methods that the document belongs to;

determining, using a natural language processing algorithm processing the at least one preprocessed document, metadata attributes of the document and a title of the document;

updating a database storing the document with the title, the category, the type and the metadata attributes;

identifying, by a planning module processing a query, the document from the database based on at least one of the title, the category, the type and the metadata attributes; and

planning a wellbore path in the geological region of interest using the earth property data comprised in the document,

wherein determining the metadata attributes of the document comprises:

generating, from the at least one preprocessed document, a set of n-grams;

obtaining from an exploration reference database, based on the category, a set of possible attributes values; and

determining the metadata attributes based on an intersection between the set of n-grams and the set of possible attributes values.

2 . The method of claim 1 , further comprising:

determining a location of a hydrocarbon reservoir in the geological region of interest using the earth property data; and

planning the wellbore path so as to cause a wellbore to penetrate the hydrocarbon reservoir based on the location.

3 . The method of claim 1 , wherein the set of trained machine-learned models comprises at least one support vector machine.

4 . The method of claim 1 , wherein determining the metadata attributes is further based on one of a set of metadata extraction exceptions and a set of metadata extraction conditions.

5 . The method of claim 1 , wherein determining the title of the document comprises:

identifying a set of sentences within the at least one preprocessed document;

determining a weight of each sentence in the set of sentences; and

identifying the title based on the weights.

6 . The method of claim 5 , wherein the weight is determined based on at least one of a list of included and excluded words and a formatting pattern of the sentence.

7 . The method of claim 5 , identifying the title based on the weights further comprises:

determining that the weight of a sentence of the set of sentences meets a predetermined criterion; and

returning the sentence as the title.

8 . The method of claim 5 , identify the title based on the weights further comprises:

determining that the weights do not meet a predetermined criterion; and

returning the metadata attributes as the title.

9 . The method of claim 1 , prior to updating the database storing the document with the title, the category, the type and the metadata attributes, the method further comprising:

determining, by a quality checking module that the title, the category, the type and the metadata attributes satisfy a set of predetermined conditions and exemptions.

10 . A system, comprising:

a machine-learned model;

a machine-readable medium storing a natural language processing (NLP) algorithm; and

a computer configured to:

obtain a document comprising earth property data regarding a geological region of interest;

preprocess the document to form at least one preprocessed document;

determine, using a set of trained machine-learned models processing the at least one preprocessed document, a category of the document and a type of the document, wherein the category is a class of subsurface exploration methods that the document belongs to;

determine, using a natural language processing algorithm processing the at least one preprocessed document, metadata attributes of the document and a title of the document;

update a database storing the document with the title, the category, the type and the metadata attributes;

identify, by a planning module processing a query, the document from the database based on at least one of the title, the category, the type and the metadata attributes; and

plan a wellbore path in the geological region of interest using the earth property data comprised in the document,

wherein the determine the metadata attributes of the document comprises:

generate, from the at least one preprocessed document, a set of n-grams;

obtain from an exploration reference database, based on the category, a set of possible attributes values; and

determine the metadata attributes based on an intersection between the set of n-grams and the set of possible attributes values.

11 . The system of claim 10 , the computer further configured to:

determine a location of a hydrocarbon reservoir in the geological region of interest using the earth property data; and

plan the wellbore path so as to cause a wellbore to penetrate the hydrocarbon reservoir based on the location.

12 . The system of claim 10 , wherein the set of trained machine-learned models comprises at least one support vector machine.

13 . The system of claim 10 , wherein determine the metadata attributes is further based on one of a set of metadata extraction exceptions and a set of metadata extraction conditions.

14 . The system of claim 10 , the computer further configured to:

identify a set of sentences within the at least one preprocessed document;

determine a weight of each sentence in the set of sentences; and

identify the title based on the weights.

15 . The system of claim 14 , wherein the weight is determined based on at least one of a list of included and excluded words and a formatting pattern of the sentence.

16 . The system of claim 14 , the computer further configured to:

determine that the weight of a sentence of the set of sentences meets a predetermined criterion; and

return the sentence as the title.

17 . The system of claim 14 , the computer further configured to:

determine that the weights do not meet a predetermined criterion; and

return the metadata attributes as the title.

18 . A non-transitory machine-readable medium comprising a plurality of machine-readable instructions executed by one or more processors, the plurality of machine-readable instructions causing the one or more processors to perform a method comprising:

obtaining a document comprising earth property data regarding a geological region of interest;

preprocessing the document to form at least one preprocessed document;

determining, using a set of trained machine-learned models processing the at least one preprocessed document, a category of the document and a type of the document, wherein the category is a class of subsurface exploration methods that the document belongs to;

determining, using a natural language processing algorithm processing the at least one preprocessed document, metadata attributes of the document and a title of the document;

updating a database storing the document with the title, the category, the type and the metadata attributes;

identifying, by a planning module processing a query, the document from the database based on at least one of the title, the category, the type and the metadata attributes; and

planning a wellbore path in the geological region of interest using the earth property data comprised in the document wherein determining the metadata attributes of the document comprises:

generating, from the at least one preprocessed document, a set of n-grams;

obtaining from an exploration reference database, based on the category, a set of possible attributes values; and

determining the metadata attributes based on an intersection between the set of n-grams and the set of possible attributes values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2024
From: AQEEL, HYDAR A.; ASFOOR, HASAN M.; DOURADO, MARIO A.; ALHARBI, DALAL A.; KOTHWAAL SHEIK, PARWEZ P.
To: SAUDI ARABIAN OIL COMPANY
Reel/Frame 068789/0620 →
Continuity (1)
Related Publication 20250384065A1 · Dec 18, 2025
References Cited (14)
US 8843815B2 · Yang et al. · 2014 [cited by applicant]
US 11435499B1 · Peredriy · 2022 [cited by examiner]
US 20120030014A1 · Brunsman · 2012 [cited by examiner]
US 20210233008A1 · Gupta · 2021 [cited by examiner]
US 20230402065A1 · Gu · 2023 [cited by examiner]
US 20240411042A1 · Bolshakov · 2024 [cited by examiner]
US 20250270921A1 · Han · 2025 [cited by examiner]
US 20250284022A1 · Sun · 2025 [cited by examiner]
CN 117271627A · 2023 [cited by applicant]
WO 2022197420A1 · 2022 [cited by applicant]
Caso, Carlo, et al., “Search and Contextualization of Unstructured Data: Examples from the Norwegian Continental Shelf,” Society of Petroleum Engineers, SPE-200750-MS, 2020 (6 pages). [cited by applicant]
Asfoor, Hasan, et al., “Applying Data Science Techniques to Improve Information Discovery in Oil And Gas Unstructured Data,” International Petroleum Conference, IPTC-20236-Abstract, 2020 (9 pages). [cited by applicant]
Wang, Xiaohong, et al., “Semantic Mediation of Metadata for Marine Geochemical Data Integration,” ACM, 2016, DOI: http://dx.doi.org/10.1145/2925995.2926025 (5 pages). [cited by applicant]
Brazier, Pearl, et al., “Geo-Seed: A Metadata Repository for Geosciences Web Service Discovery,” IEEE Computer Society, 2009 (4 pages). [cited by applicant]