Intelligent Oil and Gas Field Project of a Central State-owned Enterprise in China's Petroleum and Petrochemical Industry——Multi-type Data Storage and Computing Platform
Value Delivered
Massive Data Management: Provides a parallel processing platform for massive complex data, helping customers create a single petabyte-scale view of business data and delivering timely, efficient analytics results;
Data Security Protection: Replace foreign and open-source data storage and computing platforms with a suite of Chinese-developed database software to strengthen energy data security protection;
Unified Multi-Type Management: Achieve unified management and access to three major categories and six types of enterprise business data. Through the data governance platform, support data traceability, lineage management, and unified external services;
Solutions
The intelligent oil and gas field architecture comprises the data source layer, data aggregation layer, data storage layer, data computing layer, and data service layer.
Data Source Layer:
As the bottom layer, the data source layer includes six categories of structured, semi-structured, and unstructured data generated during daily oilfield operations: result graphics and documents (images, reports, construction documents, and reports), surveillance audio and video data, real-time IoT data, large-volume logging data, GIS imagery, and more.
Data Aggregation Layer:
The data aggregation layer aggregates data using GBaseMTK metadata synchronization tools, GBaseRTSync real-time sync tools, Kafka message queues, Sqoop, Flume, object storage, and other data transmission and processing utilities.
Data Storage Layer:
The data storage layer operates three major storage platforms: GBase8s as a high-concurrency transactional database, GBase8a as a distributed massively parallel OLAP analytical database, and a combination of HDFS, Hive, HBase (provisioned by GBaseHD), the Neo4j graph database, and object storage.
Data Computing Layer:
This layer includes computing engines and databases such as GBase8s in-database computing, GBase8a distributed computing, MapReduce, Spark, Flink, Elasticsearch, and graph computing engines. It performs computation and retrieval across the three major categories and six business data types.
Data Service Layer:
The system delivers data services to applications and external systems through structured data APIs, real-time data APIs, audio/video APIs, GIS data APIs, document and image APIs, volume data APIs, data push services, BI services, and AI services.
Requirements Analysis
The project requires a mature, stable big data platform software based on the Hadoop architecture to support storage, access, computing, and analytical processing of both structured and unstructured data. Through data virtualization, data federation, and efficient real-time data exchange channels between engines, it should integrate complex correlation analysis, stream computing, graph computing, and batch processing models to build a big data cloud platform. This platform will support OLAP, OLTP, and NoSQL workloads within the intelligent oil and gas field cloud environment.
Specific requirements include:
(1) Deployment within a central state-owned enterprise's intelligent oil and gas field cloud platform in the petroleum and petrochemical sector;
(2) Conduct big data analytics for storing unstructured data such as real-time data, volume data, and result data, along with external data sources, and perform data governance and standardization. Data originates from gas extraction plants, power dispatch, analytical laboratories, purification plants, as well as seismic data, logging curves, analytical test curves, geological model gridding data, video surveillance platforms, and selected external data sources;
(3) Establish a unified mathematical algorithm library and model repository to enable data mining analysis, including data retrieval, interactive analysis, and correlation analysis, supporting the development of intelligent oil and gas field operations;
(4) Real-time collection, storage, and analysis of streaming data;
(5) Support fast query analysis on massive datasets with sub-second query response times;
(6) Provide a unified graphical monitoring and management interface for interactive operations;
(7) Support data encryption, backup and recovery, data consistency verification, and operation auditing; unified system identity authentication and security access verification to control access to database engines and data; authorization based on users and roles.
Project Background
Aligned with the "13th Five-Year Plan" for informatization development in central state-owned oil and petrochemical enterprises, the overall master plan for the Smart Oil & Gas Field project, and the business goals of oilfield companies to enhance production efficiency and optimize operations—along with the demand for centralized integration and collaborative sharing—the construction of a pilot Smart Oil & Gas Field project was initiated.
The overarching goal of the Smart Oil & Gas Field initiative is to leverage the latest ICT technologies to build four core capabilities—comprehensive perception, integrated collaboration, early warning and prediction, and analytical optimization—around the full lifecycle management of core assets. This drives "efficient exploration and profitable development," maximizing enterprise asset value. The objective of this pilot phase is to build a Smart Oil & Gas Field cloud platform, supplement and refine the standardization and technical support frameworks, and enable dynamic reservoir management and optimization, as well as intelligent diagnostics and optimization for wells, pipelines, and equipment. The outcome will be a demonstration zone for smart oil and gas fields, serving as a national-level intelligent manufacturing benchmark project.
Given the significant impact of global trade tensions on China, the localization of critical core technologies has become increasingly crucial. Long-term reliance on foreign databases to manage petroleum exploration and production data poses risks to national energy security. Therefore, deploying domestic database software to support the Smart Oil & Gas Field cloud platform is imperative. By processing diverse massive structured, semi-structured, and unstructured data within the platform—using technologies such as data virtualization, data federation, and efficient real-time data exchange channels between engines, and integrating computational models like complex association analysis, stream computing, graph computing, and batch processing—a big data cloud platform is established. This supports OLAP, OLTP, and NoSQL business scenarios required by the Smart Oil & Gas Field cloud platform.
Since existing traditional data warehouses cannot fully meet the above business requirements, to achieve the stated objectives and better serve the development of the Smart Oil & Gas Field business, it is necessary to enhance the in-depth application of decision support systems, unlock data value, and elevate data analysis capabilities. This requires extending and expanding the current integrated data application and decision support system by adding data analysis functionalities to enable the storage, management, modeling, mining, and analysis of massive, multi-source, heterogeneous data.