Shandong Unicom Big Data Platform – Resource Integration and Data Sharing
Project Background
The Shandong Unicom big data project involved building a new big data platform to establish foundational data support and consolidate B-domain (business support domain) data. The platform delivers the ability to collect, analyze, and process diverse B-domain data sources, achieve data aggregation and standardization, and provide data services and governance. This strengthens external services and support capabilities. By building the big data platform, Shandong Unicom achieved resource integration and optimization, reduced overall investment, unified data collection and processing, enabled uniform data sharing and services, improved operational efficiency, and maximized data value. The ultimate goal is “centralized storage, unified governance, multi-point applications, and value realization.”
Requirements Analysis
The deployment of Shandong Unicom's big data platform establishes a foundational data support system that enables collection, analysis, and processing of diverse data sources from the BSS domain. It delivers data aggregation and standardization, robust data services, and governance capabilities, significantly enhancing external service and support. This is achieved through the following key requirements:
Big Data Platform Construction: Build a distributed storage and computing platform, incorporating modules for data collection, transformation, loading, real-time processing, near real-time processing, and batch processing.
Data Integration: Consolidate core BSS data by aggregating information from existing network systems—BCV, city branch data pools, front-end processors, data marts, and the cBSS—into the big data platform.
Interface Integration: Unify provincial and group-level data transmission interfaces. Provincial integration aligns interfaces between BSS and systems like Business Analysis, Grid Management, and Customer Service. Group-level integration streamlines interfaces for BSS to Group B-BSS, ECS (Enterprise Customer System), Headquarter CRM, Headquarter PRM, and from Business Analysis to Headquarter Business Analysis.
Platform Application and Management: Share computing and data capabilities across internal systems. Leverage the platform's high storage volume and processing power to enrich customer profiles within Business Analysis. Establish a data quality monitoring platform to oversee data at the collection layer, processing layer, and key metrics, achieving closed-loop data quality management.
System Architecture
The system leverages the unified BDI ETL platform to extract, cleanse, and process data. The cleansed data is then loaded into an MPP distributed database built on GBase 8a MPP. Serving as a central data hub, this MPP platform consolidates data from various business systems and supplies big data services to six vendors and 17 cities, enabling them to run their own business operations on the MPP database. Before scaling, the MPP platform handled a daily data increase of 1.6 TB, with a total data volume of 60 TB across 8 nodes and 3 loading servers. After one expansion, it now operates on 20 nodes and 3 loading servers, with the total data volume reaching 150 TB.
Beneath the unified BDI ETL platform, a Hadoop platform with cloud-native ETL capabilities stores all interface data files. BDI scans every two hours to check if data files have arrived. Once confirmed, it retrieves the data from HDFS to the GBase 8a MPP loading servers and executes loading scripts to insert the data into the database. This enables a seamless integration interface between the BDI Hadoop platform and the MPP platform.
Value Delivered
High Scalability: Leveraging the scale-out capabilities of GBase 8a MPP, we build a distributed computing and storage platform that integrates and consolidates various data sources from the BSS domain, delivering a powerful, scalable data sharing platform for equipment vendors and city-level applications.
High Integration: Through the unified BDI ETL platform and integration with the GBase 8a MPP database, we combine the processing capabilities of the MPP database with Hadoop, achieving a shared pipeline that spans data collection, transformation, loading, and processing.
High Parallelism: With GBase 8a MPP’s columnar storage, smart indexing, and other storage mechanisms designed for big data processing, together with the efficient parallel loading capabilities of the GBase 8a MPP loading engine, we achieve near-real-time data ingestion from various interface data sources into the MPP shared data platform.
High Convergence: By combining Hadoop + MPP distributed computing architectures, the platform’s computing and data storage capabilities are significantly enhanced with scalability, enabling lossless sharing of massive data sets. Leveraging the large storage capacity and powerful processing of the big data platform, this refines customer profiles for business analysis.