Beijing Mobile 2020 Independently Developed Analytical Database Expansion Project
Beijing Mobile 2020 Expansion Project for a China-Developed Analytical Database
Project Background
Since its official launch and comprehensive integration in 2004, Beijing Mobile's Business Analysis System has centrally supported the management analysis needs of all departments and branches. In 2020, a new data center was constructed, and the GBase 8a MPP database cluster was expanded to optimize the architecture of the analysis system, aiming to reduce investment costs, increase efficiency, foster diverse application development, and enhance overall system performance.
Solution
Beijing Mobile's data center system adopted a deployment model of PC Server + Linux + local disks, with a scale of nearly one hundred nodes (for the primary data warehouse), dozens (for dedicated databases), and over ten (for the self-service analytics platform). The overall system employed a hybrid architecture of multiple distributed storage processing platforms. Hadoop MapReduce and Hive handled batch processing of massive unstructured and semi-structured data; the GBase 8a MPP Cluster database processed structured massive data (including batch and near-real-time interactive processing). At the application presentation layer, a MySQL database worked alongside the GBase 8a MPP Cluster to handle some interactive processing with applications. The streaming data processing framework, including Streams, MQ, and VoltDB, enabled stream processing and complex event processing to support real-time marketing scenarios. Data transfer speed between the MPP and Hadoop clusters could reach up to approximately 30 TB per hour.
Beijing Mobile Data Center System Architecture Diagram
Within this system, the GBase 8a MPP Cluster served as the primary data warehouse for the entire enterprise data center, undertaking deep data processing and data fusion across the BOM domains—handling the most complex data processing tasks in the entire data supply chain.
The GBase 8a MPP Cluster's data sources primarily originated from upstream systems such as BOSS and CRM, which transmitted data to interface servers. At this point, data was categorized into structured and unstructured data. Unstructured data was processed in batches by Hadoop and then loaded into the MPP for further processing; structured data was loaded directly into the MPP database.
Results
Expanded data processing scope: Fully integrated the operator's B‑domain, O‑domain, and M‑domain data, establishing a data foundation for full-value-chain analysis and enabling multi-dimensional mining from perspectives such as products, customers, resources, channels, and infrastructure.
Extended data lifecycle storage, management, and processing: Supported long-term storage and management of massive data volumes, fulfilling the fundamental requirement of an enterprise data center to empower "big data."
Improved data loading speed: Raw data loading rates reached up to 20 TB per hour, representing a more than 100‑fold performance improvement over the original DB2 database.
Enhanced database operation performance: General statistical query operations saw a performance improvement of more than 1x, while update operations improved by 30% to 50%.
Increased storage space utilization: Data was further compressed to one‑quarter of its original size, significantly extending the overall data lifecycle within the system.
Reduced hardware and software construction costs: Deployment on standard x86 architecture with PC Server and open-source Linux operating system lowered both hardware and software investment costs.