Sinotruk Group Big Data Platform Replacement and Renovation Project
Project Overview
Project Background
In 2020, Sinotruk Group began building an enterprise-grade big data platform covering sales, services, human resources, connected vehicles, logistics, and manufacturing. The original platform was built using an Oracle + Hadoop architecture. The data warehouse (DW) layer primarily employed the Hive + HDFS offline data processing approach, with some workloads handled by Oracle. The data mart (DM) layer provided data services through Oracle + FineReport and Impala + Kudu + FineReport.
Over time, as application usage, data volumes, and concurrent access on the big data platform increased, querying massive amounts of structured data became a bottleneck. The Quality Department’s queries on 280 GB of indicator data were taking more than 10 seconds to return results, failing to meet business display requirements. A database product capable of handling massive structured data was urgently needed to improve the situation.
Business Requirements
Short-term: Meet the needs of business departments (sales, service, human resources, manufacturing, logistics, connected vehicles, etc.) for handling new workloads on the big data platform.
Long-term: Integrate diverse data sources and leverage real-time stream processing, in-memory computing, multi-tenancy, and container technologies. Through a next-generation converged platform architecture, gradually deliver comprehensive PaaS capabilities. This will drive the evolution from data platform construction to open data operations, fostering a thriving ecosystem of both proprietary and open services.
Construction Requirements
Support seamless migration from the existing platform, enabling a rapid transition from Oracle to a domestic MPP database.
Support storage of at least 10 TB of structured data.
Under a concurrency of at least 200, achieve query response within seconds.
Handle system load with 3,000 monthly active users and 3,000 daily active sessions.
Align with Sinotruk's future big data platform technology roadmap.
Solutions
In the first phase, General Data Technology's GBase 8a MPP Cluster database replaced Oracle to rebuild the structured data warehouse of the big data platform. The project deployed 2 nodes in the first phase, with later expansion by the customer.
The GBase 8a massively parallel distributed database cluster processes structured data, catering to OLAP computational model scenarios, enabling data querying and analysis. Leveraging the distributed computing power of the GBase 8a MPP cluster, it eliminates single points of failure and performance bottlenecks inherent in the original Oracle platform. Through its shared-nothing architecture, it enhances the customer's information processing capabilities, improves the timeliness and user experience of data analysis, and optimizes Sinotruk's big data platform architecture while boosting massive structured data storage and computing power.
Figure 1: Business Architecture Diagram
Application Results
Architecture Optimized: Phase one successfully replaced Oracle databases in Sinotruk’s big data platform, meeting massive structured data storage, analytics, and business support requirements while laying the foundation for further optimization of the big data platform’s technical architecture.
Low Cost, High Scalability: The scalable architecture runs on an x86-based, domestically manufactured server platform, delivering significantly better investment efficiency and long-term planning alignment compared to Oracle’s vertical scaling limitations.
High Performance: Data loading, aggregation, query, and processing speeds improved more than 10× over traditional databases, with storage capacity scaling to the petabyte level.
Ease of Use: GBase 8a provides a unified interface, standard SQL syntax, and comprehensive management and monitoring tools, minimizing the learning curve for developers and operations teams.