General Administration of Customs Golden Customs Project Phase II — Bridging the Past and Future to Advance Customs IT Modernization
Key Benefits
l Solidify Data Foundation and Eliminate Information Silos: First, enrich core data sources across customs operations, break down data barriers between departments, and enable seamless information interconnectivity. Second, in budget and financial management, bridge the current fragmented, standalone systems to link financial and material data, enabling cross-system comparisons and analysis across all customs information systems.
l Drastically Boost Computing Performance: Achieve 10x to 1,000x faster computation compared to legacy systems across various scenarios, meeting real-time query and analytics demands.
l Empower Customs Modernization and Efficiency: Leverage high-performance analytics and computing to overcome the scalability bottlenecks of traditional databases in OLAP scenarios. This provides a technical backbone for strengthening supervision, sharpening macro decision-making, accelerating operational efficiency, and unifying data metrics across customs.
l Simplify Data Pipelines and Reduce Costs: Support multiple analytical applications with a centralized data hub—extract and process once, query from anywhere. This streamlines data workflows and eliminates redundant hardware investments for the same business needs.
l Geo-Redundant Disaster Recovery for High Reliability: Deploy across multiple data centers in different locations to enable disaster recovery and ensure high system reliability.
Solutions
GBase 8a MPP Cluster powers a structured dynamic data warehouse subsystem that stores data from various China Customs systems. Through complex correlation calculations, in-depth analysis, and data mining, it enables data aggregation, model building and execution, and generates specific project tags, indicator libraries, and more. It provides upper-level systems with ad-hoc querying, complex computation, and data mining on massive datasets.
GBase 8a MPP Cluster adopts a Shared Nothing + MPP distributed flat architecture, offering exceptional scalability. This design not only delivers petabyte-scale data storage but also high-performance distributed data processing, achieving sub-second response for complex queries on large-scale data with high concurrency. Furthermore, a cluster-level active-active system enhances data security and elevates disaster recovery capabilities, while the multi-replica mechanism with data redundancy ensures high availability within the cluster itself.
Currently, the dynamic data warehouse subsystem has deployed a total of 124 data nodes across Beijing and Guangzhou, enabling shared underlying data and collaborative upper-layer business operations between the two sites. The Beijing center hosts one 38-node cluster (Information Resource Planning and Sharing Service Platform Data Warehouse), one 14-node cluster (DSS Decision Support System), one 6-node cluster (UDPP Unified Data Processing Platform), and one 2-node cluster (Data Center Data Warehouse). The Guangdong sub-center runs two clusters: a 38-node disaster recovery system for the Information Resource Planning and Sharing Service Platform Data Warehouse, and a 14-node disaster recovery system for the DSS Decision Support System. Additionally, the risk inspection system utilizes 4 nodes, and the tax management system uses 8 nodes. The total data volume reaches 20TB, with a daily incremental data load of 7GB. The Information Resource Planning and Sharing Service Platform Data Warehouse manages over 500 table models, while the DSS Decision Support System handles more than 800 table models.
To achieve high data security, the core systems—the Information Resource Planning and Sharing Service Platform and the DSS Decision Planning System—are deployed in a physical cluster disaster recovery configuration across Guangzhou and Beijing, with shared underlying data sources and collaborative upper-layer business operations. Golden Customs Phase II will adopt a two-site, two-center architecture to support query analysis and OLAP applications. Since OLAP data in Beijing and Guangzhou will be deployed in a cluster disaster recovery mode, synchronizing data between the two locations becomes a critical technical challenge. By leveraging the FTP push functionality of the data loading machines within the MPP database clusters at both centers, data synchronization is achieved, ensuring data consistency between the Beijing and Guangzhou clusters.
The specific data synchronization process is as follows: The Beijing center acts as the primary site, performing data extraction, cleansing, and transformation to generate new data files, which are placed on its MPP cluster’s data loading machine. This machine then uses FTP push to transfer the new data files to the data loading machine of the Guangzhou center’s MPP cluster. As the secondary site, Guangzhou processes the received new data files, thereby completing data synchronization between the two MPP database clusters.
Requirements Analysis
"Golden Customs Phase II" is a continuation and evolution of Phase I. Beyond introducing new technologies, building new frameworks, and solving emerging challenges, it must fully align with Phase I's existing systems and maximize the use of original resources. Therefore, in building the structured dynamic data warehouse subsystem, the aim was to ensure both advanced data processing technology capable of handling massive data with high performance, and system compatibility to remove barriers for data import, integration, and interfacing. For clarity, the requirements can be summarized from the following perspectives:
(1)High performance requirements for the data platform in business scenarios:
lFor large-scale data, loading speed should be ≥ 1 TB/hour;
lUpdate and delete speed should exceed 10,000 rows/second;
lSupport 500 concurrent users with an average response time within 1 minute;
lSupport concurrent read/write access; support joining multiple TB-scale tables and returning result sets of tens of millions of rows.
(2)Comprehensively plan and design customs information resources to provide a unified data platform for system interconnection, integration, and optimization, eliminating data silos and inconsistent metrics;
(3)Fully integrate business data, transforming customs systems from transaction processing-oriented to decision analysis-oriented, thereby enhancing the value of business data;
Project Background
To address the severe protection challenges at customs borders and the urgent need to improve the port clearance environment, the General Administration of Customs launched the "Golden Customs Project (Phase II)" in 2012, building on Phase I to comprehensively advance IT system construction that simultaneously enhances customs’ control and service capabilities.
In the big data era, the project’s central focus is how to fully integrate massive customs datasets, break down departmental information silos, and enable data to flow and circulate across the organization, thereby better serving upper-layer business systems. Technology selection at the data layer must lay a solid foundation, enable rational planning, and establish a forward-looking architecture.
Building a structured dynamic data warehouse subsystem is a key technical method to address these issues and achieve the expected goals. Once established, the system will support numerous statistical analysis applications, including the information resource planning system, customs monitoring and command system, enterprise credit system, anti-smuggling intelligence system, and whole-process logistics visibility system.