Inner Mongolia Development and Reform Commission Credit Reporting Phase II Project

Project Background

The Phase II Data Center of the Inner Mongolia Autonomous Region Credit Information Platform designs distinct databases for different data types within the data system, forming a Collection & Exchange Database, a Credit Internet Database, and a Credit Private Network Database, which provide a data foundation for data analysis and data services.

While ensuring the credit platform's daily data requests, it supports data service switching between the Collection & Exchange Center database and the Credit Private Network database, and delivers reliable backup and recovery mechanisms for massive historical data.

Solutions

1. System Design

Figure 1: Logical Architecture of the Large-Scale Distributed Parallel Database Cluster

The system is divided into three parts: data ingestion and processing, data storage and computing, and data sharing and application interfaces.

1) Data Ingestion and Processing

An ETL staging database is used to load incremental data from business data sources into the GBase 8a database. The primary business data source (the core database of the Phase I credit information system) is loaded incrementally via the ETL staging database.

2) Data Storage and Computing

The GBase 8a MPP Cluster parallel database system stores data from various systems and provides data computing and processing capabilities such as complex associative queries, deep analytics and mining, data aggregation, and ad-hoc queries. The GBase 8a MPP Cluster adopts a shared-nothing + MPP distributed flat architecture with outstanding scalability, enabling petabyte-scale data storage and high-performance distributed data processing. It delivers sub-second response times for highly concurrent and complex queries over massive datasets. In addition, the cluster’s multi-replica mechanism ensures high availability through data redundancy.

3) Data Sharing and Application Interfaces

Standard interfaces, including JDBC, ODBC, .NET, and C API, are provided to expose data access to upper-layer applications. The system also serves as a standard relational data source for BI tools, ETL tools, and other third-party software. For other systems requiring data sharing, a unified data extraction interface is available.

The GBase 8a cluster consists of one ETL staging server and four cluster computing nodes.

The ETL server and cluster nodes are deployed within the same LAN, interconnected via 10 Gigabit Ethernet to ensure sufficient bandwidth for data transfer.

2. Implementation

Three databases were built for this project:

  • Cloud Computing Data Center Core Analysis Database

The cloud computing center database primarily stores credit information, enterprise credit data, and individual credit data provided by peer departments across the autonomous region. It collects production data in real time from business systems in each province, with an initial designed capacity of 10 TB. GBase 8s is used for this database.

  • Cloud Computing Data Center Query and Analysis Database

The query and analysis database is a dedicated database for querying and analyzing enterprise and individual credit data within the credit information system. It houses enterprise credit libraries, individual credit libraries, public security population data, bank credit data, and housing provident fund data. With an initial capacity of 40 TB, it runs on GBase 8a.

  • Core Analysis Database of the National Development and Reform Commission Data Center

Given the total data volume and application complexity, effective backup and preservation of the overall project data became essential. Therefore, a credit information data center was established in the NDRC computer room. Its architecture is identical to that of the cloud computing center database and also uses GBase 8s. In addition to performing incremental backups of the cloud center database, it serves as a failover site to support applications in the event of an anomaly in the cloud computing center database.

Benefits

  • Enables unified data resource management, elevates data service capabilities, fully unlocks data value, and significantly improves data resource management and credit asset utilization for the NDRC's credit reporting system.

  • Delivers long-term stable operation with 24/7 uninterrupted service, ensuring business continuity for upper-layer systems. Robust backup and disaster recovery safeguard data integrity, eliminating the risk of data loss from failures—strong proof of credit data security on an all-domestic technology platform.

  • The product is mature, feature-rich, and delivers high added value, as demonstrated by:

  1. Speed: 10x–100x query and analytics performance improvement

  2. Storage Savings: 50%–90% reduction in storage space

  3. Cost Savings: 50%–90% hardware and software investment savings, 30%–50% energy savings

  4. Cloud-Native: Supports cloud architecture with horizontal scaling

  5. Full-Text: Integrated full-text search for semi-structured data (cloud files)

  6. Unstructured Data: Structured extraction and conversion from unstructured data

  7. All Data: Unified processing of structured, semi-structured, and unstructured data

  8. Visualization: GBase BI visual analytics platform