CITIC Trust Business System Project

Project Overview

CITIC Trust Co., Ltd. originally used a TDSQL database for its trust business system. It obtained accounting data from client-side leading internet companies such as Ant Group, JD.com, and Meituan. The trust loaded the data into the TDSQL database via FTP, then standardized data formats, performed aggregation, modeling, and sharding. Tables initially partitioned by date were further distributed into 16 final tables using a hash method. On these tables, it carried out accounting processing, indicator calculations, and data preparation for various regulatory reporting tasks.

Solutions

The system in this project runs on a 6-node GBase 8a MPP Cluster, housing 3–5 TB of data.
GBase 8a is a massively distributed parallel database cluster that processes structured data and fits OLAP workloads, handling data queries and analysis. To eliminate performance bottlenecks and single points of failure in the legacy system, an MPP + Shared Nothing distributed federated architecture was adopted, boosting data processing capabilities, accelerating time-to-insight, and enhancing the user experience.

Operational data is transferred to a TDSQL database via FTP, and the data requiring analysis is then loaded into the GBase 8a cluster. There, GBase 8a performs integration, cleansing, and transformation; the refined data is then delivered to upper-layer systems for consumption.

Results

  • Implementation

After deploying GBase 8a MPP Cluster, data computation and reporting in the middle tier were migrated from MySQL to the GBase cluster. Processing time dropped from 3.5 hours to just 1 hour, with results delivered promptly.

  • Benefits and Value

High-speed loading and massive storage: Enables fast loading of large tables with hundreds of millions of rows. A high compression ratio for data ingestion boosts performance and storage capacity, while massive storage simplifies consolidating data from multiple business departments.

Ad-hoc queries with sub-second response: Delivers high-speed ad-hoc queries and range queries on massive datasets, providing stable support for analytical systems.

Exceptional performance: Leverages intelligent indexing, a fully parallel architecture, and transparent compression to deliver blazing-fast query analysis, fully supporting high-performance analytical workloads.

High scalability: The cluster uses a shared-nothing architecture, supporting online, dynamic horizontal scaling on demand without service disruption, to meet your storage and computing capacity needs.