DTCC Interview | The Product Engineering Journey Behind the Innovation of GBase's Third-Generation Intelligent Distributed Database Technology
At the 15th China Database Technology Conference (DTCC 2024), GBase shared the innovative practices of its third-generation intelligent database, GBase 8c, in empowering enterprise users to tackle diverse scenarios and drive business development.
After the conference, Zhang Yi, General Manager of the GBase 8c Product Management Department, gave an interview to ITPUB, where he engaged in an in-depth discussion on hot topics in the distributed database field.
1. Could you share the changes in GBase 8c over the past year? What progress has been made?
GBase 8c is a multi-model, multi-form distributed database product launched by GBase. Over the past year, GBase 8c has made encouraging progress in both product capabilities and application depth.
In terms of product capabilities, last year GBase 8c introduced vector storage, a very important capability in the DB4AI domain in the era of large models. With vector storage, GBase 8c can process data more efficiently, particularly for data access driven by large models — a feature relatively rare among Chinese database products.
In terms of application depth, we have put GBase 8c into production in numerous core business systems in finance and telecommunications. For example, in the core credit system of a certain bank, GBase 8c successfully replaced the original database and significantly improved system processing performance. In both the O and B domains of telecom operators (OSS and BSS), GBase 8c helped customers migrate their systems onto integrated appliances, leveraging its multi-tenancy capability to significantly reduce overall O&M difficulty and costs.
2. You mentioned that in the distributed database field, competition revolves around how to achieve engineering and productization, so that products can support applications rapidly and efficiently. Could you elaborate on your engineering practices?
There is a general consensus in both the database academic and industrial communities that the theoretical framework of databases is already relatively mature. Of course, there are some subtle innovations and specific product optimizations that distinguish one product from another. However, for database vendors, standardized productization and engineering to support applications is the key to profitability and continuous iteration.
GBase has always positioned itself as a "database vendor focused on database products and services, with large-scale deployments in critical sectors such as finance and telecommunications." Over the past two decades, we have steadfastly kept our focus on database product R&D. Through this accumulation of experience, we have truly achieved standardized productization and engineering, so that customers can use our products with confidence and thus fully trust our product quality and service levels.
We strictly follow the Integrated Product Development (IPD) model and have established, within our organizational structure, a quality management department independent of the product and sales systems, ensuring that every stage of product R&D meets customer needs and expectations. This systematic R&D process enables us to respond quickly to market changes and provide customers with high-quality products and services.
3. What is your observation of the current attitude and adoption status of distributed databases across various industries? What strategies will you adopt?
When talking about distributed databases, we must also mention centralized databases. The relationship between the two has always been a focus of debate in academic and industrial circles. We believe that distributed and centralized architectures are by no means opposed or mutually exclusive. It’s just that in some scenarios distributed is more suitable, while in others centralized offers better cost-effectiveness. It is not a matter of simply using distributed for core systems and centralized for edge ones. The key is to look at the requirements: concurrency, data volume, and whether high availability capabilities can meet business needs.
For core business systems with large concurrency and massive data volumes, a distributed database may be more suitable, because it can scale out horizontally to boost processing power while better supporting elastic scaling and load balancing. For general business systems that may not require such high availability or data volumes, they can choose a distributed database like GBase 8c, leveraging its multi-tenancy capabilities to reduce management and O&M costs.
4. How do you ensure that the replacement or upgrade to a distributed database is smoother and more stable?
Migration between database products is never completely smooth — even different products on the same architecture have differences, let alone a migration from centralized to distributed. What GBase 8c does is make this differentiation process smarter.
GBase has adopted patented intelligent data distribution algorithm technology. In previous engineering practices, the customer’s business experts and our database team would jointly design a data distribution plan, then perform verification, debugging, and testing — a long and costly process.
Now, based on our self-developed intelligent data distribution technology, we automatically optimize data distribution through algorithms. After an initial coarse-grained data classification, we run tests using real business scenarios and iteratively optimize the system through cost-based evaluation. Through engineering practice, the overall time has been reduced to 20%–50% of the original, greatly lowering the costs of manual intervention and debugging.
5. What are the future plans for GBase 8c?
Our ultimate goal is to evolve GBase 8c into a Data Cloud. Previously, databases in the cloud were defined as an infrastructure component at the PaaS layer. However, when businesses move to the cloud, it becomes difficult to control the database’s disaster recovery level in terms of construction costs and data security. In the future, we want to achieve physical isolation between data and the cloud at the application layer. Data is just data — managed from the physical resource layer. Leveraging our multi-tenancy and high-security capabilities, we aim to achieve high-level high availability while meeting virtualization and cloudification requirements.