GBASE Perspective: Cloud-Native Data Warehouse Selection Focuses on Performance, Monitoring and O&M, Multi-Cloud, and Other Factors

Published on 2024-05-23

With the rapid development of cloud computing technology and the increasing demand for enterprise digital transformation, cloud-native databases have drawn growing attention from more and more enterprises. IT168 & ITPUB have launched the topic "Cloud-Native Database Selection Guide," conducting surveys and interviews with frontline experts to understand the current development status, core technical features of cloud-native databases, as well as the pain points, challenges, and practical experiences in their adoption across various industries, and to identify the key factors that enterprises and organizations focus on when selecting cloud-native databases, for industry reference.

Recently, Guan Lianpo, General Manager of the GBase 8a Product Business Department at General Data Technology Co., Ltd. (GBase), accepted an interview with ITPUB, introducing the company's definition of cloud-native databases, as well as cloud-native data warehousing application scenarios and key factors for enterprise selection.

 

What is a True Cloud-Native Database?

Cloud-native databases are the future trend of databases and a currently hot topic. However, the industry has not yet reached a unified definition of cloud-native databases.

According to Baidu Baike, a cloud-native database is a type of cloud-native data infrastructure—a database service that fully leverages the advantages of the public cloud, featuring ultimate elastic scalability, serverless capabilities, a globally-distributed architecture with high availability and low cost, and seamless integration with other cloud services.

According to Frost & Sullivan's "2023 Top 10 China Cloud-Native Database Vendor Recommendations," cloud-native databases are those designed architecturally around the characteristics of cloud computing infrastructure, fully leveraging cloud-based computing, storage, and network resources to achieve enhanced performance and expanded functionality.

Guan Lianpo pointed out that there are inconsistent definitions of cloud-native databases in the market. For example, some believe that simply migrating a database to the cloud makes it cloud-native, while others consider any database service offered by a cloud provider as cloud-native; such understandings are somewhat one-sided.

The defining characteristics of the cloud are large-scale, flexible, shared, and fully elastic. A cloud-native database must enable fully elastic scaling of all resources, support large-scale usage, and offer flexibility and convenience in deployment and use—only then can it be called cloud-native.” said Guan Lianpo. Previously, compute and storage resources were deployed together in a single box (hardware). The cloud can virtualize compute and storage separately, so a cloud-native database must support the virtualized use of both compute and storage resources.

The separation of storage and computing can be considered the foundational premise of a cloud-native data warehouse. Traditional MPP data warehouses have a coupled storage-compute architecture; when scaling compute resources, data needs to be redistributed and migrated. Even when deployed on the cloud, they cannot achieve flexible elastic scaling, which has become a bottleneck hindering further development. Cloud-native data warehouses such as Snowflake and GBase GCDW adopt a decoupled storage-compute architecture, decoupling storage and computing to fully leverage the flexible elastic advantages of cloud-native, enable pay-as-you-go models, and represent a breakthrough in data warehouse technology.

Guan Lianpo further pointed out that, taking GBase GCDW as an example, the management of compute, storage, and metadata resources can be fully elastic, and many components are fully containerized for easy operations and management. These are some technical indicators for judging a cloud-native database.

Overall, from the perspective of supply and demand, databases as a middleware layer must evolve in response to changes in upstream application requirements and underlying infrastructure. They need to be re-architected at the architecture and kernel levels based on cloud-era storage, compute, and network resources, to fully harness the advantages of the cloud. Simply moving a database to the cloud does not make it a cloud-native database.

 

Cloud-Native Data Warehouse Application Scenarios and Requirements

Guan Lianpo explained that as businesses evolve rapidly, higher demands are placed on database scalability, elasticity, and operations management. This may lead to the adoption of cloud-native data warehouses. Requirements vary across different businesses. According to his observations, the main scenarios for cloud-native data warehouses include:

First, agile business scenarios requiring elastic resource scaling with high stability demands. For example, in the financial industry, report calculation and precise regulatory reporting must meet compliance requirements without any latency; batch processing and report generation times must be strictly controlled. Since cloud resources are shared and prone to contention, to avoid performance fluctuations impacting batch runs, cloud data warehouses are often deployed using fully isolated bare-metal servers within the cloud.

Second, analyst business and real-time analytics for internet finance scenarios, which require full data sharing and on-demand use of storage and compute resources. This is a relatively typical cloud-native data warehouse scenario. Particularly, agilely developed ToC (To Consumer) businesses are well-suited for cloud-native data warehouses.

Third, government cloud business, where timeliness requirements are not as high as financial reporting. Government cloud environments heavily rely on office automation systems; using a cloud-native data warehouse is more friendly for agile application development, offering greater flexibility and elasticity.

Through exchanges with customers, Guan Lianpo found that, after years of cloud computing development, database migration to the cloud has been accepted by most head enterprises. However, critical industries like finance remain relatively cautious. Although they are interested in cloud-native data warehouses, most are still in a wait-and-see or trial phase.

These financial customers have some concerns, including the actual adoption of cloud-native data warehouses by other large banks. Additionally, cloud-native data warehouses differ from traditional MPP architectures in schema design, algorithms, and operations. For example, containerization brings significant differences in log collection and viewing. The industry needs to cultivate more personnel with combined skills in both cloud and data warehouse product operations to meet the staffing requirements of cloud-based infrastructure.

 

Key Selection Factors: Performance, Monitoring & O&M, Cost, and Multi-Cloud

Database selection has never been an easy task. Guan Lianpo explained that when enterprises and organizations select a cloud-native data warehouse, they mainly consider the following factors:

First, monitoring and O&M — problem localization. When issues occur, many customers care about distinguishing whether the problem lies with the cloud or with the data warehouse. This requires more detailed metric monitoring in the cloud-native data warehouse. In traditional data warehouse technology stacks, the warehouse runs on operating systems and hardware—a relatively mature and reliable environment with proven monitoring capabilities and low failure rates. In a cloud-native data warehouse, network, CPU, memory, and storage are all virtualized, increasing technology stack complexity and making problem localization more challenging. Once network fluctuations occur, rapid problem identification becomes critical.

Second, version maintenance. Traditional MPP data warehouses are deployed directly on physical environments, with full isolation between businesses. For example, in a bank, each set of business applications is deployed separately, making upgrades and maintenance relatively simple. However, after moving to the cloud, all components are in a flexible state. As a unified cloud-native data warehouse with many shared components, the question of how to upgrade components as business requirements change arises—this relates to the degree of product standardization. GBase GCDW has the capability of grayscale online upgrades, enabling non-disruptive upgrades.

Third, performance — whether it can guarantee no performance degradation with the same resource allocation and even deliver better performance.

Fourth, cost management. In traditional deployment models, resource consumption costs on existing hardware are relatively controllable and easy to assess. On the cloud, how to evaluate and keep costs under control becomes a concern for enterprises.

Fifth, after full containerization, the issue of container volatility — whether it can be resolved.

Sixth, multi-cloud support. To avoid cloud lock-in and mitigate risks while leveraging the strengths of different clouds, enterprises often adopt a multi-cloud strategy. They will focus on whether the cloud-native data warehouse has been adapted to mainstream clouds. Different clouds grant different permissions for external components, requiring significant adaptation work at the database level.

Guan Lianpo pointed out that a data warehouse should not dictate the infrastructure; the cloud is just another form of infrastructure. As a cloud-neutral database vendor like General Data Technology (GBase), on one hand, it must adapt to mainstream cloud providers to meet enterprises' multi-cloud strategic needs. At the same time, the data warehouse should break free from the constraints of all infrastructures, including clouds, abstracting away the underlying complexity. Whether it is a single cloud or multiple clouds, public, private, or hybrid cloud, or even traditional hardware deployment, it should be able to support them all to cater to various business scenarios.

Cloud-native databases are part of the overall cloud ecosystem. Their future development requires the joint efforts of the entire cloud ecosystem—cloud infrastructure, databases, and applications.