CAICT Interview | The New Era of Data Processing: GBASE and Lakehouse Integration
In January 2023, the China Academy of Information and Communications Technology (CAICT) released the results of the 15th batch of “Trusted Big Data” evaluations. General Data Technology Co., Ltd. (GBASE) participated in and passed the assessment for a cloud-native data lakehouse platform. The evaluation was based on the Technical Requirements for Cloud-Native Data Lakehouse Platforms, covering five capability domains: lakehouse data integration, lakehouse storage, lakehouse computing, lakehouse data governance, and other lakehouse capabilities. Recently, Zhang Shaoyong, Chief Engineer of the GBase 8a Product Division, was interviewed by the CAICT Cloud Computing and Big Data Research Institute to discuss what a data lakehouse is, why to build one, its technical features, and deployment practices.
CAICT Cloud and Big Data Research Institute: Mr. Zhang, could you introduce what a data lakehouse is and how it relates to existing data tools like data warehouses and data lakes?
Zhang Shaoyong: A data lakehouse is an organic combination of a data lake and a data warehouse. It is a new architecture that leverages the advantages of both to efficiently process massive enterprise data, including structured, semi-structured, and unstructured data, and to handle both non-real-time batch processing and real-time stream processing. By adopting a storage-compute separation architecture, it consolidates all data into a single, low-cost storage system that supports unlimited scaling. It provides multiple compute engines to meet the performance requirements of various upper-layer applications for batch and stream data, enabling data value mining.
CAICT Cloud and Big Data Research Institute: Why build a data lakehouse, and what are its key technical features?
Zhang Shaoyong: The data lakehouse is a natural evolution of database technology and enterprise big data platform needs. As businesses grow, data volumes increase year over year. To simultaneously handle large amounts of low-value-density data and high-value-density data, enterprises often end up with siloed data processing architectures consisting of a data lake and multiple data warehouses. This increasing complexity has driven the demand for modernization, giving rise to the data lakehouse. Its technical features include at least storage-compute separation, open data formats, and support for multiple computing workloads. Storage-compute separation allows storage and compute to scale independently, supporting virtually unlimited storage and multiple compute clusters in the future. Open data formats effectively open data channels between the lake and warehouse, enabling cross-system data operations. Support for multiple computing workloads meets diverse needs such as batch computing, stream computing, and graph computing.
CAICT Cloud and Big Data Research Institute: What are the application scenarios for a data lakehouse?
Zhang Shaoyong: The data lakehouse architecture evolves naturally with customers’ data businesses. GBASE’s database products have already been adopted at scale in the financial and telecom sectors. Through close collaboration with customers in these industries, we have early insights into how data lakehouses are being applied there:
Financial Industry
At financial institutions, each customer’s data platform typically comprises a data lake, multiple data warehouses, and data marts. Their data processing chains often span the data lake, warehouses, and marts. In such scenarios, it is essential to further improve data processing efficiency. The data lakehouse is the optimal technical approach to address this. It effectively unifies the lake and warehouses, leveraging the unique strengths of each to make enterprise data processing more efficient and resource-effective.
Telecom Industry
In the telecom sector, data lakes are widely used to process data from BSS and OSS domains, transforming low-value-density data into high-value-density data. Meanwhile, data warehouses are used for analytics, correlating high-value-density data to generate insights for decision support systems. Given this landscape, adopting data lakehouse technology in telecom significantly improves data processing efficiency. It provides a single system for all data processing capabilities—unified data integration, storage, computing, scheduling, security, and governance.
CAICT Cloud and Big Data Research Institute: How does GBASE implement a data lakehouse, and what is its architecture?
Zhang Shaoyong: GBASE’s data lakehouse solution is built on its own big data products, including the cloud data warehouse GCDW, the data warehouse GBase 8a MPP, and the data platform GBase UP. GBASE is a specialist data warehouse vendor. The cloud data warehouse GCDW is the core product for delivering data lakehouse solutions. It supports key data lakehouse technologies such as storage-compute separation, extreme elasticity, open data formats, multi-model compute engines, and unified stream-batch processing. This enables unified big data storage, scheduling, language, interfaces, metadata management, and security. It fulfills end-to-end data lifecycle management needs, providing various tools, compute engines, and job scheduling software for different data processing stages—from data collection, integration, and storage to computing, governance, and tiered management—helping enterprises build an efficient data lakehouse platform.