CAICT Interview: GBASE's Perspective on Lakehouse and Its Implementation Path
In January 2023, the China Academy of Information and Communications Technology (CAICT) released the results of the 15th "Trusted Big Data" evaluation. General Data Technology Co., Ltd. (GBASE) participated in and passed the evaluation of a cloud-native lakehouse data platform. The evaluation was based on the "Technical Requirements for Cloud-Native Lakehouse Data Platforms", covering five capability domains: lakehouse data integration, lakehouse storage, lakehouse computing, lakehouse data governance, and other lakehouse capabilities. Recently, Zhang Shaoyong, Chief Engineer of the GBase 8a Product Division, had a discussion with the Cloud Computing and Big Data Research Institute of CAICT, exploring what a lakehouse is, why to build it, its technical characteristics, and how to implement it.
CAICT Cloud Big Data Institute: Mr. Zhang, could you please introduce what a lakehouse is and how it relates to traditional data tools such as data warehouses and data lakes?
Zhang Shaoyong: A lakehouse is an organic combination of a data lake and a data warehouse into a new architecture that leverages the advantages of both. It efficiently handles massive enterprise data, including structured, semi‑structured, and unstructured data, and supports both non-real-time batch processing and real-time stream processing. By adopting a storage-compute separation architecture, it unifies all types of data in a low-cost storage system with unlimited scalability, while providing various computing engines to meet the performance requirements of upper-layer applications for both batch and stream data processing, enabling data value mining.
CAICT Cloud Big Data Institute: Why build a lakehouse, and what are its technical characteristics?
Zhang Shaoyong: The lakehouse is an inevitable result of database technology evolution and the demands of enterprise big data platforms. As enterprises grow, data volumes increase year after year. To simultaneously process large amounts of low-value-density data and high-value-density data, organizations often end up with siloed architectures containing a data lake alongside multiple data warehouses. This growing complexity has driven enterprises to seek reform, giving rise to the lakehouse. Its technical characteristics include at least storage-compute separation, open data formats, and support for diverse workloads. Storage-compute separation allows storage and compute to scale independently, supporting unlimited storage and multiple compute clusters in the future. Open data formats bridge the gap between data lakes and data warehouses, enabling cross-lake and cross-warehouse data jobs. Support for multiple computing workloads meets the demands for batch, stream, graph, and other computing paradigms.
CAICT Cloud Big Data Institute: What are the application scenarios for the lakehouse?
Zhang Shaoyong: The lakehouse architecture evolves naturally with customers’ data businesses. GBASE’s database products have been widely deployed in the financial services and telecommunications industries. Through close collaboration with customers in these sectors, we recognized the lakehouse application scenarios early on:
Financial Services Industry
For financial services customers, each data platform typically consists of a data lake, multiple data warehouses, and multiple data marts. Their data processing pipelines often span these components, making it essential to further improve processing efficiency. The lakehouse is the best technical solution to this challenge, effectively merging the data lake and data warehouse, leveraging the strengths of each to increase efficiency and reduce resource consumption for enterprise data operations.
Telecommunications Industry
In the telecommunications industry, data lakes are widely used to process B-domain and O-domain data, transforming low-value-density data into high-value-density data. Meanwhile, data warehouses are used for analytics, deriving decision-support insights from high-value-density data. In this context, adopting lakehouse technology significantly improves data processing efficiency, delivering all capabilities within a single system: unified data integration, unified storage, unified computing, unified scheduling, unified security, and unified governance.
CAICT Cloud Big Data Institute: Could you talk about how GBASE implements the lakehouse and what its architecture looks like?
Zhang Shaoyong: GBASE’s lakehouse solution is built on its own big data products, including the cloud data warehouse GCDW, the data warehouse GBase 8a MPP, and the data platform GBase UP. As a professional data warehouse vendor, GBASE offers GCDW as the core product that delivers a lakehouse solution. GCDW supports the key lakehouse technologies—storage-compute separation, extreme elasticity, open data formats, multi-model computing engines, and stream-batch unified processing—achieving unified big data storage, unified scheduling, unified language, unified interface, unified metadata management, and unified security. It meets the full lifecycle management needs for all enterprise data, providing the necessary tools, computing engines, and business scheduling software across different phases—from data collection and integration to storage, computing, governance, and tiered management—helping enterprises build an efficient lakehouse data platform.