Data Management Expert: Big Data
On June 26, 2019, General Data Technology Co., Ltd. (GBase) held its 2019 Partner Conference under the theme “Phoenix Reborn, a New Voyage” at the Millennium Hotel Beijing. More than 100 customers and partners attended, including the State Information Center, General Administration of Customs, State Grid, Agricultural Bank of China, CAAC Information Center, Huawei, Digital China, Ronglian, H3C, and Weiwang Technology. Many attendees left the event highly inspired and eager for deeper insights, prompting us to share this in-depth introduction to GBase’s big data products today.
In large-scale data processing, GBase 8a MPP clusters have been deployed with over 6,500 total nodes, managing more than 100 PB of data. The solution is widely used across key sectors including finance, telecommunications, government, security, and defense, serving users in 17 countries and all 32 provincial-level administrative regions across China.
The GBase 8a MPP Cluster V9 released at the conference (hereafter “V9”) represents the culmination of GBase’s years of big data expertise, delivering world‑class complex data processing capabilities that align with enterprises’ next‑generation big data platform requirements for AI, Big Data, Cloud, and embedded Devices (ABCD).
V9 is a massively parallel processing (MPP) cluster designed for heterogeneous data processing. Its core value can be summarized in four aspects: virtual cluster support for ultra‑large‑scale deployments, heterogeneous data fusion, in‑database AI, and data security technologies.
Virtual cluster support for ultra‑large‑scale deployment
V9 leverages virtual cluster technology to achieve resource isolation and data sharing in large‑scale cluster deployments, and introduces the industry’s first cluster mirroring capability. This dramatically improves manageability of massive database clusters and, for the first time in the industry, enables MPP clusters to scale beyond 1,000 nodes.
Virtual cluster technology enables coarse‑grained resource partitioning (at the node level), providing physical resource isolation between tenants while maintaining unified management and allowing data sharing across tenants.
Concept
A virtual cluster is a resource isolation mechanism that physically partitions a large cluster into multiple logical sub‑clusters. Each logical sub‑cluster can independently plan and scale its cluster size and computing resources according to the storage and computation requirements of different workloads.
Virtual clusters provide all logical sub‑clusters with a unified access interface, unified metadata view, unified resource management, unified execution scheduling, and unified authentication and authorization.
Virtual clusters provide capabilities for data exchange, data migration, and data correlation between clusters.
Virtual clusters support cluster mirroring, with real‑time table‑level synchronization between mirrored clusters, enabling active‑active data real‑time processing and T+0 high availability.
Use cases
Virtual clusters are ideal for environments where multiple independent clusters need to be deployed and managed under a single system, with each cluster serving independent business areas.
Virtual clusters include data management, user management, and cluster version management.
Transparent data migration, data correlation, and data sharing are supported between all logical sub‑clusters of a virtual cluster.
Value
Virtual clusters are suitable for big data platforms, integrated BI systems, data warehouses, and data mart systems that contain relatively independent business domains or different types of analysis. Different application scenarios run in separate logical sub‑clusters, all under unified management. This solves the high cost of managing, monitoring, and maintaining multiple physical clusters, while accommodating the differentiated characteristics of various business scenarios. It maximizes resource utilization and enhances both the scalability and maintainability of the cluster.
Heterogeneous data fusion
V9 introduces the logical data warehouse concept, enabling heterogeneous data fusion. The logical data warehouse goes beyond structured data to include unstructured data such as video, audio, and documents. Logically, it functions as a large data warehouse while underlying it can encompass various data sources for correlated processing. Regardless of whether data resides on‑premises, in the cloud, on a device, or anywhere else, it can be correlated without moving the data to a specific location. Through correlation, it automatically discovers data, uses machine‑driven awareness to identify valuable data, determines data value, analyzes data, automatically applies appropriate security measures, shares data, and optimizes data.
V9 spans OLTP, OLAP, NoSQL, and other large‑scale structured and unstructured data processing technologies. Through data virtualization and data federation, it establishes efficient data exchange channels between engines, integrating complex correlation analysis, stream computing, graph computing, and batch processing models. The result is a big data platform product that is outwardly unified and inwardly extensible, capable of adapting to diverse business scenarios and serving as essential infrastructure for enterprise big data platforms.
In‑database AI
V9 provides users with the GBase Machine Learning Library (GBMLlib) — an in‑database machine learning algorithm library that enables deep data mining capabilities.
In‑database big data analysis
Eliminates the need to move data from the database to external analysis nodes via API or ODBC.
Various analysis and mining algorithms run natively inside the database as UDF/UDAF functions.
The value of in‑database big data analysis
Scheduling is done through the database execution plan, fully leveraging the parallel computing resources of the distributed database.
Data movement is minimized, reducing network and disk I/O overhead.
Data security technologies
V9 further strengthens data security, offering column‑level encryption and decryption that allows users to encrypt data based on the sensitivity level of individual data fields. The encryption/decryption process is transparent to users, happening automatically in the background, with a performance impact of less than 5%.
Dynamic data masking technology ensures that, without altering the actual stored data, authorized users see plaintext while unauthorized users only see transformed, non‑plaintext data.
GBase 8a MPP has already powered world‑class big data platforms for hundreds of high‑end users across finance, telecommunications, government, security, and other sectors. Tempered by years of intense business pressure, the product’s functionality, performance, stability, and reliability — as well as the capabilities of our R&D, technical support, and after‑sales service teams — have reached world‑class levels. We remain wholeheartedly committed to providing the optimal solutions and the best service to every user, and to becoming the most trusted data management expert for customers in China.