GCDW Ops Assistant | Monitoring System
What is GCDW Monitoring System?
The monitoring system serves as the operations support system for GCDW cloud data warehouse, delivering comprehensive monitoring coverage of key components including cloud warehouse services, FDB services, and ELK services. It features real-time data collection and visualization, dynamically tracking health status, resource consumption, and performance metrics, and provides precise alerts when anomalies occur, helping ops teams quickly pinpoint issues. Moreover, the monitoring system enables top-level monitoring and management across all tenants, offering unified, global visualization so that operations teams can centrally control tenant allocation and management.
The system supports the configuration of nearly 100 monitoring metrics. Through a two-tier structure — global and per-tenant — thresholds, collection intervals, and alert rules can be set independently for each metric according to actual business loads and usage patterns. This granular policy management ensures stability for critical tenants while offering flexibility in monitoring strategy configuration. Ultimately, with features such as visualized historical data review and SNMP integration with third-party systems, the monitoring system empowers ops teams to achieve truly fine-grained and observable closed-loop operations, fully safeguarding the reliable running of the GCDW cloud data warehouse system.
Features
Monitoring Policy Configuration
Policy configuration is divided into two levels: global policies and tenant-specific policies, with each metric independently configurable. This allows precise monitoring tailored to different tenants. Each monitoring metric can be set up with the following options:
- Alert Severity: Critical, High, Medium, Info
- Metric Monitoring Status: whether metric collection is enabled (On, Off)
- Metric Alert Status: whether alerting is enabled for the metric (On, Off)
- Threshold: the alert threshold value
- Condition: the judgment condition for threshold evaluation. For numeric metrics: =, <>, >=, <=, >, <. For status metrics: =, <>.
- Consecutive Breach Count: the number of consecutive times the metric exceeds the threshold before triggering an alert. A positive integer from 1 to 50.
- Alert Silence Count: effective only when Continuous Alerting is enabled. The number of alert cycles to suppress before a new alert is triggered. An integer from 0 to 50.
- Notify on Recovery: whether to generate a recovery notification when the metric returns to normal. Options: Yes, No.
- Continuous Alerting: whether to send repeated alerts as long as the breach persists. Options: Yes, No.
- Collection Interval: global metrics do not support per-tenant collection interval settings.
Monitoring & Alerting
The GCDW cloud data warehouse monitoring and operations system offers three core functional modules to help operations teams fully grasp system status.
The monitoring history query displays collected metric data in detail, supports filtering by conditions, and allows access to historical data for trend analysis and issue tracing. The system also provides an automatic data purge function: you can set a retention period for metrics, after which expired data is automatically deleted. This lifecycle management can be flexibly enabled or disabled.
The alert information module presents all generated alert records, supports filtering, and features a clear-alerts function, helping ops personnel quickly remove handled or unnecessary alerts and keep the list clean and effective. Alerts can also be pushed to third-party systems via SNMP for cross-system interoperability.
The operation log module records detailed configuration changes, including modifications to global policies, tenant-specific policies, and connection settings. These complete logs make every change auditable and traceable, supporting troubleshooting and compliance. Together, these modules deliver integrated operations capabilities covering runtime monitoring, alert management, and audit trails for the GCDW cloud data warehouse.
Tenant Management
The GCDW tenant management module provides comprehensive and flexible lifecycle operations. Administrators can quickly locate target tenants via search and filtering.
When creating a new tenant, the following core information is required: Tenant Name (unique identifier in the K8s system), Tenant Password, Company Name, Phone Number, Email, Nickname, Description, and the super admin username and password for the tenant database (similar to MySQL root account). The underlying storage system must also be selected, supporting S3 or HDFS. If S3 is chosen, additional configuration of Access Key (AK), Secret Key (SK), Endpoint, Region, and Bucket is needed. If HDFS is selected, configure HDFS URL and NameNodes, choose the authentication method (simple or Kerberos), and, for Kerberos, upload the conf file and keytab and provide the Principal information.
After creation, you can add or remove Coordinator nodes to elastically adjust compute resources. Tenant information can be edited, and tenants can be deleted. The entire module, with its simplicity and efficiency, meets the fine-grained operations needs of multi-tenant scenarios.
Health Checks
GCDW cloud data warehouse includes a comprehensive health check mechanism covering task management, result viewing, and threshold configuration.
Health check task management supports flexible lifecycle operations. Users can create new tasks, providing the task name, schedule rules (None, Daily, Weekly, Monthly, or Cron custom expression), scheduling status (On or Off), and scope (ELK, FDB, and the tenant information completed on the configuration page).
Once created, tasks can be edited, manually triggered, or deleted when no longer needed.
Health check records track the execution of each task. Users can view detailed results, filter records by criteria, and export health check reports. Single-record deletion and the option to clear all records are provided for easy maintenance. The threshold configuration allows flexible modification of policy information and various thresholds. For status-type thresholds, the alert severity can be adjusted; for numeric thresholds, the threshold value, judgment condition (e.g., greater than, less than), and corresponding alert severity can be modified independently. This fine-grained control over sensitivity and alert behavior helps operations teams keep all GCDW components healthy.
Use Cases
As GCDW has been adopted by commercial banks, the cloud data warehouse often hosts data warehouse workloads for multiple branches and various lines of business. Faced with numerous tenants, diverse workload profiles, and strict regulatory requirements, the monitoring system has become the core tool for ensuring stable operations. The operations team has shifted from reactive firefighting to proactive prevention and control, with the monitoring system truly serving as the “cockpit” for stable cloud warehouse operations.
Independent Monitoring Policy Configuration
Balancing Critical Business and Resource Efficiency
Given that different tenants have varying sensitivity to anomalies, they generally fall into high-sensitivity and lower-sensitivity categories. Highly sensitive tenants are extremely sensitive to certain metrics. We can set dedicated thresholds for these metrics on a per-tenant basis, lowering the thresholds and shortening the collection interval so that anomalies are spotted as early as possible. For less sensitive tenants, we configure more relaxed thresholds and longer collection intervals, and can use additional settings like consecutive breach counts to eliminate sporadic noise. This tenant-level, metric-by-metric independent configuration capability allows us to keep critical business running smoothly while avoiding alert storms from other tenants, achieving fine-grained operations.
Alerts and Historical Analysis
Rapid Troubleshooting and Risk Mitigation
During GCDW operation, the monitoring system triggered an alert: “S3 storage capacity usage exceeded 80%” (synchronously pushed via SNMP to the bank’s unified alert platform). Operations personnel immediately filtered data for that time period in the monitoring history module, combined the alert timestamp logs, and quickly determined that the corresponding tenant’s S3 space was about to run out. They promptly coordinated to add storage for that tenant, timely averting business impact and resolving the system risk.
Full Tenant Lifecycle Management and Health Checks
When onboarding a new tenant, the tenant management module makes creation straightforward. After the tenant goes live, by scheduling an overnight health check task, the module automatically runs the predefined checks in the early morning, generates a report the next day for export and compliance auditing. At the same time, the operation log records every policy modification and tenant change made through the monitoring system, meeting audit requirements.
Summary & Outlook
The GCDW cloud data warehouse monitoring system provides users with an integrated closed-loop operations capability covering infrastructure, tenant management, and health checks. In the future, as the system continues to iterate, we will further explore more intelligent monitoring technologies. We believe it will evolve from a monitoring tool into an intelligent ops brain, providing solid assurance for GCDW’s stable operation.