Observing System Resources with Wazuh
Learn to monitor server CPU, RAM, and disk space using Wazuh! Set up custom scripts and rules for proactive system health checks.
This article details a method for monitoring critical system resources—CPU, memory, and disk space—on a server hosting the Wazuh manager service. By leveraging a custom bash script and Wazuh's built-in capabilities, administrators can proactively detect and alert on potential system health issues.
System Resource Monitoring Script
A bash script, metric.sh, was developed to collect system resource utilization. This script identifies the server's IP address, current memory consumption, disk space usage for the root partition, and CPU utilization.
The script's core functionalities include:
- Host Identification: It determines the IP address associated with the network interface that has a route to
8.8.8.8. Alternatively, a hostname can be used. - Memory Usage: It captures the percentage of consumed memory using a command like
free -mh. For instance, a value of68.73indicates that 68.73% of the allocated memory is in use. - Disk Space Usage: It reports the percentage of disk space used on the root partition, as shown by
df -sh. The script is designed to target the root partition but can be modified to monitor other mount points, such as NFS shares, if logs are written elsewhere. - CPU Utilization: It measures the current CPU usage. In a low-activity environment with one Wazuh agent and minimal Elasticsearch load, CPU usage might appear low, such as
0.08.
The collected data is then formatted into a JSON string with fields for host, ram, cpu, and disk. This JSON output is written to /tmp/health.json.
Integrating with Wazuh Manager
To ensure continuous monitoring, the metric.sh script is configured to run every 30 seconds directly by the Wazuh manager. This is achieved by adding a configuration block to the Wazuh manager's configuration file.
The configuration specifies:
- Tag:
metric - Command:
/var/ossec/bin/wazuh-execd /opt/metric.sh(assuming the script is placed in/opt/) - Interval:
30s(seconds) - Run on start:
yes
The output of the script is directed to the /tmp/health.json file, and the Wazuh manager is configured to collect this file as a JSON source. This allows Wazuh to parse the resource metrics.
Detection Rules and Alerts
Custom detection rules are created within Wazuh to trigger alerts when resource utilization crosses predefined thresholds. These rules are designed to parse the JSON data from /tmp/health.json.
Key rule configurations include:
- Log Format: Explicitly set to JSON.
- Grouping Rule: A rule with
signature_id 1000014is used to group all metric health check events. This rule checks for the existence of thecpufield, acting as a wildcard for metric collection logs. - Conditional Alerts: Based on the grouped metric events, specific conditions are evaluated:
- Memory Usage: Alerts are triggered if the
ramfield starts with8(80-89%),9(90-99%), or equals100. The alert message is "Memory usage is high," followed by the actual consumed value. - CPU Utilization: Alerts are triggered if the
cpufield starts with8(80-89%),9(90-99%), or equals100. The alert message is "CPU usage is high," followed by the current CPU utilization. - Disk Space: Alerts are triggered if the
diskfield starts with7(70-79%),8(80-89%),9(90-99%), or equals100. The alert message is "Disk space is running low," followed by the current disk usage percentage.
- Memory Usage: Alerts are triggered if the
These alerts are assigned a severity level of 12, indicating a higher priority. The system can then route these alerts to appropriate channels such as Slack or PagerDuty for timely resolution.
This approach enables Wazuh to monitor not only its own processes but also the underlying server infrastructure, providing a comprehensive view of system health.
Introduction to System Resource Monitoring with Wazuh
Introduction to monitoring server health with Wazuh, focusing on CPU, memory, and disk space as crucial resources. Explains the impact of high consumption on server performance and Wazuh's alerting capabilities.
- Wazuh can monitor system resources on the server where the manager is installed.
- Key resources to monitor are CPU, memory, and disk space.
- High CPU can slow down the server.
- Running out of memory can be catastrophic.
- Running out of disk space prevents Wazuh from writing alerts, leading to missed events.
Creating the System Metrics Bash Script
Details the creation of a bash script (`metric.sh`) to collect CPU, memory, and disk usage. The script outputs this data into a JSON format in `/tmp/health.json`.
- A bash script (
metric.sh) is used to gather system metrics. - The script collects host IP, RAM usage, disk usage percentage, and CPU usage.
- The collected data is formatted into a JSON string.
- The JSON output is written to
/tmp/health.json.
Configuring Wazuh to Run and Collect Metrics
Explains how to configure the Wazuh manager to execute the `metric.sh` script every 30 seconds and collect the resulting `/tmp/health.json` file.
- The Wazuh manager configuration (
ossec.conf) is modified to run the script. - The script is set to run every 30 seconds using the
intervalparameter. - The output of the script is ignored as it's written to a file.
- The
run_on_startoption ensures the script runs when the manager starts. - Wazuh is configured to collect the
/tmp/health.jsonfile usinglocalfileconfiguration.
Creating Wazuh Rules for Resource Thresholds
Demonstrates how to create custom Wazuh rules to detect and alert on high CPU, memory, or low disk space based on the collected metrics in `health.json`.
- Custom rules are created in Wazuh to analyze the
health.jsondata. - Rules are designed to trigger alerts when CPU reaches 80%, RAM reaches 80-90%, or disk space drops below 70%.
- Specific signature IDs are used to group and identify metric health check rules.
- Alerts include the resource name and its current value.
- The severity level for these alerts is set to 12.
Testing and Verifying Alerts
Tests the configured rules by simulating high resource usage and verifying that Wazuh generates the expected alerts in the security events.
- Rule set tests are performed to validate the configuration.
- Simulating high memory usage triggers a 'memory usage is high' alert.
- Simulating low disk space triggers a 'disk space is running low' alert.
- The generated alerts appear in Wazuh's security events, showing the resource values.
- The system successfully detects and alerts on resource threshold breaches every 30 seconds.
