This guide details the configuration of Zabbix for monitoring MongoDB instances. Zabbix provides built-in templates for common databases, and MongoDB is no exception. By following the steps outlined, you can effectively monitor your MongoDB deployments.
- Importing the MongoDB Template
Begin by importing the official MongoDB template into your Zabbix environment. The specific steps for template import can be found in the Zabbix documentation.
- Configuring Macros
Once the template is imported, you'll need to configure the necessary macros to allow Zabbix to connect to your MongoDB instances. These typically include:
{$MONGODB.CONNSTRING}: The connection string for your MongoDB instance (e.g.,tcp://192.168.xx.xx:27017).{$MONGODB.USER}: The username for authentication.{$MONGODB.PASSWORD}: The password for authentication.
- Testing the Connection
After configuring the macros, you can test the connectivity from the Zabbix server or agent to the MongoDB instance. Use the zabbix_get utility with the appropriate keys:
# Test general availability
zabbix_get -s <zabbix_agent_host> -k 'mongodb.ping["{$MONGODB.CONNSTRING}","{$MONGODB.USER}","{$MONGODB.PASSWORD}"]'
# Example with hardcoded values
zabbix_get -s mongos.node -k 'mongodb.ping["tcp://192.168.xx.xx:27017","root","your_password"]'
- Considerations for MongoDB 5.0+
For MongoDB versions 5.0 and later, the default Zabbix agent might not be compatible. You may need to install a specific MongoDB plugin or driver for Zabbix agent 2. Ensure you are using a compatible version, such as zabbix-agent2-plugin-mongodb-6.4.14-release1.el7.x86_64.rpm if applicable to your environment.
- Custom Monitoring Scripts (UserParameter)
For more granular or speciifc monitoring needs, you can define custom UserParameters in the Zabbix agent configuration. This involves creating shell scripts that collect the desired metrics.
5.1. Adjusting Zabbix Agent User Permissions
Ensure the Zabbix agent process has the necessary permissions to execute the custom scripts and access required files. This might involve modifying the Zabbix agent service file (e.g., /usr/lib/systemd/system/zabbix-agent2.service) and potentially setting up environment variables for the MongoDB client tools.
# Example: Setting PATH for MongoDB binaries
echo 'export PATH=/data/iuap/middleware/mongodb-34970/bin/:$PATH' >> /home/mongodb/.bash_profile
source /home/mongodb/.bash_profile
5.2. Defining UserParameters
Create a configuration file for your custom parameters (e.g., /etc/zabbix/zabbix_agent2.d/monitor_mongodb.conf) and define your UserParameters. Each parameter will map a Zabbix item key to a shell script execution.
[global]
# Example UserParameters
UserParameter=mongo.alter,sh /etc/zabbix/scripts/jkalert.sh
UserParameter=mongo.connect,sh /etc/zabbix/scripts/jkcur_connect.sh |grep cur_connect
UserParameter=mongo.delay,sh /etc/zabbix/scripts/jkdelay.sh
UserParameter=mongo.alive,sh /etc/zabbix/scripts/jkmongo_alive.sh
UserParameter=mongo.oplog,sh /etc/zabbix/scripts/jkoplog.sh
UserParameter=mongo.replic,sh /etc/zabbix/scripts/jkreplic.sh
These UserParameters can monitor various aspects of MongoDB, such as:
mongo.alter: Basic alerts, errors.mongo.connect: Connection counts.mongo.delay: Replication lag.mongo.alive: MongoDB service status.mongo.oplog: Oplog size.mongo.replic: Replica set health.
5.3. Script Permissions and Ownership
Pay close attention to file permissions. Ensure scripts are executable by the user running the Zabbix agent (often the zabbix user). Also, ensure log files that scripts read have appropriate read permissions.
# Example: Change ownership of scripts to zabbix user
chown zabbix:zabbix /etc/zabbix/scripts/*
# Example: Grant read access to a log file
chmod a+r /path/to/mongodb.log
5.4. Creating Zabbix Items
Finally, in the Zabbix frontend, create new monitoring items associated with the host you want to monitor. Use the keys defined in your UserParameter configuration (e.g., mongo.alive, mongo.connect).
- Example Scripts
Below are example scripts that can be used with the UserParameter method. These scripts typically rely on the mongo shell and a configuration file.
6.1. db_monitor.conf
This file contains configuration variables used by the monitoring scripts.
#!/bin/sh
# Export MongoDB path if not already in system's PATH
export PATH=/data/iuap/middleware/mongodb-34970/bin/:$PATH
export PATH
# MongoDB connection details
PASSWORD='your_password'
USER='admin'
IP='192.xx.xx.xx'
PORT=27017
# Construct the mongo command string (adjust for authentication if needed)
MONGOCMD="mongo ${IP}:${PORT} -u${USER} --port=${PORT} -p${PASSWORD}"
# Log file path and name
LOGPATH=/data/iuap/middleware/mongodb-34970/log
LOGFILE=mongod.log
# Other monitoring thresholds
REP_DELAY='Y'
DELAY_ALTER=1000
CUR_CONNECT_ALERT=500
6.2. jkalert.sh
Monitors MongoDB logs for specific error patterns.
#!/bin/bash
# Define patterns to search for in the MongoDB logs
ALERT_LOG_ALERT="ERROR|Assertion|failed|Error|error|SSL|lastErrorObject|getLastErrorDefaults|SocketException|LogicalSessionCacheRefresh"
# Source configuration and set current directory
curdir=$(dirname "$0")
. "${curdir}/db_monitor.conf"
cd "${curdir}"
# Filter log file for alerts and save to alert.log
cat "${LOGPATH}/${LOGFILE}" | grep -E "${ALERT_LOG_ALERT}" | grep -v 'SSL' | grep -v lastErrorObject | grep -v getLastErrorDefaults | grep -v SocketException | grep -v LogicalSessionCacheRefresh > "${curdir}/alert.log"
# Compare current alerts with previous state to show new errors
if [ -s "${curdir}/alert.log" ]; then
# Use diff to find new entries (this part of the script might need refinement based on exact logging behavior)
errinf=$(diff -a "${curdir}/alert.log" "${curdir}/temp.log" | grep '^<' | sed 's/^/')
fi
echo "$errinf"
# Update temp.log with current alerts for the next run
cat "${curdir}/alert.log" > "${curdir}/temp.log"
6.3. jkcur_connect.sh
Checks the current number of active connections.
#!/bin/bash
# Source configuration and set current directory
curdir=$(dirname "$0")
. "${curdir}/db_monitor.conf"
cd "${curdir}"
# Get current connection count from serverStatus
echo "print('cur_connections#', db.serverStatus().connections.current)" | ${MONGOCMD} > "${curdir}/cur_connections.log"
CUR_CONNECT=$(cat "${curdir}/cur_connections.log" | grep cur_connections | awk '{print $2}')
# Log connection count over time (optional)
echo "$(date)" >> "$curdir/tmpfile/connect_number.log"
echo "$CUR_CONNECT" >> "$curdir/tmpfile/connect_number.log"
echo "$CUR_CONNECT"
# Trigger alert if connections exceed threshold
if [ "$(awk -v a=${CUR_CONNECT} -v b=${CUR_CONNECT_ALERT} 'BEGIN{print(a>b)?a:b}')" = "${CUR_CONNECT}" ]; then
alert_info="cur_connect:${CUR_CONNECT}"
echo "$alert_info"
fi
6.4. jkdelay.sh
Monitors replicatino lag between primary and secondary nodes.
#!/bin/bash
# Source configuration and set current directory
curdir=$(dirname "$0")
. "${curdir}/db_monitor.conf"
cd "${curdir}"
# Get replication delay information
echo "rs.printSlaveReplicationInfo()" | ${MONGOCMD} | grep 'behind the primary' > "${curdir}/cur_delay.log"
while read hang; do
cur_delay=$(echo "${hang}" | awk '{print $1}')
# Log delay over time (optional)
echo "$(date)" >> "$curdir/tmpfile/cur_delay.log"
echo "$cur_delay" >> "$curdir/tmpfile/cur_delay.log"
# Trigger alert if delay exceeds threshold
if [ "$(awk -v a=${cur_delay} -v b=${DELAY_ALTER} 'BEGIN{print(a>b)?a:b}')" = "${cur_delay}" ]; then
alert_info="Current delay is:${cur_delay} second"
echo "$alert_info"
fi
done < "${curdir}/cur_delay.log"
6.5. jkmongo_alive.sh
Checks if the MongoDB instance is running and accessible.
#!/bin/bash
# Source configuration and set current directory
curdir=$(dirname "$0")
. "${curdir}/db_monitor.conf"
cd "${curdir}"
# Attempt to run a simple command like rs.status()
alert_info=$(echo "rs.status()" | ${MONGOCMD})
# Check the exit status of the mongo command
if [ $? -ne 0 ]; then
alert_info='The mongoDB instance is down!'
fi
echo "$alert_info"
6.6. jkoplog.sh
Monitors the remaining size of the operation log (oplog).
#!/bin/bash
# Source configuration and set current directory
curdir=$(dirname "$0")
. "${curdir}/db_monitor.conf"
cd "${curdir}"
# Get oplog length
oplog_length=$(echo "rs.printReplicationInfo()" | ${MONGOCMD} | grep "log length" | awk -F: '{print $2}' | awk -Fs '{print $1}')
# Trigger alert if oplog length is below a critical threshold (e.g., 2 hours = 7200 seconds)
if [ "${oplog_length}" -le 7200 ]; then
alert_info="The oplog length is ${oplog_length}S"
echo "${alert_info}"
fi
6.7. jkreplic.sh
Checks the health status of all nodes in a replica set.
#!/bin/bash
# Source configuration and set current directory
curdir=$(dirname "$0")
. "${curdir}/db_monitor.conf"
cd "${curdir}"
# Get replica set status and filter for unhealthy nodes
alert_info=$(echo "rs.status()" | ${MONGOCMD} | grep 'health' | grep -v '"health" : 1')
# Trigger alert if any node is reported as unhealthy
if [ -n "$alert_info" ]; then
alert_info='The replica set has an unhealthy node. Please investigate.'
echo "$alert_info"
fi